Research

AI-generated text

Nine benchmark-driven questions Epoch AI uses to probe the societal and economic effects of advancing AI

Epoch AI’s research team outlines nine broad questions they consider central to understanding how advances in AI capabilities will affect economies and society.

Nine benchmark-driven questions Epoch AI uses to probe the societal and economic effects of advancing AI

This piece is drawn from a post in Epoch AI’s Gradient Updates newsletter; the views presented reflect the authors’ opinions and not necessarily those of Epoch AI as a whole. The author lays out nine high-level questions they consider central to how AI capability improvements will translate into real-world economic and social effects. These questions motivate much of Epoch’s benchmarking work: the goal is to build benchmarks and perform analyses that help answer them.

1. Can AI do my job?

More specifically: can AI move from narrowly scoped tasks to handling messier, open-ended jobs? Epoch’s polling shows that when people use AI at work, they typically use it for parts of tasks rather than whole jobs. A shift toward end-to-end task completion by AI would presage larger labor-market disruption and sustained fast revenue growth for model developers.

Most benchmarks so far focus on narrow tasks (bug fixing, math problems, report writing). Some benchmarks extend this: MirrorCode asks for full large software packages to be implemented from scratch in a structured setting. The Remote Labor Index extracts real freelancing projects and has humans grade AI outputs against human references. A concrete, real-world experiment is Andon Café, a real café run by an AI agent — the sort of case study-style benchmarking that may become more common.

2. Is AI making progress where we expect big economic impact?

The “ChatGPT moment” arrived when AI could hold coherent conversations across topics; the “Claude Code moment” when AI could credibly tackle many coding tasks. We can’t predict the next big moment precisely, but we can identify capabilities likely to unlock near-term impact and benchmark them.

A salient domain today is cybersecurity: when will AI be able to exploit many vulnerable systems? Cybersecurity benchmarks were already signaling risks even prior to the Hugging Face incident. Another domain is computer use: if AI becomes fast and reliable at using computers (not just chat interfaces), that could unlock further adoption. Physical industries are also worth tracking — not necessarily robotics per se, but whether AI can guide a less-skilled technician through repairing a broken factory machine.

3. How consistent is the gap between frontier and trailing models across domains?

This question covers open vs. closed weights, U.S. vs. China, and frontier U.S. firms (OpenAI, Anthropic) vs. near-frontier (xAI, Meta). The economic question is whether leading developers can capture enough value to justify heavy infrastructure investment. Geopolitical questions appear in the US/China frame.

Closed-weight developers could capture value if the open/closed capability gap is larger than it seems — for instance if open-weight teams concentrate scarce resources on a few central domains (like coding) and neglect a long tail of economically valuable domains. An extreme is “benchmaxxing,” where models are optimized to score well on benchmarks even while underperforming on the real capabilities those benchmarks intend to measure.

Well-designed benchmarks can shed light on these dynamics, though they may not fully resolve them. Even small capability gaps can translate into large value capture under strong winner-take-all dynamics.

4. Why are benchmark scores so correlated?

Epoch Capabilities Index (ECI) combines many benchmark scores into a single general capability measure because benchmark results tend to be highly correlated across domains. Interpreting the single underlying dimension that statistical analysis reveals is important.

One explanation is that AI teams work to improve every benchmark independently (for example, by acquiring domain-specific training data) with no deeper unifying cause. A different possibility is a general capability factor — analogous to human IQ. Some correlations are notable: before it saturated, the logarithm of METR’s Time Horizons measurement correlated strongly with ECI. Another candidate driver is the maximum context length over which models can sustain coherent reasoning.

ECI growth is especially useful for detecting acceleration in capability progress. If ECI can be shown to map more directly to real-world impact (e.g., marginal ECI points translating to revenue), interpreting acceleration would be simpler.

5. Can AI do AI R&D?

This classic recursive self-improvement question asks whether automated AI R&D could produce runaway capability growth. A comprehensive suite of AI R&D benchmarks would serve as a leading indicator for such dynamics.

AI R&D is complex and automating it likely requires many skills. Benchmarks could cover many parts of the process or focus on a single metric (or on the ability to generate larger-grained research outputs). Two key benchmarking challenges are realism — the most consequential R&D happens inside frontier firms that are partially opaque — and cost, since realistic-scale R&D uses expensive resources (e.g., many GPUs) that are costly to provision for benchmarking.

6. Can AI learn on the fly?

Related to continual learning: can an AI improve at a task through repeated attempts, typically via context management rather than weight updates? If so, pre-deployment testing may not bound capabilities, especially for economically valuable or dangerous tasks.

Epoch’s EBR-bench probes this by having AI repeatedly play a long, strategically rich, campaign-style board game. AI isn’t terrible at the game out of the box, but so far it has shown little if any positive learning trajectory from repeated play. The benchmark would reveal a continual-learning capability if one were to emerge.

7. What are the returns to inference scaling?

A close relative of the prior question asks whether AI can solve any task if it is allowed to ‘‘think’’ long enough. Experiments tend to show at best logarithmic gains, meaning required scale may be impractical for many tasks. However, capabilities that currently cost an impractical number of tokens may become feasible with falling inference costs or further training.

Two challenges for benchmarking are: we don’t know the optimal way to scale inference compute (serial thinking time, multi-agent architectures, or specialized harnesses may differ in efficiency), and returns may vary by domain. Some domains may be effectively hill-climbable with steady logarithmic gains, while others plateau early.

8. How well does reinforcement learning generalize across domains and out of distribution?

Capability gains likely owe something to more and better training data. But to what extent does that fully explain improvements? Are systems weak on tasks far from their training distribution?

Not in the strongest sense: AI systems have improved on a benchmark of “puzzles” drawn from an undisclosed game (analogous to chess puzzles), despite it seeming unlikely they were post-trained on that game. MirrorCode also shows near-equal performance in low-resource languages versus high-resource languages. Nevertheless, generalization is less well understood the further one moves from training data and for tasks that are harder to verify. It remains possible that models have crossed a threshold and are ‘‘getting better at everything at once,’’ but this is uncertain.

Benchmarks that truly lie far from training distributions, without confounding factors that explain poor performance, could help determine the extent of generalization.

9. Can AI generate "new ideas"?

Human technological progress depends on people coming up with and testing novel ideas. AI could plausibly increase the rate and value of new ideas, producing unprecedented impact — though other bottlenecks (for example physical testing) might slow adoption.

So far, we do not have a ‘‘data center of geniuses.’’ Mathematical innovation is the area closest to this, but mathematicians’ assessments suggest that AI math insights so far resemble ideas humans could and do produce. Benchmarking can help detect departures from this pattern. FrontierMath: Open Problems is explicitly designed to collect problems likely to require what humans would recognize as genuinely new ideas, so that any AI solution would merit further human scrutiny.

Conclusion

Epoch AI’s nine questions highlight where benchmarks and empirical measurement can most illuminate how AI capability improvements translate into economic and societal effects. The author argues for a broad mix of benchmarks and repeated real-world case studies to better understand labor impacts, generalization, scaling returns, AI-driven R&D, and whether AI can produce genuinely novel ideas. If these questions interest you, Epoch AI invites applications to join their work.