Research

Epoch launches EBR-bench, finds little evidence of learning-from-experience; publishes new AI capability and vulnerability analyses

Epoch AI introduced EBR-bench, a new benchmark that measures whether models improve on a complex board game through repeated play and reports little sign so far that models learn from experience.

Epoch launches EBR-bench, finds little evidence of learning-from-experience; publishes new AI capability and vulnerability analyses

Epoch AI’s latest briefing introduces EBR-bench, two short Data Insights, a new Gradient Update, and a list of open roles as the organisation expands.

Model evaluations: EBR-bench

EBR-bench is a new benchmark designed to test whether AI systems can improve by learning from repeated experience. The benchmark uses a complex board game called Earthborne Rangers: models play the game repeatedly so evaluators can observe whether their performance improves and whether they learn from past mistakes.

Epoch’s current findings show little evidence that the evaluated models substantially improve through repeated attempts. EBR-bench is intended to be a continuing tool to detect if and when models begin to demonstrate learning-from-experience.

Expanding benchmarking work

EBR-bench is one of several recent additions to Epoch’s benchmarking suite. Two weeks earlier, Epoch launched MirrorCode, developed jointly with METR, which asks models to autonomously write code for weeks at a time to rebuild real-world programs from scratch — some tasks involved programs with tens of thousands of lines of code. The best model so far scores 56% on MirrorCode.

Epoch has also increased the number of benchmarks it tracks: nine were added last month and another thirteen were added most recently. These cover agentic work, cybersecurity, algorithm engineering, forecasting, and research-level physics. Seven of the latest additions feed into the Epoch Capabilities Index (ECI), the organisation’s aggregate measure of model capability.

Data Insights: two new short analyses

Epoch published two new Data Insights that highlight notable trends in concise form.

  • GPT-4 led the ECI far longer than any other model

    OpenAI’s GPT-4 topped the Epoch Capabilities Index for roughly a year after its release in March 2023, and no subsequent model has held the lead for as long. The second-longest leadership period, held by OpenAI’s o1, lasted a little over three months — less than a third of GPT-4’s run.

  • Cyber vulnerability disclosures spiked around Claude Mythos Preview

    Researcher Luke Emberson reports that notable organisations disclosed about 1,500 high- and critical-severity CVEs in June, more than 3.5× the previous monthly record before the release of Claude Mythos. The spike follows Anthropic’s April announcement that Claude Mythos Preview could autonomously discover software vulnerabilities and that partners in the company’s Project Glasswing program had already been using it to find and fix bugs before the model’s public release.

Gradient Update: the missing half of AI futurism debates

In the latest Gradient Update, senior researcher JS Denain and researcher Anson Ho argue that many bold predictions about automated AI research rapidly enabling advanced technologies (such as nanotech, Dyson swarms, or near-light-speed spacecraft) rest primarily on continued advances in model capability and often do not include careful analysis of how difficult those futuristic technologies would be to build. They propose applying exploratory engineering with explicit assumptions about AI capabilities to ground such predictions more rigorously.

Gradient Updates represent the authors’ views and do not necessarily reflect the views of Epoch AI as a whole.

Other updates: careers

Epoch is hiring across research, engineering, and operations. Open roles mentioned include:

  • Finance Specialist / Manager — to run accounting and finance operations
  • Researcher (Benchmark Reviews) — to develop and publish critiques and reviews of AI benchmarks
  • Researcher (Evaluations) — to evaluate frontier models on hard-to-grade tasks
  • Software Engineer, Benchmarking — to build and maintain benchmarking infrastructure
  • Talent Scout — to help find and recruit exceptional people
  • Senior Product Designer — to lead UI/UX and data visualization
  • Data Scientist (Contract) — to assist with literature review and data analysis

Applications are rolling and Epoch encourages candidates to apply soon.