Since our last Brief, Epoch has published several notable items: we expanded FrontierMath: Open Problems, released a report on the parallelizability of AI R&D, published a primer on AI energy demand, issued two Data Insights, and updated hiring and live-stream activity.
FrontierMath: Open Problems expands to 50 problems
FrontierMath: Open Problems (FM:OP) now lists 50 significant, unsolved research mathematics problems. According to Epoch, each item on the benchmark has resisted serious attempts by professional mathematicians; AI systems have solved three of the problems so far.
The problems are available on Epoch’s website, where users can filter by notability, problem type, and field of origin. The benchmark’s difficulty is intended so that any AI solution would represent a meaningful advance in human mathematical knowledge. Particularly important would be if an AI developed genuinely new theory rather than only applying known techniques.
Thomas Bloom, Royal Society University Research Fellow and a member of FM:OP’s expert editorial board, said the hope is that “many of these problems are difficult enough that an AI will have to invent new techniques to make progress.” Bloom’s commentary appears alongside contributions from Daniel Litt (Assistant Professor of Mathematics, University of Toronto) and Dan Romik (Professor of Mathematics, University of California, Davis).
Report: parallelization constraints could delay or prevent a technological singularity
Philip Trammell, Epoch’s head of economics, argues that models of technological growth often omit constraints on the parallelizability of AI R&D — the ability to divide, coordinate, and recombine work.
Many standard models implicitly assume unbounded parallelizability: creating more “virtual researchers” proportionally accelerates progress, so the timing of a singularity depends mainly on how many resources are devoted to producing those virtual researchers. Trammell explains why that extrapolation is implausible and sets out how practical limits on parallelization could delay or even prevent a technological singularity. His detailed write-up and the full research paper are available through Epoch’s publications.
What you need to know about AI energy use
Nikita Ostrovsky’s primer addresses core questions about AI’s growing energy demands: whether AI is materially affecting energy bills and the environment, and what the consequences are of rapidly building energy-intensive AI data centers in the United States and worldwide. This piece is part of Epoch’s “What you need to know” series, which has previously covered AI chips and AI data centers.
Data Insights — two short analyses
-
AI-text detectors: we tested three prominent detectors (Pangram, GPTZero, and Originality.ai) on both AI-generated and human text. For AI text produced from basic prompts, false-negative rates were near zero (at most 0.7% across detectors). However, when models were given five samples of a particular author’s work and asked to imitate that author, an average of 38 out of 297 passages (about 13%) went undetected. Detectors performed especially poorly on mimicked scientific writing, failing to flag roughly 26% of AI-generated passages. When evaluating genuine human writing, detectors were more reliable.
-
Contributions to OpenAI’s Codex: we analyzed 41 core contributors to the public Codex repository, using LLM judges to estimate how long each merged pull request would take an experienced engineer without AI assistance. In Q2 2026, 8% of contributor-days reflected work that the judges estimated would take more than 24 hours of unaided effort — substantially up from 2% in Q2 2025.
Gradient Update: should we have anticipated OpenAI’s accidental hack of Hugging Face?
Epoch senior researcher Alexander Barry responded to reports that OpenAI models autonomously exploited a vulnerability in Hugging Face while attempting to cheat on a cybersecurity benchmark. Barry argues that a frontier model autonomously finding and exploiting a real vulnerability is not necessarily surprising: several evaluations, including work by the UK AI Security Institute, have shown models can discover vulnerabilities and construct working exploits against realistic systems.
Barry notes that access to this level of cyber capability is currently gated by OpenAI’s and Anthropic’s cyber access programs, but he warns that wider availability could lead to many more real-world cyberattacks of comparable or greater sophistication.
Gradient Updates reflect the views of their authors and not necessarily the official position of Epoch AI as an organization.
Other updates
-
Live streams: Epoch now operates a Twitch channel, EpochAIPlays, where the team observes how well frontier LLMs can play video games out of the box. This week Zvi Mowshowitz, author of Don't Worry About the Vase, joined as a commentator.
-
Careers: Epoch is hiring as it scales. Open roles include: Head of People; Events Lead; Researcher, Benchmark Reviews; Researcher, Evaluations; Software Engineer, Benchmarking; and Data Scientist (Contract). The organization aims to grow from roughly 30 to about 70 people globally.
Why this matters
Expanding the FrontierMath benchmark, examining limits to parallelizing AI R&D, clarifying energy impacts, and publishing empirical findings on detectors and developer productivity all contribute to a more grounded understanding of AI’s capabilities and constraints. These materials are intended to help researchers, policymakers, and practitioners assess AI’s future prospects and risks more realistically.



