Epoch AI has launched MirrorCode, a long-horizon coding benchmark co-developed with METR, designed to probe the largest software projects an AI can complete autonomously. MirrorCode tasks models with rebuilding 25 real-world programs spanning bioinformatics tools, Unix utilities, cryptography components, interpreters, and more. Models have no access to the original source code and no human-in-the-loop assistance is allowed.
The benchmark provides a much larger inference budget than typical software engineering (SWE) benchmarks: many existing SWE benchmarks limit inference to roughly $1–$10 per task and runs that last minutes or at most hours, whereas some MirrorCode tasks require orders of magnitude more resources. One of the largest MirrorCode tasks cost $2,600 for a single run and involved the model working for 19 days without human intervention. MirrorCode estimates that the hardest tasks would take an unaided human engineer weeks to months to complete.
As of the release, Claude Opus 4.7 leads other models on the benchmark with a 56% solve rate, indicating substantial headroom for improvement. Full results and analysis are available in the release.
Financial insight: hyperscaler capex may outpace operating cash flows by end of 2026
Senior researcher Isabel Juniewicz finds that the world’s largest hyperscalers — Microsoft, Amazon, Alphabet, Meta, and Oracle — are increasing cash capital expenditures faster than cash inflows from operations. Many hyperscalers have already turned to external financing to fund growing investments in AI infrastructure or are considering doing so. Juniewicz's analysis suggests that, on current trends, capex could overtake operating cash flows by the end of 2026.
Gradient Updates: two new short analyses
Epoch published two new Gradient Updates in which researchers and guest authors present more opinionated or informal takes on major questions in AI progress. Gradient Updates reflect the views of their authors and do not necessarily represent Epoch AI's official position.
What we learned from 1,604 Chinese AI job postings
Cheryl Wu, JS Denain, and Anson Ho scraped more than 1,600 job postings from six major Chinese firms to better understand their strategies. Their analysis finds that, similar to US firms, Chinese labs do not follow a single playbook and display distinct “personalities” and strategic differences.
Toward an O*NET for AI R&D
How close is AI to automating AI research and development? Joe Kwon, together with Epoch researchers Jean-Stanislas Denain and Anson Ho, propose a finer-grained tool: a taxonomy of more than 60 tasks involved in frontier AI research. They argue that existing economic tools for tracking automation are too blunt to determine which parts of R&D are automatable and that a research-specific taxonomy would provide sharper measurement.
Other updates and hiring
Epoch AI is hiring across several fully remote roles:
- Senior Product Designer and Product Designer: convert complex research into intuitive, engaging, high-impact designs, dashboards, and visualizations.
- Researchers and Senior Researchers: lead new projects across expanding teams.
- Data Scientist (Contract): support AI research through technical literature review, benchmark tracking, and analysis of AI models, data centers, and companies.
Applications are rolling, so interested candidates are encouraged to apply soon.



