Prime Intellect reported its largest autonomous AI research experiment to date: 18 frontier models were placed in isolated, offline sandboxes for eight days and ran 153 fully autonomous research trials. The test focused on a nanoGPT optimizer speedrun, where the objective is to reach a fixed validation loss in as few training steps as possible.
What happened
- Over an eight-day period the models operated without human intervention in disconnected sandboxes.
- In total, they executed 153 fully autonomous experiments on the nanoGPT optimizer speedrun challenge.
- Human engineers hold a record of 2,600 training steps for this task, achieved after months of iterative work.
- Model performance varied: according to Prime Intellect, Fable 5 closed 82% of the gap toward a 2,726-step result; Opus 5 reached 54%; and GPT-5.5 managed 8%.
- Crucially, none of the models invented a technique that was not already present in the published literature.
Real-world implications
The experiment directly tested a claim made by Jack Clark, co-founder of Anthropic, in May: that genius is one percent inspiration and ninety-nine percent perspiration, and that AI consuming the perspiration could self-improve by 2028. On the ‘‘perspiration’’ side—the repetitive, grind-heavy work of running experiments, abandoning dead ends, and re-testing failed approaches—the top model reproduced about 82% of a human-built record. On the ‘‘inspiration’’ side—the single novel idea that produces a breakthrough—the models scored zero: every technique they used was drawn from existing publications.
Takeaway
The results suggest current autonomous AIs are exceptionally tireless research assistants that can close a large portion of the gap to human engineering effort, but they do not (yet) produce genuinely novel scientific or methodological ideas. In other words, we have automated much of the labor of research, but not the decisive creative leap that turns labor into discovery.



