A multi‑institution research team tested whether current frontier AI agents can originate genuinely novel research ideas that would advance the field of AI. Participants in the project included Princeton University, Cornflower Labs, UK AI Security Institute, University of Toronto, UC Berkeley, Georgetown University (CSET), Johns Hopkins University, the Golden Gate Institute for AI, AI Digest, and Stanford University.
Method: shadow evaluation using unpublished NeurIPS 2026 submissions
The study used a “shadow evaluation” design. Researchers partnered with the authors of two papers submitted to NeurIPS 2026 that were not yet public and extracted the central research questions. They then tasked a well‑resourced frontier agent—Claude Opus 4.8 running in the OpenClaw harness—with answering those questions. The original paper authors graded the agent outputs as if they were conference submissions. This approach is conceptually similar to the earlier “First Proof” experiment that tested whether AI systems could contribute to ongoing, unpublished mathematical problems.
The two research lines attempted by the agents were:
- the structure and controllability of LLM personas (referred to here as the Personas experiment), and
- designing a distribution shift detector for tabular foundation models (a TabPFN‑style task).
Findings: strong engineering, weak creative research
The investigators found that while agents could complete the engineering tasks required to run experiments and produce implementations, they failed to deliver original research at the standard expected at top machine‑learning conferences.
Concrete outcomes:
- The human authors rejected both agent-produced papers. The Personas paper received a score of 2 ("Reject"), and the TabPFN paper received a score of 1 ("Strong Reject").
- Reviewer comments emphasized the same shortcomings: poorly motivated choice of data and experiments, no novel contribution, and opaque prose.
Observed failure modes included early commitment to a narrow set of research paths, insufficient responsiveness to synthetically generated feedback intended to improve experimental design, and difficulty abandoning unpromising approaches to explore alternatives.
Why this matters for timelines of automated AI development
These results matter for forecasts about recursive self‑improvement and the pace at which AI could automate its own advancement. If current AI agents lack the capacity for tasteful, intuitive scientific creativity—despite their engineering competence—then scenarios that assume rapid, fully automated leaps in capability become less likely or at least more uncertain.
The findings are consistent with earlier work (for example, Anthropic’s experiments) showing that automating parts of scalable oversight or research often requires human researchers to supply strong priors or promising directions; without that seeding, agents may make technical progress but fail to explore sufficiently creative ideas to drive large performance gains.
Conclusion
This shadow evaluation suggests that contemporary, well‑resourced AI agents are excellent at engineering tasks needed to carry out research but do not reliably originate novel, high‑quality scientific contributions on their own. That distinction is significant when assessing the speed at which AI might become capable of substantially automating its own research and development.
The authors cite the arXiv paper “Can AI agents conduct open‑ended AI research? Early evidence from two case studies” as the detailed report of this work.



