OpenAI reports that an internal, unreleased model and an agent-based workflow produced a purported solution to the Navier–Stokes existence and smoothness problem, one of the Millennium Prize Problems. The company says it began a targeted evaluation on September 1 and reached a resolution on September 5, with formal verification completed on September 6. Mathematicians Tristan Buckmaster (New York University) and Levent Alpöge (currently at Anthropic) dispute elements of the timeline and raise concerns that their months-long work, including draft sessions stored in Codex, may have influenced model behavior or training.
Timeline and key events
- Tristan Buckmaster and Levent Alpöge say they worked on related problems for nearly a year and frequently used Claude and Codex (mainly GPT‑5.6 Sol). They report a breakthrough on August 15 and subsequently published a rapid account of their results.
- According to OpenAI, on Tuesday, September 1, after hearing rumors that two Millennium Prize problems had been resolved, they launched an internal effort to evaluate their improved model on all open Millennium Prize problems and several other high‑impact problems.
- OpenAI states their agents reached a resolution in the Navier–Stokes case on Saturday, September 5, about 88 hours after the first agents were launched; a lean formalization and verification step took an additional 17 hours using GPT‑6 Astra, concluding on September 6.
Computational scale reported by OpenAI
OpenAI reported that across all attempted problems the agents exchanged approximately 4.9 million messages and produced around 300 billion output tokens. For only the Navier–Stokes resolution, they report 2.7 million messages and roughly 130 billion output tokens. Using public API price levels for GPT‑6 Astra, OpenAI notes that 300 billion output tokens would correspond to an estimated $15,000,000, although the company has not disclosed the exact cost structure of their internal models.
Points of contention
- Buckmaster asked OpenAI when the first prompt had been sent; he says the question was not answered directly for some time and was later told it had been sent in the days after information about Buckmaster and Alpöge's work reached OpenAI.
- Buckmaster and Alpöge say they kept drafts and session material in Codex. He asked whether the OpenAI model had been trained on, or accessed, those sessions. OpenAI maintains that their researchers and agents did not see the duo's work until it was publicly released, and that no specific user data was accessed to solve the problem. However, OpenAI also wrote that they cannot completely rule out that de‑identified data derived from users’ use of their products may have helped improve their models.
- OpenAI reportedly offered to coordinate a concurrent release acknowledging priority, or to have Buckmaster author a paper about OpenAI’s result, but indicated that Levent Alpöge would not be invited as a co‑author because of OpenAI’s competitive relationship with his employer, Anthropic.
OpenAI’s technical note and differences
OpenAI emphasized that their proofs and approach differ significantly from the Buckmaster–Alpöge approach, and that precise results diverge in related Euler cases (forced versus unforced settings). They frame their effort as a rapid internal evaluation spurred by rumors and by a perceived step change in their internal model's performance.
Why this matters beyond this case
The dispute raises broader questions about:
- Data use and transparency: What does it mean in practice when companies say user interactions are used to improve model performance? How can researchers be confident their in‑progress work did not influence a later model’s capabilities?
- Research ethics and competition: When corporate groups mobilize extensive compute to reproduce or preempt unpublished academic work, how should priority, authorship and fair collaboration be handled?
The situation has been compared to security research dynamics, where a rumor of a vulnerability can prompt automated systems to discover and exploit it quickly; here, a rumor of an unpublished solution may motivate rapid, high‑cost model runs aimed at reaching the same result first.
Documentation and next steps
Buckmaster published a quick PDF describing his and Alpöge's work and the chronology from their perspective. OpenAI has published its own account with technical details and resource usage numbers. The technical and ethical questions raised by both parties — especially around whether and how user data contributed to model improvements — remain subjects for further clarification and community discussion.
Conclusion
Beyond the substantive claim about Navier–Stokes, this episode spotlights how interactions between academic researchers and commercial AI labs can create conflicts about data provenance, model training, and authorship. The case underscores unresolved issues in transparency and governance as powerful models are applied to frontier mathematical research.
(Reporting here relies on the parties' public statements and the technical figures they have shared.)



