Research

AI-generated text

Anthropic's Claude Rapidly Raised the Proven Proportion of Zeta Zeros on the Critical Line

Anthropic ran an unreleased model of Claude on the Riemann zeta zeros and significantly increased the proportion provably on the critical line, from 41.6% to 67.2% in about 36 hours.

Anthropic's Claude Rapidly Raised the Proven Proportion of Zeta Zeros on the Critical Line

Anthropic ran an unreleased Claude model on the Riemann zeta zeros. This did not produce a proof of the Riemann Hypothesis — which remains unproven after 167 years — but it moved a numerical benchmark that human mathematicians had only nudged slowly.

Over the past 37 years human progress had raised the proportion of nontrivial zeta zeros proven to lie on the critical line (Re(s)=1/2) by just 0.8 percentage points. Anthropic's Claude increased that proportion by 25.6 percentage points in roughly a day and a half, from 41.6% to 67.2%.

How the model worked

The breakthrough was not presented as a flash of mathematical insight but as sustained computational effort. Several components used in the approach were already known: for example, Hugh L. Montgomery's results from 1973 and work by Goldston and Suriajaya that identified an exact missing step in certain strategies. Humans had not previously combined these parts successfully.

According to Anthropic, Claude carried out hundreds of experimental attempts: 650 failed runs, the involvement of 60 separate agents, and about 31 million tokens processed — roughly half of which produced no useful outcome. Where human teams tended to stop after around thirty dead ends, the model continued calmly to the 650th attempt.

Significance and limits

This result is not a proof of the Riemann Hypothesis. The technique used may still fall short of yielding a general proof. Nevertheless, the contrast — 37 years of modest human progress versus a 36-hour computational campaign — highlights a shift: some future field-defining advances may come not from who out-thinks a field, but from who outlasts it computationally.

Across 2026, AI systems have repeatedly set records on diverse mathematical problems, from Erdős-style questions to decades-old conjectures, often on a near-monthly cadence. Two years ago, ranking AIs by olympiad-style medals on problems with known answers was a common benchmark; now AI is setting records in areas where answers were previously unknown.

Closing note

The Claude outcome emphasizes that AI's role in mathematical research can lie as much in endurance and large-scale trial-and-error as in singular creative insights.

Cited figures (from the reported account)

  • Human progress over 37 years: +0.8 percentage points.
  • Claude's improvement: +25.6 percentage points in ~36 hours (41.6% → 67.2%).
  • Experiment scale: 650 failed attempts, 60 agents, ~31 million tokens, about half nonproductive.
  • Referenced prior work: Montgomery (1973); Goldston and Suriajaya (identification of the missing step).