Anthropic says one of its unreleased Claude models stumbled into a mathematical breakthrough while trying to solve a problem that has baffled researchers since 1859.
The company said an internal research version of Claude was asked to “take a real stab” at the Riemann hypothesis, one of the Clay Mathematics Institute’s Millennium Prize Problems that carries a $1 million reward for a proof.
Claude did not prove the hypothesis. Instead, Anthropic says the model discovered a better lower bound for a closely related question involving the zeros of the Riemann zeta function. According to the company, the model increased the known lower bound from 41.6% to 67.2%.
Anthropic wrote, “We don’t expect that the techniques Claude used will lead to proving the Riemann hypothesis. But its work serves as the latest example of the speed of progress in AI models’ mathematical capabilities.”
The result was reviewed by Anthropic mathematicians Levent Alpöge and Ralph Furman, and the company said it also shared the paper with outside experts Brian Conrey and Dan Goldston. Claude also produced a formally verifiable proof using the Lean proof assistant.
How the model reached the result
Anthropic said the research model worked across two Claude Code sessions and generated 31 million output tokens. It first explored 650 unsuccessful ideas before coordinating about 60 subagents that ran thousands of numerical checks, executed 2,400 shell commands, and wrote hundreds of Python scripts.
The work was initiated by an Anthropic staff member without significant mathematical training, who simply prompted the model to “take a real stab” at the hypothesis and then allowed it to continue working autonomously for about a day and a half.
Anthropic said the breakthrough came from combining several existing research frameworks developed by mathematicians including Baluyot, Goldston, Suriajaya, Turnage-Butterbaugh, and Bombieri.
The claim is likely to attract attention because it suggests AI systems may be becoming more useful as research collaborators rather than just problem solvers. The model did not replace decades of mathematical work; it built on that work and produced a result that experts could inspect and verify.
Anthropic itself emphasized that the achievement is narrower than solving the Riemann hypothesis. The company presented it as an unexpected byproduct of a failed attempt at the larger problem.
What this could mean next
If the proof holds up under broader mathematical scrutiny, the result could become another sign that advanced AI models are moving beyond answering questions and toward participating in the research process itself.
That matters far beyond mathematics. Researchers across science and technology spend enormous amounts of time testing ideas, checking calculations, writing code, and eliminating approaches that go nowhere. Claude’s experiment suggests AI systems could increasingly take on parts of that exploratory work, helping human researchers investigate more possibilities faster.
Important limits remain. Anthropic’s result has not gone through the normal academic peer-review process, and Claude did not solve the Riemann hypothesis. But for anyone watching the evolution of AI, the larger question is whether systems like Claude are becoming capable not just of finding answers, but of helping produce genuinely new knowledge.
Related reading: For another look at how governments are putting AI models to the test, see how Japan is benchmarking its homegrown LLMs against Claude and Amazon Nova in a major public-sector AI pilot.


