Mathematics used to be AI’s safest trophy case: benchmarks, Olympiad medals, and problems with known answers.
Well, OpenAI says internal Astra, its next major model, generated ten results on long-standing geometry, cryptography, quantum-computing, and pure-math problems. Some settle conjectures; others (supposedly) improve best-known limits.
Here's what happened
- Astra made the first improvement since 1978 to a major high-dimensional sphere-packing limit: how tightly equal objects fit in many dimensions.
- It constructed a “non-sofic group,” an object researchers wondered existed, disproving Connes’s rigidity conjecture.
- Successful solution tokens would cost roughly $2,000 at Sol API rates, OpenAI says.
- Humans prepared papers with Astra; the model converted every proof into Lean, whose checker verifies each logical step.
- OpenAI published a 249-page paper, Lean files, and model-written reconstructions of the ideas’ development (paper, reasoning walkthrough).
In his first analysis, Gary Marcus applauded Astra but warned that checkable math does not prove universal scientific reasoning. Math offers right-or-wrong feedback and endless synthetic practice; cancer research, military strategy, and most real-world decisions do not.
Then Marcus’s follow-up: Anthropic mathematician Levent Alpöge said public Claude Fable reproduced roughly half the results within 24 hours with a generic prompt, no internet, and full autonomy. That needs an apples-to-apples public comparison but weakens Astra’s singular-threshold claim.
Why this matters
The breakthrough may exceed one secret model.
Frontier AI appears to be entering a reusable research loop: Humans select verifiable problems; models search huge idea spaces (with not-literal but close gigawatts of compute); proof software checks answers. That could compress years of mathematical trial and error into days… and humanity benefits.
Our take
The missing number is the denominator. Noam Brown acknowledged OpenAI tried other major problems unsuccessfully, but it has not disclosed how many, how much human guidance it received, or the total cost of failed attempts. The next credible test is a pre-committed open-problem set where every failure counts.
That said, if Astra keeps this hit rate, AI has not “solved” science per se but has permanently changed how some science gets done. Someone pray for the academic paper reviewers who have to verify all this stuff…
Editor’s note: This article originally appeared on our sister publication, The Neuron.


