In brief
- OpenAI published 722 math manuscripts in 372 result families on GitHub from an unreleased internal model.
- Only 162 of the 722 papers have a Lean-formalized main result, and OpenAI warns that some unformalized results could have issues.
- MIT’s Andrew Sutherland says the one-prompt, single-agent claim is unverified until the model is released, and the release omits the prompts that an advisory group at the Institute for Advanced Study recommended disclosing.
OpenAI published 722 math manuscripts on GitHub on Tuesday, all produced by an internal model the company has not released. An OpenAI spokesperson said almost everything came from a single prompt handed to a single AI agent, though some may have taken multiple attempts.
It’s a bold claim and a potentially significant breakthrough in the field of mathematics. But not everyone is a fan, or buying the hype.
“Until and unless they release the model and people can replicate their results, I think you should treat any claims about one-shotting problems with a single agent as unverified,” Andrew Sutherland, a mathematician at MIT, told Scientific American. “We should ask for receipts,” he said.
The papers are grouped into 372 “families” of related results, and a family can bundle a main theorem with companion arguments, consequences or alternative proofs. That makes 722 a count of manuscripts, not of solved problems. OpenAI says it posed roughly 4,000 problems to the model and kept the outputs it judged significant enough to publish.
The average result used the equivalent of roughly three hours of ChatGPT Pro thinking compute, per OpenAI. The Navier-Stokes claim last month looked very different, with 10,000 coordinating agents working for 88 hours.
OpenAI released abridged reasoning summaries for 10 of the results. That said, only 162 of the 722 papers come with a computer-checked main result, according to a formalization catalog in the repository. That is about 22% of the collection, translated into Lean, software that checks every logical step mechanically.
OpenAI itself says not all manuscripts have Lean formalizations and that “some of the unformalized results could have issues.” In other words, a lot of what they published could be wrong.
A passing Lean check confirms only that the proof follows from the statement as written in Lean. It does not show that the statement matches the original problem, or that the result is new or important, which is the part mathematicians now have to judge.
And this is where researchers raise their eyebrows.
I was trying to read the OpenAI proof that chromatic number of a plane is >= 6. But it is totally unbelievable alien math?
Somehow the model found that any K-coloring <=> “weakly measurable” K-coloring, which seems out of nowhere
1/2 pic.twitter.com/W13ScT3nel
— Dmitry Rybin (@DmitryRybin1) October 7, 2026
The openai/math repo has Issues turned off and has never accepted a pull request. That’s disappointing. If you publish 722 manuscripts and ask for Lean formalisations, you need somewhere for people to send them.
I’m formalising OpenAI’s Saxl’s Conjecture proof in Lean 4 against…
— Keith Adler (@keithadler) October 7, 2026
“It is now the case that AI can output mathematical arguments in situations without the human who prompted it being able to understand the arguments, verify them, or take responsibility for them,” The Institute for Advanced Study in Princeton, New Jersey, said in a statement. “We believe that human understanding of mathematics remains of paramount importance. How, in this new era, can we work towards a new paradigm that includes human understanding of mathematics as part of responsible scholarly output?”
Others, though, like Professor Abhishek Saha, are pretty excited. “It is a very big day for mathematics,” he wrote, but noted that most of the problems fit in the categories of “exceptional advances within an existing program” of “surprising breakthroughs.”
This means most of the problems in the set are interesting, but not impossible or game changing like the millennium problems. That spot is reserved for exactly one problem out of the 722: the Quasi-Riemann Hypothesis.
Some further thoughts on the 372 results released by OpenAI today, across 722 manuscripts.
If I were to classify theorems that mathematicians prove and publish according to their groundbreaking nature, I would (very roughly) divide them into four categories:
A) Non-breakthrough… https://t.co/BA10SxBlx7
— Abhishek Saha (@ObhishekSaha) October 7, 2026
The release also falls short of what an advisory group at the Institute for Advanced Study recommended on September 29: the model name, the prompts, a summarized chain of thought, the time taken and the compute cost for every result. OpenAI published average compute figures and 10 reasoning summaries but no prompts, and says it is still working to release the model responsibly.
Daniel Litt, a mathematician at the University of Toronto, took the opposite view, arguing there is no reason to ask the company to keep the answers to these math questions secret.
Anthropic took a different route with its Lean-checked Fermat’s Last Theorem proof last month, posting all 13 million lines publicly on GitHub. That proof formalized a theorem Andrew Wiles published in 1995, rather than claiming new results.
OpenAI says it will add Lean formalizations as it obtains them; for now, 162 of the 722 manuscripts have one.
Daily Debrief Newsletter
Start every day with the top news stories right now, plus original features, a podcast, videos and more.


