Key facts
- OpenAI published 722 math manuscripts on GitHub from an unreleased internal AI model.
- Most results were reportedly generated from a single prompt to a single AI agent.
- Only 162 of the 722 papers have a Lean-formalized main result.
- Some unformalized results may have issues, according to OpenAI.
- Mathematicians are demanding the model's release to verify the claims.
- The Institute for Advanced Study stated that human understanding of mathematics remains paramount.
OpenAI has released 722 math manuscripts on GitHub, generated by an unreleased internal AI model. The company claims that most of these results were achieved with a single prompt given to a single AI agent, though some may have required multiple attempts. This announcement has been met with skepticism from the mathematical community, who are demanding the model's release to verify the claims.
Mathematicians like Andrew Sutherland of MIT have stated that claims of "one-shotting problems with a single agent" should be treated as unverified until the model can be replicated. The published papers are grouped into 372 families of related results, with 722 manuscripts in total. OpenAI stated it posed approximately 4,000 problems to the model and retained the outputs it deemed significant.
According to OpenAI, the average result used compute equivalent to about three hours of ChatGPT Pro thinking. However, only 162 of the 722 papers, about 22%, come with a computer-checked main result formalized in Lean software. OpenAI itself acknowledges that some unformalized results could have issues, implying potential errors. A Lean check confirms proof validity but not necessarily the originality or importance of the result, which mathematicians must now assess.
The Institute for Advanced Study in Princeton expressed concern that AI can now output mathematical arguments that humans cannot understand, verify, or take responsibility for, emphasizing the continued importance of human understanding in mathematics. Professor Abhishek Saha, however, expressed excitement, categorizing most problems as "exceptional advances within an existing program" rather than "game changing" breakthroughs, with the Quasi-Riemann Hypothesis being a notable exception.
An advisory group at the Institute for Advanced Study had recommended disclosing the model name, prompts, chain of thought, time taken, and compute cost for each result. OpenAI provided average compute figures and 10 reasoning summaries but has not released the prompts, citing responsible release efforts. Daniel Litt of the University of Toronto argued there is no reason to keep the answers secret.
