Key facts
- OpenAI released 719 claimed solutions to complex math problems.
- Only 10 of OpenAI's 719 manuscripts included the model's chain of thought.
- 42% of OpenAI's released proofs had not undergone formalization.
- A paper documented discrepancies between OpenAI's natural language proof and its Lean code for a problem derived from the Navier-Stokes equations.
Mathematicians have expressed concerns that OpenAI's recent release of hundreds of claimed solutions to complex mathematical problems falls short of established standards for rigor and human understanding. The Advisory Group on Mathematics and Artificial Intelligence (AGMAI), a group of prominent researchers, stated that OpenAI did not fully adhere to its guidelines, particularly concerning the formalization of proofs and the transparency of the AI's reasoning process.
AGMAI, hosted by Princeton University's Institute for Advanced Studies, released guidelines at the end of September emphasizing the need for human understanding of AI-generated mathematical results. One of their primary requests was to cease testing advanced mathematical problems on proprietary models, a guideline OpenAI appears to have disregarded by evaluating its proprietary models on open research problems.
While OpenAI did release results promptly and included some information on how models reached conclusions, only ten of the 719 manuscripts provided the model's chain of thought. Furthermore, only 42% of the released proofs had undergone formalization, a process suggested by AGMAI for proofs that are not easily understood by humans. AGMAI also suggested that OpenAI should help fund human mathematicians to make the AI's solutions meaningful.
Terence Tao, a mathematician critical of OpenAI's approach, noted on social media that AI prompters solving problems autonomously may lack interest in the broader field and may not fully understand the AI's output. This issue is highlighted by a paper from the University of Cambridge and King's College in London, which found discrepancies between OpenAI's natural language explanations and the formal Lean code for a problem derived from the Navier-Stokes equations. The authors of this paper concluded that such auto-formalized proofs should not be trusted without the same peer review and scrutiny as human-generated proofs.
Melanie Wood, a mathematics professor at Harvard University, commented that when AI models release solutions, there is an initial lack of human understanding, and the work of interpretation begins afterward, contrasting with the process of human discovery where responsibility and engagement with the community are immediate.
