OpenAI's recent release of hundreds of claimed solutions to difficult mathematical problems has not fully met the standards outlined by an advisory group of mathematicians. The Advisory Group on Mathematics and Artificial Intelligence (AGMAI), hosted by Princeton University’s Institute for Advanced Studies, had issued guidelines for frontier labs solving math problems in late September.
OpenAI's approach diverged from AGMAI's recommendations, particularly regarding the need for human understanding of mathematical results. The advisory group's first request was to stop testing advanced mathematical problems on proprietary models, yet OpenAI's release explicitly states it is evaluating its proprietary models using open research problems.
While OpenAI did follow some principles, such as releasing results promptly and including information on how models reached conclusions, it did not adhere to all. Only 10 of the 719 manuscripts included the model's chain of thought, and 42% of the proofs had not undergone formalization, a process suggested for papers that people do not understand.
Concerns have also been raised about the translation of natural language proofs into formal code. A new paper by mathematicians from the University of Cambridge and King’s College in London documents at least two discrepancies between the natural language proof and the Lean code for an OpenAI solution to a problem derived from the Navier-Stokes equations.
Mathematicians, including Terence Tao, have criticised the lack of human understanding of AI-generated solutions at the point of release. AGMAI had suggested that OpenAI should help fund human mathematicians to make the lab's solutions meaningful.