OpenAI releases 377 mathematical results, raising questions about review

OpenAI has published hundreds of results from an internal AI model tackling open mathematical problems. Researchers welcome access to the work but are debating how independently the proofs have been checked and whether the model drew on mathematicians’ ideas.

The Sydney Morning Herald reported 377 new results spanning algebra, number theory, theoretical computer science, mathematical logic and topology. OpenAI posted the work to GitHub on October 6. The release includes formalizations of many proofs in Lean, a programming language that lets computers check the logic of proofs.

The company also shared 10 summaries of the model’s reasoning and information about the computing resources used. OpenAI said the average result used computing equivalent to about three hours of ChatGPT Pro thinking. The model that produced the work remains internal and has not been publicly released.

Researchers are discussing how independent the findings are from existing human work. Tristan Buckmaster of New York University, who has worked on the Navier-Stokes problem, said some proofs may have extended ideas from mathematicians using AI. He also questioned whether such a large collection of results could have been properly checked so quickly.

OpenAI said it consulted an independent group of mathematicians at the Institute for Advanced Study in Princeton. The group recommends publishing the prompts given to AI agents and their reasoning chains. On September 29, it called on AI labs to stop testing closed models on difficult problems. Its members stressed that releasing the work is only the beginning of the process of reviewing it and incorporating it into mathematical knowledge.

Last month, OpenAI said its model had solved the Navier-Stokes equation, one of the Clay Mathematics Institute’s seven Millennium Prize Problems, each carrying a $1 million award. The company says internal testing is needed to build scientific tools and that the proofs were a byproduct of that work. The release has renewed debate over how results produced with closed models should be checked and credited.