AI News Feed
Market watch
Large Language Models

‘Pure Insanity’: Mathematicians Face Years of Work After OpenAI Releases Nearly 400 Math Results

OpenAI released nearly 400 AI-generated mathematical results across more than 700 manuscripts, according to The Verge. Mathematicians described the drop as overwhelming and said understanding it could take years, while fewer than half the manuscripts have been formally verified.

The results cover combinatorics, several branches of geometry, number theory, theoretical computer science, algebra, topology, probability and statistical mechanics, and mathematical physics. The collection is so vast that OpenAI felt the need to publish guidance on how to navigate the sprawling GitHub repository. Even a preliminary assessment proved difficult. Most mathematicians who spoke with The Verge said they were still struggling to digest everything, and several said that simply working through the roughly 40-page table of contents and abstracts took them the better part of an hour. “Just going over the entire list of abstracts is overwhelming,” said Álvaro Lozano-Robledo, a professor of mathematics at the University of Connecticut.

Some of the manuscripts include formalizations in Lean, a programming language and proof assistant that allows results to be verified computationally. Such formalizations have helped assess some of OpenAI’s previous mathematical claims, giving researchers confidence that a claim is logically correct even if they do not fully understand the argument behind it. But the degree to which each result had been verified varied widely. On GitHub, OpenAI acknowledged that the results are “at different stages of verification” and that “many, but not all, of the manuscripts have been formalized.” As of writing, fewer than half the manuscripts in the collection appeared to have been described formally. OpenAI said only 300 top-line results out of 719 manuscripts had been formalized, around 42 percent, and that it “will update the repository with more formalizations as we obtain them.”

Several mathematicians complained to The Verge about the lack of formalization, particularly given the sheer number of results, and stressed that even when Lean code accompanies a result, evaluation is not instantaneous. Researchers must check that the formalization actually proves what the result claims, another time-consuming process. Several digging through the papers said that even where computer-verifiable proofs had been provided, the quality was inconsistent and the statements they verified did not always appear to map neatly onto the claims in accompanying manuscripts.

Kevin Buzzard, a mathematics professor at Imperial College London, said he had identified numerous theorems in his area of work, algebraic number theory, of which only around six immediately “stood out.” Few, if any, of those appeared to be formally verified in Lean. “Hence, I either have to read possibly-not-correct slop, or wait for others to do the same, or wait for someone to formalise them before I can say for sure that the results are even correct,” Buzzard said. His concerns were echoed by numerous other researchers.

Buzzard was far from alone in worrying about AI “slop.” The term is a shorthand for low-quality, frequently erroneous AI-generated material that increasingly crops up online, in the real world, and in academic papers. Mathematicians told The Verge they have seen a huge uptick in such material produced with tools like ChatGPT and Claude in recent years. Much of it is confusing, hard to read, and demonstrates little understanding of the subject; it is especially shoddy when it comes to crediting other researchers. OpenAI’s previous mathematical write-ups were widely criticized by experts for their sloppy nature, particularly their poor or nonexistent attribution. The release also left academics facing careers upended overnight, with researchers forced to separate solutions from slop while OpenAI moves on.