OpenAI has handed the world's mathematicians a huge reading assignment that nobody asked for. On Oct. 6 the company posted 722 manuscripts written by an AI model it hasn't released, and its own notes admit that some of the results without machine-checked proofs "could have issues". This is an analysis piece, and my argument is simple. Writing these proofs was the cheap part. The expensive part is checking them, and OpenAI has left that job to people who don't work for it.
Cheap to write, slow to check
Start with what it cost to produce them. OpenAI says the model worked through about 4,000 problems, and the average result used roughly three hours of ChatGPT Pro thinking time. The results it kept were grouped into 372 "families" of related work. For a company that runs data centers, that bill is small.
Now look at what checking costs. When OpenAI announced just ten results on Aug. 1, two researchers, Mikołaj and Krzysztof Sienicki, set out to audit them. Their write-up covered 18 chapter-by-chapter reviews. It was first posted on Aug. 3 and was still being revised on Sept. 9. One specialist asked for major revisions to an argument that was too compressed to follow, and in another chapter a sign error only came to light after formatting lost in a PDF conversion was recovered. The good news is that, in the reviews they examined, no confirmed serious error in a main result remained. The catch is scale. That audit covered ten results, and this new batch is 72 times bigger.
Doesn't a computer check the proofs?
Partly. Many of the manuscripts come with versions written in Lean, a programming language that lets a computer verify every step of a proof. OpenAI says "many, but not all" of the papers have been formalized this way, and that it will add more over time. That's real progress, and it's more than most human-written papers offer.
But Lean only answers one question: does this proof really follow from this statement? It can't tell you whether the statement typed into the computer means what the paper claims in plain English. A preprint by Maher Kallel and Mohamed El Louadi is about exactly that gap. They studied OpenAI's earlier batch, the ten results it published in August, not the new 722. In that set, the computer-checked proofs came to about 20.6 MB. The formal statements a human has to read took up only 55.6 KB, but they relied on 218 custom definitions instead of the field's shared standard ones. Their conclusion is that making checking free "shifts the burden to layers dependent on scarce expert attention". The machine handles the tedious part. A person still has to confirm the right question was asked.
Who picks up the tab
So who does that work? OpenAI's own advisers answer that plainly. The Advisory Group on Mathematics and Artificial Intelligence, nine mathematicians including Edward Witten and Timothy Gowers whom OpenAI consulted on the release, called it "the beginning, not the completion, of the process of human understanding." They added that "only the mathematical community can undertake the assessment that is needed". The group says its members take no payment for the work and operate independently of any AI company. It also has no say over how fast OpenAI moves. The company said the group won't advise it on "how to pace our internal progress".
That's where the bill lands. The "mathematical community" the advisers point to is mostly university mathematicians, postdocs and journal referees. OpenAI's release invites people to report problems so it can fix them, but we found nothing in it, or in OpenAI's announcement of the advisory group, about paying anyone for that review. Those people now have two choices. They can spend months checking claims from a model they can't use themselves, or they can ignore the claims and risk spending a year on a problem that may already be solved. Either way, they pay.
Outsiders can't rerun the work either. MIT mathematician Andrew Sutherland told Engadget that until the model is released and people can reproduce the results, claims of one-shot solutions should be treated "as unverified". OpenAI also held back exact prompts and compute figures for individual problems, even though its advisers had pushed for that kind of disclosure.
This complaint isn't new. In September, 25 Fields Medal winners signed a letter warning that labs can "spend tens of millions of dollars using LLMs to beat the original researchers to a proof," a race they said would push researchers toward secrecy. The letter came right after OpenAI claimed a solution to the famous Navier-Stokes problem, a claim that is still unverified.
If you don't follow math, the pattern should still look familiar. A company ships fast, gets the headline, and other people do the slow cleanup afterward. Here the cleanup is careful reading by experts, the kind that filled 18 chapter-by-chapter reviews for just ten results, and there's no shortcut for it.
The best case for OpenAI
To be fair, publishing beats hoarding. OpenAI could have kept these results private. Instead it put them online under an open license, attached Lean proofs to many of them and promised to fix problems "quickly". The one outside audit of its earlier batch found the main results held up. If most of these 722 papers turn out to be right, checking them may be less painful than feared, and the payoff, hundreds of answered questions, could be very large.
That's a serious argument. But "it will probably be fine" isn't a verification process, and OpenAI's own notes concede that some unformalized results could be wrong. Its advisers were blunt that only the wider community can do the assessment. Someone has to do the reading before anyone can honestly call these results correct, and the company getting the headline isn't the one doing that reading.
What would settle it
My argument gets weaker if OpenAI starts paying for independent review: funding referees, postdocs or formalization work at universities, publishing the prompts and per-problem compute its advisers already asked for, or letting outside researchers test the model. It also gets weaker if mathematicians work through the backlog quickly and find few errors, as the auditors of the first ten results did. It holds if, a year from now, most of these papers are still unread and unconfirmed while the number 722 keeps showing up in OpenAI's pitches.
So here's the question. When an AI can produce work faster than people can check it, should the company that built it have to pay for the checking?
Sources
- 1.openai/math repository README · OpenAI (GitHub)
- 2.Advisory Group on Mathematics and Artificial Intelligence · AGMAI / Institute for Advanced Study
- 3.OpenAI just posted hundreds more results on major math problems · Engadget
- 4.OpenAI forms math advisory group as its AI resolves more than 100 open problems · TechCrunch
- 5.OpenAI's feud with mathematicians is only escalating · TechCrunch
- 6.A Human Audit of OpenAI's AI-Generated Mathematical Proofs · arXiv (Mikołaj Sienicki, Krzysztof Sienicki)
- 7.Verification abundance, adjudication scarcity: what happens to mathematical knowledge when proof checking becomes free · arXiv (Maher Kallel, Mohamed El Louadi)
Reported by the WattsUpNext desk from the sources linked below. Spot an error? Tell us at corrections@wattsupnext.com.
The WattsUpNext Brief
Get stories like this in your inbox.
One email each weekday, only the topics you choose. Real news, sourced, no fluff. Unsubscribe in one click.




