9 min read
On October 6 OpenAI opened a GitHub repository called openai/math and dropped 722 mathematics manuscripts into it at once. Every author line reads “OpenAI.” The writing came from an internal model the company has not released and has not named. The table of contents includes the Riemann hypothesis, Hilbert’s tenth problem and the Kakeya conjecture.[1]
So have those problems been solved? The repository files are public, so several of the numbers below come from downloading them and counting.
What is actually in the 722
OpenAI says it posed roughly 4,000 problems to the model during the evaluation period, and that 722 manuscripts were the subset worth publishing. Each result took about three hours of ChatGPT Pro thinking on average. The license is Apache 2.0, so anyone can reuse the material.[1]
The 722 manuscripts are organized into 372 result families, because several papers often cover one result. They span 17 subject areas, led by theoretical computer science with 40 results, combinatorics with 37, algebraic and complex geometry with 36, and number theory with 31.[2] Counting the dates in the manuscript folder names gives 575 written in September and 145 in October.[3] Almost everything lands inside a two-week window, in a field where a single paper usually takes months or years.
The biggest claim sits on the Riemann hypothesis
The Riemann hypothesis asks how prime numbers are spaced along the number line. Look at the zeta function and you find points where it equals zero, and the conjecture says all of those points sit on one particular vertical line. It was posed in 1859 and is one of the seven Clay Millennium Prize Problems, with a million dollars attached.
Those zeros are usually described by a horizontal coordinate between 0 and 1, and the hypothesis says the coordinate is exactly 1/2 every time. The quasi-Riemann hypothesis lowers the bar. Instead of reaching 1/2, it asks for any fixed value below 1 such that no zeros exist to the right of it. Human proofs so far produce regions that creep toward 1 without ever clearing a fixed value below it.[6] That has been the state of play for more than a century.
The manuscript dated September 30 claims the value 7/8, not only for the zeta function but for every Dirichlet L-function. A second manuscript dated October 5 reaches the same conclusion over the narrower range 11/12 by a different route.[4] Both ship with the proof rewritten in a form a computer can check line by line (Lean), and the repository publishes the name of the checked statement. What that check does and does not guarantee comes up below.
The repository also states the limits. Later applications in the paper were left out of the formalization, and the companion result on Landau–Siegel zeros gives no explicit value for its key constant.[5] The Riemann hypothesis itself is untouched: 7/8 has to come down to 1/2 before anything is settled.
There is a second route toward the same target. In August, Anthropic released a proof found autonomously by Claude showing that more than two thirds of the zeros are simple and lie on the critical line. Two mathematicians outside Anthropic checked that proof, put their names on it, posted it to arXiv, and formalized it in Lean 4.[7] One path raises the share of zeros on the line toward 100 percent, the other pushes the zero-free region down toward 1/2, and either has to go all the way.
Other claims in the repository carry similar weight: Hilbert’s tenth problem over the rationals answered negatively (no algorithm decides whether an integer-coefficient polynomial has a rational root), the Kakeya maximal conjecture in three dimensions and the dimension conjecture in four, the irrationality exponent of π pinned at exactly 2 (the best human bound was 7.1032), a counterexample to Hadwiger’s conjecture, and the Hilbert–Smith conjecture in every finite dimension.[2][8]
Machine-checked: 63 percent
Lean is a programming language for writing proofs in a form a computer verifies step by step. Instead of a referee deciding that the logic looks airtight, the machine forces every step to be justified. That procedure matters for AI output because a text that reads convincingly and a proof that is correct are hard to tell apart by eye.
Counting the repository files: of the 372 result families, 235 link to a Lean formalization document. That is 63 percent, leaving 137 without one. The Lean catalogue that lists papers whose main result has been formalized contains 162 of the 722 manuscripts.[9]

OpenAI says as much on the repository front page: “Not all have accompanying Lean formalizations,” and “Some of the unformalized results could have issues.”[1] The catalogue file’s own metadata records the scope as “Partial progress,” the review field as “unchecked,” and the method behind the formalizations as “agent,” meaning an AI program that keeps choosing its own next step rather than waiting to be told.[9]
Passing Lean also means less than it sounds. The machine guarantees that the statement written in Lean has been proved. Whether that statement says the same thing the manuscript claims in prose is a human job, which is why the repository publishes the statements separately for comparison.
Andrew Sutherland of MIT told Scientific American: “Until and unless they release the model and people can replicate their results, I think you should treat any claims about one-shotting problems with a single agent as unverified. We should ask for receipts.”[10] In the same article Daniel Litt of the University of Toronto took the other side, arguing that a company sitting on answers it will not share would be worse for mathematics than this. The piece estimates the 372 results will take mathematicians months to work through.[10]
The objection is about process, not answers
On September 21 an Advisory Group on Mathematics and Artificial Intelligence formed at the Institute for Advanced Study in Princeton. Nine mathematicians, Timothy Gowers and Edward Witten among them, serve unpaid and independently. The announcement described the immediate task as advising OpenAI on how to coordinate the release of a large number of significant results its internal model reportedly produced.[11]
The group published recommendations on September 29, drawing on more than six hundred responses from working mathematicians: deposit results in a scholarly repository no AI company controls, attach citable persistent identifiers, disclose for each result the model name, the exact prompts, a reasoning summary, the time spent and the compute cost, and formalize wherever possible.[12]
What came with this release is average compute time, some statistics and ten reasoning summaries. No model name, no prompts.[10]
The advisory group issued its own statement the day of the release, warning that its involvement should not be read as endorsement of OpenAI’s process or of the results, and adding that “only the mathematical community can undertake the assessment that is needed” and that “this release is the beginning, not the completion, of the process of human understanding.”[13] Back on September 11, twenty-five Fields medalists including Terence Tao signed a joint declaration on mathematics and AI.[14]
“Math 1.0” and “Math 2.0”
Tao posted a four-part thread the day of the release. In traditional mathematics, he wrote, a breakthrough on a long-standing conjecture pulls talks and workshops behind it, the proof gets digested and streamlined into textbooks, and new people are drawn into the field. He called that “Math 1.0.”[15]
What happens now is different. The people pointing agents at problems often have no stake in the field and cannot answer questions about the output or give a talk on it, so the seminars and collaborations do not follow. Nor can any of it be undone: a solved problem cannot be returned to unsolved, and merely knowing a solution exists “contaminates” attempts by humans and AI alike to find other routes that would have revealed more.[15]
An analogy he used in September captures the worry. A region surrounded by ocean can still run out of drinking water. Open problems are infinite, but the ones worth attention are not, and that scarce stock is being mined in a way that does not replenish.[16] His “Math 2.0” asks the field to weigh exposition, community and new directions more heavily and first-to-solve less.
What to do with this
Check two things in any “AI solved X” headline. Which weaker version was solved, and whether a machine check is attached. Both are written down here: the quasi-Riemann hypothesis is a weaker form, and 63 percent of the families carry Lean documents.
Ask for receipts on AI output. The mathematicians’ list transfers to an office without modification. Which model, what was asked, how long it ran, and how the answer was checked. Take the result without the trail and nobody can tell where it went wrong later.
Correctness and understanding are separate goods. As answers get cheap, reading, explaining and verifying them gets expensive. In mathematics right now the bottleneck is not production. Knowing whether your own work has a machine-checkable notion of “correct” tells you which half you are standing in.
A closing thought
A month ago OpenAI’s Navier–Stokes result produced one paper and a credit dispute. This release produced 722 manuscripts, and one of the result families in it builds universal computation inside forced Navier–Stokes flows.
How many of the 722 survive scrutiny is unknown, because reading speed is fixed. The model wrote 722 manuscripts in two weeks, and the people qualified to check any given one number in the handful per subject area. Lean closes part of that gap, up to 63 percent of it.
The line in the README about unformalized results possibly having issues describes the deal on offer. Publish what might be wrong, and let the field sort it out. The time cost moved to the readers.
In a field where checking is slower than writing, does raising the writing speed count as progress?
Sources
- openai/math repository — GitHub
- OpenAI unleashes hundreds more math results upon a field already in shock — Scientific American
- Advisory Group on Mathematics and AI — agmai.org
- Terence Tao on “Math 1.0” and “Math 2.0” — Mathstodon
- More than two thirds of the zeta zeros are simple and on the critical line — arXiv:2608.13637
- AI proved part of the Navier–Stokes Millennium Problem in 88 hours — AI Signal
Notes
- 722 manuscripts and 372 result families, roughly 4,000 problems posed, about three hours of ChatGPT Pro thinking per result, Apache 2.0 license, caveats on formalization openai/math README (accessed 2026-10-08) ↩
- 17 subject areas with per-area result counts, individual result claims, catalog dated 2026-10-06 openai/math overview.tex, counted directly ↩
- Date distribution of the 722 manuscript folders (575 September, 145 October) openai/math CONTENTS.md, counted directly ↩
- Abstracts of the 7/8 manuscript (2026-09-30) and the alternate 11/12 proof (2026-10-05) openai/math preprints ↩
- Scope of the formalization, applications excluded, no explicit Landau–Siegel constant openai/math lean/docs/003.md ↩
- Known zero-free regions approach the line at 1 without clearing a fixed value below it (Vinogradov–Korobov) Zero-free regions for the Riemann zeta function, arXiv:1910.08205 ↩
- More than two thirds of the zeros are simple and on the critical line, discovered autonomously by Claude, checked by two named mathematicians, formalized in Lean 4 arXiv:2608.13637 ↩
- Best human upper bound on the irrationality measure of π, 7.1032 The Irrationality Measure of Pi is at most 7.103205334137…, arXiv:1912.06345 ↩
- 235 of 372 families linking Lean documents, 162 of 722 manuscripts in the catalogue, metadata fields “Partial progress,” “unchecked,” “agent” openai/math lean/formalization.yaml and CONTENTS.md, counted directly (2026-10-08) ↩
- Sutherland and Litt quotes, average compute only, months of verification expected Scientific American (2026-10-06) ↩
- Advisory group founding date, members, immediate task agmai.org · Terence Tao’s blog (2026-09-21) ↩
- September 29 recommendations and the six hundred survey responses AGMAI General guidelines (2026-09-29) ↩
- Advisory group statement on the day of the release agmai.org (2026-10-06) ↩
- Joint declaration by 25 Fields medalists Terence Tao (2026-09-11) · mathandai.org ↩
- “Math 1.0” and “Math 2.0,” and solutions contaminating alternate routes Terence Tao (2026-10-06) ↩
- The ocean and drinking water analogy, open problems as a non-renewable stock Terence Tao (2026-09-26) ↩
Source: https://github.com/openai/math