OpenAI Posts 722 Math Manuscripts From an Unreleased Model, and 162 Carry a Lean Check
OpenAI posted 722 manuscripts from an unreleased model on Oct. 6, and its own catalogue shows 162 with a formalized main result; mathematicians' advisory group asked labs to stop, and capability moves up 0.4.
By Ryan Elliott Dennis · 7 sources · 7 min read
OpenAI posted 722 mathematics manuscripts to a public GitHub repository at 6 P.M. EDT on Tuesday, Oct. 6, and attributed them to an unreleased internal model 2. The papers fall into 372 families, drawn from roughly 4,000 problems posed to the model, at about three hours of ChatGPT Pro thinking compute per result 1. Lean, a programming language that checks every step of a proof mechanically, backs the main result of 162 of the papers, which is 22 percent of the collection 6.
OpenAI's 722 manuscripts lift the reading 0.4 points
The moveThe data5 rows · sources
| Measure | Value |
|---|---|
| Reading before this day | 2.6 |
| This piece's move | +0.4 (Capability, reported) |
| Band for reported evidence | 0.3 to 0.8 |
| Reading after the day | 3.0 |
| Distance to 100 | 97.0 |
What did OpenAI publish on Oct. 6?
The README describes the method in plain terms. OpenAI evaluates its models on open research problems, widened those evaluations after its existing math tests saturated, and kept the outputs that met "an appropriate level of significance" 1. An OpenAI spokesperson told Scientific American that the model produced almost every result from one prompt handed to one AI agent, with some results taking several attempts 2.
Compare that with the Navier-Stokes proof of September. That one came from 10,000 coordinating agents working 88 hours and cost millions of dollars in computing power 25. If the spokesperson's account holds, the unit of work shrank from a swarm to a single agent running about three hours.
Claims in the collection run wide. Joseph Howlett of Scientific American lists a solution to the four-dimensional Kakeya conjecture, improvements to important algorithms, and progress toward the Riemann hypothesis 2. Mike Pearl of Gizmodo opened one paper at random and found a claim of "the full Birch-Swinnerton-Dyer leading term formula," which holds only for elliptic curves over the rationals under a condition on the corank of a Selmer group 4. "I'm not remotely qualified to evaluate any of this," Pearl wrote 4. OpenAI's spokesperson added that many of the new results still await understanding by the company's own mathematicians 2.
Dates ride in the repository's directory names. Of the 722 manuscripts, 565 bear a date from Sept. 23 to Sept. 27, another 112 bear Oct. 5, and the remaining 45 are spread across other days in September and October 7. A reader can take those stamps as the dates of the PDFs, which leaves the model's run time unreported. They show a catalogue assembled in about two weeks, and they place nearly four fifths of it before the mathematicians' advisory group published its recommendations.
How much of the collection has a machine check?
The repository ships a file named formalization.yaml, a catalogue of papers with a formalized main result. It holds 162 entries 6. Decrypt reached the same count from the same file and put it at "about 22%" of the collection 5. That leaves 560 manuscripts whose main result lacks a Lean proof in the catalogue.
OpenAI says so itself. "Some of the unformalized results could have issues. We will endeavor to fix any such issues quickly," the README states 1. The sentence reads as a disclosure and a promise at once, and the verb "endeavor" promises effort, a result left to later.
A passing Lean check also has a boundary. Decrypt notes that it shows the logic holds for the formal statement, and a person still has to confirm the formal statement matches the problem mathematicians posed 5. The 162 are therefore the strongest rows in the collection, and the remaining question about each is whether the theorem Lean proved is the theorem the title announces.
Andrew Sutherland, a mathematician at the Massachusetts Institute of Technology, drew the line at the model itself. "Until and unless they release the model and people can replicate their results, I think you should treat any claims about one-shotting problems with a single agent as unverified," Sutherland told Scientific American 2. His sentence is conditional on release, which OpenAI told Scientific American it is working toward as quickly and responsibly as possible 2. Until then, the single-prompt claim rests on a spokesperson's word.
162 of 722 OpenAI papers have a Lean-formalized main result
ComparedWhich of the advisory group's requests did the release meet?
OpenAI announced an independent advisory group of mathematicians on Sept. 21 2. The Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study published its recommendations eight days later, drawing on more than 600 replies from the mathematical community 3. Its list for each released result names five items: the model, the prompts, a summarized chain of thought, the time taken and the cost of computation 3.
In reply, OpenAI supplied average compute, ten reasoning summaries, and some attempt statistics, and held back the prompts and the model name 125. The spokesperson told Scientific American the company is taking the guidelines seriously and doing its best to comply, and described them as advisory 2.
The group's opening paragraph reads as a request addressed to this exact release. "We want to state clearly from the start: we do not endorse this practice, and we ask them to stop testing advanced mathematical problems on proprietary models," the Advisory Group wrote 3. The release arrived seven days later, with the model still proprietary. In an email to Gizmodo, OpenAI spokesperson Lindsay McCallum Rémy replied, "AGMAI's advice and public recommendations have informed how we're sharing the results" 4. Both sentences can be true at once: advice informs a release while the release proceeds.
565 manuscripts carry Sept. 23 to 27 dates, before the advisory group spoke
TimelineThe data5 rows · sources
Should a lab hold results until mathematicians can read them?
Two serious answers exist. Daniel Litt, a mathematician at the University of Toronto, argued for release. "If we want to know the answers to these math questions, I see no reason why we should ask the company to keep them secret from us," Litt told Scientific American 2. Litt's argument is about access. A result on GitHub can be read, attacked and formalized by anyone, and a result in a company's private repository can only be believed.
From the opposite direction, the Advisory Group on Mathematics and Artificial Intelligence made its case about what a release does to the field. "We strongly recommend that AI labs refrain from treating the release of mathematical results as marketing vehicles to promote their models, ignoring the substantial negative externalities that this practice inflicts on the mathematical community," the Advisory Group on Mathematics and Artificial Intelligence wrote 3. It also warns that proprietary internal models risk "a two-tier system where labs outrun the rest of the field" 3.
Practitioners pushed on the mechanics. Keith Adler, writing on X on Oct. 7, observed that the repository has its Issues tab switched off, and added, "If you publish 722 manuscripts and ask for Lean formalisations, you need somewhere for people to send them" 5. Abhishek Saha, a mathematician who read the catalogue the same day, wrote that "It is a very big day for mathematics," and sorted most of the results into "exceptional advances within an existing program" rather than breakthroughs 5.
Read together, the two sides share a premise. Both treat the results as worth reading. Litt wants them public now, and the Advisory Group wants them released under conditions that let a mathematician reproduce the path. The disagreement is over who pays for the checking.
Daniel Litt backs release; the advisory group asks labs to stop
Both sidesThe data2 rows · sources
| Side | Who | Claim | Source |
|---|---|---|---|
| For | Daniel Litt | Mathematicians gain from learning the answers, so he favors publishing the results over keeping them secret. | [2] |
| Against | Advisory Group on Mathematics and Artificial Intelligence | Labs should stop testing advanced problems on proprietary models, and should stop treating result releases as marketing for their models. | [3] |
Why does the step stop at 0.4?
The method sizes a reported event between 0.3 and 0.8 points, and it sizes a confirmed one, such as a measured result on an uncontaminated test, between 0.8 and 1.5. This release is a company's own account of an unreleased model, partly checked by software. The 162 Lean-backed papers justify a step above the floor of the reported band. Four facts hold it in the lower third.
First, the model is internal, so the single-prompt claim awaits outside replication 2. Second, 560 papers lack a formalized main result, and OpenAI concedes some "could have issues" 1. Third, referees and journals have yet to report on any manuscript, and OpenAI's own mathematicians have yet to digest many of them 2. Fourth, the prompts stay private, which removes the one artifact that would let a reader price the human effort behind each result 23.
Capability carries a weight of 0.20 in the reading, and the component already stood at 3.2 before this piece. The piece adds 0.4 and leaves the question of how many of the 372 families survive review to the months ahead.
What the repository shows, and what rests on OpenAI's word
The recordThe data8 rows · sources
| Column | Item | Source |
|---|---|---|
| On the record | 722 manuscripts in 372 families sit on public GitHub | [1] |
| On the record | 162 papers have a Lean-formalized main result | [6] |
| On the record | Ten abridged reasoning summaries are published | [1] |
| On the record | Average compute is about three hours per result | [1] |
| Claimed or open | One prompt to one agent rests on a spokesperson's account | [2] |
| Claimed or open | 560 papers lack a formalized main result | [6] |
| Claimed or open | Prompts and the model name stay private | [3] [5] |
| Claimed or open | Referee reports on any family have yet to appear | [2] |
By the numbers
- 722 manuscripts in 372 families went on GitHub at 6 P.M. EDT on Oct. 6 12.
- Roughly 4,000 problems were posed to the model, and each result used about three hours of ChatGPT Pro thinking compute on average 1.
- Of the 722 papers, 162 have a Lean-formalized main result in the repository's catalogue, which is 22 percent 6.
- Directory dates place 565 manuscripts between Sept. 23 and Sept. 27 and 112 on Oct. 5 7.
- The Advisory Group's Sept. 29 list asks for five disclosures per result, and OpenAI published averages and ten reasoning summaries and withheld the prompts 35.
- Over 600 mathematicians replied to the group's survey before it wrote its recommendations 3.
- The Navier-Stokes proof of September used 10,000 agents over 88 hours 25.
What to watch
Watch the formalization catalogue first. Growth from 162 toward the 722 papers shows whether the unformalized results hold up, and the README promises more Lean files as OpenAI obtains them 1. A referee's named report on any single family, the Birch-Swinnerton-Dyer claim being the most scrutinized, would move the piece from reported toward confirmed. Release of the model to outside mathematicians would let Sutherland's condition be tested. Retraction of any headline family, or a Lean statement found to differ from its title, would pull the step back down.
Sources
- 1openai/math: README, OpenAI (GitHub), OpenAI, Oct. 6, 2026
- 2OpenAI unleashes hundreds more math results upon a field already in shock, Scientific American, Joseph Howlett, Oct. 6, 2026
- 3Responsible Release of AI-Generated Mathematics, Advisory Group on Mathematics and Artificial Intelligence, AGMAI, Sept. 29, 2026
- 4OpenAI Dumps 377 New Math Results on GitHub, Publishes Hand-Wringing Blog Post, Gizmodo, Mike Pearl, Oct. 6, 2026
- 5OpenAI Says a Secret AI Model Cracked Hundreds of Open Math Problems in One Prompt, Mathematicians Want Receipts, Decrypt, Jose Antonio Lanz, Oct. 7, 2026
- 6lean/formalization.yaml: catalogue of papers with a formalized main result, OpenAI (GitHub), OpenAI, Oct. 6, 2026
- 7CONTENTS.md: manuscript map, OpenAI (GitHub), OpenAI, Oct. 6, 2026