SuperIntelligence Infrastructure 3.0%reading
researchOpenAILayer lab

OpenAI Posts 722 Math Manuscripts From an Unreleased Model, and 162 Carry a Lean Check

OpenAI posted 722 manuscripts from an unreleased model on Oct. 6, and its own catalogue shows 162 with a formalized main result; mathematicians' advisory group asked labs to stop, and capability moves up 0.4.

▲ +0.4 Capability reported Reading after Oct. 7, 2026: 3.0

By Ryan Elliott Dennis · 7 sources · 7 min read

OpenAI posted 722 mathematics manuscripts to a public GitHub repository at 6 P.M. EDT on Tuesday, Oct. 6, and attributed them to an unreleased internal model 2. The papers fall into 372 families, drawn from roughly 4,000 problems posed to the model, at about three hours of ChatGPT Pro thinking compute per result 1. Lean, a programming language that checks every step of a proof mechanically, backs the main result of 162 of the papers, which is 22 percent of the collection 6.

OpenAI's 722 manuscripts lift the reading 0.4 points

The move
Capability steps up 0.4 at reported confidence: OpenAI released 722 manuscripts from an internal model, 162 with a Lean-formalized main result. An unreleased model and an unreviewed majority keep the step in the lower third of its band.
The data5 rows · sources
MeasureValue
Reading before this day2.6
This piece's move+0.4 (Capability, reported)
Band for reported evidence0.3 to 0.8
Reading after the day3.0
Distance to 10097.0

What did OpenAI publish on Oct. 6?

The README describes the method in plain terms. OpenAI evaluates its models on open research problems, widened those evaluations after its existing math tests saturated, and kept the outputs that met "an appropriate level of significance" 1. An OpenAI spokesperson told Scientific American that the model produced almost every result from one prompt handed to one AI agent, with some results taking several attempts 2.

Compare that with the Navier-Stokes proof of September. That one came from 10,000 coordinating agents working 88 hours and cost millions of dollars in computing power 25. If the spokesperson's account holds, the unit of work shrank from a swarm to a single agent running about three hours.

Claims in the collection run wide. Joseph Howlett of Scientific American lists a solution to the four-dimensional Kakeya conjecture, improvements to important algorithms, and progress toward the Riemann hypothesis 2. Mike Pearl of Gizmodo opened one paper at random and found a claim of "the full Birch-Swinnerton-Dyer leading term formula," which holds only for elliptic curves over the rationals under a condition on the corank of a Selmer group 4. "I'm not remotely qualified to evaluate any of this," Pearl wrote 4. OpenAI's spokesperson added that many of the new results still await understanding by the company's own mathematicians 2.

Dates ride in the repository's directory names. Of the 722 manuscripts, 565 bear a date from Sept. 23 to Sept. 27, another 112 bear Oct. 5, and the remaining 45 are spread across other days in September and October 7. A reader can take those stamps as the dates of the PDFs, which leaves the model's run time unreported. They show a catalogue assembled in about two weeks, and they place nearly four fifths of it before the mathematicians' advisory group published its recommendations.

How much of the collection has a machine check?

The repository ships a file named formalization.yaml, a catalogue of papers with a formalized main result. It holds 162 entries 6. Decrypt reached the same count from the same file and put it at "about 22%" of the collection 5. That leaves 560 manuscripts whose main result lacks a Lean proof in the catalogue.

OpenAI says so itself. "Some of the unformalized results could have issues. We will endeavor to fix any such issues quickly," the README states 1. The sentence reads as a disclosure and a promise at once, and the verb "endeavor" promises effort, a result left to later.

A passing Lean check also has a boundary. Decrypt notes that it shows the logic holds for the formal statement, and a person still has to confirm the formal statement matches the problem mathematicians posed 5. The 162 are therefore the strongest rows in the collection, and the remaining question about each is whether the theorem Lean proved is the theorem the title announces.

Andrew Sutherland, a mathematician at the Massachusetts Institute of Technology, drew the line at the model itself. "Until and unless they release the model and people can replicate their results, I think you should treat any claims about one-shotting problems with a single agent as unverified," Sutherland told Scientific American 2. His sentence is conditional on release, which OpenAI told Scientific American it is working toward as quickly and responsibly as possible 2. Until then, the single-prompt claim rests on a spokesperson's word.

162 of 722 OpenAI papers have a Lean-formalized main result

Compared
OpenAI's own formalization catalogue lists 162 papers, 22 percent of the 722, with a formalized main result. The other 560 rest on the README's warning that some unformalized results could have issues.Sources [6]
The data2 rows · sources
Unit: manuscripts
ItemValueSource
Formalized main result
22 percent
162[6]
Main result without a Lean proof
722 minus 162
560[6]

Which of the advisory group's requests did the release meet?

OpenAI announced an independent advisory group of mathematicians on Sept. 21 2. The Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study published its recommendations eight days later, drawing on more than 600 replies from the mathematical community 3. Its list for each released result names five items: the model, the prompts, a summarized chain of thought, the time taken and the cost of computation 3.

In reply, OpenAI supplied average compute, ten reasoning summaries, and some attempt statistics, and held back the prompts and the model name 125. The spokesperson told Scientific American the company is taking the guidelines seriously and doing its best to comply, and described them as advisory 2.

The group's opening paragraph reads as a request addressed to this exact release. "We want to state clearly from the start: we do not endorse this practice, and we ask them to stop testing advanced mathematical problems on proprietary models," the Advisory Group wrote 3. The release arrived seven days later, with the model still proprietary. In an email to Gizmodo, OpenAI spokesperson Lindsay McCallum Rémy replied, "AGMAI's advice and public recommendations have informed how we're sharing the results" 4. Both sentences can be true at once: advice informs a release while the release proceeds.

565 manuscripts carry Sept. 23 to 27 dates, before the advisory group spoke

Timeline
OpenAI named its advisory group Sept. 21. Directory dates put 565 manuscripts before the group's Sept. 29 request to stop testing proprietary models, and the release followed Oct. 6, seven days after.Sources [1] [2] [3] [7]
The data5 rows · sources
DateEventSource
OpenAI announces an independent advisory group of mathematicians[2]
First of 565 manuscripts dated Sept. 23 to 27[7]
Advisory group asks labs to stop testing on proprietary models[3]
112 manuscripts carry this date[7]
OpenAI posts 722 manuscripts at 6 P.M. EDT[1] [2]

Should a lab hold results until mathematicians can read them?

Two serious answers exist. Daniel Litt, a mathematician at the University of Toronto, argued for release. "If we want to know the answers to these math questions, I see no reason why we should ask the company to keep them secret from us," Litt told Scientific American 2. Litt's argument is about access. A result on GitHub can be read, attacked and formalized by anyone, and a result in a company's private repository can only be believed.

From the opposite direction, the Advisory Group on Mathematics and Artificial Intelligence made its case about what a release does to the field. "We strongly recommend that AI labs refrain from treating the release of mathematical results as marketing vehicles to promote their models, ignoring the substantial negative externalities that this practice inflicts on the mathematical community," the Advisory Group on Mathematics and Artificial Intelligence wrote 3. It also warns that proprietary internal models risk "a two-tier system where labs outrun the rest of the field" 3.

Practitioners pushed on the mechanics. Keith Adler, writing on X on Oct. 7, observed that the repository has its Issues tab switched off, and added, "If you publish 722 manuscripts and ask for Lean formalisations, you need somewhere for people to send them" 5. Abhishek Saha, a mathematician who read the catalogue the same day, wrote that "It is a very big day for mathematics," and sorted most of the results into "exceptional advances within an existing program" rather than breakthroughs 5.

Read together, the two sides share a premise. Both treat the results as worth reading. Litt wants them public now, and the Advisory Group wants them released under conditions that let a mathematician reproduce the path. The disagreement is over who pays for the checking.

Daniel Litt backs release; the advisory group asks labs to stop

Both sides
Litt argues public release serves mathematics, while the advisory group argues proprietary models and marketing-driven release harm the field. The evidence holds level until outside referees report.Sources [2] [3]
The data2 rows · sources
SideWhoClaimSource
ForDaniel LittMathematicians gain from learning the answers, so he favors publishing the results over keeping them secret.[2]
AgainstAdvisory Group on Mathematics and Artificial IntelligenceLabs should stop testing advanced problems on proprietary models, and should stop treating result releases as marketing for their models.[3]

Why does the step stop at 0.4?

The method sizes a reported event between 0.3 and 0.8 points, and it sizes a confirmed one, such as a measured result on an uncontaminated test, between 0.8 and 1.5. This release is a company's own account of an unreleased model, partly checked by software. The 162 Lean-backed papers justify a step above the floor of the reported band. Four facts hold it in the lower third.

First, the model is internal, so the single-prompt claim awaits outside replication 2. Second, 560 papers lack a formalized main result, and OpenAI concedes some "could have issues" 1. Third, referees and journals have yet to report on any manuscript, and OpenAI's own mathematicians have yet to digest many of them 2. Fourth, the prompts stay private, which removes the one artifact that would let a reader price the human effort behind each result 23.

Capability carries a weight of 0.20 in the reading, and the component already stood at 3.2 before this piece. The piece adds 0.4 and leaves the question of how many of the 372 families survive review to the months ahead.

What the repository shows, and what rests on OpenAI's word

The record
Solid tiles come from the repository itself. Dashed tiles rest on a spokesperson, a missing prompt or a referee who has yet to report, which is why the step stays reported.Sources [1] [2] [3] [5] [6]
The data8 rows · sources
ColumnItemSource
On the record722 manuscripts in 372 families sit on public GitHub[1]
On the record162 papers have a Lean-formalized main result[6]
On the recordTen abridged reasoning summaries are published[1]
On the recordAverage compute is about three hours per result[1]
Claimed or openOne prompt to one agent rests on a spokesperson's account[2]
Claimed or open560 papers lack a formalized main result[6]
Claimed or openPrompts and the model name stay private[3] [5]
Claimed or openReferee reports on any family have yet to appear[2]

By the numbers

  • 722 manuscripts in 372 families went on GitHub at 6 P.M. EDT on Oct. 6 12.
  • Roughly 4,000 problems were posed to the model, and each result used about three hours of ChatGPT Pro thinking compute on average 1.
  • Of the 722 papers, 162 have a Lean-formalized main result in the repository's catalogue, which is 22 percent 6.
  • Directory dates place 565 manuscripts between Sept. 23 and Sept. 27 and 112 on Oct. 5 7.
  • The Advisory Group's Sept. 29 list asks for five disclosures per result, and OpenAI published averages and ten reasoning summaries and withheld the prompts 35.
  • Over 600 mathematicians replied to the group's survey before it wrote its recommendations 3.
  • The Navier-Stokes proof of September used 10,000 agents over 88 hours 25.

What to watch

Watch the formalization catalogue first. Growth from 162 toward the 722 papers shows whether the unformalized results hold up, and the README promises more Lean files as OpenAI obtains them 1. A referee's named report on any single family, the Birch-Swinnerton-Dyer claim being the most scrutinized, would move the piece from reported toward confirmed. Release of the model to outside mathematicians would let Sutherland's condition be tested. Retraction of any headline family, or a Lean statement found to differ from its title, would pull the step back down.

Sources

  1. 1openai/math: README, OpenAI (GitHub), OpenAI, Oct. 6, 2026
  2. 2OpenAI unleashes hundreds more math results upon a field already in shock, Scientific American, Joseph Howlett, Oct. 6, 2026
  3. 3Responsible Release of AI-Generated Mathematics, Advisory Group on Mathematics and Artificial Intelligence, AGMAI, Sept. 29, 2026
  4. 4OpenAI Dumps 377 New Math Results on GitHub, Publishes Hand-Wringing Blog Post, Gizmodo, Mike Pearl, Oct. 6, 2026
  5. 5OpenAI Says a Secret AI Model Cracked Hundreds of Open Math Problems in One Prompt, Mathematicians Want Receipts, Decrypt, Jose Antonio Lanz, Oct. 7, 2026
  6. 6lean/formalization.yaml: catalogue of papers with a formalized main result, OpenAI (GitHub), OpenAI, Oct. 6, 2026
  7. 7CONTENTS.md: manuscript map, OpenAI (GitHub), OpenAI, Oct. 6, 2026