SuperIntelligence Infrastructure 2.6%reading
researchAnthropicLayer lab

Anthropic's Internal Model Formalizes Fermat's Last Theorem in Lean in 11 Days

Anthropic said on Sept. 4 that an internal model formalized Fermat's Last Theorem in Lean in eleven days; Kevin Buzzard vouches for the check, Daniel Kokotajlo says AI runs behind his scenario, and capability moves up 0.6.

▲ +0.6 Capability reported Reading after Sept. 4, 2026: 2.2

By Ryan Elliott Dennis · 6 sources · 6 min read

Anthropic published a research note on Sept. 4 describing an internal model that formalized Fermat's Last Theorem in the Lean proof assistant, start to finish, in eleven days 1. The run began on Aug. 7 and ended on Aug. 18. It produced roughly 6 billion output tokens, 13 million lines of Lean, and 29,500 intermediate theorems 1. The whole structure compiles under Lean's kernel, with the axioms of mathematics as its only assumptions 1. Kevin Buzzard reviewed the result and put his name to it 1. He is the Imperial College mathematician who has led the human effort to formalize the same proof since 2024. Anthropic describes the model as roughly comparable to Claude Fable 5.1, which it released to the public three days earlier.

Why does a proof of a theorem Andrew Wiles settled in 1995 belong on a ledger about superintelligence? Because the ledger's definition asks for unsupervised work that runs for weeks on tasks experts find hard. This run is the longest single autonomous mathematical job on the public record. The question is how much of a step it earns.

Anthropic's Lean proof of Fermat's Last Theorem lifts capability 0.6 points

The move
Capability steps up 0.6 on reported evidence: a Lean kernel checked a full research-grade proof after eleven days of autonomous work. An internal model and known mathematics hold the step inside the reported band.
The data5 rows · sources
MeasureValue
Reading before this day1.6
This piece's move+0.6 (Capability, reported)
Band for reported evidence0.3 to 0.8
Reading after the day2.2
Distance to 10097.8

Lean's kernel checked every step of the model's proof

Formalization is translation under a strict grader. A human proof of Fermat's Last Theorem runs through modularity lifting, Galois representations and a great deal of algebraic number theory. Mathematicians accept much of it on trust in one another. Lean demands each step in full. Every step has to reduce to earlier steps until the chain reaches the axioms, and the kernel either compiles the whole thing or it rejects it 1.

That grader is the reason the result carries weight on this ledger. A benchmark score arrives from the vendor's own harness, and independent boards have disagreed with vendor harnesses by fifteen to thirty points across the summer. Here the checker ships with the proof. Anthropic states that a human set the goal and the environment, and the proof assistant was the only judge along the way 1.

Buzzard's role matters for the same reason. His team spent two years building the Lean infrastructure for the proof. He is the person best placed to say whether the model's 29,500 lemmas amount to the theorem or to a pile of technicalities that dodge it. "This extraordinary autoformalization achievement, which Anthropic researchers say only took 11 days, proves Fermat's Last Theorem with no assumptions other than the axioms of mathematics," Buzzard said in the note 1. Read the sentence for its hedge. Eleven days is Anthropic's number, and Buzzard marks it as such. The proof itself he vouches for outright.

Lean's kernel checked every step of the proof, and Buzzard vouched for it

Who connects
A human set the goal and the environment, the model worked for eleven days, and Lean's kernel judged every step against the axioms. Buzzard, whose group spent two years on the same target, vouched for the result.Sources [1]
The data7 rows · sources
FromLinkToSource
Anthropicset the goal and the environmentInternal model[1]
Internal modelroughly comparable, in Anthropic's accountClaude Fable 5.1[1]
Internal model13 million lines in eleven daysLean kernel[1]
Lean kernelreduces every step to themAxioms of mathematics[1]
Buzzard's Lean grouptwo years of Lean infrastructureLean kernel[1]
Kevin Buzzardleads the human effort since 2024Buzzard's Lean group[1]
Kevin Buzzardreviewed the finished proof and vouched for itInternal model[1]

The case for the move

Buzzard's sentence is the strongest case that the reading should move up. A research-grade proof, fully checked, in eleven days, is a long-horizon result. Static benchmarks stopped measuring that kind of work once frontier models cleared them. An outsider verified this one, which is the bar the method sets for the higher steps.

Three features of the run push toward a larger step. Eleven days of continuous work is a horizon well past anything the public time-horizon measurements have recorded for a single task. A program checked the output, in place of a vendor grading it. Human experts had already invested two years on the same target, which gives the comparison a real baseline.

Anthropic's model worked 11 days on one prompt, graded by one checker

The number
Eleven days of continuous work is a horizon well past anything public time-horizon measurements have recorded for one task. Anthropic reports the duration and Buzzard marks it as Anthropic's number, one reason the move stays reported.Sources [1]
The data1 row · sources
Unit: days
MeasureValueSource
Continuous autonomous run, Aug. 7 to Aug. 18, with Lean's kernel as the only judge11 days[1]

The case for a smaller step

Three reasons shrink the step.

First, the mathematics was known. Wiles published in 1995, Buzzard's group had mapped much of the route into Lean, and the model's job was to close the remaining distance under a checker. That is a hard translation task and a long one. Finding a theorem is a different thing. The definition's first clause asks for economically valuable work across occupations, and formalizing a settled proof is one narrow instance of it.

Second, the model is internal. Anthropic describes it as comparable to Fable 5.1, and only Anthropic has run it. A result from a system the public can use would earn a confirmed step. One from a system only the vendor has run earns a reported one, whatever the checker says about the output.

Third, the forecasting record moved the other way over the same months. Daniel Kokotajlo wrote AI 2027, the fastest widely read timeline of 2025, and he said this year that the field is running behind it 2. "Things seem to be going somewhat slower than the AI 2027 scenario," Kokotajlo said 2. He put a number on the revision in a later interview. "Our timelines were longer than 2027 when we published and now they are a bit longer still," he said, adding that around 2030 is what he now tells people, with a great deal of uncertainty 3. So Kokotajlo lengthened his scenario while machine proofs got longer. That pairing is the tell. Capability results of this kind were already priced into his 2027 story. The rest of the story, the part about systems automating research at scale, has arrived slower than he wrote it.

Gary Marcus makes the sharper version of the same argument. Writing in July about a panel he shared with three other researchers, he said: "None of the four of us believe that LLMs on their own are sufficient for AGI" 4. A Lean proof from a language model answers part of that claim, since the proof is real and the model produced it. It leaves the rest standing. Eleven days of formalization say little about reliability on open-ended work where the grader is missing.

What Lean and Buzzard confirmed, and what keeps the step reported

The record
A program and an outside mathematician confirmed the left column. The right column holds the reasons the result stays reported: an internal model, a theorem settled since 1995, and a duration Anthropic alone measured.Sources [1] [3] [4]
The data10 rows · sources
ColumnItemSource
Checked by the kernel and BuzzardWhole proof compiles under Lean's kernel[1]
Checked by the kernel and BuzzardAxioms of mathematics are its only assumptions[1]
Checked by the kernel and Buzzard13 million lines of Lean and 29,500 intermediate theorems[1]
Checked by the kernel and BuzzardRoughly 6 billion output tokens across the run[1]
Checked by the kernel and BuzzardBuzzard reviewed the result and vouched for the proof[1]
Holding the step at reportedModel is internal, and Anthropic alone has run it[1]
Holding the step at reportedTheorem settled by Wiles in 1995, its Lean route partly mapped[1]
Holding the step at reportedEleven days is Anthropic's figure, marked as such by Buzzard[1]
Holding the step at reportedKokotajlo moved his central estimate to around 2030[3]
Holding the step at reportedMarcus: all four panelists judge LLMs alone insufficient for AGI[4]

Where this sits on the ledger

The reading measures four clauses: unsupervised expert work across occupations, ten gigawatts acting as one machine, measured self-improvement, and third-party proof of execution. This result touches the first clause and, in a small way, the fourth, since a Lean kernel is a form of proof of what ran. It leaves the compute clause untouched. Self-improvement stays untouched too, because the model formalized mathematics instead of its own successor.

Epoch AI's trend data gives the scale of the background against which one result lands. Frontier training compute has doubled roughly every 5.2 months since 2020, and Epoch expects the pace to slow from 2026 5. In mid-July, FutureSearch's tracker put the Metaculus community median for general AI at January 2033, with a 25 percent chance by 2029 6. A ledger that moved a full point on every impressive result would have crossed fifty by now. This one moves 0.6, in the capability component, at reported confidence. It will move further on the day an outside team runs a public model against a comparable target and publishes what happened.

Buzzard vouches for the proof; Kokotajlo says AI runs behind his 2027 scenario

Both sides
Buzzard vouches for a proof whose only assumptions are the axioms. Kokotajlo, author of the fastest widely read timeline, says the field runs slower than his scenario, so the scale tips up by a reported step.Sources [1] [2] [3]
The data2 rows · sources
SideWhoClaimSource
ForKevin BuzzardAn extraordinary autoformalization proves Fermat's Last Theorem with the axioms of mathematics as its only assumptions; the eleven days are Anthropic's figure.[1]
AgainstDaniel KokotajloThe field runs somewhat slower than the AI 2027 scenario; his timelines, already longer than 2027 at publication, have lengthened again, to around 2030.[2] [3]

By the numbers

  • 11 days from Aug. 7 to Aug. 18 for the full formalization run 1
  • Roughly 6 billion output tokens generated across the run 1
  • 13 million lines of Lean and 29,500 intermediate theorems in the finished proof 1
  • Two years of human effort on the same Lean target under Buzzard's group before the run 1
  • 5.2 months per doubling of frontier training compute since 2020, with a slowdown expected from 2026 5
  • January 2033 as the Metaculus community median for general AI, with 25 percent by 2029 6
  • Around 2030 as Kokotajlo's revised central estimate 3

What to watch

An independent group running a public model against a comparable formalization target, with the run logged, would turn this reported step into a confirmed one. New mathematics, checked the same way, would engage the definition's first clause directly. Kokotajlo's next timeline update is the other marker. So is the first outside measurement of how much frontier research the labs' own agents now perform, which would settle whether the forecasters or the lab notes read the record correctly.

Sources

  1. 1Formalizing Fermat's Last Theorem, Anthropic, Sept. 4, 2026
  2. 2Things seem to be going somewhat slower than the AI 2027 scenario: Daniel Kokotajlo, OfficeChai, OfficeChai Staff, 2026
  3. 3AI 2027, six months later, FutureSearch, 2026
  4. 4Gary Marcus on X: None of the four of us believe that LLMs on their own are sufficient for AGI, X, Gary Marcus, July 2026
  5. 5Key trends and figures in machine learning, Epoch AI, 2026
  6. 6AGI timeline tracker, FutureSearch, July 2026