SuperIntelligence Infrastructure 2.6%reading
researchOpenAILayer lab

Epoch AI Declares FrontierMath Tier 4 Saturated After GPT-6 Astra Solves the Last Problem

Epoch AI declared FrontierMath Tier 4 saturated on Sept. 10 after GPT-6 Astra solved Jay Pantone's last standing problem; Gary Marcus says OpenAI reports its successes and keeps its misses, and capability moves up 0.6.

▲ +0.6 Capability confirmed Reading after Sept. 10, 2026: 1.8

By Ryan Elliott Dennis · 12 sources · 9 min read

Epoch AI said on Sept. 10 that AI models have now solved every problem in FrontierMath Tier 4, and that OpenAI's GPT-6 Astra solved the last one 1. Tier 4 is the hardest set the evaluator publishes: 43 research-level problems, written by professional mathematicians, of the kind a specialist expects to spend days on 2. When the tier opened on July 11, 2025, the top score was 5 percent 8. OpenAI's launch page for Astra reported 97.6 percent on Sept. 3, with GPT-5.6 Sol at 83.0 4. Epoch rounds Astra to 98 and calls the benchmark saturated 3.

Why does one exhausted math test belong on a ledger built around megawatts and proofs of execution? Because capability carries a fifth of the reading, and the definition asks for tests that stay uncontaminated. FrontierMath was written to be that test: private problems, expert authors, answers checked by code. Its exhaustion is a dated measurement by an outside evaluator, which is the evidence class the method calls confirmed. How large a step it earns runs through two facts on the same page: who wrote the last problem, and who paid for the test.

Epoch AI's Tier 4 measurement lifts capability 0.6 points

The move
Capability steps up 0.6 on confirmed evidence: Epoch AI measured a public model solving the last problem on its private, expert-written Tier 4 set. OpenAI's funding of the test and a 2 of 68 Erdős reading hold the step under the 0.8 floor.
The data5 rows · sources
MeasureValue
Reading before this day1.2
This piece's move+0.6 (Capability, confirmed)
Band for confirmed evidence0.8 to 1.5
Reading after the day1.8
Distance to 10098.2

What Epoch measured

FrontierMath's private set runs under fixed rules. A model submits a Python function that returns its answer, with a hard limit of 1,000,000 tokens per attempt, a forced submission at 660,000, and 30 seconds of runtime for the function itself 2. Version 2 arrived on June 12, 2026, after an audit found errors in 42 percent of the whole benchmark; in Tier 4 it corrected 12 problems and removed 7, which left the 43 that Astra faced 2. The audit had its own trigger. OpenAI reported during testing that FrontierMath held more errors than expected, and Epoch screened the set with GPT-5.5 and Claude Opus 4.7 before handing each flagged problem to a mathematician for review 9.

Epoch AI announced the result in three sentences. "Every FrontierMath Tier 4 problem has now been solved by AI, with GPT-6 Astra solving the last problem standing," the evaluator wrote 1. It then turned to method. "Mathematicians often commented that AI found unintended shortcuts when solving their Tier 4 problems. Not so for this last one, which was created by Jay Pantone," the post continued 1. Read the order of the clauses. The score comes first, and the news Epoch chose to add is about the route the model took, because on earlier problems the route had been the complaint.

Pantone is an associate professor of mathematics at Marquette University, and his name has been on this benchmark's record before. When Epoch reported GPT-5.2 Pro's jump to 31 percent in January, 15 of 48 problems, its write-up noted that two of Pantone's problems had fallen and that the model had "used numerical shortcuts that he didn't intend" 11. Eight months later 36Kr reported his verdict on the last problem: Astra's solution ran close to his own, and he could hardly be surprised anymore 9.

The case for the move

Elliot Stewart writes the Epoch Brief, the evaluator's own newsletter, and his Sept. 12 edition put the measurement in the plainest form it takes. "It also scored 98% on FrontierMath: Tier 4, solving the only remaining unsolved problem, leading us to consider the benchmark saturated," Stewart wrote 3. He set the score beside a wider result: "Astra set multiple new records, including on the Epoch Capabilities Index, taking the top spot among the 247 models we track" 3. The verb to weigh is "consider". OpenAI declares; Epoch considers, and Epoch considered the test finished.

Epoch's own forecasts are the sharpest measure of how fast this went. In October 2025 Greg Burnham, who runs benchmark work at Epoch AI, published a ceiling for the whole of FrontierMath. "We estimate that even repeatedly running all of these models would cap out below 70%," Burnham wrote 10. Eleven months later the hardest tier alone stands at 98. His estimate, built from repeated runs of every model then available, lasted under a year.

Three features push the step toward the confirmed band. An outside organization ran the test on a private set it controls, so the score belongs to Epoch and to OpenAI both. Working mathematicians wrote the problems, which Terence Tao, on reading them, called "extremely difficult" 9. And the machine that did it is public: Astra shipped on Sept. 3, and Aidan Clark, OpenAI's vice president of research, tied it to the compute clause the ledger tracks. "It's the first time we've pretrained on more than 100,000 GPUs at our Stargate site in Texas," Clark told Fortune 5. Greg Brockman, OpenAI's president, drew the larger conclusion at the same launch. "It's not unreasonable to feel that we are now in the AGI era," Brockman said 5.

FrontierMath Tier 4 went from a 5 percent top score to 98 percent in 14 months

Timeline
Tier 4 opened at a top score of 5 percent in July 2025 and closed at 98 percent in September 2026. Burnham's ceiling of 70 percent stood for eleven months, a measure of how fast the labs learned to train for the test.Sources [1] [2] [3] [4] [8] [10] [11]
The data7 rows · sources
DateEventSource
Tier 4 opens with a top score of 5 percent[8]
Burnham puts the FrontierMath ceiling below 70 percent for every model then available[10]
GPT-5.2 Pro sets a record at 15 of 48 problems, 31 percent[11]
Version 2 corrects 12 problems and removes 7, leaving 43[2]
OpenAI's launch page reports 97.6 percent for GPT-6 Astra[4]
Astra solves Jay Pantone's problem, the last one standing[1]
Stewart's Epoch Brief records 98 percent and calls the benchmark saturated[3]

The case for a smaller step

Three reasons shrink the step, and the first sits on Epoch's own page.

FrontierMath was funded by OpenAI, and OpenAI holds exclusive access to a subset of the benchmark 2. Epoch discloses this in a conflict of interest statement linked from the Tier 4 hub 2. The disclosure changes the reading of a score and leaves the score itself in place. A test OpenAI paid for, with a slice only OpenAI can see, is an independent measurement with a footnote, and the footnote is why the step sits at the bottom of its band.

Gary Marcus makes the second case, and he made it a month before the tier fell, about Astra's earlier claim to have solved ten open problems. Marcus, the NYU professor emeritus who has spent a decade contesting benchmark claims, wrote on Aug. 3 that OpenAI reported its successes and kept its misses to itself. "They have given us a numerator without a denominator, always a worrisome sign," Marcus wrote 6. He quoted OpenAI's Noam Brown conceding that other major problems had been tried and had held 6. His conclusion carried into September: "Astra is an obviously an advance to some degree, but that degree may well turn out to be merely incremental relative to other recent models" 6. On launch day he made the same point about benchmarks as a class. "Success on ARC-AGI is great and impressive, but not—despite the name of the task—proof of AGI," Marcus wrote 7. Swap the benchmark's name and the sentence holds for Tier 4. A saturated test says what the model does on that test. It stays silent on the problems the vendor chose to leave out of the announcement.

The third reason is the arithmetic around the score. Fortune's Emily Forlini tracked OpenAI's launch page through three versions on Sept. 3 and 4. Astra's 97.6 stayed fixed, while Anthropic's Fable 5.1 moved from 87.8 to 78 and back to 83 on the same test, and GPT-5.6 Sol went from 83 to 80.5 and back to 83 4. An OpenAI spokesperson told Fortune the changes were corrections. "Most evaluations have noise within a few percentage points based on the exact checkpoint, scaffold, and evaluation run used in reporting," the spokesperson said 4. Vincent Sunn Chen, who leads benchmark and evaluation research at Snorkel AI, read it as ordinary. "It's usually a function of final launch logistics," Chen said 4. Anka Reuel and Mike Hardy, researchers at Stanford University's Intelligent Systems Laboratory and Trustworthy AI Lab, read the timing differently. "This can be done in a very tight timeframe, and it's better for their marketing," they told Fortune 4. Ten points of movement on a rival's score, inside 36 hours, on this benchmark.

OpenAI's launch page held Astra at 97.6 while rival scores moved 10 points

Compared
Astra's 97.6 held across three versions of OpenAI's launch page on Sept. 3 and 4, while Fable 5.1 went from 87.8 to 78 to 83 and GPT-5.6 Sol from 83 to 80.5 to 83. Ten points of drift keep the step small.Sources [4]
The data7 rows · sources
Unit: percent on FrontierMath Tier 4, OpenAI launch page
ItemValueSource
Fable 5.1, version one87.8[4]
GPT-5.6 Sol, version one83[4]
Fable 5.1, version two78[4]
GPT-5.6 Sol, version two80.5[4]
Fable 5.1, version three83[4]
GPT-5.6 Sol, version three83[4]
GPT-6 Astra, all three versions97.6[4]

Astra solves 2 of 68 problems on the next test

Saturation on one tier arrived beside weak results on the tier behind it. FrontierMath Erdős, Epoch's set of 68 open Erdős problems, gave Astra 2 solutions in its official run and 5 with repeated attempts that cost more than $220,000 in compute 8. On Humanity's Last Exam, Astra's 57.2 percent trails its predecessor's 65.0 8. Stewart's brief records the Erdős result as a first, since Astra is the first model to solve any problem on that set 3. Two of 68 is also the opening reading on the test that replaces the one that just closed.

Epoch's own researchers had already written the caution that applies. In February, Florian Brand and Greg Burnham studied the benchmarks that claim to measure economic value and found each of them covering a narrow slice of a job. "High scores on the benchmarks would imply a shift in how jobs are done, not replacement," Brand and Burnham wrote 12. A research-level math test is narrower still. The definition's first clause asks for unsupervised expert work across occupations, and a 43-problem answer set, however hard, engages one occupation for one afternoon per problem.

What Epoch AI measured, and what holds the step at 0.6

The record
Epoch AI, an outside evaluator, dated the left column on a private set it controls. The right column carries the reasons the step sits at 0.6: a vendor-funded test, weak readings on neighbouring tests, and misses the announcement left out.Sources [1] [2] [3] [6] [8] [9]
The data9 rows · sources
ColumnItemSource
Measured by Epoch AIPrivate set, run under fixed token and runtime limits[2]
Measured by Epoch AILast problem, written by Jay Pantone, fell to GPT-6 Astra[1]
Measured by Epoch AIAstra tops the Epoch Capabilities Index among 247 models[3]
Measured by Epoch AIPantone judged Astra's solution close to his own[9]
Measured by Epoch AITao called the problems extremely difficult[9]
Holding the step smallOpenAI funded FrontierMath and holds exclusive access to a subset[2]
Holding the step smallErdős set: 2 of 68 in the official run, 5 after more than $220,000[8]
Holding the step smallHumanity's Last Exam: Astra at 57.2 percent, its predecessor at 65.0[8]
Holding the step smallMarcus: a numerator with its denominator hidden on open problems[6]

Where this sits on the ledger

The reading measures four clauses: unsupervised expert work across occupations, ten gigawatts acting as one machine, measured self-improvement, and third-party proof of execution. A saturated static benchmark touches the first, and the definition states in its own text that a perfect score on one moves capability a little. Clark's 100,000 GPUs speak to the second clause as a training run; the ledger counts energized sites 5. Self-improvement and proof of execution stay where they were.

So the move is 0.6 points, up, in capability, at confirmed confidence. Confirmed because Epoch AI, an outside evaluator, measured a public model on a private set and put the date on it. Small because OpenAI funded the test its own model exhausted, because the next test in the series reads 2 of 68, and because a benchmark that took fourteen months from 5 percent to saturation records, above all, how fast the labs learned to train for it. Burnham's ceiling of 70 percent stood for eleven months. The number this ledger shows will stand for one day, which is the point of the method.

Epoch's Stewart calls the test finished; Marcus says OpenAI hides its misses

Both sides
Stewart, writing for Epoch AI, considers the test finished; Marcus reads Astra's record as successes reported and misses kept private. An outside measurement tips the scale up, and OpenAI's funding keeps the step small.Sources [3] [6]
The data2 rows · sources
SideWhoClaimSource
ForElliot StewartAstra scored 98 percent on Tier 4 and solved the only remaining unsolved problem, so Epoch considers the benchmark saturated.[3]
AgainstGary MarcusOpenAI reports its successes and keeps its misses, a numerator with the denominator hidden; Astra's advance may prove incremental relative to recent models.[6]

By the numbers

  • 43 problems in FrontierMath Tier 4 after the June 12, 2026 revision, with 12 corrected and 7 removed 2
  • Five percent was the top score when the tier opened on July 11, 2025 8
  • GPT-6 Astra reported 97.6 percent and GPT-5.6 Sol 83.0 on OpenAI's launch page 4
  • Fifteen of 48 problems, 31 percent, was GPT-5.2 Pro's record in January 2026 11
  • Under 70 percent: Burnham's October 2025 estimate of the ceiling for every model then available, run repeatedly 10
  • 1,000,000 tokens per attempt, with forced submission at 660,000 and 30 seconds of runtime 2
  • 2 of 68 Erdős problems in Astra's official run, rising to 5 after more than $220,000 in repeated attempts 8
  • 247 models sit on the Epoch Capabilities Index, with Astra at the top 3

What to watch

A second evaluator, one funded by a different lab or by neutral money, rerunning Astra on the Tier 4 set would firm up this step. Release of the subset OpenAI holds exclusively, with a rerun on the whole, would do more. The next reading on FrontierMath Erdős is the forward measure: movement past 5 of 68 engages the first clause, and a stall there says the tier that closed was measuring training as much as mathematics. Marcus's denominator is the last marker, and OpenAI publishing the problems Astra tried and missed would settle it.

Sources

  1. 1Every FrontierMath Tier 4 problem has now been solved by AI, X, Epoch AI, Sept. 10, 2026
  2. 2FrontierMath Tier 4 (v2), Epoch AI, June 12, 2026
  3. 3The Epoch Brief - September 12, 2026, Epoch AI, Elliot Stewart, Sept. 12, 2026
  4. 4OpenAI quietly boosts some of Astra's evaluation metrics, and continues to change others post-launch, Fortune, Emily Forlini, Sept. 4, 2026
  5. 5OpenAI launches GPT-6 Astra, its most powerful model yet, and touts its ability to use your computer, Fortune, Emily Forlini, Sept. 3, 2026
  6. 6Two critical updates re: Astra and mathematics, Marcus on AI, Gary Marcus, Aug. 3, 2026
  7. 7Hot take on GPT-6 Astra, Marcus on AI, Gary Marcus, Sept. 3, 2026
  8. 8OpenAI's GPT-6 Astra Cracks Epoch's Hardest Math Benchmark in 14 Months, AlphaSignal, Sept. 10, 2026
  9. 9AI Math Field's Final High Wall Falls: GPT-6 Astra Fully Conquers All FrontierMath Tier 4 Tasks, 36Kr, Sept. 11, 2026
  10. 10Less than 70% of FrontierMath is within reach for today's models, Epoch AI, Greg Burnham, Oct. 17, 2025
  11. 11New record on FrontierMath Tier 4, Epoch AI, Jan. 23, 2026
  12. 12What do 'Economic Value' Benchmarks Tell Us?, Epoch AI, Florian Brand and Greg Burnham, Feb. 13, 2026