The ARC Prize Foundation Launches ARC-AGI-3, and Frontier AI Scores Under 1 Percent
The ARC Prize Foundation launched ARC-AGI-3 on March 25 with 135 hand-built games every human tester solved and every frontier model scored under 1 percent on; capability moves down 0.3.
By Ryan Elliott Dennis · 8 sources · 8 min read
The ARC Prize Foundation published ARC-AGI-3 on March 25, and the two sentences at the top of Greg Kamradt's announcement carry the whole story: "Humans score 100%. Frontier AI scores 0.51%." 1. Kamradt, the foundation's president, described "hundreds of original turn-based environments, each handcrafted by a team of human game designers" 1. He put $2,000,000 in prizes behind any system that matches untrained people 1. François Chollet, who created the ARC series in 2019, marked the launch on stage at Y Combinator's San Francisco office in a fireside conversation with Sam Altman of OpenAI 1. One day later The Decoder printed the per-model figures: Gemini 3.1 Pro Preview at 0.37 percent, GPT 5.4 at 0.26 percent, Opus 4.6 at 0.25 percent and Grok-4.20 at 0.00 percent 3.
Why does a benchmark launch move a ledger about superintelligence down? Because the reading measures public evidence against a definition. A fresh test that places every public frontier system under 1 percent on work every tested human completed is evidence of distance. The step is small. Half of this piece explains why it stays small.
ARC-AGI-3's launch scores pull the capability reading down 0.3 points
The moveThe data5 rows · sources
| Measure | Value |
|---|---|
| Reading before this day | 0.0 |
| This piece's move | −0.3 (Capability, inferred) |
| Band for inferred evidence | 0.1 to 0.5 |
| Reading after the day | 0.0 |
| Distance to 100 | 100.0 |
The ARC Prize Foundation swaps grid puzzles for games
Static ARC is finished as a frontier test. Anthropic's Opus 4.5 scored 37.6 percent on ARC-AGI-2 at $2.20 a task by December 4. A refinement pipeline from Poetiq, built on Gemini 3 Pro, reached 54 percent at $30 a task 4. Mike Knoop, the foundation's co-founder, wrote in the 2025 results post that the third version "marks the first major format change since ARC was introduced in 2019" 4. Puzzles on a grid gave way to games.
The foundation's own paper, posted to arXiv on March 24, states the design in one sentence: "We introduce ARC-AGI-3, an interactive benchmark for studying agentic intelligence through novel, abstract, turn-based environments in which agents must explore, infer goals, build internal models of environment dynamics, and plan effective action sequences without explicit instructions." 2. Read the verb list. Explore, infer, build, plan: each names a thing the ledger's definition asks a system to do unsupervised for weeks. The paper then gives the calibration: "Our testing shows humans can solve 100% of the environments, in contrast to frontier AI systems which, as of March 2026, score below 1%." 2.
Scoring is where the number gets its meaning. The Decoder described the metric, relative human action efficiency, with per-level credit computed as the square of human actions divided by AI actions 3. A system that wins a game in twice the moves a person needs earns a quarter of the credit for that level. Learning speed is the score, and a win by brute search earns a sliver of credit. Human testers solved all 135 environments with, in the foundation's words, "no prior knowledge and no instructions," and 25 of them are public 3.
ARC-AGI-2 climbed from 1% to 54% in nine months; ARC-AGI-3 opened at 0.51%
TimelineThe data7 rows · sources
| Date | Event | Source |
|---|---|---|
| ARC-AGI-2 launches; reasoning models score 1 to 1.3 percent | [8] | |
| Results post puts Poetiq's pipeline at 54 percent on ARC-AGI-2 | [4] | |
| Foundation posts the ARC-AGI-3 paper to arXiv | [2] | |
| Launch: humans 100 percent, frontier AI 0.51 percent | [1] | |
| The Decoder prints per-model scores, Gemini 3.1 Pro Preview first at 0.37 | [3] | |
| Chollet tells Semafor a rapid climb would mark real progress | [5] | |
| Goertzel reports 32.58 percent on the 25 public games | [7] |
The case for the move
Chollet made the argument himself on X the day after launch, in answer to the question of whether a hand-built harness should count: "true AGI shouldn't need task-specific human guidance, especially when ordinary humans can handle the same tasks without any help" 3. Read the phrase he chose for the baseline. "Ordinary humans" sets the bar at anyone, and "without any help" sets the condition at cold. A model that clears the games with a scaffold a researcher wrote for those games has cleared the scaffold.
His standard for the whole series predates this launch. In the December results post he wrote: "You'll know AGI is here when the exercise of creating tasks that are easy for regular humans but hard for AI becomes simply impossible." 4. The foundation then spent the winter creating hundreds of such tasks, and every one of them worked. By his own rule the games answer the question, and the answer is the gap.
Kamradt had set out the efficiency principle a year earlier, when ARC-AGI-2 launched. "Intelligence is not solely defined by the ability to solve problems or achieve high scores. The efficiency with which those capabilities are acquired and deployed is a crucial, defining component," he told TechCrunch in March 2025 8. Version 3 turns that sentence into arithmetic. The squared penalty is the efficiency clause written as a formula.
Chollet told Semafor two weeks after launch that even top models "score below 1%" and that any rapid improvement "would represent real progress towards AGI" 5. So the foundation has published the number, the method, the human baseline and the price of a claim. It has said in advance what a climb would mean. That is a measurement, in the method's terms, from an evaluator outside the labs it measures. The definition's first clause asks for unsupervised expert work across occupations. Here is a test of unsupervised work at the level of a game a person learns in minutes, and the public systems sit under 1 percent on it.
Human testers solved 100 percent; Gemini 3.1 Pro Preview scored 0.37
ComparedThe case for a smaller step
Four reasons shrink the step.
First, the ARC record says launch-day floors move. TechCrunch reported at the ARC-AGI-2 launch in March 2025 that reasoning models scored between 1 and 1.3 percent 8. OpenAI's o3 had reached 75.7 percent on ARC-AGI-1 at $200 a task and 4 percent on the new set, while human panels averaged 60 percent 8. Nine months on, the refinement pipeline stood at 54 percent 4. Twice now an ARC floor near 1 percent has lasted under a year.
Second, the same launch data shows how much of the gap is scaffolding. Opus 4.6 scored 97.1 percent on a known environment using a hand-crafted harness and dropped to 0 percent on an unfamiliar one 3. The model can play a game it has been told about. It struggles to find the game. Which of those two facts the ledger should weigh is the open question, and the 0.51 percent headline reports only the second.
Third, Ben Goertzel, the AGI researcher who leads SingularityNET, ran his own systems against the 25 public games and published the results on May 8. "On the 25 public ARC-AGI-3 games, our system fully solved 7 and reached near-human performance on another 6. Aggregating per-game scores gives a mean human-normalized score of 32.58%," he wrote, and he noted that Symbolica had reported around 36 percent 7. His reading of what that proves is the sharpest objection on the record: "ARC-AGI-3 is an interesting benchmark but it seems quite plausible that systems with no real general-intelligence pretensions can be engineered to score well on it." 7. Read the verb. "Engineered" says the score will rise while the meaning stays flat. That cuts against the benchmark from both sides at once: the launch number under-measures what a built system can do, and a later high number may over-credit it. Either way the 0.51 percent is a reading of March, and the March reading was already moving by May.
Fourth, the record from launch week contradicts itself. Jensen Huang, Nvidia's chief executive, answered Lex Fridman's question about an AI that could build a billion-dollar business by saying: "I think it's now. I think we've achieved AGI." 6. Shane Legg, who co-founded DeepMind and coined the term, gave Fortune the opposite bar: "I define an AGI to be an artificial agent that can do the kinds of cognitive things that people can typically do. I see this as the natural minimum bar." 6. Altman told the same reporter the term "has become a very sloppy term" 6. Five days after Huang declared AGI arrived, a test built to Legg's minimum bar put the best model at 0.37 percent. Both statements sit in the record. The ledger scores the test.
Opus 4.6 scored 97.1% on a known game with a harness and 0% on a new one
The numberThe data1 row · sources
| Measure | Value | Source |
|---|---|---|
| Opus 4.6 with a hand-crafted harness on a known environment; 0 on an unfamiliar one | 97.1 percent | [3] |
Where this sits on the ledger
The reading measures four clauses: unsupervised expert work across occupations, ten gigawatts acting as one machine, measured self-improvement, and third-party proof of execution. This launch touches the first clause alone. Compute, energy and fabric stay untouched, and proof of execution rests on the foundation's own testing protocol.
So the move is 0.3 down, in capability, at inferred confidence. Down, because the definition's authors wrote that static tests had stopped measuring the thing that matters. This is the first public test since then that measures it, and the public systems sit under 1 percent. Inferred, because a benchmark launch is a demo-class event in the method's table. Day one of a fresh set of games gives a day-one score, and the ARC record twice shows such scores climbing fifty points inside a year. A larger step would need a held-out result that stays low after the labs have had a season to try.
Chollet calls the gap the measurement; Goertzel says built systems will climb
Both sidesThe data2 rows · sources
| Side | Who | Claim | Source |
|---|---|---|---|
| For | François Chollet | True AGI clears tasks ordinary humans clear unaided; a harness written for these games scores the harness, and every public model sits under 1 percent. | [3] |
| Against | Ben Goertzel | His system reached 32.58 percent on the 25 public games; systems with little claim to general intelligence can plausibly be engineered to score well. | [7] |
By the numbers
- Human testers solved every environment, and frontier AI scored under 1 percent as of March 2026 2
- Headline frontier score in the foundation's announcement: 0.51 percent 1
- Per-model figures on March 26 ran from Gemini 3.1 Pro Preview at 0.37 percent through GPT 5.4 at 0.26 and Opus 4.6 at 0.25 down to Grok-4.20 at 0.00 3
- Launch set: 135 environments, 25 of them public 3
- Prize pool for ARC Prize 2026: $2,000,000 1
- Opus 4.6 with a hand-crafted harness scored 97.1 percent on a known environment and 0 on an unfamiliar one 3
- Poetiq's refinement pipeline reached 54 percent on ARC-AGI-2 at $30 a task, nine months after that test's launch 4
- Goertzel's system on the 25 public games: 32.58 percent, reported May 8 7
What to watch
The foundation's verified leaderboard on the hidden environments is the marker. An OpenAI, Anthropic or Google model, run by the foundation on a plain agent loop, above 10 percent on the private set with a published cost per task would begin to reverse this step. One paid-out prize would reverse it outright. Public frontier models holding under 5 percent on the private set through the summer, while engineered systems climb on the public 25, would confirm it. That pattern would also confirm Goertzel's reading. The other marker is the foundation's own cost data: a score that arrives at $200 a task measures search, and a score that arrives at human action counts measures learning.
Sources
- 1Announcing ARC-AGI-3, ARC Prize, Greg Kamradt, March 25, 2026
- 2ARC-AGI-3: A New Challenge for Frontier Agentic Intelligence, arXiv, ARC Prize Foundation, March 24, 2026
- 3ARC-AGI-3 offers $2M to any AI that matches untrained humans, yet every frontier model scores below 1%, The Decoder, Maximilian Schreiner, March 26, 2026
- 4ARC Prize 2025 Results and Analysis, ARC Prize, Mike Knoop, Dec. 5, 2025
- 5AI research foundation releases test that will warn when AGI arrives, Semafor, Reed Albergotti, April 8, 2026
- 6Nvidia's Jensen Huang says 'we've achieved AGI.' But no one can agree on what that means, Fortune, Jeremy Kahn, March 30, 2026
- 7Playing Around with the ARC-AGI-3 Benchmark, Ben Goertzel (Substack), Ben Goertzel, May 8, 2026
- 8A new, challenging AGI test stumps most AI models, TechCrunch, Maxwell Zeff, March 24, 2025