SuperIntelligence Infrastructure 2.6%reading
modelsAnthropicLayer lab

Anthropic Sells Opus 5 at Half Fable 5's Price; ARC Prize Scores It 30 Percent

Anthropic released Claude Opus 5 on July 24 at half of Fable 5's price, ARC Prize confirmed a 30 percent ARC-AGI-3 score, and Gennaro Cuofano argues the harness carries most of each score; capability moves up 0.4.

▲ +0.4 Capability reported Reading after July 24, 2026: 1.6

By Ryan Elliott Dennis · 9 sources · 7 min read

Anthropic released Claude Opus 5 on July 24 at $5 per million input tokens and $25 per million output tokens, the Opus 4.8 price and half the rate of Claude Fable 5 1. Anthropic's launch page says the model "comes close to the frontier intelligence of Claude Fable 5 at half the price" 1. ARC Prize tested it the same day and posted 30.16 percent on ARC-AGI-3 at high reasoning effort, the top score on that board, against a prior best of 7.8 percent held by a GPT-5.6 variant 38. Russell Brandom, AI editor at TechCrunch, noted that the model arrived two months after Opus 4.8, which shipped May 28 2. Anthropic's own page carries the caveat: the model "still shows important limitations on long-running, autonomous research tasks" 1.

Why does a price cut belong on a ledger about superintelligence? Because the definition's first clause asks for expert work across most occupations, and price per task decides which occupations can afford the work. A frontier score at half the price widens the set of jobs a system can take. How much of a step does that earn? The evidence splits three ways: one independent number, several vendor numbers, and one admission.

ARC Prize's release-day score lifts capability 0.4 points

The move
Capability steps up 0.4 on reported evidence: ARC Prize posted the ARC-AGI-3 score on release day and the price halves the cost of frontier work. Internal harnesses and a conceded long horizon hold the step short of confirmed.
The data5 rows · sources
MeasureValue
Reading before this day1.2
This piece's move+0.4 (Capability, reported)
Band for reported evidence0.3 to 0.8
Reading after the day1.6
Distance to 10098.4

What ARC Prize measured

The test is interactive. ARC Prize describes it as a benchmark that "challenges AI agents to explore novel environments, acquire goals on the fly, build adaptable world models, and learn continuously," and a full score means "AI agents can beat every game as efficiently as humans" 9. Opus 5 posted 30.16 percent at high reasoning effort on July 24 3. Only one effort setting was run, because the testing window was short 3. ARC Prize wrote that the model "completed five additional Public Demo environments that no model had previously beaten, demonstrating strong logical reasoning" 3. On the older boards it scored 97.5 percent on ARC-AGI-1 and 90.4 percent on ARC-AGI-2 at maximum effort, results ARC Prize called "competitive with previous frontier leaders, though at slightly higher cost" 3.

Read the two ARC sentences together. Five environments fell that had stood against every earlier model, and the price of the older scores went up. OfficeChai reported the prior ARC-AGI-3 best as 7.8 percent, held by GPT-5.6 Sol at maximum effort 8. So ARC Prize agrees with Anthropic's "three times" claim on the new test and adds a cost note on the old ones.

Opus 5 scored 30.16% on ARC-AGI-3, almost four times GPT-5.6 Sol's 7.8%

Compared
ARC Prize ran Opus 5 at high effort on release day and posted 30.16 percent, against 7.8 percent for the prior leader. The board ran one effort setting inside a short window, which keeps the score reported.Sources [3] [8]
The data2 rows · sources
Unit: percent on ARC-AGI-3
ItemValueSource
GPT-5.6 Sol, prior best
maximum effort
7.8[8]
Claude Opus 5
high effort, the one setting run
30.16[3]

Anthropic's own runs supply every other number

Every other claim on the launch page comes from Anthropic's own runs. On Frontier-Bench v0.1, "Opus 5 surpasses all other models, and more than doubles Opus 4.8's performance at a lower cost per task" 1. A footnote states the setup: "These results are from an internal run of Frontier-Bench v0.1, on the mini-SWE-agent harness and a GKE backend, mean reward over 5 attempts per task" 1. On OSWorld 2.0, the page says Opus 5 "outperforms every other model at any given cost, surpassing Fable 5's best result at just over a third of the cost" 1. Zapier AutomationBench shows a pass rate "around 1.5× the next-best model for the same cost" 1.

Nicolas Zeeb of Vellum read the same page on launch day and marked where the numbers stop. "Without published head-to-head numbers in text form, the exact margins stay on the chart," Zeeb wrote 7. He added a line on price: "Opus 5 doesn't undercut anyone on raw price. It's still above Sol and well above K3" 7. The half-price claim is measured against Fable 5, and against Fable 5 alone.

Sualeh Asif, co-founder of Cursor, supplied the customer voice on Anthropic's page. "Claude Opus 5 delivers near Fable 5 intelligence at Opus speed and cost," Asif said 1. His sentence carries the same structure as Anthropic's: intelligence near the flagship, cost at the tier below. Simon Willison, who publishes the weblog that tracks each release, wrote on launch day that the model was "currently leading the Artificial Analysis leaderboard, in front of even Fable 5" 4. An aggregate index placing the cheaper model above the flagship is the strongest outside support for the price-per-capability claim, and it rests on tests the vendors chose to run.

ARC Prize ran one outside test; Anthropic ran every other number

The record
ARC Prize measured its own benchmarks on release day; every figure in the right column came from Anthropic's runs on harnesses Anthropic chose. Two outside aggregates disagree on the ordering, placing Opus 5 first and third.Sources [1] [3] [4] [6]
The data11 rows · sources
ColumnItemSource
Measured outside AnthropicARC-AGI-3: 30.16 percent at high effort, posted July 24[3]
Measured outside AnthropicFive Public Demo environments completed that had stood against every earlier model[3]
Measured outside AnthropicOlder boards: 97.5 percent on ARC-AGI-1, 90.4 on ARC-AGI-2, maximum effort[3]
Measured outside AnthropicArtificial Analysis leaderboard puts Opus 5 ahead of Fable 5, per Willison[4]
Measured outside AnthropicEpoch Capabilities Index places Opus 5 behind Fable 5 and GPT-5.6[6]
Anthropic's own runsFrontier-Bench v0.1: more than double Opus 4.8, internal run on mini-SWE-agent[1]
Anthropic's own runsMean reward over 5 attempts per task, on a GKE backend[1]
Anthropic's own runsOSWorld 2.0: Fable 5's best result at just over a third of the cost[1]
Anthropic's own runsZapier AutomationBench: around 1.5 times the next-best model at equal cost[1]
Anthropic's own runsClassifiers expected to intervene around 85 percent less often than Fable 5's[1]
Anthropic's own runsLong-running autonomous research: important limitations, conceded on the same page[1]

The case for the move

Price is the argument. The definition's first clause asks for expert work across most occupations, and a task priced at $25 per million output tokens reaches occupations that a $50 task leaves out. ARC Prize's number is the second argument, since an outside board ran the test and posted the score on the day of release. Brandom drew the practical conclusion for buyers. "While smaller than Fable 5, the model will be both cheaper and less restrictive than Fable, likely making it preferable in most use cases," he wrote 2. Less restrictive has a number attached: Anthropic expects its cyber classifiers to "intervene around 85% less often" than Fable 5's 1, and Brandom reported that Opus 5 sits outside the 30-day data retention policy that covers Fable and Mythos 2. Opus 5, cheaper, less filtered and free of the retention rule, will run more of the work, and more work on the record is what the ledger counts.

Opus 5 costs $25 per million output tokens, half of Fable 5's rate

Compared
Opus 5 keeps the Opus 4.8 price of $25 per million output tokens, half of Fable 5's rate. A cheaper frontier task reaches more occupations, the first clause of the definition and the core case for the move.Sources [1]
The data3 rows · sources
Unit: USD per million output tokens
ItemValueSource
Claude Opus 4.825[1]
Claude Opus 5
close to Fable 5, in Anthropic's account
25[1]
Claude Fable 550[1]

The case for a smaller step

Gennaro Cuofano, who writes FourWeekMBA, made the harness argument in September after scaffolded ARC results arrived. "A benchmark score is not a property of a model. It is a property of a triple: the model, the harness wrapped around it, and the evaluation set it was scored on," Cuofano wrote 5. He put a number on the gap: base Opus 5 at 30.16 percent on the ARC-AGI-3 public demo set against a scaffolded system at 100 percent on the same set, close to 70 points apart 5. His conclusion followed: "If the reported gap between base-model performance and harness-augmented performance on the same public set is close to 70 percentage points, the scaffolding is not a footnote to the model result — it is most of the result" 5. Apply that reading to the launch page. Frontier-Bench ran on mini-SWE-agent, a harness Anthropic chose, with five attempts per task and the mean reported 1. The definition's own rule states that a vendor harness counts once an independent board reproduces it. Only ARC has.

Luis Chavez-Mattos, director of product at MindStudio, found the cost curve inside Anthropic's own charts. "More reasoning effort doesn't always mean a better score. Anthropic's own data shows Opus 5's performance on Frontier code peaks at medium reasoning effort (around 53%) and actually gets worse, and more expensive, at higher settings," he wrote two days after launch 6. Price per capability, then, is a curve, and the buyer picks the point. He also placed the model on an aggregate: "The Epoch Capabilities Index, which aggregates dozens of benchmarks using item response theory, places Opus 5 behind both Fable 5 and GPT-5.6 in overall composite capability" 6. Two aggregates, two orderings. Artificial Analysis puts Opus 5 first; Epoch puts it third.

Then the admission. Anthropic wrote that the model "still shows important limitations on long-running, autonomous research tasks, which is where we expect AI models to pose the most substantial biology-related risks" 1. Long-running autonomous work is the definition's first clause, word for word. Anthropic, clearing ARC-AGI-3 and conceding the long horizon in the same document, has told the ledger which component to leave alone. Autonomy stays where it was.

Cursor's Asif praises the price; Cuofano says the harness carries the score

Both sides
Asif makes the price-per-capability case; Cuofano reports a gap near 70 points between base and scaffolded runs on the same set. ARC Prize's release-day score tips the balance up, and the harness gap holds the step at reported.Sources [1] [5]
The data2 rows · sources
SideWhoClaimSource
ForSualeh AsifOpus 5 delivers intelligence near Fable 5's at Opus speed and Opus cost.[1]
AgainstGennaro CuofanoA score belongs to a model, its harness and its evaluation set together; base Opus 5 at 30.16 percent against a scaffolded 100 makes the harness most of the result.[5]

Where this sits on the ledger

The reading measures four clauses. Opus 5 touches the first, expert work across occupations, through price and through one independent score. It leaves the ten-gigawatt clause, the self-improvement clause and the proof-of-execution clause where they were. Capability moves 0.4, up, at reported confidence. Why reported, when ARC Prize measured the number itself? Because ARC Prize ran one effort setting on a public demo set inside a short window, and every other claim on the page came from an internal harness. Why 0.4 and above the floor of the band? Because the one outside number is almost four times the prior best, and the price halves the cost of reaching it. A confirmed step waits on a second outside board and a longer run.

By the numbers

  • $5 and $25 per million input and output tokens, the Opus 4.8 price and half of Fable 5's 1
  • 30.16 percent on ARC-AGI-3 at high effort, tested July 24 3
  • GPT-5.6 Sol held the prior ARC-AGI-3 best at 7.8 percent, at maximum effort 8
  • Five public demo environments completed that had stood against every earlier model 3
  • 90.4 percent on ARC-AGI-2 and 97.5 percent on ARC-AGI-1 at maximum effort 3
  • Around 85 percent fewer classifier interventions than Fable 5, by Anthropic's estimate 1
  • 57 days between Opus 4.8 on May 28 and Opus 5 on July 24 2
  • 2.3 on Anthropic's automated audit of misaligned behaviour, the lowest of its recent models 1

What to watch

ARC Prize running Opus 5 at maximum effort, and publishing cost per task beside the score, would show whether the 30 percent holds when the model spends more. An independent run of Frontier-Bench v0.1 or OSWorld 2.0, on a harness the outside board chose, would turn the cost claims into measured ones and lift the step toward the confirmed band. Epoch's index and Artificial Analysis agreeing on one ordering would settle which aggregate to trust. A METR time-horizon measurement on Opus 5 would test the long-running limitation Anthropic conceded, and a poor result there would hold autonomy flat while capability rose.

Sources

  1. 1Introducing Claude Opus 5, Anthropic, July 24, 2026
  2. 2Anthropic launches Opus 5, TechCrunch, Russell Brandom, July 24, 2026
  3. 3Claude Opus 5 - ARC-AGI Results, ARC Prize, July 24, 2026
  4. 4Introducing Claude Opus 5, Simon Willison's Weblog, Simon Willison, July 24, 2026
  5. 5Claude Opus 5 Scores 97.5% and 30.16% Simultaneously, and Both Numbers Are Correct, FourWeekMBA, Gennaro Cuofano, Sept. 26, 2026
  6. 6Claude Opus 5 Benchmarks: The Numbers Anthropic Didn't Headline, MindStudio, Luis Chavez-Mattos, July 26, 2026
  7. 7Claude Opus 5 Benchmarks Explained, Vellum, Nicolas Zeeb, July 24, 2026
  8. 8Claude Opus Scores 30% On ARC-AGI 3, Triples Previous Best Score By Any Model, OfficeChai, OfficeChai Team, July 24, 2026
  9. 9ARC-AGI-3, ARC Prize, 2026