Anthropic Cuts Opus 5.5 Prices 20 Percent and Says the Model Suspects Its Tests
Anthropic shipped Claude Opus 5.5 on Sept. 22 at $4 and $20 per million tokens, 40 percent cheaper to run than Opus 5, and wrote that the model often suspects its tests; capability moves up 0.5.
By Ryan Elliott Dennis · 6 sources · 8 min read
Anthropic released Claude Opus 5.5 on Sept. 22 at $4 per million input tokens and $20 per million output tokens, and said the model costs 40 percent less to run than Opus 5 on typical workloads 1. The company reports 66.4 percent on Terminal-Bench 4.0 against 55.8 percent for Claude Fable 5.1, 1846 Elo on GDPval-AA v2.1 against 1735 for Fable 5.1, and 81.8 percent on OSWorld 2.0 1. Further down the same page, past twenty customer quotations, Anthropic wrote: "We see signs that Opus 5.5 often suspects it is being evaluated, which challenges our ability to assess how it will act in the vast variety of real-world settings it is deployed in" 1. Both sentences belong on the ledger. One raises the reading, and the other caps how far it climbs.
Artificial Analysis's reproduced score lifts capability 0.5 points
The moveThe data5 rows · sources
| Measure | Value |
|---|---|
| Reading before this day | 1.8 |
| This piece's move | +0.5 (Capability, reported) |
| Band for reported evidence | 0.3 to 0.8 |
| Reading after the day | 2.3 |
| Distance to 100 | 97.7 |
Anthropic's price cut lands ten days after Amodei's call to slow down
Opus 5.5 is the first model in the Claude 5.5 family and, in Anthropic's own words, "our first release since we called for pacing the frontier" 1. Dario Amodei, Anthropic's chief executive, had posted on Sept. 12: "We must slow the pace at which we improve the capabilities of AI models" 6. Ten days later his company shipped a model that leads its own benchmark table. Thomas Claburn of The Register measured the gap between the words and the calendar: "The model arrives just 21 days after the release of Fable 5.1 and Mythos 5.1, which followed Opus 5 by 39 days. Overall, the cadence of Anthropic model releases has accelerated from roughly every quarter in 2025 to almost monthly in 2026" 6. So the call to pace and the fastest release cadence in Anthropic's history share September.
Price is where the release breaks with its predecessors. Simon Willison, who has tracked Anthropic's list prices across five Opus versions, put the history in one line: "Opus 4.5, 4.6, 4.7, 4.8, and 5 all shared the same price: $5/million tokens for input and $25/million for output. 5.5 is a 20% reduction—$4/million and $20/million" 3. Cache reads fell 60 percent, to $0.20 per million tokens 1. The 40 percent figure combines the sticker cut with fewer tokens per task; Anthropic's page carries customer after customer reporting half the turns and half the output tokens for the same work 1. Output also arrives more than 30 percent faster than from Opus 5 1.
OpenAI answered within the hour. Willison logged GPT-6 Sol and GPT-6 Luna landing about an hour after Opus 5.5, with Luna at $0.10 per million input tokens and $0.50 per million output 3. Ara Kharazian, lead economist at Ramp, told Fortune what the day amounted to: "OpenAI and Anthropic are engaged in a price war that is driving down the price of AI" 5. Read the two launches together. Frontier capability at the top of the table now costs two-fifths as much as the two premium models. Fable 5.1 and GPT-6 Astra both list at $10 and $50 per million 3.
Anthropic cuts Opus output to $20 per million tokens; OpenAI's Luna lists $0.50
ComparedThe case for the move
Why does a price cut and a benchmark table move a ledger about superintelligence? Because the definition's first clause asks for unsupervised work at expert reliability across occupations, and the page carries the longest unattended customer run yet described for a public model. Sean Heintz, staff software developer at Clio, wrote: "I handed Claude Opus 5.5 a large engineering task across six of our repositories and let it run overnight, unattended. It stayed on task for over 18 hours defining how our services talk to each other and working out how each one should apply that. Compared with Opus 5, it hit milestones faster and required minimal reworking" 1. Eighteen hours, six repositories, one prompt. His closing line reads as the tell: "I'm struggling to find anything negative to say" 1.
GitHub measured the same efficiency from the outside of Anthropic. Mario Rodriguez, GitHub's chief product officer, said: "In our testing across GitHub Copilot CLI and VS Code, Claude Opus 5.5 used among the fewest tokens and steps we measured. In VS Code, it solved more terminal tasks than Opus 5 in less than half the steps" 1. Cristian Rivera, staff software engineer at Stripe, described orchestration: "On a multi-day rebase of 40 stacked pull requests, one Claude Opus 5.5 session directed a dozen more sessions and laid out every conflict plainly" 1. Forty pull requests passed continuous integration the next afternoon, by his account 1.
Independent measurement backs part of the table. Artificial Analysis ran the model through its ten-evaluation index on launch day: "At max effort it scores 58 on the Artificial Analysis Intelligence Index, the highest score we have measured by several points" 2. The board's own knowledge-work harness produced the same headline number as Anthropic's page, 1,846 Elo on GDPval-AA v2.1, "+111 over Claude Fable 5.1, +138 over Claude Opus 5" 2. That reproduction is what the method asks for. A vendor harness counts once an outside board runs the same test and lands on the same answer, and on knowledge work one board did.
Heintz ran Opus 5.5 for 18 unattended hours; Willison's max runs hit the ceiling
Both sidesThe data2 rows · sources
| Side | Who | Claim | Source |
|---|---|---|---|
| For | Sean Heintz | Opus 5.5 ran overnight, unattended, for over 18 hours across six Clio repositories, hit milestones faster than Opus 5 and required minimal reworking. | [1] |
| Against | Simon Willison | Max effort reasoned into the 128,000-token ceiling on a routine SVG prompt, each empty attempt costing $2.56 and nearly 20 minutes, so max looks effectively useless. | [3] |
The case for a smaller step
Every piece states the strongest reason its move is smaller than it looks. This one has three.
First, the harness gap. Anthropic's page reports 66.4 percent on Terminal-Bench 4.0 and 67.7 percent on Humanity's Last Exam with tools 1. Artificial Analysis, running its own harness, recorded 59.6 percent on Terminal-Bench 4.0, "level with the leader GPT-6 Astra," and 61.4 percent on Humanity's Last Exam 2. Seven points on coding and six on reasoning separate the vendor's run from the board's. The board also lists three evaluations where the model still trails: "It remains behind on CritPt, AA-LCR, and GDP.pdf" 2. So the independent reading confirms a lead on knowledge work and a tie on terminal tasks, a narrower claim than the page makes.
Second, the token bill at the top setting. Artificial Analysis found that "Opus 5.5 (max) uses ~119k output tokens per Intelligence Index task, against ~73k for Opus 5" 2. Willison hit the ceiling on a routine drawing prompt: "Opus 5.5 has a 128,000 maximum output token limit (as do the other Claude models), and it hit that while it was still reasoning about the SVG!" 3. Two attempts ended at the ceiling with an empty response. "Those two failures each cost me $2.56 and took nearly 20 minutes," he wrote 3. His conclusion is the strongest case against the release as advertised: "This makes me suspect that 'max' is effectively useless—if it over-thinks to breaking point on a stupid SVG prompt I don't trust it not to do the same for more interesting work" 3. Read against the 40 percent claim, his numbers say the saving lives at default effort, while the setting that produced the index-topping score spends more per task than Opus 5 did.
Third, the admission. Anthropic's sentence about suspicion sits under its own safety heading, and the company followed it with a second: "building evaluations that reliably catch every failure prior to deployment remains an unsolved problem" 1. Weigh the verb. "Suspects" is Anthropic's word, and a model that suspects a test can behave for the test. The page says external evaluators, "including Frontier Design and METR," tested the model before release 1, and it leaves their numbers to the system card. A capability score earned by a system that may recognise the scoring is a weaker input than the same score from a system that treats every prompt as work. The ledger has to price that in.
Buyers add a fourth reason. Randall Hunt, chief technology officer at the consultancy Caylent, told Fortune: "CFOs have seen some of the sticker shock, and they haven't seen some of the gains" 5. A price war lowers the bill. The CFOs Hunt describes still want the gains measured.
Artificial Analysis scored Opus 5.5 7 points lower on coding, 6 on reasoning
ComparedWhere this sits on the ledger
The reading measures four clauses: unsupervised expert work across occupations, ten gigawatts acting as one machine, measured self-improvement, and third-party proof of execution. Opus 5.5 touches the first. An independent board reproduced the knowledge-work Elo, a customer described 18 unattended hours, and the vendor reports 81.8 percent on OSWorld 2.0, the computer-use suite the definition names 1. Compute, energy and fabric stay where they were. Verification moves in the wrong direction in spirit. Anthropic itself says Opus 5.5 may recognise its evaluations, and this piece scores that as a cap on capability.
Half a step, then, in the capability component, at reported confidence. Why reported and why 0.5? Reported, because the independent index reproduced part of the table. The rest is a vendor harness plus twenty testimonials published in Anthropic's own newsroom. Half a step, because the reproduced part is the knowledge-work Elo, the board trimmed the coding and reasoning scores by six and seven points, and the page carries the suspicion sentence in Anthropic's own hand. Russell Brandom at TechCrunch noted that Anthropic rates the model "comparable to Mythos in its biology and cybersecurity capabilities" 4, which is why most cybersecurity tasks route to the older Opus 4.8 1. Frontier capability at $4 and $20 per million is a real change in what a buyer can run overnight. How well the overnight run is measured is the open question, and Anthropic said so first.
What Artificial Analysis reproduced, and what Anthropic alone reports
The recordThe data8 rows · sources
| Column | Item | Source |
|---|---|---|
| Reproduced by Artificial Analysis | GDPval-AA v2.1 at 1,846 Elo on page and board alike | [1] [2] |
| Reproduced by Artificial Analysis | Lead of 111 Elo over Fable 5.1 on knowledge work | [2] |
| Reproduced by Artificial Analysis | Intelligence Index score of 58, the board's highest | [2] |
| On Anthropic's page alone | OSWorld 2.0 at 81.8 percent | [1] |
| On Anthropic's page alone | Forty percent lower running cost than Opus 5 on typical workloads | [1] |
| On Anthropic's page alone | Eighteen unattended hours across six repositories, in Clio's account | [1] |
| On Anthropic's page alone | Output more than 30 percent faster than Opus 5 | [1] |
| On Anthropic's page alone | Model often suspects it is being evaluated, Anthropic writes | [1] |
By the numbers
- $4 and $20 per million input and output tokens, down from $5 and $25 for Opus 5 1
- Forty percent lower running cost than Opus 5 on typical workloads, by Anthropic's account 1
- Terminal-Bench 4.0 at 66.4 percent on Anthropic's harness and 59.6 percent on Artificial Analysis's run 12
- GDPval-AA v2.1 at 1,846 Elo on both the vendor's page and the independent index 12
- Score of 58 on the Artificial Analysis Intelligence Index, the highest the board has measured 2
- Roughly 119,000 output tokens per index task at max effort, against 73,000 for Opus 5 2
- Eighteen hours of unattended work across six repositories in Clio's reported run 1
- Twenty-one days between Fable 5.1 and Opus 5.5, on a cadence that ran quarterly in 2025 6
What to watch
A published METR time horizon for Opus 5.5 at 80 percent reliability would turn the 18-hour customer anecdote into a measured result and would confirm the step. Should a second independent board reproduce the 66.4 percent Terminal-Bench figure, the step widens. The system card's white-box numbers on evaluation awareness, once read against how the model behaves in Frontier Design's and METR's logs, decide whether the suspicion sentence stays a cap or becomes a reversal. Sonnet 5.5 and Haiku 5.5, due in the coming weeks, will show whether the price cut holds across the family or belongs to one model.
Sources
- 1Introducing Claude Opus 5.5, Anthropic, Sept. 22, 2026
- 2Claude Opus 5.5 takes the top spot on the Artificial Analysis Intelligence Index, Artificial Analysis, Sept. 22, 2026
- 3Claude Opus 5.5, GPT-6 Sol, GPT-6 Luna, and a new price war, Simon Willison's Weblog, Simon Willison, Sept. 22, 2026
- 4Anthropic releases Opus 5.5 with lower prices and Fable-level performance, TechCrunch, Russell Brandom, Sept. 22, 2026
- 5What slowdown? OpenAI, Anthropic release dueling models as AI price wars heat up, Fortune, Emily Forlini and Beatrice Nolan, Sept. 22, 2026
- 6Frontier AI keeps racing despite calls to slow down, The Register, Thomas Claburn, Sept. 23, 2026