SuperIntelligence Infrastructure 2.6%reading
modelsGoogle DeepMindLayer lab

Google DeepMind Matches GPT-6 Astra With Gemini 4 Argon and Gives Cyber Defenders First Access

Google DeepMind released Gemini 4 Argon to cyber defenders on Sept. 30; Artificial Analysis scored it 53, level with GPT-6 Astra, and capability holds flat while Matthias Bastian says Anthropic still leads.

● 0.0 Capability reported Reading after Sept. 30, 2026 (7 pieces that day): 2.5

By Ryan Elliott Dennis · 7 sources · 8 min read

Google DeepMind released Gemini 4 Argon on Sept. 30 to a small group of cyber defenders in its Fairwind Program 1. Artificial Analysis ran the model through its ten-evaluation Intelligence Index the same day and scored it 53 2. That ties OpenAI's GPT-6 Astra at 53. It sits five points under Anthropic's Claude Opus 5.5, which the same board scored 58 on Sept. 22 7. So Google has a frontier model again, seven months after its last one, and the frontier line on the board stays at 58. Which of those two facts belongs on a ledger about superintelligence?

Gemini 4 Argon ties GPT-6 Astra, and capability holds flat

The move
Capability holds at 0.0 on reported evidence: Artificial Analysis scored Gemini 4 Argon 53, level with GPT-6 Astra, while Claude Opus 5.5 keeps the top score of 58 and Argon stays limited to cyber defenders.
The data5 rows · sources
MeasureValue
Reading before this day1.6
This piece's move0.0 (Capability, reported)
Band for reported evidence0.3 to 0.8
Reading after the day (with 6 other pieces that day)2.5
Distance to 10097.5

Artificial Analysis scores Argon 53, level with GPT-6 Astra

Start with the outside number, because the method weighs it first. Artificial Analysis called Argon "Google DeepMind's first proprietary model above the Flash class in over 7 months" and placed it at 53 on high reasoning, the top setting Google offers 2. GPT-6.1 Sol sits one point lower at 52 2. Gemini 3.1 Pro Preview, Google's previous model above the Flash class, scored 30. That is a 23-point jump in one generation 2.

Matthias Bastian, who covered the launch for The Decoder, filled in the rest of the table from the same board. "Anthropic's models still lead," he wrote. "Claude Opus 5.5 sits at 58 points and Claude Sonnet 5.5 at 56" 3. Read the order. Two Anthropic models sit above Argon, and Argon shares third place with Astra and Claude Fable 5.1 3.

Agentic results carry the clearest gain. Artificial Analysis put Argon first on its own version of AutomationBench at 78 percent, seven points ahead of Claude Sonnet 5.5 2. Terminal Bench 4 came in at 57 percent, behind Sonnet 5.5 at 64, Opus 5.5 at 60 and GPT-6 Astra at 59 2. Hallucination is the standout. Argon's rate on AA-Omniscience came in at 15 percent, against 51 percent for GPT-6 Astra, and the board wrote that "Argon is much more likely to acknowledge when it does not know an answer rather than guess incorrectly" 2.

Artificial Analysis puts Argon at 53, five points under Claude Opus 5.5

Compared
Artificial Analysis scored Gemini 4 Argon 53, tied with GPT-6 Astra and 23 points above Gemini 3.1 Pro Preview, while Claude Opus 5.5 at 58 and Sonnet 5.5 at 56 still sit above it.Sources [2] [3] [7]
The data6 rows · sources
Unit: Artificial Analysis Intelligence Index score
ItemValueSource
Claude Opus 5.5
max effort
58[7]
Claude Sonnet 5.556[3]
Gemini 4 Argon
high reasoning
53[2]
GPT-6 Astra
max
53[2]
GPT-6.1 Sol
max
52[2]
Gemini 3.1 Pro Preview
Google's previous model
30[2]

The case for the move

Koray Kavukcuoglu, senior vice president of Google DeepMind and Google's chief AI architect, signed the announcement. His opening claim is about work, and the benchmarks come second. "Built to sustain deep reasoning across complex, long-horizon workflows, Argon is fundamentally changing the way we work and build at Google," Kavukcuoglu wrote 1. The verb is "changing," present tense. He describes a model already in use, and the page lists what it did.

Three internal results anchor the claim. A team of Argon agents read fleet-wide profiling data and applied memory fixes across Google's data centers, "freeing up over 300 TiB of memory once rolled out" 1. Argon agents replaced 32,000 lines of SIMD code in libgav1, Google's open video decoder, and the result runs 2.7 times faster than the earlier Rust port with identical output 1. Other agents are moving C and C++ code to Rust, up to more than 800,000 lines for the Fuchsia Zircon kernel 1. Google adds that those large rewrites are still in audits and review before production 1.

Vendor benchmarks point the same way. Google reports 77.9 percent on DeepSWE v1.1, a test of long software engineering tasks 1. Abner Li of 9to5Google read Google's comparison table: "Claude Opus 5.5 comes in at 74.2% followed by GPT-6 Astra's 74.1%" 5. Google also raised the output limit from 64,000 tokens to one million, so a single run can reason for hundreds of thousands of tokens 1.

Why would any of this move a ledger built on clause one of the definition, unsupervised expert work across occupations? Because memory savings and kernel migrations are economically valuable work, done by agents, inside Google, which runs the data centers and can measure the savings. The low hallucination score adds weight here. A model that admits a gap in its knowledge is a model a person can leave alone longer.

Google's harness puts Argon at 77.9 percent on DeepSWE v1.1

Compared
Google reports 77.9 percent for Gemini 4 Argon on DeepSWE v1.1, ahead of Claude Opus 5.5 at 74.2 and GPT-6 Astra at 74.1; all three figures come from Google's own comparison table.Sources [1] [5]
The data3 rows · sources
Unit: percent of DeepSWE v1.1 tasks, Google's table
ItemValueSource
Gemini 4 Argon77.9[1]
Claude Opus 5.5
in Google's table
74.2[5]
GPT-6 Astra
in Google's table
74.1[5]

Kavukcuoglu says Argon changes Google's work; Bastian says Anthropic still leads

Both sides
Kavukcuoglu points to 300 TiB of memory freed and kernel migrations inside Google; Bastian points to the 58 that Opus 5.5 still holds. The evidence holds level, and so does the reading.Sources [1] [3]
The data2 rows · sources
SideWhoClaimSource
ForKoray KavukcuogluArgon sustains deep reasoning across long workflows and is already changing how Google works, freeing over 300 TiB of data-center memory.[1]
AgainstMatthias BastianArgon puts Google back among the top three labs, though Anthropic likely still holds the lead, and the Vals ranking deserves a grain of salt.[3]

The case for a smaller step

Bastian makes the strongest case that the frontier stayed put. "It puts the ad giant back among the top three AI labs, though Anthropic likely still holds the lead," he wrote in The Decoder 3. His headline says the same thing in fewer words: Argon "closes the gap" and "doesn't take a clear lead" 3. A closed gap is a story about Google. The ledger asks a different question, about how far the best public system has gone, and the best score on the board is still 58.

Cost is the second reason. Artificial Analysis measured a promotional price of $1.99 per index task, 60 percent of GPT-6 Astra's $3.26 2. The board then explained where the saving comes from. "This cost efficiency is driven by lower token prices, rather than reduced token use, with Gemini 4 Argon averaging 62k output tokens per task, compared with 27k for GPT-6 Astra (max)," Artificial Analysis wrote 2. Argon spends more than twice the tokens for the same score. Once the 50 percent launch discount ends, the cost per task rises to $3.98, about 1.2 times Astra's 2. Google has yet to confirm when the discount ends 2.

Third, Google's own scoreboard leans on a test Bastian distrusts. Argon leads the Vals Index at 68.9 percent 3. "In this test, though, the mid-tier model Sonnet 5.5 also outranks Anthropic's top model Opus 5.5, so take it with a grain of salt," Bastian wrote 3. Lucas Ropek of TechCrunch placed the launch inside a pattern. "Top AI labs have rushed to release increasingly powerful models, attempting to outdo each other even as those very same firms warn that AI could spin out of control," Ropek wrote 4. His report relays Google's benchmark claims and the Vals ranking, and leaves the checking of Google's own numbers to others 4.

Fourth, people inside Google have doubts. Julia Love and Davey Alba of Bloomberg reported that some Google employees find the model strong on benchmarks and weaker when staff put it to real work, according to The Next Web's account of their story 6. Google told Bloomberg it would be inaccurate to say Gemini 4 underperforms in areas such as coding 6. That dispute stays open until outsiders run the coding tasks themselves.

Argon spends 62,000 output tokens per task to Astra's 27,000

Compared
Artificial Analysis found Gemini 4 Argon averages about 62,000 output tokens per index task against 27,000 for GPT-6 Astra, so its cheaper cost per task comes from a launch discount on token prices.Sources [2]
The data2 rows · sources
Unit: thousand output tokens per index task
ItemValueSource
Gemini 4 Argon
high reasoning
62[2]
GPT-6 Astra
max
27[2]

Google hands Argon to cyber defenders before anyone else

Access is the fifth reason, and it shapes how the ledger counts the release. Argon is "currently being rolled out to selected users and is not publicly available," Artificial Analysis noted 2. Google set the order. "For trusted defenders and our own internal teams at Google, we'll be releasing Argon without cyber guardrails so they can leverage its full frontier-level cybersecurity defense capabilities," Kavukcuoglu wrote 1. Paid API customers and Google AI Ultra subscribers come next, on a date Google left open 1.

The cyber claims are the sharpest on the page. Google says Argon can find, validate and patch critical vulnerabilities on its own, and ties for first on CWE-bench v1 at 68 percent 1. Wiz runs the model through its Scan for Good program, and Google says Argon found a critical flaw in healthcare software used by hospitals worldwide 1. The announcement leaves the software unnamed. Each of these results comes from Google's tests or a partner's internal ones.

Read the release order as strategy. Google is taking part in the U.S. government's voluntary pre-release access process while it widens access 1. By shipping its most capable model to defenders first, with guardrails removed for them, Google is treating offensive cyber skill as the capability to manage. The definition's fourth clause, proof of what a system ran, gets one small line in Google's favour: the page describes monitors that watch Argon's chain of thought and stop execution when needed 1.

What Artificial Analysis measured, and what rests on Google's word

The record
Artificial Analysis measured the index score, the token use and the hallucination rate on launch day; the cyber results, the DeepSWE lead and the internal savings rest on Google's tests, which keeps the move reported.Sources [1] [2]
The data8 rows · sources
ColumnItemSource
Measured by Artificial AnalysisIntelligence Index score of 53, tied with GPT-6 Astra[2]
Measured by Artificial AnalysisAbout 62,000 output tokens per index task[2]
Measured by Artificial AnalysisHallucination rate of 15 percent on AA-Omniscience[2]
Measured by Artificial AnalysisTerminal Bench 4 at 57 percent, behind Opus 5.5 at 60[2]
On Google's page aloneDeepSWE v1.1 at 77.9 percent[1]
On Google's page aloneCWE-bench v1 at 68 percent, tied for first[1]
On Google's page aloneCritical healthcare-software flaw found through Wiz's Scan for Good[1]
On Google's page aloneOver 300 TiB of data-center memory freed by Argon agents[1]

Where this sits on the ledger

The reading measures four clauses: unsupervised expert work across occupations, ten gigawatts acting as one machine, measured self-improvement, and third-party proof of execution. Argon touches the first. An outside board confirmed it on release day, and the confirmation places it level with a model OpenAI shipped earlier and below one Anthropic shipped eight days before.

Capability holds flat, at reported confidence. Reported, because Artificial Analysis measured the index and every other figure rests on Google's harness or its partners. Flat, because the top score stayed at 58 and the public can still buy Opus 5.5 while Argon waits behind the Fairwind Program. Three labs at the frontier is real news for buyers and for pricing. For the ledger, it shows the frontier spreading across companies, and the reading moves only when the frontier itself moves.

By the numbers

  • 53 on the Artificial Analysis Intelligence Index, tied with GPT-6 Astra 2
  • Five points separate Argon from Claude Opus 5.5 at 58 27
  • Gemini 3.1 Pro Preview scored 30, so Argon gained 23 points in one generation 2
  • Output tokens per index task run about 62,000 for Argon against 27,000 for GPT-6 Astra 2
  • Hallucination on AA-Omniscience fell to 15 percent, against 51 percent for Astra 2
  • DeepSWE v1.1 at 77.9 percent on Google's harness, with Opus 5.5 at 74.2 percent 15
  • $1.99 per task at the launch discount, rising to $3.98 at standard pricing 2
  • Over 300 TiB of memory freed in Google's data centers by Argon agents, by Google's account 1

What to watch

An outside run of DeepSWE v1.1 or Terminal Bench 4 that reproduces Google's lead would turn a vendor table into a measured result and could earn capability a small step up. A public release date for paid API customers, with standard pricing, would let buyers test the 62,000-token habit on real work. The Bloomberg dispute settles when developers outside Google post coding results of their own. A score above 58 from any lab is the number that moves this component.

Sources

  1. 1Gemini 4 Argon: our next era of frontier intelligence, Google, Koray Kavukcuoglu, Sept. 30, 2026
  2. 2Gemini 4 Argon: Google is back as one of the top three labs in intelligence achieved, Artificial Analysis, Sept. 30, 2026
  3. 3Google Gemini 4 Argon closes the gap with OpenAI and Anthropic but doesn't take a clear lead, The Decoder, Matthias Bastian, Sept. 30, 2026
  4. 4Google releases Gemini 4 Argon, called its most powerful model yet, TechCrunch, Lucas Ropek, Sept. 30, 2026
  5. 5Google announces Gemini 4 Argon as its new frontier model, 9to5Google, Abner Li, Sept. 30, 2026
  6. 6Google unveils Gemini 4 Argon, and cyber defenders get it first, The Next Web, Ana Maria Constantin, Sept. 30, 2026
  7. 7Claude Opus 5.5 takes the top spot on the Artificial Analysis Intelligence Index, Artificial Analysis, Sept. 22, 2026