SuperIntelligence Infrastructure 2.6%reading
modelsDeepSeekLayer lab

NIST's CAISI Puts DeepSeek V4 Pro Eight Months Behind the US Frontier

NIST's Center for AI Standards and Innovation scored DeepSeek V4 Pro at GPT-5's level on May 1, eight months behind the frontier and cheaper than GPT-5.4 mini on five benchmarks of seven; the ledger holds flat.

● 0.0 Capability confirmed Reading after May 1, 2026: 0.0

By Ryan Elliott Dennis · 10 sources · 9 min read

The Center for AI Standards and Innovation, the NIST unit that tests frontier models for the U.S. government, published its evaluation of DeepSeek V4 Pro on May 1 1. Nine benchmarks across five domains, two of them held out from the public, put the model at an estimated 800 Elo against 1,260 for OpenAI's GPT-5.5 and 999 for Anthropic's Opus 4.6 1. "CAISI evaluations indicate that DeepSeek V4's capabilities lag behind the frontier by about 8 months," the center wrote 1. On the same page it found the model cheaper than GPT-5.4 mini, its chosen U.S. reference, on five of seven benchmarks, at prices that ranged from 53 percent below to 41 percent above 1. DeepSeek had released the weights one week earlier, on April 24, with a technical report that placed its own gap at three to six months 24.

Why does a Chinese model's distance from the American frontier belong on this ledger? Because the reading counts public systems measured on tests that stay clean, and an open-weight model anyone can download is the most public system there is. CAISI ran it on two held-out tests and published the numbers. That is the kind of evidence the method calls confirmed. The question is what the numbers move.

CAISI's DeepSeek measurement leaves the capability reading flat

The move
Capability holds flat on confirmed evidence. CAISI measured DeepSeek V4 Pro on held-out tests at GPT-5's level, a level the public record has carried since August 2025, so the finding adds provenance and a price to the reading.
The data5 rows · sources
MeasureValue
Reading before this day0.0
This piece's move0.0 (Capability, confirmed)
Band for confirmed evidence0.8 to 1.5
Reading after the day0.0
Distance to 100100.0

What CAISI ran

CAISI served the model from cloud H200 and B200 GPUs at DeepSeek's recommended settings, with maximum thinking on, and reproduced DeepSeek's own GPQA-Diamond score before it trusted anything else 1. The evaluation covered cyber, software engineering, natural sciences, abstract reasoning and mathematics. Two benchmarks stayed private: the semi-private ARC-AGI-2 set from the ARC Prize Foundation, and PortBench, a CAISI-built test that asks a model to port a command-line tool from one language to another 1. Agent runs used Inspect's ReAct agent under fixed token budgets 1.

The results split by domain. DeepSeek V4 Pro scored 97 percent on OTIS-AIME-2025, above Opus 4.6's 92, and matched GPT-5.5 on PUMaC 2024 at 96 percent 1. It scored 74 percent on SWE-Bench Verified, beside GPT-5.4 mini's 73 and under GPT-5.5's 81 1. Then the held-out tests arrived. PortBench came in at 44 percent against 78 for GPT-5.5 and 60 for Opus 4.6. ARC-AGI-2 gave 46 percent against 79, and the cyber archive gave 32 against 71 1. Public mathematics looked frontier. Private agentic work looked like August 2025.

CAISI's method turns those scores into one number by treating each model as a student and each task as an exam question, the item-response approach of psychometrics 1. Equal weight goes to each benchmark inside a domain and to each of the five domains. That weighting is a choice, and it is the choice that produces eight months: the two held-out agentic tests carry two fifths of the score between them, and they are where the gap is widest.

CAISI scores DeepSeek V4 Pro at 800 Elo and GPT-5.5 at 1,260

Compared
CAISI's item-response estimate folds nine benchmarks across five domains into one score. DeepSeek V4 Pro lands 460 points under GPT-5.5, and the two held-out agentic tests carry two fifths of the weight.Sources [1]
The data4 rows · sources
Unit: Elo, CAISI item-response estimate
ItemValueSource
GPT-5.51260[1]
Opus 4.6999[1]
DeepSeek V4 Pro800[1]
GPT-5.4 mini749[1]

The case for the finding

CAISI's page is the advocate for its own number, and it says so plainly. "DeepSeek V4 scores better on DeepSeek's self-reported evaluations than on CAISI evaluations," the center wrote, adding that on DeepSeek's data the model sits beside Opus 4.6 and GPT-5.4, released two months before, while on CAISI's held-out tests it sits beside GPT-5, released eight months before 1. Both halves of that sentence come from measurements, and the second half comes from tests kept away from the vendor.

DeepSeek's own report drew the gap narrower. The model "falls marginally short of GPT-5.4 and Gemini 3.1 Pro, suggesting a developmental trajectory that trails state-of-the-art frontier models by approximately three to six months," the company wrote, as Fortune quoted it on launch day 4. TechCrunch carried the same claim on the same day 3. One week earlier the gap had been three to six months by DeepSeek's count. CAISI measured eight.

One footnote on the page argues for the finding more than the headline does. "CAISI scores on SWE-Bench Verified tend to be lower than those of other evaluators, likely due to system prompt, scaffolding, and token budget differences," the center wrote 1. CAISI, by publishing where its harness runs low, shows it has examined that harness. Every model in the table ran under one budget and one scaffold, so the ranking survives the footnote.

DeepSeek claimed a three-to-six-month gap; CAISI measured eight months

The record
DeepSeek's own report placed the model three to six months behind the frontier. CAISI's held-out tests placed it eight months back, level with GPT-5, and the vendor's claim shrank under outside measurement.Sources [1] [4]
The data8 rows · sources
ColumnItemSource
Measured by CAISIGap to the U.S. frontier: about 8 months[1]
Measured by CAISIEstimated 800 Elo against 1,260 for GPT-5.5[1]
Measured by CAISIPortBench, held out: 44 percent against 78 for GPT-5.5[1]
Measured by CAISIARC-AGI-2 semi-private: 46 percent against 79[1]
Measured by CAISIClosest peer: GPT-5, released eight months before[1]
Claimed by DeepSeekThree to six months behind state-of-the-art frontier models[4]
Claimed by DeepSeekMarginally short of GPT-5.4 and Gemini 3.1 Pro[4]
Claimed by DeepSeekBeside Opus 4.6 and GPT-5.4 on self-reported evaluations[1]

Kabanov says the gap is the wrong number

Ilya Kabanov, who writes The Weather Report's research section, read the same page and reached the opposite emphasis. "If you want frontier capability, operational sovereignty, and control over post-training, you have NO other choice but to use a Chinese model," he wrote on May 4, after listing Kimi K2.6, Qwen 3.5, GLM and DeepSeek as the open-weight frontier and placing Llama 4 behind them on capability and on licence terms 6. His argument accepts the eight months and moves past it. A lab that downloads the weights and fine-tunes them on its own hardware buys something a closed API sells at a premium, and CAISI's cost table shows the premium.

That table gives Kabanov his strongest support. GPT-5.4 mini was the one U.S. model that survived CAISI's filter for similar capability at similar price, and DeepSeek V4 Pro beat it on cost on five benchmarks of seven 1. Developer prices on the page put DeepSeek at $1.74 per million uncached input tokens and $3.48 per million output tokens, against $0.75 and $4.50 for the mini model 1. The Register priced GPT-5.5 at $5 per million input tokens and $30 per million output on the same day DeepSeek launched 7. So a model at GPT-5's capability level ships at roughly one ninth of GPT-5.5's output price and undercuts the smallest model OpenAI sells on most of the government's own tests.

Grace Shao, an AI analyst who writes the AI Proem newsletter, told Fortune in July where that price comes from. "Labs are so compute-constrained, capital-constrained, and talent-constrained that a lot of them are being cautious in how they use their resources," Shao said 5. Paul Triolo, a partner at DGA-Albright Stonebridge Group, put the same point in terms of the export controls. "The idea that Meituan could train a 1.6 trillion-parameter model on domestic hardware would have been inconceivable in October 2022," Triolo said, naming the month the U.S. restricted chip sales to China 5. The Register reported that V4 Pro trained on 33 trillion tokens and that its attention design holds a million-token context in 9.5 to 13.7 times less memory than V3.2 7. Huawei said its Ascend processors would carry the model in full, and Fortune reported DeepSeek expected further price cuts as Ascend 950 production scaled 4.

Decrypt collected the third-party leaderboards for its May 4 piece. "The Artificial Analysis Intelligence Index v4.0—a rating system tracking frontier model intelligence across 10 evaluations—shows OpenAI near 60 points and DeepSeek in the low 50s as of May 2026, compressed far tighter than a year ago," Jose Antonio Lanz wrote 8. Public indexes show a closing gap because public tests are what they hold. CAISI's held-out tests show a wider one because vendors train toward what they can see.

DeepSeek launched at one ninth of GPT-5.5's output price, then raised it

Compared
Output prices per million tokens across spring and summer. V4 Pro launched at $3.48 against $30 for GPT-5.5, sat near $0.87 by late July, and rose to $3.96 at peak hours from Aug. 16.Sources [1] [5] [7] [9]
The data5 rows · sources
Unit: USD per million output tokens
ItemValueSource
OpenAI GPT-5.5
April 24
30[7]
GPT-5.4 mini4.5[1]
DeepSeek V4 Pro at launch3.48[1]
Late-July rate for V4 Pro
Fortune's figure
0.87[5]
Peak-hour rate from Aug. 16
as Bloomberg reported
3.96[9]

Why the ledger holds flat

The method sizes a confirmed step at 0.8 to 1.5 points and asks for a measured result on an uncontaminated test. CAISI supplied one. What did it measure? That the best open-weight model in the world sits where GPT-5 sat in August 2025. The ledger's capability component already carried GPT-5's measured level, and it carried GPT-5.5's when that model arrived. A second system reaching a level the public record already held adds provenance and adds a price.

A step down was the other option, and the case for it is real. DeepSeek claimed three to six months and CAISI measured eight. Vendor benchmarks placed the model beside GPT-5.4; held-out tests placed it beside GPT-5. On the skeptical stance the method takes, a claim that shrinks under outside measurement is ordinary cause for a down step. Two things argue against taking it here. The ledger booked the measurement alone, so the record holds a single figure and it stands. And the confirmed band starts at 0.8, a size that would penalise the reading for a model that CAISI itself calls the most capable Chinese system it has tested 1. A move of 0.3 down would sit in the reported band for a confirmed measurement, which misdescribes the evidence. Flat, at confirmed confidence, describes it.

The strongest reason the finding is smaller than it looks is the one Kabanov named. Capability at eight months' remove, at one ninth the price, with the weights in hand, is a different product from the frontier, and the definition this ledger scores counts price only through the capital component. The price then moved. By late July, Fortune priced a million output tokens from V4 Pro at about $0.87 5. On Aug. 13, Bloomberg reported that DeepSeek would raise peak-hour rates to $3.96 per million from Aug. 16, more than four times the current level, with off-peak at $1.98, as founder Liang Wenfeng balanced expansion against capital demands ahead of a possible listing 9. The efficiency story survives that increase. Cheapness took a cut.

Then CAISI's leadership changed. Chris Fall, the former Energy Department official who took over CAISI in late April, resigned in July after three months, TechCrunch reported 10. The center had "released a few reports on the capabilities of Chinese open-weight models Z.ai's GLM-5.2 and DeepSeek V4 Pro," Julie Bort wrote, and TechCrunch's questions to Commerce and NIST about the evaluation process were still open 10. CAISI promised on May 1 to publish a fuller description of PortBench and of its item-response method 1. The ledger treats a measured result as confirmed on the day it is measured, and an evaluator's method as confirmed on the day an outsider reruns it.

CAISI measures an eight-month gap; Kabanov says open weights still win

Both sides
CAISI measured a gap wider than DeepSeek claimed. Kabanov accepts the eight months and prices what open weights buy. Capability holds flat, and price counts through the capital component.Sources [1] [6]
The data2 rows · sources
SideWhoClaimSource
ForCenter for AI Standards and InnovationHeld-out tests place DeepSeek V4 Pro beside GPT-5, about 8 months behind the U.S. frontier, a wider gap than DeepSeek's own evaluations show.[1]
AgainstIlya KabanovBuyers who want frontier capability, operational sovereignty and control over post-training have to choose a Chinese open-weight model, and CAISI's cost table prices the premium.[6]

By the numbers

  • 8 months: CAISI's estimate of the gap between DeepSeek V4 Pro and the U.S. frontier 1
  • 800 Elo for DeepSeek V4 Pro against 1,260 for GPT-5.5, 999 for Opus 4.6 and 749 for GPT-5.4 mini 1
  • Five of seven benchmarks on which DeepSeek V4 Pro undercut GPT-5.4 mini, in a range from 53 percent cheaper to 41 percent dearer 1
  • PortBench at 44 percent and ARC-AGI-2 semi-private at 46, against 78 and 79 for GPT-5.5 1
  • OTIS-AIME-2025 at 97 percent, second in the table to GPT-5.5 at 100 1
  • 1.6 trillion parameters, 49 billion active, one-million-token context, weights released April 24 2
  • Launch output price of $3.48 per million tokens against $30 for GPT-5.5 17
  • Peak-hour price of $3.96 per million from Aug. 16, more than four times the July rate 9

What to watch

CAISI's promised write-up of PortBench and the item-response method would let an outside lab rerun the eight-month figure, which would turn the method itself into a confirmed instrument. A second held-out evaluation, of DeepSeek's next release or of GLM-5.2 on the same tests, would show whether the gap is closing at the vendor's rate or the government's. On the price side, the ledger watches whether DeepSeek's Ascend-backed cuts arrive as Fortune reported or whether the August increase holds. An open-weight model clearing a held-out agentic test at the frontier's level would move capability up by a confirmed step; a second overstatement caught by an outside evaluator would move it down.

Sources

  1. 1CAISI Evaluation of DeepSeek V4 Pro, NIST, Center for AI Standards and Innovation, May 1, 2026
  2. 2DeepSeek-V4 Preview Release, DeepSeek API Docs, DeepSeek, April 24, 2026
  3. 3DeepSeek previews new AI model that 'closes the gap' with frontier models, TechCrunch, Ram Iyer, April 24, 2026
  4. 4DeepSeek unveils V4 model, with rock-bottom prices and close integration with Huawei's chips, Fortune, Nicholas Gordon, April 24, 2026
  5. 5China's Moonshot, Z.AI, and DeepSeek are challenging U.S. AI labs, and beating them on cost, Fortune, Nicholas Gordon, July 26, 2026
  6. 6CAISI Evaluation of DeepSeek V4 Pro, The Weather Report, Ilya Kabanov, May 4, 2026
  7. 7DeepSeek's new models offer big inference cost savings, The Register, Tobias Mann, April 24, 2026
  8. 8US Government Says China's Best AI Models Lag Behind. Experts Aren't So Sure, Decrypt, via Yahoo Tech, Jose Antonio Lanz, May 4, 2026
  9. 9DeepSeek increases prices for AI services by multiple times, Fortune, via Bloomberg, Saritha Rai, Aug. 13, 2026
  10. 10Trump's latest AI czar has already resigned, TechCrunch, Julie Bort, July 20, 2026