SuperIntelligence Infrastructure 2.6%reading
computeDeepSeekLayer L0

DeepSeek Open-Sources Huawei Ascend Kernels and Claims 99.8% of Chip Peak

DeepSeek released DeepGEMM-Ascend on Sept. 30 under the MIT license, reporting up to 99.8 percent of Ascend 950 peak; Venkat Somala counters that Huawei ships under 4 percent of Nvidia's compute.

▲ +0.3 Compute reported Reading after Sept. 30, 2026 (7 pieces that day): 2.5

By Ryan Elliott Dennis · 6 sources · 7 min read

DeepSeek published DeepGEMM-Ascend on GitHub on Sept. 30, a matrix-multiply kernel library ported to Huawei's Ascend 950, and released it under the MIT license 1. The library ships API-compatible with DeepSeek's CUDA version, so a developer installs one package and keeps the same workflow on Chinese silicon 1. DeepSeek reports that dense matrix multiplication reaches up to 99.8 percent of the hardware limit on an Ascend 950DT, with BF16 at 99.8 percent, FP8 at 99.5 percent, and FP4 at 98.3 percent 1. The company says it co-optimized the kernels with Huawei for a 128-chip Ascend 950 supernode 6. Bloomberg reports DeepSeek plans to deploy at least 160,000 Huawei accelerators in Inner Mongolia 2.

Why does a software release belong on a ledger that measures megawatts and energized compute? Because the compute component asks whether frontier compute is being built and used at scale, and the single largest tax on using Huawei silicon has been software. CUDA is the moat. A kernel library that approaches peak on a domestic chip chips at that moat, and the question is how much.

What DeepSeek shipped on Sept. 30

DeepGEMM-Ascend is one of six tools DeepSeek open-sourced that day, alongside TileLang, DeepEP, FlashMLA, and more, all aimed at programming Huawei's accelerators 4. TileLang is the piece Bloomberg frames as China's answer to CUDA, the language layer a developer writes against 2. FlashMLA reached 95 percent of theoretical performance on the Ascend 950 for some operations, by DeepSeek's account 4. Each tool arrives with the kernels DeepSeek uses in its own training and inference, with the production path intact 6.

DeepSeek fixed the port on one hardware target. The company validated the library on the Ascend 950 series using Huawei's CANN 9.20 toolkit, and the benchmark table lists exact shapes drawn from the DeepSeek model series 1. Dense FP4 reached 1,701 TFLOPS against a 1,730 TFLOPS hardware limit, and BF16 reached 431 against 432 1. Those are the numbers that produce the 99.8 percent figure. Peng Zhang, writing for Geopolitechs on Sept. 30, recorded DeepSeek's own summary of the work: "Every TileLang operator currently used in DeepSeek training now has a corresponding high-performance implementation on Ascend" 6.

The case for the move

DeepSeek states the claim plainly in the README. "Despite its lightweight codebase, DeepGEMM Ascend can achieve peak hardware performance across a wide range of matrix shapes," the company wrote 1. Read the sentence for what it concedes and what it asserts. Lightweight is the concession, since a thin kernel library is easier to audit and port than a sprawling one. Peak hardware performance across many shapes is the assertion, and it is the whole value of the release: a chip a developer could use at a fraction of its silicon before now runs matrix multiplication near its ceiling.

A second DeepSeek line, carried by Bloomberg, states the strategy behind the code. "A sophisticated programming language with a simple coding process is essential for building an effective AI ecosystem," the company said in a WeChat post 2. That sentence names the real target. An ecosystem, meaning the developers, the libraries, and the habits that keep workloads on one vendor's chips, is what CUDA defends and what TileLang attacks. DeepSeek acknowledged Huawei's engineers in the README and described their help as "unreserved and vigorous support" during the research, with more collaboration promised 12.

The move matters because the actor is DeepSeek itself, moving its own production stack. DeepSeek post-trained its V4-Pro model on Huawei's older Ascend 910C in June 2026, so the port follows real use ahead of any press cycle 4. Every Chinese developer who trains on domestic silicon now inherits open kernels, which lowers the cost to follow. That is a compute-component story, told in software.

What DeepSeek and Huawei tied together on Sept. 30

Who connects
DeepSeek released six open-source tools built on Huawei's CANN toolkit and co-optimized for a 128-chip Ascend 950 supernode, moving its matrix-multiply stack off CUDA onto domestic silicon.Sources [1] [3] [4] [6]
The data6 rows · sources
FromLinkToSource
DeepSeekreleased Sept. 30DeepGEMM-Ascend[1]
DeepSeekopen-sourcedTileLang[4]
DeepSeekco-optimized supernodeHuawei[6]
DeepGEMM-Ascendnear-peak kernelsAscend 950[1]
DeepGEMM-Ascendbuilt on toolkitCANN 9.20[1]
Huaweimakes the chipAscend 950[3]

A vendor benchmark earns only the reported step

The 99.8 percent figure arrives from DeepSeek's own harness, which is the reason the step stays small. Sahil Khanna, writing for Lapaas Voice on Sept. 30, read the release the way the ledger reads it. "This is a development milestone with a defined deployment gap, rather than evidence of a fully interchangeable chip ecosystem," Khanna wrote 5. He noted that DeepSeek listed its pipeline-parallel and remote-memory features as experimental, and its reduce-scatter and all-reduce kernels as still being built 5. Khanna also flagged the benchmark itself as "vendor-reported results" from "a company benchmark" that outside parties have yet to verify 5.

Open code and a running fleet are two different milestones, and the Sept. 30 release cleared the first 5. Peak utilization measured against the Ascend chip says how well the kernel uses that chip. It stays silent on how the chip compares with the part it would replace, a distinction one Geopolitechs reader drew sharply 6. So the reading moves on a real engineering result, held to the reported band because a vendor measured it, the hardest communication kernels remain unfinished, and the recommended commercial kit publishes around Oct. 15 5.

What the release delivered, and what rests on DeepSeek's word

The record
Open code and the Ascend port are on the record; the 99.8 percent peak, the undelivered chips, and the unfinished communication kernels rest on DeepSeek's account, which holds the step to reported.Sources [1] [2] [4] [5]
The data6 rows · sources
ColumnItemSource
On the recordDeepGEMM-Ascend released under the MIT license on Sept. 30[1]
On the recordSix tools open-sourced, including TileLang and FlashMLA[4]
On the recordSupports BF16, FP8 and FP4 on the Ascend 950[1]
Claimed or open99.8 percent of peak is DeepSeek's own benchmark[5]
Claimed or openPipeline-parallel and remote-memory features marked experimental[5]
Claimed or openAscend 950 deliveries and energized sites stay uncounted[2]

Why Huawei still ships under 4 percent of Nvidia's compute

Venkat Somala of Epoch AI set the ceiling on how far any kernel library can carry Huawei this year. In a report dated Sept. 4 and updated Sept. 24, Somala wrote: "Its most powerful chip, the Ascend 950, delivers roughly half the performance of Nvidia's H100, which began shipping in 2022, and Huawei will produce less than 4% of Nvidia's compute output this year" 3. Half an H100, three to four years late, is the chip these kernels run near peak on.

Volume is the harder wall. Epoch estimates Huawei will produce 1.5 million units in 2026 against Nvidia's 5.9 million, and on a common scale that reads as 880,000 H100-equivalents of Huawei compute against 23,000,000 for Nvidia 3. Somala's conclusion sets the horizon for the whole effort: "Huawei will substantially improve its AI compute solutions by 2030, but export controls constrain its most important scaling levers, making it unlikely to catch up with Nvidia" 3. Software efficiency multiplies the chips a buyer holds. Multiplying 880,000 by a near-peak kernel still leaves a figure far under 23 million, and the gap in usable high-bandwidth memory ran to 73 times in 2026 by Epoch's count 3.

The two sides agree on more than they contest. DeepSeek argues that the software tax on domestic silicon just fell, which is true. Somala argues that the silicon under the software stays scarce and a generation behind, which is also true. One result raises how much of each Ascend chip a developer can reach; the other caps how many Ascend chips exist to reach.

Huawei's 2026 compute against Nvidia's, on one scale

Compared
Epoch estimates 880,000 H100-equivalents of Huawei compute in 2026 against 23,000,000 for Nvidia, the ceiling a near-peak kernel library leaves in place.Sources [3]
The data2 rows · sources
Unit: H100-equivalents
ItemValueSource
Huawei
Epoch 2026 estimate
880000[3]
Nvidia
Epoch 2026 estimate
23000000[3]

Where this sits on the ledger

The reading measures four clauses: unsupervised expert work across occupations, ten gigawatts acting as one machine, measured self-improvement, and third-party proof of execution. This release touches the second clause, the physical one, from the software side. A kernel library that reaches near peak lets more of a chip's rated compute turn into useful work, so the effective supply of domestic compute rises even while the chip count holds. Ten gigawatts acting as one machine still needs the chips, the interconnect, and the power, and a kernel delivers software alone.

So the step is 0.3, in the compute component, at reported confidence. A larger step waits on an independent benchmark of DeepGEMM-Ascend on delivered hardware, or on a count of Ascend 950s energized in Inner Mongolia. Awarding a full point on a vendor benchmark would read the record wrong. The step adds 0.3 to the day's reading, and the reading moves further the day an outside team runs these kernels on chips it can see.

A small step up in compute, held to the reported band

The move
Compute moves up 0.3 at reported confidence, inside the 0.3 to 0.8 band: DeepSeek open-sourced near-peak Ascend kernels, yet a vendor measured the peak and the chips stay undelivered.
The data5 rows · sources
MeasureValue
Reading before this day1.5
This piece's move+0.3 (Compute, reported)
Band for reported evidence0.3 to 0.8
Reading after the day (with 6 other pieces that day)2.5
Distance to 10097.5

By the numbers

  • 99.8 percent of the Ascend 950 hardware limit for BF16 dense matrix multiplication, by DeepSeek's benchmark 1.
  • Six open-source tools released Sept. 30, including DeepGEMM-Ascend, TileLang, DeepEP, and FlashMLA 4.
  • 128 chips in the Ascend 950 supernode DeepSeek and Huawei co-optimized 6.
  • At least 160,000 Huawei accelerators DeepSeek plans to deploy in Inner Mongolia 2.
  • Under 4 percent of Nvidia's 2026 compute output is Epoch's estimate for Huawei 3.
  • Roughly 880,000 H100-equivalents of Huawei compute in 2026, against 23,000,000 for Nvidia 3.
  • 73 times the gap in usable high-bandwidth memory between the two in 2026, by Epoch's count 3.
  • Around Oct. 15 is when the recommended commercial kit reaches public availability 5.

What to watch

An independent benchmark of DeepGEMM-Ascend on an Ascend 950DT, run by a party other than DeepSeek, would turn the 99.8 percent claim into a measured result and earn a confirmed step. A public count of Ascend 950s energized in the Inner Mongolia site would engage the compute clause directly. Huawei's delivery of the commercial kit near Oct. 15, and the completion of the reduce-scatter and all-reduce kernels DeepSeek marked experimental, would show whether the stack reaches production. A revised Epoch estimate of Huawei's 2026 output is the marker that would move the ceiling either way.

Sources

  1. 1DeepGEMM Ascend, DeepSeek, Sept. 30, 2026
  2. 2DeepSeek unveils Huawei AI chip tools that may replace Nvidia's, The Star (Bloomberg), Luz Ding, Sept. 30, 2026
  3. 3Will Huawei catch up to Nvidia by 2030?, Epoch AI, Venkat Somala, Sept. 24, 2026
  4. 4DeepSeek has partnered with Huawei to open-source six AI development tools, Gigazine, log1d_ts, Oct. 1, 2026
  5. 5Huawei AI Chips Get DeepSeek Code, But Face a Deployment Gap, Lapaas Voice, Sahil Khanna, Sept. 30, 2026
  6. 6DeepSeek Builds for Huawei Ascend, Geopolitechs, Peng Zhang, Sept. 30, 2026