DeepSeek CEO Names Huawei Training Chips as Company's Biggest Bet, Q4 Deliveries Expected
A June post-training run on 1,000 Ascend 910C chips cleared the way for a much larger commitment

DeepSeek founder and CEO Liang Wenfeng told investors that moving AI training onto Huawei's domestic accelerators is now one of the company's largest strategic bets, with Huawei expected to begin delivering dedicated training hardware as early as the fourth quarter of 2026, according to The Information. The disclosure — made in a closed investor meeting and reported for the first time on Monday — marks a qualitatively different commitment than what was publicly known before. Where DeepSeek had previously moved Huawei silicon into its serving stack, Liang is now describing the hardware company as essential to where models get built.
The shift matters because training and inference are not interchangeable workloads. Running a trained model to generate responses — inference — is a relatively bounded task that Huawei's Ascend chips have handled for DeepSeek since the R1 era. Teaching a model from scratch on hundreds of billions of tokens, coordinating tens of thousands of processors in tight synchronization for weeks at a time — that is a fundamentally different kind of computation, and one that Huawei's chips failed to complete reliably as recently as August 2025. That history makes Liang's investor statement the most significant signal yet that something in DeepSeek's domestic hardware picture has actually changed.
What Changed Between the R2 Failure and Today
To understand why this announcement carries weight, the August 2025 failure needs to be on the table. After DeepSeek's R1 model rattled markets in January 2025, Chinese government officials encouraged the lab to train its next model, R2, on Huawei's Ascend accelerators rather than Nvidia hardware. The attempt ran into persistent problems documented by the Financial Times and Reuters: unstable multi-chip performance during extended training runs, inter-chip communication slower than Nvidia's interconnect, and Huawei's CANN software framework — Huawei's equivalent of Nvidia's CUDA — lacking the mature tooling and debugging infrastructure that production-scale training requires. Even with Huawei engineers working on-site at DeepSeek's offices for weeks, no successful training run was completed. The lab reverted to Nvidia for R2's pretraining and kept Huawei's chips for the less demanding inference workload.
What changed, according to reporting, is a June 2026 post-training run of V4 Pro — DeepSeek's 1.6-trillion-parameter mixture-of-experts model — completed using approximately 1,000 Ascend 910C chips without instability. Post-training is less demanding than frontier pretraining: it refines an already-completed model rather than building it from scratch, involves a smaller and more bounded compute job, and is more tolerant of the kinds of multi-chip synchronization issues that crashed the R2 runs. The June success does not prove that Huawei silicon can handle a 2-trillion-parameter pretraining run — it proves it can handle the lighter finishing phase. But for Liang, it appears to have provided the evidence needed to commit.
Ascend 950DT: The Chip at the Center of the Training Bet
The hardware Liang is betting on is not the same Ascend 910C used in the June proof-of-concept. The chip at the center of the new training commitment is the Ascend 950DT, scheduled for commercial delivery in Q4 2026. Huawei designed the 950DT specifically for training and decoding workloads that demand high memory capacity and bandwidth — the workload profile that broke the older 910C at scale in 2025.
On paper, the 950DT's memory specifications are notable. It integrates 144 gigabytes of Huawei's proprietary HiZQ 2.0 high-bandwidth memory, delivers 4 terabytes per second of memory bandwidth, and provides 2 terabytes per second of chip-to-chip interconnect. At FP8 precision, it reaches 1 petaflop of compute; at MXFP4, it can reach 2 petaflops. The interconnect bandwidth — 2.5 times that of the predecessor Ascend 910C — is particularly relevant for training, where the bottleneck is often the speed at which thousands of chips can exchange gradients during backpropagation.
For comparison, Nvidia's H200 carries 80 gigabytes of HBM3e memory at 3.35 terabytes per second of bandwidth. On memory capacity and raw bandwidth, the 950DT's published specs are competitive with or ahead of the H200. Where the gap remains severe is in the higher-precision and overall compute tiers: Nvidia's Vera Rubin VR200, due in the second half of 2026, is rated at 35 petaflops of FP4 training compute versus the 950DT's 2 petaflops at MXFP4 — a substantial difference that matters most for the largest and most demanding training runs. Huawei has acknowledged it operates on a "one generation per year, doubling compute with each release" cycle, with the Ascend 960 series targeting 2027 aiming to double the 950's capabilities. The company's own Atlas 950 SuperCluster specification links together more than 520,000 of these chips with a claimed aggregate compute of 524 exaflops.
The CANN Gap and the Software Challenge Nobody Announces
Chip specifications alone do not determine whether a training run succeeds. The R2 failure was partly a hardware story — slow interconnects, instability at scale — but it was equally a software story. Nvidia's CUDA platform carries roughly 15 years of development in debugging tools, optimization libraries, and distributed training frameworks that the entire industry has built against. Huawei's CANN (Compute Architecture for Neural Networks) framework is functionally capable but historically lacking in the depth of that ecosystem.
DeepSeek has been co-engineering CANN adaptations with Huawei since at least early 2026, fine-tuning the software layer that maps DeepSeek's training code to Huawei's accelerator architecture. That work is not publicly detailed, but the June post-training success suggests the software integration has improved substantially from the R2 era. What remains unknown is how far that improvement extends — whether it covers the full sweep of a multi-week, tens-of-thousands-of-chips pretraining run that a 2-trillion-parameter model would require, or whether it covers only the narrower post-training workload that was tested. Nvidia CEO Jensen Huang addressed the threat directly on the Dwarkesh Podcast in May 2026, saying that DeepSeek optimizing for Huawei hardware would be "a horrible outcome for America."
Read more: DeepSeek V4 Flash Vision model opens multimodal agent workflows at text-model prices
DeepSeek's Scaling Ambitions and Why Domestic Silicon Matters
The context for why Liang is making this bet now is the scale of what DeepSeek intends to build next. The lab is currently training a model with 2 trillion parameters and has plans for an eventual 8-trillion-parameter model, according to the investor disclosures as reported by HuggingNews. Those numbers exist in a different category than anything DeepSeek has shipped. V4 Pro, the current flagship, is a 1.6-trillion-parameter mixture-of-experts model in which only 49 billion parameters activate per token — meaning its actual inference compute cost resembles a 49-billion-parameter dense model, not a 1.6-trillion-parameter one.
A genuine 2-trillion-parameter pretraining run would require substantially more raw compute than V4 Pro's training — and an 8-trillion-parameter model would represent a scale of training ambition comparable to the largest runs publicly announced anywhere in the world. At that scale, access to Nvidia hardware through China's constrained supply channels — a combination of stockpiled pre-ban chips, limited H200 allocations under a December 2025 US-China deal that has yet to produce significant deliveries, and undisclosed sourcing — cannot be assumed to scale indefinitely. Domestic silicon, if it works, removes that supply-chain vulnerability.
Nvidia's Position and the Significance of the Forecast
The backdrop against which Liang made his investor statement is a semiconductor market that has shifted with unusual speed. Nvidia's FY2026 10-K, filed with the SEC, states that the company is "effectively foreclosed from competing in China's data center computing/compute market" — the company's own language for the outcome of three years of tightening export controls. Bernstein Research estimates Nvidia's share of China's AI chip market has fallen from approximately 40 percent in 2025 to roughly 8 percent by the end of 2026, while Huawei's share has risen to approximately 50 percent.
That shift was accelerated by specific regulatory moves: the April 2025 ban on even the downgraded H20 chip, which Nvidia had built specifically to comply with existing controls; the September 2025 Chinese government directive telling domestic companies to halt Nvidia GPU purchases; and the limited reprieve of a December 2025 arrangement allowing H200 sales to roughly 10 approved Chinese companies under a 75,000-unit-per-customer cap and a 25-percent US government revenue share, which as of mid-2026 had resulted in very few actual deliveries. The net effect is that China's major AI labs — including DeepSeek — are operating on finite stockpiles of Western silicon and must plan their future training infrastructure around hardware they can actually obtain.
What Needs to Go Right
Several conditions must hold for Liang's bet to pay off. The Ascend 950DT must arrive at scale. Huawei's production capacity for the 950DT is constrained in 2026 by limited availability of the advanced high-bandwidth memory components the chip requires, with estimates placing 2026 output at roughly 200,000 to 300,000 units. DeepSeek's 160,000-chip order for its Ulanqab inference cluster — a 1-gigawatt data center under construction in Inner Mongolia, about 350 kilometers northwest of Beijing — already accounts for a substantial share of that production. Training deliveries will compete with inference demand from DeepSeek and other Chinese labs simultaneously adopting Huawei silicon.
CANN software must mature from post-training reliability to pretraining stability at scale. The R2 failure was specifically a pretraining failure; the June V4 Pro success was a post-training success. Proving Huawei chips can sustain a weeks-long, multi-thousand-chip frontier pretraining run without instability is a different engineering challenge than the post-training job that cleared the path for Liang's announcement.
And the models themselves must be worth training. DeepSeek's 2-trillion-parameter target is more than double V4 Pro's total parameter count. Scaling in MoE architectures does not follow linear rules — routing stability, load balancing across experts, and training convergence at multi-trillion scale on non-Nvidia hardware remain open engineering questions.
What Comes Next
Huawei's accelerated chip roadmap creates a specific near-term test window. The Ascend 960DT, announced at Huawei Connect 2026 in Shanghai on September 17, is now scheduled for Q1 2027 — pulled forward from a prior Q3 2027 target — with a 960PR variant following in Q3 2027. The 960 series is positioned to roughly double the compute and memory capabilities of the 950. If DeepSeek can begin training workloads on 950DT hardware in Q4 2026, the results will be the earliest real-world evidence of whether Liang's bet is paying off — and will shape how aggressively other Chinese frontier labs follow the same path.
The Ulanqab inference cluster, expected to bring its first compute capacity online in late 2027 or early 2028, will also provide a measurable signal: not whether Huawei chips can train models, but whether they can serve them at the throughput levels DeepSeek's approximately 130 million monthly active users in China and its growing global API user base actually require. Inference at that scale on domestic silicon would give other Chinese labs a data point that no prior deployment has provided.
Liang told investors that "within one year, real-world deployment will demonstrate that the ecosystem works." That statement, reported from an investor meeting in mid-2026, now has an implicit deadline: Q3 2027. By then, either Huawei chips will have carried a DeepSeek frontier training run to completion — or the gap between the bet and the outcome will be visible in the compute stack of whatever model DeepSeek ships next.
Read more: Nvidia Vera Rubin NVL72 makes MLPerf debut with 3.7x throughput lead over Blackwell