DeepSeek Open-Sources Six Huawei Ascend Kernel Libraries as CUDA-Free AI Stack Takes Shape
TileLang, DeepGEMM, DeepEP, and FlashMLA give Chinese AI labs a production-tested migration path off Nvidia

On September 30, DeepSeek published Ascend NPU support for six software infrastructure components built for Huawei's Ascend 950 neural processing unit — releasing not just an additional model or benchmark, but the programming language, matrix computation library, distributed communication library, operator collection, attention kernel, and data-selection tooling that together form a coherent alternative to Nvidia's CUDA software ecosystem. The release arrives as Huawei's Ascend platform is set to receive its first dedicated training chip shipments in Q4 2026, and as Chinese enterprises place large Ascend 950 orders following DeepSeek's successful deployment of its V4.1 model on domestic hardware.
The timing is deliberate. DeepSeek has structured the release to coincide with the moment Chinese AI labs move from evaluating Ascend hardware to committing production workloads to it. The software stack now in the open — TileLang, DeepGEMM-Ascend, DeepEP-Ascend, TileKernels, FlashMLA-Ascend, and DeepSelect — mirrors its Nvidia-facing counterparts API for API. Developers who already use DeepSeek's Nvidia tools are intended to switch backends, not rewrite codebases.
Read more: DeepSeek CEO bets big on Huawei training chips
Why Nvidia's Real Moat Is Software, Not Silicon
Understanding the significance of this release requires understanding why China's AI labs have remained dependent on Nvidia hardware even as alternatives have improved. Huawei's Ascend 950DT chip, with one petaflop of FP8 throughput and two petaflops at MXFP4, is not trivially inferior to Nvidia's H100 on raw specifications. The problem has been software.
Nvidia's Compute Unified Device Architecture, introduced in 2007, accumulated nearly two decades of investment in surrounding infrastructure. cuBLAS handles matrix operations, cuDNN provides deep learning primitives, NCCL manages collective communication across GPU clusters, and frameworks including PyTorch and TensorFlow were built assuming CUDA as the substrate. Switching accelerator hardware does not just mean recompiling — it means porting every custom kernel, every communication primitive, and every profiling and debugging workflow to a new target. Historically, that represented six to twelve months of engineering per major operation.
DeepSeek's strategy is to eliminate that friction through direct API mirroring. Each of the six open-sourced components maps onto a specific layer of the Nvidia ecosystem, providing the same calling convention and roughly equivalent functionality on Ascend. The goal, documented across DeepSeek's open infrastructure repositories, is to make migration a backend swap rather than a rewrite.
TileLang: A Multi-Backend Compiler Built for Hardware Independence
TileLang is the keystone of the stack. It is a domain-specific language using Python syntax, with a compiler built on Apache TVM, that allows developers to write high-performance AI compute kernels without targeting a specific hardware instruction set directly. The compiler generates code for the target platform — Nvidia GPUs from SM70 through SM120, AMD ROCm, Apple Metal, and now Huawei Ascend 950 — from a single source.
For the Ascend backend, TileLang wraps Huawei's Ascend C, a C-variant designed for NPU programming. Rather than requiring engineers to manage Ascend's fractal matrix layouts, alignment constraints, and address calculation manually, TileLang handles the translation automatically. DeepSeek says the resulting code targets the hardware's performance ceiling rather than trading efficiency for portability — a claim the company supports with benchmark numbers: TileLang-compiled GEMM kernels reportedly reach 0.98 times the performance of hand-written Ascend C, with attention and vector kernels at 0.95–0.96 times. Those are project-published figures, not independently reproduced.
The deeper significance of TileLang is architectural. The analogy to LLVM is apt: just as LLVM provided a hardware-neutral intermediate representation for CPU compilers that eventually supported dozens of architectures, TileLang aims to serve the same function for AI accelerators. DeepSeek has explicitly stated that its goal is for TileLang to become a reference for other Chinese AI chip makers to adopt — positioning the language not as a Huawei-exclusive tool but as common infrastructure that any NPU vendor could support.
DeepGEMM and DeepEP: Getting to the Hardware Limits
Matrix multiplication is the computational core of large model training and inference, so DeepGEMM-Ascend's stated performance numbers are the most consequential in the release. On the Ascend 950DT, DeepSeek claims Dense GEMM achieves 99.8 percent of the hardware's theoretical BF16 peak, with FP8 at 99.5 percent and FP4 at 98.3 percent.
These figures, if they hold independently, would indicate that the library is not just functional but competitive at extracting hardware throughput — historically the hardest engineering challenge when porting from one architecture to another. The implementation abstracts Huawei's matrix accumulate and distribute (MAD) primitives behind the same API surface as DeepSeek's Nvidia-targeting DeepGEMM, and adds coroutine-based pipeline optimization specific to Ascend's memory hierarchy. It supports the BF16, FP8, and FP4 precision formats that DeepSeek's V4 model family uses across training and inference, as well as the multi-query attention logits (MQA logits) and MegaMoE operators that appear in current DeepSeek model architectures.
DeepEP-Ascend addresses a different bottleneck: collective communication in Mixture-of-Experts clusters. MoE architectures like DeepSeek V4, which routes tokens to a small subset of 256 expert sub-networks per forward pass while keeping 1.6 trillion parameters in total, require every accelerator in a training cluster to exchange token activations with every other after each routing step. At scale, this all-to-all communication pattern routinely becomes the binding constraint rather than compute. DeepEP's Ascend version is implemented on Huawei's HCCL, HCOMM, UBMEM, and URMA communication primitives, with kernels compiled at runtime via DeepJIT. In tests on the Ascend 950DT's 128-card supernode configuration, dispatch bandwidth reaches 90 to 95 percent of the physical network ceiling at expert-parallel scales up to 32. Performance above 32 expert-parallel groups is not reported.
FlashMLA and DeepSelect: Sparse Computation on Long Contexts
FlashMLA-Ascend handles DeepSeek's Multi-head Latent Attention architecture, the compression scheme introduced in V2 that dramatically reduces KV cache memory by projecting keys and values into a low-dimensional latent space before attention is computed. The Ascend version adds DeepSeek Sparse Attention (DSA), a technique that uses a component called Lightning Indexer to pre-select the top-K most relevant tokens for each query before the attention operation runs, cutting the effective computation on long sequences substantially.
The performance numbers for FlashMLA-Ascend are striking as headline figures. Prefill throughput reaches 410 teraflops on the Ascend 950, which the company characterizes as 95 percent of the chip's theoretical peak. Decoding throughput hits 360 teraflops, or 83 percent of peak. Both figures are company-reported. DeepSeek has also published a technical report documenting the optimization techniques behind the Ascend prefill implementation. The FP8 KV cache format for V4.1 consumes 528 bytes per token; the FP4 variant compresses this to 288 bytes — a reduction that compounds across the long contexts where sparse attention matters most.
DeepSelect, the smallest of the six components, optimizes the TopK operation that appears at two critical points in DeepSeek's inference pipeline: Lightning Indexer's token selection, where up to 512 tokens are chosen from a very long sequence per query in V4, and vocabulary sampling, where the inference system selects from a 128,000-token vocabulary at each generation step. Both are hot-path operations that run once per generated token. According to DeepSelect's GitHub repository, the library's Ascend NPU support was added on September 30, 2026. DeepSeek claims DeepSelect runs two to twenty times faster than PyTorch's native torch.topk on the Ascend platform. The wide performance range reflects the sensitivity of TopK efficiency to input size and distribution.
The Supernode Ambition: 128 Chips Working Together
Individual chip performance is only one part of the infrastructure challenge. DeepSeek and Huawei have also co-designed an Ascend 950 supernode configuration that links 128 chips into a single computing unit, with inter-chip interconnect and collective communication architecture built to support the workloads these libraries produce.
The 128-chip scale matters for MoE training in particular, because expert-parallel routing creates all-to-all traffic patterns that stress interconnect rather than compute. Designing DeepEP around the specific topology of a 128-card cluster — using Huawei's external Clos network and optimizing dispatch for that traffic pattern — means the library's bandwidth numbers are benchmarked against a real production cluster rather than a synthetic test. The collaboration also signals intent: Huawei and DeepSeek are not publishing tooling for individual developers experimenting with a few chips. The infrastructure target is cluster-scale AI training and inference.
DeepSeek's June 2026 post-training run on approximately 1,000 Ascend 910C chips — the first to complete without instability after an August 2025 failure on Huawei hardware — demonstrated that this was achievable, but the software stack available then was internal and proprietary. The open-source release makes the approach reproducible for any organization with Ascend hardware access.
Read more: DeepSeek CEO bets big on Huawei training chips
What the Gap Still Looks Like
The performance numbers DeepSeek has released are impressive as stated. But several important limitations remain that the announcement does not address.
Every benchmark figure in this release is self-reported by the developer and has not been independently reproduced. The GEMM efficiency numbers (99.8 percent of BF16 hardware peak), the FlashMLA throughput figures, and the DeepSelect speedup claims await external verification. This is not a reason to dismiss the numbers, but it is a reason not to treat them as established facts — particularly given that hardware vendors and AI labs routinely measure efficiency under favorable conditions.
Huawei's CANN software development kit, which sits beneath DeepSeek's tooling layer and provides the low-level compiler and profiling infrastructure, remains substantially less mature than Nvidia's CUDA tooling ecosystem. Debugging novel Ascend kernels is harder than debugging CUDA kernels because fewer developers have done it, fewer debugging tools exist, and community knowledge is more limited. DeepSeek's open-source release addresses the applications layer; it does not solve the deeper tooling maturity gap that cost DeepSeek its first Huawei training run in August 2025.
The community gap is stark. CUDA developers number in the millions globally, with decades of accumulated answers on developer forums, open-source repositories, and documentation. The Ascend kernel development community is orders of magnitude smaller. Open-sourcing TileLang and the operator libraries accelerates growth in that community, but it does not instantly create the accumulated expertise that CUDA has built since 2007.
DeepSeek as China's Default AI Platform Maintainer
The most consequential implication of this release is not the performance of any individual library. It is the platform position that DeepSeek is constructing.
By open-sourcing tools that API-mirror Nvidia's stack, DeepSeek is not just helping Huawei compete with Nvidia. It is positioning itself as the de facto maintainer of the software layer that sits between Chinese AI chips and Chinese AI applications — a position analogous to what Meta occupies globally with PyTorch. Organizations that adopt TileLang for kernel development, DeepGEMM for matrix computation, and DeepEP for cluster communication will build engineering workflows around DeepSeek's design choices, versioning decisions, and API evolution. DeepSeek stated explicitly that its goal is for TileLang to become a model for other Chinese chip makers, which means that if successful, the language's authors effectively set the interface standard for China's emerging domestic AI silicon market.
This platform dynamic has compounding value independent of the specific efficiency numbers. Nvidia's sustainable advantage in AI compute comes not from being the only chip that can train a neural network, but from the accumulated software investment — in frameworks, debugging tools, performance profilers, and developer habit — that makes switching costs prohibitive. DeepSeek is building the same dynamic on the Ascend side, and doing so openly, which removes the argument that domestic chip software will always lag behind a proprietary ecosystem.
Liang Wenfeng, DeepSeek's founder, set a one-year window — implicitly pointing to Q3 2027 — for real-world Ascend deployment to demonstrate that the ecosystem is viable. With six production-tested open-source components, a jointly specified 128-card supernode, and V4.1 already running in production on Ascend hardware, the remaining test is whether a broader developer community can validate and extend these tools beyond DeepSeek's own workloads. The Huawei Ascend 950DT training chip deliveries expected in Q4 2026 will bring that community its most direct test yet.