Nvidia's moat was never really the chip. It was the four million developers who write CUDA. That sentence has been the quiet foundation of the entire AI hardware debate for a decade: you can out-build the silicon, but you can't out-develop the ecosystem, and every rival — AMD first, now everyone else — has learned that the expensive way. So when DeepSeek and Huawei just announced open-source programming tools for Huawei's Ascend chips, with a language called TileLang at the core, the story isn't "China shipped another chip tool." The story is that the most consequential software war in computing now has a new front, and both sides just moved their flagships into it.
Let's walk through what actually happened, why the language matters more than the chips, and why even the analysts who think CUDA's moat is "potentially dead" would still hand Nvidia the next round.
What DeepSeek and Huawei Shipped
Per DeepSeek's official channel and Reuters' reporting, the two companies teamed up to build programming tools for Huawei's Ascend AI chips — libraries for computation and for moving data between chips, with everything made open source. The centerpiece is TileLang, an open-source programming language for AI chips that was originally developed by researchers at Peking University and that DeepSeek has been using for roughly a year. In DeepSeek's own framing, anyone who wants to build an independent software ecosystem for AI chips first needs a universal language that is easy to program but still gets full hardware performance out of the device — and TileLang, they argue, offers a simpler programming model than CUDA. The two companies also optimized a "supernode," a cluster of 128 Ascend 950 chips, as part of the release.
Two details in that paragraph carry most of the weight. First, TileLang isn't a government-mandated tool; it's a university-born language that a frontier lab adopted on its own and has now been running its most advanced work on — The New York Times reports it is now DeepSeek's main tool for its AGI work. Second, the release is open source, which means the ecosystem it's building is, by design, not a walled garden. That is the exact opposite of how CUDA built its moat, and it's the most interesting part of the whole story.
Why the Language Is the Frontier
Here's the structural problem Huawei and the Chinese chip industry have been stuck on: the models are already winning, the software is the bottleneck. The New York Times notes that Chinese model makers like Z.ai and Moonshot AI have moved faster than the country's chipmakers — the frontier is in the model layer, and the hardware layer is racing to catch up with it. Every time a new Chinese frontier model ships, the question isn't "can we run it?" it's "can we run it fast enough, cheaply enough, on our own silicon, at scale?" That question is answered by the software stack, not the chip. CANN, Huawei's existing stack, is already the only one besides CUDA to have supported DeepSeek V4 on day one, by SemiAnalysis' accounting. But a day-one-support stack and a developer-native language are different things. One carries the model; the other teaches the world how to write the next one.
And the context around the release is deliberately full. Two weeks before DeepSeek's announcement, Huawei unveiled new AI processors and supernode systems it says will be widely used for model training next year. Huawei's rotating chairman Eric Xu was unusually blunt about the strategic point: the company plans to sell fewer chips abroad, and, pointing at U.S. export controls, "can't accept a future that hinges on whether others are willing to sell chips to China." That is not a supply-chain statement. That's a sovereignty statement with a road map attached — and the language release is the layer of that roadmap that decides whether the whole thing holds.
The Western Countermove: The Moat Is Cracking, But Not Where You'd Think
This is where it gets genuinely interesting, because the strongest evidence that CUDA's moat is weakening doesn't come from China at all. Research firm SemiAnalysis, after testing OpenAI's in-house inference chip Jalapeño, called the CUDA moat "potentially dead" — OpenAI gets new models running on its own hardware astonishingly fast, and Jalapeño beat Nvidia's Blackwell on performance-per-watt across most of the scenarios tested. OpenAI's models, in a nice twist, helped design the chip that's now running them.
But the caveats are doing real work, and they matter. SemiAnalysis only tested relatively easy-to-optimize scenarios — around 8,000 input tokens and 1,000 output tokens. They haven't run AgentX, the benchmark for multi-step agent workloads, and that's exactly where their August analysis found Nvidia still well ahead: with AMD's current software stack, Nvidia would come out cheaper per token even if AMD gave its hardware away. The authors' conclusion is precise — Nvidia's lasting advantage isn't the silicon. It's the software that links many chips into one system. And Huawei's chips weren't part of the AgentX comparison at all.
So the honest state of play is a three-way split: the simple inference cases are commoditizing fast (Jalapeño, and a dozen others), the agent workloads still crown the incumbent, and the Chinese stack — TileLang plus CANN plus a 128-chip supernode — is the only challenger being built as a coherent, open, nationally-sustained alternative rather than as one more closed ecosystem. That's the framing the "CUDA moat is dead" headlines don't give you.
What It Means
1. The open-source move is the cleverest part, and it's a deliberate mirror of the West's best trick. CUDA won by being a de-facto standard that a four-million-developer community built on. PyTorch won by being open and being the best. China's answer to the "you can't out-develop the ecosystem" thesis is: don't out-develop it, out-share it. A university-born language, a frontier lab's production workloads, a national chipmaker's supernode, and an open license — that's an ecosystem play, not a tools release. The 128-chip supernode optimization is the hint that the next step is the multi-chip system layer, which is exactly where SemiAnalysis says Nvidia's real moat lives.
2. "Simpler than CUDA" is a marketing claim until the compiler wins the war. Every rival has said "easier than CUDA" for a decade, and the reason the moat held is that "simpler" and "full performance" have been in tension — you either abstract the hardware and lose its peak, or you stay close to the metal and lose your developers. TileLang's claim is that the tile-based model resolves that tension. The only way to test it is the ugly, unglamorous benchmark: take a frontier model, port it, and measure tokens-per-dollar on the hardest workloads — the agent ones, the long-context ones — not the 8,000-token demos. Watch for that number. It's the whole story in one metric.
3. The real contest is now between three software stacks, not two. Pre-Jalapeño, the story was "CUDA vs. everything." Post-Jalapeño and post-AgentX, it's "CUDA still leads the agent layer, in-house silicon owns the inference layer it serves, and the Chinese stack is the only one building all three layers — chip, language, system — as a single national program." That asymmetry is what "closing ranks" actually means. It's not an alliance. It's a stack. And a stack that's open on the outside and coordinated on the inside is a genuinely new species of AI infrastructure.
4. The four-million number is the clock, and it's running on both sides. Ecosystems are measured in developers, and developers are measured in time. CUDA's four million is a stock, not a flow. TileLang's open license and a national training mandate are building a flow. The question that decides the chip war of the next decade isn't "which chip is faster" — it's "which language do the next ten million AI engineers learn first?" Every university that teaches TileLang, every model that ships with it as the default, every supernode that's optimized for it, is a deposit in that future account. The chips are the headline. The language is the balance sheet.
🔥 Hot Takes
1. "Potentially dead" is the most dangerous phrase in the whole chip war, and it's coming from the most credible voice in the room. When SemiAnalysis says the CUDA moat is potentially dead — on the strength of a single in-house chip beating Blackwell on perf-per-watt in easy scenarios — that's not a eulogy. That's a warning shot to Nvidia that the moat's outer ring has already been breached by its own customer's customer. The moat wasn't a wall around Nvidia; it was a wall around a workflow. The workflow is now being rebuilt in three places at once: OpenAI's silicon, AMD's stack, and China's open language. The question isn't whether the moat dies. It's whether it dies fast enough for Nvidia to still be the center of gravity when the dust settles.
2. The open license is the actual strategic weapon, and it's aimed at Western labs as much as Western chips. Think about who TileLang's open source really targets. It's not just Chinese developers. It's every lab on Earth that's quietly worried about being locked to one vendor's toolchain — which is, this week, a lot of labs, given how fast the in-house silicon wave is moving. A Peking University language that runs Huawei, runs old Nvidia chips, and is the main tool for one of the world's frontier model programs is a universal option with a free tier. That's not a national project. That's a global developer bet, funded by a national one. The West should be scared of the free tier.
3. The 128-chip supernode is the tell that this is a systems war, not a language war. Nobody optimizes a 128-chip cluster for a language. You do it because you know the next workload — the long-horizon, multi-step, many-chip agent workload that SemiAnalysis says still crowns CUDA — is the next battlefield, and you want to be the one who's already wired for it. Huawei is not shipping a compiler. It's shipping the answer to the AgentX question before AgentX has finished being asked. That's a two-years-ahead move wearing a release-notes disguise, and it's the single most important data point in the entire story.
The Bottom Line
Strip the press-release gloss and this is a clean statement about where the AI hardware war actually is: not in the die, not in the benchmark, but in the language. DeepSeek and Huawei just open-sourced a simpler, universal, national-scale answer to the four-million-developer moat, wired it to a 128-chip supernode, and pointed it at the agent workloads everyone quietly knows is the next prize. The West's own in-house silicon wave has cracked CUDA's outer ring; China's coordinated open stack is the only player rebuilding the entire layer from language to system. The chip war's center of gravity just moved up one abstraction, and both sides felt it the same week. The developers will decide it — and for the first time in the CUDA era, the developers have two open doors to walk through.