🐾 LIVE
Chinese Tech Workers Are Training Their AI Replacements — And Fighting Back Xiaomi miclaw Becomes China's First Government-Approved AI Agent OpenAI's Quiet Acquisitions Signal Existential Questions About Its Future Google Gemini Launches Native Mac App: The Desktop AI Wars Are On Cerebras Files for IPO at $23B, Backed by $10B OpenAI Partnership DeepSeek Raising $300M at $10B Valuation — While Remaining Profitable ByteDance vs Alibaba vs Tencent: China's AI Video War Heats Up Chinese Tech Workers Are Training Their AI Replacements — And Fighting Back Xiaomi miclaw Becomes China's First Government-Approved AI Agent OpenAI's Quiet Acquisitions Signal Existential Questions About Its Future Google Gemini Launches Native Mac App: The Desktop AI Wars Are On Cerebras Files for IPO at $23B, Backed by $10B OpenAI Partnership DeepSeek Raising $300M at $10B Valuation — While Remaining Profitable ByteDance vs Alibaba vs Tencent: China's AI Video War Heats Up
Infra

AMD's Helios Just Declared War on Nvidia — And the AI World Should Care

The first rack-scale system that beats Vera Rubin on paper, backed by $5B Anthropic deal and a customer list that reads like Silicon Valley's greatest hits

2026-07-24 By AgentBear Editorial Source: TechCrunch + The Register + CNBC + The Decoder 12 min read
AMD's Helios Just Declared War on Nvidia — And the AI World Should Care

At AMD's sold-out Advancing AI conference in San Francisco this week, CEO Dr. Lisa Su unveiled Helios — the company's first true rack-scale AI system, and the most serious challenge Nvidia has faced in the datacenter since the GPU revolution began.

This isn't just another chip launch. Helios is a complete infrastructure play: 72 GPUs per rack, liquid-cooled, using AMD's own CPUs, networking, and software stack. And according to AMD's own specs, it outperforms Nvidia's Vera Rubin in several key metrics while being significantly larger and more powerful.

But the real story isn't just about specs. It's about what Helios represents: the first time a single company has assembled enough pieces — GPUs, CPUs, networking, software — to offer a credible alternative to Nvidia's vertically integrated empire. And the customer list proves it's not vaporware.

The Specs That Matter

Helios is built around the new Instinct MI455X GPU, based on CDNA 5 architecture and fabricated on TSMC's bleeding-edge 2nm process. Each GPU uses a "silicon sandwich" design with 24 chiplets — 8 compute dies stacked on 2 fabric/cache dies, all interconnected with advanced packaging. The result? A chip that delivers up to 4x higher floating-point performance for AI workloads compared to last year's MI355X.

The MI455X makes some bold architectural choices. It drops FP64 entirely — no double-precision math — dedicating every transistor to AI-centric datatypes like MXFP4 and MXFP8. It also ditched its last-level "Infinity" cache in favor of a larger shared L2 cache that delivers 1.5x the bandwidth of the entire Infinity cache on the previous generation.

Here's what makes Helios interesting:

The rack measures 1.2 meters wide and stands 44U tall — nearly twice the size of Nvidia's NVL72. It packs 72 MI455X GPUs across 18 liquid-cooled compute blades, each containing four GPUs, a single 96-core Venice EPYC CPU clocking up to 5 GHz, and networking gear from AMD's Pensando acquisition.

For scale-up communication between GPUs, Helios uses UALoE (Ultra Accelerator Link over Ethernet) with 12 Broadcom Tomahawk 6 switch ASICs spread across six switch trays. Each provides 512 lanes of 200 Gbps connectivity, giving each MI455X 3.6 TB/s of bidirectional bandwidth. This means any GPU can talk to any other while keeping latency minimal — and critically, it uses standard Ethernet switches instead of Nvidia's proprietary NVLink.

For scale-out networking, each MI455X gets three 800 Gbps Pensando Vulcano network cards for a total of 2.4 Tbps — significantly beefier than Nvidia's single 1.6 Tbps ConnectX-9 superNIC per GPU.

All of this is powered by a 50V liquid-cooled DC bus bar that can deliver between 225 and 245 kW under load. If AMD's power estimates hold, Helios would be both faster than Vera Rubin and more power-efficient.

The Customer List Is Uncomfortable for Nvidia

AMD didn't just announce Helios — they announced who's buying it, and the names are staggering:

Eight of the top ten AI companies now run workloads on AMD Instinct GPUs. That's not a trickle — that's a flood.

Futurum Group estimates Helios costs $5-5.5 million per rack vs. $3.5-4 million for Nvidia's Vera Rubin. But AMD's pitch is about total cost of ownership and lowest cost per token — not upfront hardware price. Forrest Norrod, AMD's data center head, called it "our baby" and emphasized: "We're very focused on providing the best total cost of ownership, the lowest cost per token, all in."

The Cerebras Partnership: Inference Speed Redefined

On the same day, AMD announced a collaboration with Cerebras Systems to create a disaggregated AI inference platform. The concept is elegant: AMD Instinct GPUs handle compute-heavy prompt processing, while Cerebras' SRAM-based Wafer Scale Engines handle memory-intensive token generation.

Cerebras CEO Andrew Feldman was blunt about why this matters: "Nvidia is just an arms dealer." Cerebras' wafer-scale chips don't need HBM4 — their on-chip SRAM is orders of magnitude faster. The result? Inference speeds exceeding 2,000 tokens per second, with up to 5x better tokens-per-watt compared to GPU-only approaches.

This directly competes with Nvidia's $20 billion Groq acquisition, which aimed to solve the same inference problem with LPUs (Language Processing Units). Where Nvidia needs 2,000 Groq LPUs to serve a trillion-parameter model like Kimi K3, AMD+Cerebras needs at most a few dozen WSE chips.

Su said during a press conference: "You can expect that we're going to do more workload disaggregation going forward." This isn't a one-off deal — it's a strategic direction.

The Bigger Picture: Why This Matters Beyond Specs

Nvidia controls approximately 95% of the datacenter GPU market. AMD holds roughly 4.5%. But those numbers don't tell the full story. Under Lisa Su's 12-year leadership, AMD went from a company whose market share "rounded to zero percent" in datacenters to one that now counts eight of the top ten AI companies as workload customers.

Daniel Newman, CEO of Futurum Group, told CNBC: "There's a serious case that AMD can get to 20-25%, and that's hundreds of billions of dollars in revenue."

AMD's Q1 2026 data center revenue jumped 57% year-over-year, driven almost entirely by AI accelerators. The company expects to book tens of billions in data center AI revenue starting in 2027, with Helios as the primary driver.

Meanwhile, the competitive landscape is shifting in ways that benefit AMD. Google's TPUs reportedly saved OpenAI 30% on Nvidia pricing simply by existing as an alternative. SemiAnalysis found that TPUs offer a total cost of ownership roughly 44% lower than comparable Nvidia GB200 systems. When Anthropic — one of AMD's biggest new customers — is simultaneously securing up to one million Google TPUs, you can see the pattern: the hyperscalers are actively diversifying away from Nvidia.

This isn't just about AMD vs Nvidia. It's about whether the AI industry can sustain a single-vendor dependency model when every major lab is racing to build alternatives.

What Comes Next

Helios ships in H2 2026. Alongside it, AMD is working on several other CDNA 5-based GPUs: the MI440X (a cut-down version for enterprise), the MI430X (optimized for HPC with FP64), and the Venice-X EPYC CPU launching in 2027 with up to 1152 MB of 3D V-Cache, 96 cores, and 5.15 GHz boost clock.

Su projected that by 2030, the AI accelerator market will reach $1.4 trillion — approaching the size of the entire semiconductor market today. "By the end of the decade," she said, "GPUs are going to make up the vast majority of that market because the algorithms are still very much in their infancy, and we're still continuing to see the workloads change, and that favors programmability."

Whether that prediction holds depends largely on whether AMD can convert its spec-sheet advantages into real-world dominance. The first shipment will tell us everything.

🔥 Hot Takes

1. The "CUDA moat" is more porous than Silicon Valley admits. Everyone keeps saying Nvidia's software ecosystem is unbeatable. But Helios runs on ROCm — and major labs are already deploying it at gigawatt scale. When OpenAI, Anthropic, and Meta are all running production workloads on ROCm, the question isn't whether CUDA is good. It's whether vendors will accept a world where their biggest customer also has alternatives. The moment Nvidia realizes they've lost pricing power on enterprise deals, the moat won't feel so deep. Remember: OpenAI negotiated a 30% discount on Nvidia chips simply by credibly threatening to use Google TPUs. Pricing power evaporates fast when alternatives exist.

2. AMD's circular funding strategy is genius — and dangerous. Here's the loop: AMD gives OpenAI 10% equity → OpenAI buys $billions of AMD GPUs → AMD's revenue goes up → AMD invests $5B in Anthropic → Anthropic buys more AMD GPUs → repeat. It's capital recycling at scale. The risk? If AI lab spending slows or these models don't monetize fast enough, the whole structure wobbles. We're already seeing signs of strain: Tesla missed earnings with free cash flow turning negative, Alphabet and Tesla are testing Wall Street's patience with massive AI capex. But if it works, AMD becomes the de facto infrastructure layer for the entire AI industry — not just a chip vendor. That's a valuation re-rating waiting to happen.

3. UALoE over Ethernet is the Trojan horse that kills Nvidia's hardware lock-in. This is the insight most people miss. Nvidia's real advantage isn't the GPU — it's the interconnect. NVLink + custom switches + CUDA = sticky as hell. Helios says: use standard Ethernet + Broadcom switches + UALoE protocol instead. Suddenly hyperscalers can source switches from any vendor, cable from any vendor, and build clusters without Nvidia's permission. That's not competition — that's ecosystem decoupling. And once the supply chain adapts, there's no going back. Think about it: Broadcom is already shipping 102.4 Tbps Tomahawk 6 ASICs. The infrastructure exists. AMD just needs to convince enough buyers to use it.

4. The real winner might be the customers, not AMD. Every dollar AMD takes from Nvidia flows back to hyperscalers in the form of lower prices and better terms. OpenAI already saved 30% by playing Google against Nvidia. Now with Helios, they have another lever. Meta, Microsoft, Oracle — they're all playing the same game: keep Nvidia honest by funding alternatives. The question isn't whether AMD will win. It's whether Nvidia will be forced to share margin, and whether the industry can sustain multiple competing stacks without fragmenting into incompatibility hell.

Enjoyed this analysis?

Share it with your network and help us grow.

More Intelligence

Infra

The CXMT Shock: A $489 Billion Chinese Chip IPO Just Rewrote the Memory Market

Infra

China's GPU IPO Frenzy Meets New York's Data Center Ban — The AI Infrastructure Crisis Nobody's Talking About

Back to Home View Archive