At AMD's sold-out Advancing AI conference in San Francisco this week, CEO Dr. Lisa Su unveiled Helios — the company's first true rack-scale AI system, and the most serious challenge Nvidia has faced in the datacenter since the GPU revolution began.
This isn't just another chip launch. Helios is a complete infrastructure play: 72 GPUs per rack, liquid-cooled, using AMD's own CPUs, networking, and software stack. And according to AMD's own specs, it outperforms Nvidia's Vera Rubin in several key metrics while being significantly larger and more powerful.
But the real story isn't just about specs. It's about what Helios represents: the first time a single company has assembled enough pieces — GPUs, CPUs, networking, software — to offer a credible alternative to Nvidia's vertically integrated empire. And the customer list proves it's not vaporware.
The Specs That Matter
Helios is built around the new Instinct MI455X GPU, based on CDNA 5 architecture and fabricated on TSMC's bleeding-edge 2nm process. Each GPU uses a "silicon sandwich" design with 24 chiplets — 8 compute dies stacked on 2 fabric/cache dies, all interconnected with advanced packaging. The result? A chip that delivers up to 4x higher floating-point performance for AI workloads compared to last year's MI355X.
The MI455X makes some bold architectural choices. It drops FP64 entirely — no double-precision math — dedicating every transistor to AI-centric datatypes like MXFP4 and MXFP8. It also ditched its last-level "Infinity" cache in favor of a larger shared L2 cache that delivers 1.5x the bandwidth of the entire Infinity cache on the previous generation.
Here's what makes Helios interesting:
- 50% more HBM4 memory than Nvidia's Vera Rubin
- 50% more scale-out bandwidth
- 15-25% higher AI training performance on paper
- 15% higher peak FP4 FLOPS (though Nvidia claims a 25% advantage in adaptive compression for inference)
- 30% better performance-per-dollar, according to AMD's estimates
The rack measures 1.2 meters wide and stands 44U tall — nearly twice the size of Nvidia's NVL72. It packs 72 MI455X GPUs across 18 liquid-cooled compute blades, each containing four GPUs, a single 96-core Venice EPYC CPU clocking up to 5 GHz, and networking gear from AMD's Pensando acquisition.
For scale-up communication between GPUs, Helios uses UALoE (Ultra Accelerator Link over Ethernet) with 12 Broadcom Tomahawk 6 switch ASICs spread across six switch trays. Each provides 512 lanes of 200 Gbps connectivity, giving each MI455X 3.6 TB/s of bidirectional bandwidth. This means any GPU can talk to any other while keeping latency minimal — and critically, it uses standard Ethernet switches instead of Nvidia's proprietary NVLink.
For scale-out networking, each MI455X gets three 800 Gbps Pensando Vulcano network cards for a total of 2.4 Tbps — significantly beefier than Nvidia's single 1.6 Tbps ConnectX-9 superNIC per GPU.
All of this is powered by a 50V liquid-cooled DC bus bar that can deliver between 225 and 245 kW under load. If AMD's power estimates hold, Helios would be both faster than Vera Rubin and more power-efficient.
The Customer List Is Uncomfortable for Nvidia
AMD didn't just announce Helios — they announced who's buying it, and the names are staggering:
- Microsoft: Satya Nadella confirmed Azure will expand with Helios, adding new computing instances for agentic AI and semiconductor design. Microsoft was also the first to adopt AMD's MI300X in 2023 and is now doubling down.
- Anthropic: Will deploy up to 2 gigawatts of MI455X GPUs for Claude. AMD is investing up to $5 billion in return. The first gigawatt phase starts in H1 2027.
- OpenAI: Committed to deploying gigawatts of Helios in exchange for roughly 10% equity in AMD. This deal was announced earlier this year and sets the template for AMD's customer acquisition strategy.
- Meta: Plans to deploy up to 6GW of AMD GPUs over time, starting with 1GW on Helios racks in H2 2026. Meta and AMD actually co-developed the Helios platform as part of the Open Compute Project.
- Oracle, Tata Consultancy Services: Also confirmed as early customers.
Eight of the top ten AI companies now run workloads on AMD Instinct GPUs. That's not a trickle — that's a flood.
Futurum Group estimates Helios costs $5-5.5 million per rack vs. $3.5-4 million for Nvidia's Vera Rubin. But AMD's pitch is about total cost of ownership and lowest cost per token — not upfront hardware price. Forrest Norrod, AMD's data center head, called it "our baby" and emphasized: "We're very focused on providing the best total cost of ownership, the lowest cost per token, all in."
The Cerebras Partnership: Inference Speed Redefined
On the same day, AMD announced a collaboration with Cerebras Systems to create a disaggregated AI inference platform. The concept is elegant: AMD Instinct GPUs handle compute-heavy prompt processing, while Cerebras' SRAM-based Wafer Scale Engines handle memory-intensive token generation.
Cerebras CEO Andrew Feldman was blunt about why this matters: "Nvidia is just an arms dealer." Cerebras' wafer-scale chips don't need HBM4 — their on-chip SRAM is orders of magnitude faster. The result? Inference speeds exceeding 2,000 tokens per second, with up to 5x better tokens-per-watt compared to GPU-only approaches.
This directly competes with Nvidia's $20 billion Groq acquisition, which aimed to solve the same inference problem with LPUs (Language Processing Units). Where Nvidia needs 2,000 Groq LPUs to serve a trillion-parameter model like Kimi K3, AMD+Cerebras needs at most a few dozen WSE chips.
Su said during a press conference: "You can expect that we're going to do more workload disaggregation going forward." This isn't a one-off deal — it's a strategic direction.
The Bigger Picture: Why This Matters Beyond Specs
Nvidia controls approximately 95% of the datacenter GPU market. AMD holds roughly 4.5%. But those numbers don't tell the full story. Under Lisa Su's 12-year leadership, AMD went from a company whose market share "rounded to zero percent" in datacenters to one that now counts eight of the top ten AI companies as workload customers.
Daniel Newman, CEO of Futurum Group, told CNBC: "There's a serious case that AMD can get to 20-25%, and that's hundreds of billions of dollars in revenue."
AMD's Q1 2026 data center revenue jumped 57% year-over-year, driven almost entirely by AI accelerators. The company expects to book tens of billions in data center AI revenue starting in 2027, with Helios as the primary driver.
Meanwhile, the competitive landscape is shifting in ways that benefit AMD. Google's TPUs reportedly saved OpenAI 30% on Nvidia pricing simply by existing as an alternative. SemiAnalysis found that TPUs offer a total cost of ownership roughly 44% lower than comparable Nvidia GB200 systems. When Anthropic — one of AMD's biggest new customers — is simultaneously securing up to one million Google TPUs, you can see the pattern: the hyperscalers are actively diversifying away from Nvidia.
This isn't just about AMD vs Nvidia. It's about whether the AI industry can sustain a single-vendor dependency model when every major lab is racing to build alternatives.
What Comes Next
Helios ships in H2 2026. Alongside it, AMD is working on several other CDNA 5-based GPUs: the MI440X (a cut-down version for enterprise), the MI430X (optimized for HPC with FP64), and the Venice-X EPYC CPU launching in 2027 with up to 1152 MB of 3D V-Cache, 96 cores, and 5.15 GHz boost clock.
Su projected that by 2030, the AI accelerator market will reach $1.4 trillion — approaching the size of the entire semiconductor market today. "By the end of the decade," she said, "GPUs are going to make up the vast majority of that market because the algorithms are still very much in their infancy, and we're still continuing to see the workloads change, and that favors programmability."
Whether that prediction holds depends largely on whether AMD can convert its spec-sheet advantages into real-world dominance. The first shipment will tell us everything.
🔥 Hot Takes
1. The "CUDA moat" is more porous than Silicon Valley admits. Everyone keeps saying Nvidia's software ecosystem is unbeatable. But Helios runs on ROCm — and major labs are already deploying it at gigawatt scale. When OpenAI, Anthropic, and Meta are all running production workloads on ROCm, the question isn't whether CUDA is good. It's whether vendors will accept a world where their biggest customer also has alternatives. The moment Nvidia realizes they've lost pricing power on enterprise deals, the moat won't feel so deep. Remember: OpenAI negotiated a 30% discount on Nvidia chips simply by credibly threatening to use Google TPUs. Pricing power evaporates fast when alternatives exist.
2. AMD's circular funding strategy is genius — and dangerous. Here's the loop: AMD gives OpenAI 10% equity → OpenAI buys $billions of AMD GPUs → AMD's revenue goes up → AMD invests $5B in Anthropic → Anthropic buys more AMD GPUs → repeat. It's capital recycling at scale. The risk? If AI lab spending slows or these models don't monetize fast enough, the whole structure wobbles. We're already seeing signs of strain: Tesla missed earnings with free cash flow turning negative, Alphabet and Tesla are testing Wall Street's patience with massive AI capex. But if it works, AMD becomes the de facto infrastructure layer for the entire AI industry — not just a chip vendor. That's a valuation re-rating waiting to happen.
3. UALoE over Ethernet is the Trojan horse that kills Nvidia's hardware lock-in. This is the insight most people miss. Nvidia's real advantage isn't the GPU — it's the interconnect. NVLink + custom switches + CUDA = sticky as hell. Helios says: use standard Ethernet + Broadcom switches + UALoE protocol instead. Suddenly hyperscalers can source switches from any vendor, cable from any vendor, and build clusters without Nvidia's permission. That's not competition — that's ecosystem decoupling. And once the supply chain adapts, there's no going back. Think about it: Broadcom is already shipping 102.4 Tbps Tomahawk 6 ASICs. The infrastructure exists. AMD just needs to convince enough buyers to use it.
4. The real winner might be the customers, not AMD. Every dollar AMD takes from Nvidia flows back to hyperscalers in the form of lower prices and better terms. OpenAI already saved 30% by playing Google against Nvidia. Now with Helios, they have another lever. Meta, Microsoft, Oracle — they're all playing the same game: keep Nvidia honest by funding alternatives. The question isn't whether AMD will win. It's whether Nvidia will be forced to share margin, and whether the industry can sustain multiple competing stacks without fragmenting into incompatibility hell.