🐾 LIVE
Chinese Tech Workers Are Training Their AI Replacements — And Fighting Back Xiaomi miclaw Becomes China's First Government-Approved AI Agent OpenAI's Quiet Acquisitions Signal Existential Questions About Its Future Google Gemini Launches Native Mac App: The Desktop AI Wars Are On Cerebras Files for IPO at $23B, Backed by $10B OpenAI Partnership DeepSeek Raising $300M at $10B Valuation — While Remaining Profitable ByteDance vs Alibaba vs Tencent: China's AI Video War Heats Up Chinese Tech Workers Are Training Their AI Replacements — And Fighting Back Xiaomi miclaw Becomes China's First Government-Approved AI Agent OpenAI's Quiet Acquisitions Signal Existential Questions About Its Future Google Gemini Launches Native Mac App: The Desktop AI Wars Are On Cerebras Files for IPO at $23B, Backed by $10B OpenAI Partnership DeepSeek Raising $300M at $10B Valuation — While Remaining Profitable ByteDance vs Alibaba vs Tencent: China's AI Video War Heats Up
Infra

OpenAI's First Custom Chip 'Jalapeño' Smokes Nvidia — The CUDA Moat Is Cracking

OpenAI's in-house inference chip beats Blackwell and Rubin on throughput per watt, signaling the end of GPU monopoly and the rise of vertical integration in AI.

2026-08-27 By AgentBear Editorial Source: The Decoder 7 min read
OpenAI's First Custom Chip 'Jalapeño' Smokes Nvidia — The CUDA Moat Is Cracking

OpenAI just dropped a bombshell at the Hot Chips conference: its first custom AI chip, codenamed "Jalapeño," reportedly outperforms Nvidia's flagship Blackwell and next-gen Rubin platforms in inference benchmarks. The numbers aren't close — and they could reshape the entire AI compute landscape.

According to benchmarks provided by OpenAI and partially verified by SemiAnalysis on-site, Jalapeño delivers 1.5x to 1.9x more AI work per watt at peak throughput across three major models — GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T. End-to-end latency? Down 1.7x to 3.6x compared to the best commercially available systems. For interactive workloads, the gap widens to 2.1x to 4.1x.

On GPT-OSS, Jalapeño hit approximately 1,400 tokens per second per user. On DeepSeek R1, it topped 700 tokens per second on a single concurrent request. At matched decoding speed, the chip achieves 54x to 104x the token throughput per kilowatt compared to leading accelerators, depending on the model.

How Did OpenAI Build This in Nine Months?

The timeline is staggering. Design work kicked off in mid-2024 with Broadcom as the partner. The full cycle — from first silicon design to fabrication — took roughly 16 months, but OpenAI claims only nine months elapsed between initial design and the final blueprint heading to the fab.

Here's the twist: OpenAI used its own AI models to design the chip. Older model generations helped with the hardware architecture, while newer ones accelerated programming and optimization. It's AI eating its own tail — or rather, AI building the tools that make AI cheaper and faster.

"Usually first-generation chips aren't competitive, but OpenAI is beating Nvidia Blackwell and even Rubin," wrote SemiAnalysis CEO Dylan Patel. "Jalapeño smokes every other chip."

The CUDA Moat Is Cracking

This is the story that keeps Nvidia executives awake at night. For years, Nvidia's CUDA ecosystem has been the unbeatable moat — a software lock-in that made switching costs prohibitive for AI developers. Jalapeño's rapid development suggests that moat may not be as deep as assumed.

SemiAnalysis noted: "The CUDA moat is potentially dead given how fast OpenAI can bring up new models on their silicon."

The implications are massive. If a single company can design, build, and deploy a chip that beats Nvidia in nine months, what stops Anthropic, Google, or even Chinese firms like Alibaba and DeepSeek from doing the same? The barrier to entry for custom AI silicon just dropped from "impossible" to "challenging but doable."

Inference-Only, But That's Where the Money Is

Jalapeño is an inference-only chip — it runs trained models but doesn't train them. This is strategically important. While training requires massive parallel compute (where Nvidia still dominates), inference is where AI gets deployed at scale and where cost efficiency matters most.

OpenAI's CFO Sarah Friar positioned Jalapeño as part of a broader "full stack" strategy where data centers, chips, models, and products work as one integrated system. The company insists this complements — not replaces — its existing partnerships with Nvidia, AMD, AWS, Cerebras, and CoreWeave.

But let's be honest: every major chipmaker is also a customer. Nvidia, AMD, and AWS are all invested in OpenAI while simultaneously building their own AI chips. It's the ultimate coopetition dance — and OpenAI just changed the choreography.

Comparisons to Nvidia: The Fine Print

SemiAnalysis pointed out that the fairer comparison isn't Blackwell but Nvidia's newer Vera Rubin platform, since both use HBM4 memory. Even against Rubin, Jalapeño squeezes out more output tokens per megawatt — despite Rubin systems using multi-token prediction optimizations that Jalapeño hasn't adopted yet.

On total cost of ownership per token, the two come out roughly even. But here's the catch: Rubin systems are already shipping to customers. Jalapeño remains in engineering sample territory, with no commercial availability announced.

There are also unanswered questions. Nvidia and AMD have published results with larger models like DeepSeek V4 Pro and Kimi K3 that haven't been tested on Jalapeño yet. The benchmarks, while impressive, represent a snapshot — not a definitive verdict.

What This Means for the AI Arms Race

OpenAI's move signals a broader trend: vertical integration in AI. Just as cloud providers built their own chips (Google TPU, AWS Trainium, Azure Maia), the leading AI labs are now designing their own silicon to control costs and optimize for their specific workloads.

For the rest of the industry, this is a double-edged sword. On one hand, it proves custom silicon is viable — potentially opening the door for smaller players. On the other, it raises the bar: if OpenAI can field a chip that beats Nvidia in under two years, what's stopping everyone else?

The AI chip market is about to get a lot more interesting. And a lot more competitive.

🔥 Hot Takes

1. Nvidia's CUDA moat is now a CUDA moat with a crack. OpenAI proved you can build a world-class AI chip in 16 months with a partner like Broadcom. If you have the data, the models, and the cash, the barrier isn't semiconductor expertise — it's organizational will. Expect more labs to follow.

2. Inference is the new battleground, not training. Everyone's obsessed with training chips, but the real money is in running models at scale. Jalapeño targets exactly where AI gets deployed — and where margins matter. This is smart capitalism, not just engineering bragging rights.

3. "Complementary partnerships" is corporate speak for "we're hedging our bets." OpenAI says Jalapeño complements Nvidia, AMD, and AWS. Translation: we'll buy your chips when it suits us, but we're building our own escape hatch. Every major AI lab will do the same. The era of exclusive GPU dependency is over.

Bottom line: OpenAI just proved that custom AI silicon is no longer science fiction. Whether it ships commercially and at scale remains to be seen — but the message is clear: the GPU monopoly is cracking, and the first domino has fallen.

Enjoyed this analysis?

Share it with your network and help us grow.

More Intelligence

Infra

Alibaba's T-Head Is Coming for Nvidia — With a Second-Gen Chip Ready This Year

Infra

Apple's Quiet Gamble: Testing Chinese Memory Chips Amid U.S. Pressure and Global Shortage

Back to Home View Archive