OpenAI Unveils Jalapeño Inference Chip, Challenging Nvidia's Power Efficiency at Hot Chips 2026
The AI compute build-out accelerates with major GPU deployments and new partnerships, as liquid cooling gains traction to manage increasing power densities.
The story
OpenAI detailed its custom-designed Jalapeño inference chip at Hot Chips 2026 this week, marking a significant entry into the competitive AI silicon market. Co-developed with Broadcom, the 700-watt ASIC is designed for low-latency inference workloads and claims to outperform Nvidia's GB200 and GB300 in performance-per-watt.
OpenAI plans to deploy Jalapeño in its own data centers later this year at gigawatt scale, potentially reducing its reliance on third-party GPUs for its AI models. The chip features a NUMA-style architecture with 64 memory/core slices and 216 GB of HBM4 memory, capable of up to 13.4 MXFP4 PFLOPS. This move underscores the growing trend of hyperscalers developing custom silicon to optimize for specific AI workloads and manage escalating operational costs and power consumption within their expanding infrastructure.
Silicon
Jalapeño
Maker: OpenAI (co-developed with Broadcom)
What: 700W ASIC with 216 GB HBM4, up to 13.4 MXFP4 PFLOPS, designed for low-latency inference, claiming superior performance-per-watt over Nvidia's GB200/GB300.
For Whom: OpenAI's own data centers for AI model deployment.
Vera Rubin Platform
Maker: Nvidia
What: A full-stack computing platform comprising seven chips, including the Rubin GPU and Vera CPU, now in full production for agentic AI workloads.
For Whom: Hyperscalers and cloud providers building AI factories.
The build-out
| Project | Who | Scale | Where |
|---|---|---|---|
| AI Compute Deployment | ChronoScale and Microsoft | 50 megawatts, utilizing Nvidia GB300 NVL72 systems with liquid cooling. | North America |
| GPU Deployment Expansion | AWS and Nvidia | 2 million additional Nvidia GPUs | Across AWS's global infrastructure |
| Data Center Campus | OpenAI (contract with Georgia Power) | 3.2 gigawatts of power supply | Georgia, USA |
| Two Data Centers | Undisclosed developers | $700 million each | Near Austin, Texas |
| Atlas One Initiative | AZIO AI Holdings Inc. | Combines South Texas property, behind-the-meter natural gas power, dedicated fiber, and modular computing. | South Texas, USA |
Supply & policy signals
AI chip demand continues to squeeze semiconductor capacity, leading to longer lead times for integrated circuits and passive components.
Implication: Data Image has secured a one-year supply for some industrial control products, with materials for Q1 and Q2 2027 already in inventory, indicating persistent supply chain tightness.
Proposals for gas power to supply U.S. data centers have doubled in the last six months.
Implication: This reflects a significant increase in planned data center capacity and the associated energy needs, highlighting the ongoing challenge of securing sufficient power infrastructure.
Cisco Systems has expanded its partnership with Nvidia to deliver AI infrastructure solutions, emphasizing liquid cooling as a new standard for AI chip deployments.
Implication: The shift toward liquid cooling infrastructure underscores the thermal and power challenges posed by current-generation AI accelerators.
What we'll be watching
- Microsoft's planned unveiling of its Maia 300 AI chip, expected as soon as September, targeting production of over a million units for Azure cloud.
- OpenAI's commencement of deploying its Jalapeño inference chip in its own data centers later this year at gigawatt scale.
- Data Center World Power 2026, taking place September 21-23 in Dallas, focusing on energy solutions for AI data centers.
Reporting + analyst voices: grounded via Google Search at publish time.