NVDA 220.78 ▲1.48%GOOGL 339.35 ▼2.09%MSFT 507.29 ▼1.22%AMD 470.72 ▲1.10%INTC 89.51 ▲0.04%TSMC 415.32 ▼0.53%AMZN 259.77 ▼2.50%META 572.34 ▼0.98%AAPL 316.85 ▼0.89%PLTR 186.38 ▲0.05%
Markets at last close

OpenAI · Chips

OpenAI’s Jalapeño shows strong early inference performance

·1 min read

OpenAI has detailed Jalapeño, a self-designed inference ASIC developed with Broadcom and built from a blank slate for LLM inference. Design work began in the middle of 2024 and reached manufacturing tape-out in ~16 months, an unusually fast cycle for a first-generation chip.

Early InferenceX results show Jalapeño ahead of Nvidia Blackwell across most performance-per-watt scenarios and competitive with Vera Rubin, even without multi-token prediction, speculative decoding, or prefill-decode disaggregation. The chip reached over 700 tokens per sec per user at concurrency 1 on DeepSeek R1, while Kimi-K2.5 and GPT-OSS ran at approximately 1,400 tok/sec/user. The benchmark runs were verified in OpenAI’s lab, but OpenAI supplied the numbers, and the full InferenceX suite and AgentX workloads have not yet been tested.

Jalapeño uses HBM4, a simplified memory system, and a homogenous architecture aimed at reducing latency and data movement across varied inference workloads. Production is scheduled to ramp over 2027, with most output expected near the end of next year. At rack scale, the design reaches 128 Jalapeño ASICs per rack and can connect up to 2,048 XPUs, while OpenAI’s next deployment target is 100MW.

Originally reported by newsletter.semianalysis.comRead the source →
Related coverage
All OpenAI news →