NVDA 230.36 ▲0.84%GOOGL 338.46 ▼1.17%MSFT 499.70 ▼2.04%AMD 477.57 ▲4.69%INTC 95.80 ▲4.51%TSMC 428.91 ▲2.85%AMZN 258.51 ▼0.15%META 616.77 ▲1.00%AAPL 319.97 ▼2.51%PLTR 174.33 ▼4.49%
Markets at last close

OpenAI · Chips

OpenAI’s Jalapeño shows strong early inference performance

·1 min read

OpenAI has detailed Jalapeño, a self-designed inference ASIC developed with Broadcom and built from a blank slate for LLM inference. Design work began in the middle of 2024 and reached manufacturing tape-out in ~16 months, an unusually fast cycle for a first-generation chip.

Early InferenceX results show Jalapeño ahead of Nvidia Blackwell across most performance-per-watt scenarios and competitive with Vera Rubin, even without multi-token prediction, speculative decoding, or prefill-decode disaggregation. The chip reached over 700 tokens per sec per user at concurrency 1 on DeepSeek R1, while Kimi-K2.5 and GPT-OSS ran at approximately 1,400 tok/sec/user. The benchmark runs were verified in OpenAI’s lab, but OpenAI supplied the numbers, and the full InferenceX suite and AgentX workloads have not yet been tested.

Jalapeño uses HBM4, a simplified memory system, and a homogenous architecture aimed at reducing latency and data movement across varied inference workloads. Production is scheduled to ramp over 2027, with most output expected near the end of next year. At rack scale, the design reaches 128 Jalapeño ASICs per rack and can connect up to 2,048 XPUs, while OpenAI’s next deployment target is 100MW.

Originally reported by newsletter.semianalysis.comRead the source →
Related coverage
All OpenAI news →