News
OpenAI Publishes First Benchmark Results for Its Jalapeño Inference Chip

OpenAI has released the first benchmark results for Jalapeño, its inaugural AI inference chip designed with Broadcom, showing up to 1.9 times more work per watt and up to 3.6 times lower latency than comparable systems.
Contents
OpenAI has published the first performance benchmark results for Jalapeño, its first in-house chip designed for large language model inference. The company says the chip, co-designed with Broadcom, outperforms comparable systems in both speed and power efficiency, with deployment across OpenAI's infrastructure set to begin later this year.
What the Benchmarks Show
OpenAI used the InferenceX benchmark, developed by analytics firm SemiAnalysis, to evaluate performance. In tests involving workloads that require very low latency, Jalapeño achieved 2.1 to 4.1 times higher performance than competing solutions. For the Kimi K2.5 1T model, the company recorded 1.5 times higher peak performance per watt and 3.4 times lower end-to-end latency.
OpenAI compared its chip against systems built on the latest generation of Nvidia and AMD chips. According to the company, those chips offer 1.46 to 2.1 times more raw compute power, but only about 85 percent of the memory bandwidth per rack compared with OpenAI's solution. A single Jalapeño chip is said to deliver 13.4 petaflops in MXFP4 format, 216 GB of HBM4 memory, and 15.4 TB/s of memory bandwidth, with a rated power draw of 700 watts, though measured consumption in the tested workloads stayed at or below 550 watts.
Architecture Built for Inference
Jalapeño is a chip designed exclusively for inference, not for training models. OpenAI emphasizes that the architecture is meant to minimize data movement during the prefill phase and inter-chip communication, which the company identifies as typical bottlenecks in serving large language models. A large SRAM cache is meant to keep key-value cache data local during response generation, cutting latency without sacrificing throughput.
Jalapeño can handle more AI work per unit of power, while also returning answers faster - Richard Ho, OpenAI
Partnership with Broadcom
OpenAI and Broadcom unveiled the Jalapeño project in June 2026 as the first result of a partnership announced in October 2025. Nine months passed from the initial design to tapeout, the final production-ready version, which the companies describe as the result of close collaboration between OpenAI's engineering teams and Broadcom's chip manufacturing expertise. Broadcom CEO Hock Tan had previously spoken of cost savings of around 50 percent compared with typical GPU-based systems.
The chip design process itself was reportedly aided in part by OpenAI's own models. The company said AI-generated implementations of portions of the chip were 1.5 to 1.8 times faster than existing approaches, which shortened the project's development time.
Why It Matters for the Market
OpenAI's own inference chip reduces the company's dependence on Nvidia, its main supplier of AI accelerators, and fits into a broader trend of leading AI companies building their own hardware, alongside similar efforts from Google (TPU) and Amazon. For enterprise customers using OpenAI's models, lower inference costs could translate into cheaper API pricing over the long run, though the first deployments are expected to cover only small volumes for now.
OpenAI cautions that competitors may significantly improve their chips before Jalapeño reaches wider deployment, and benchmark results produced by the company itself typically require independent verification. Still, publishing concrete performance numbers rather than marketing claims alone suggests the project has entered an advanced silicon-testing phase rather than remaining at the concept stage.
Deployment in data centers is expected to happen through partners, including Microsoft, with the first small production volumes arriving later in 2026 and a significant scale-up in 2027.

