
OpenAI releases Jalapeño benchmark data showing efficiency gains over Nvidia GB300
Benchmark figures released at the Hot Chips conference indicate OpenAI's custom inference processor delivers up to 1.9 times more work per watt than Nvidia GB300 systems.
Benchmark results and inference performance
At the Hot Chips conference on Tuesday, OpenAI presented the first performance benchmarks for Jalapeño, its custom artificial intelligence inference processor. Testing conducted on the SemiAnalysis InferenceX benchmark evaluated the processor against Nvidia GB200 and GB300 systems across three large language models: GPT-OSS 120B, DeepSeek R1, and the 1-trillion-parameter version of Kimi K2.5. Jalapeño delivered between 1.5 and 1.9 times more AI work per watt than the comparison hardware at peak performance. Response latency decreased by a factor between 1.7 and 3.6 across the evaluated models, with the widest performance margin observed on Kimi K2.5. Richard Ho, the vice president of hardware at OpenAI, stated that the chip achieved both higher throughput and lower latency simultaneously.
The bottom line is that the results show a very, very significant performance advance over state of the art.
Architecture and power specifications
Jalapeño is an application-specific integrated circuit developed in collaboration with Broadcom and Celestica. OpenAI used its own artificial intelligence models during development, completing the design process from initial concept to final silicon in nine months. The processor operates at a nominal power draw of 700 watts, while sustained power remained below 550 watts during benchmark testing. In comparison, Nvidia GB200 systems consume 1,200 watts and GB300 systems draw 1,400 watts. Jalapeño is engineered exclusively for inference tasks rather than model training, optimizing memory management and local key-value caching to reduce data transmission delays during prefill and communication phases. Independent analysis firm SemiAnalysis noted that tests against Blackwell-generation hardware precede comparisons with Nvidia's newer Rubin architecture.
- OpenAI Jalapeño
- 700 W
- Nvidia GB200
- 1200 W
- Nvidia GB300
- 1400 W
Deployment schedule and multi-generational roadmap
OpenAI plans to deploy Jalapeño in limited volumes within its internal data center infrastructure before the end of 2026, followed by expanded volume in 2027 and large-scale production in the fourth quarter of 2027. The company is currently developing a second-generation chip scheduled to tape out in the coming months, while a third iteration is in the conceptual design phase. OpenAI confirmed it has no plans to sell Jalapeño chips to external customers, reserving all production to run services such as ChatGPT, Codex, the public API, and agent products. The multigenerational platform allows OpenAI to co-develop algorithms, memory systems, and physical processors to match future model requirements.
- Initial development announced
- OpenAI presents the Jalapeño processor
- Performance benchmarks revealed at Hot Chips conference
- Initial small-volume deployment in OpenAI infrastructure
- Scheduled start of large-scale production
Supplier relations and industry context
Because Jalapeño cannot train new models, OpenAI remains reliant on external semiconductor suppliers for its broader computing infrastructure. Company leadership affirmed that commercial accelerators from other vendors will continue to power training workloads alongside the proprietary inference silicon.
We're going to need a lot of compute. And Jalapeño is part of that. Cerberus is part of that. Nvidia is part of that. AMD is part of that.
The custom hardware effort follows similar developments across the industry, with Anthropic designing in-house silicon, Google negotiating with Marvell for custom inference chips, and Amazon and Microsoft operating proprietary processors. Meanwhile, the European Union has committed 20 billion euros toward AI gigafactories that purchase merchant hardware on the open market, coinciding with server price increases of more than 15% announced by Nvidia.


