OpenAI's Jalapeño Chip Beats Nvidia Blackwell in Early Inference Benchmarks

OpenAI presents the first benchmark results for its Jalapeño inference chip at the Hot Chips conference, showing significant performance gains over current state-of-the-art processors. Tested on SemiAnalysis' InferenceX benchmark, Jalapeño delivers more tokens per user and more throughput per kilowatt than an Nvidia Blackwell system. Richard Ho, OpenAI's head of hardware, says the chip serves more AI work per unit of power while returning responses more quickly, combining high efficiency with low latency.

Developed in collaboration with Broadcom and first announced last October, Jalapeño benefits from OpenAI's full-stack approach, in which AI models assist in development and products, models, chips, and memory evolve together across multiple generations. This design lets OpenAI target specific phases of inference that often cause friction, particularly the prefill and communication stages that frequently act as bottlenecks.

OpenAI says Jalapeño minimizes data movement and communication delays by keeping model state, including the KV cache used during response generation, local while activating the right combination of compute, memory, and networking for each inference phase. The chip deploys in very small volumes at the end of 2026, with more significant rollout coming in 2027 — though competitors like Nvidia may advance their own offerings in the meantime.

Read More at the original source →