OpenAI says Jalapeño chip outperforms Nvidia on AI inference
OpenAI's new chip, Jalapeño, delivers up to 1.9x more work per watt and 3.6x lower latency than Nvidia's superchips in benchmark tests.
Jalapeño is an application-specific integrated circuit (ASIC) developed with Broadcom, designed for AI inference. It was first introduced in June, according to The Verge, though TechCrunch reported that the chip was first announced last October.
In OpenAI's tests using the InferenceX benchmark, Jalapeño delivered 1.5 to 1.9 times more AI work per watt than Nvidia's GB200 or GB300 superchips across models including GPT-OSS 120B, DeepSeek R1 and Kimi K2.5 1T, while end-to-end latency was 1.7 to 3.6 times lower.
OpenAI hardware vice president Richard Ho said: "The bottom line is that the results show a very, very significant performance advance over state of the art." He also described Jalapeño as offering the "best of both worlds" with lower latency and higher throughput.
OpenAI designed Jalapeño to minimize delays during the prefill and communication phases of processing, which often act as bottlenecks. The chip can explicitly place and keep model state, including the KV cache, local while activating the right combination of compute, memory and networking for each phase.
Ho said OpenAI plans to deploy Jalapeño in small volumes by the end of this year, with more significant deployment in 2027. The company did not say how many chips it plans to deploy next year. OpenAI expects to continue working with partners like Nvidia and will develop second and third generations of Jalapeño.