OpenAI's 'Jalapeño' Chip Outperforms Nvidia Blackwell in Benchmark Tests
OpenAI's first self-designed chip, Jalapeño, beat Nvidia's Blackwell in SemiAnalysis benchmarks, achieving up to 1.9x per-watt throughput and 3.6x lower latency. Mass production starts in 2027.
The tests were conducted at OpenAI's lab using SemiAnalysis' InferenceX benchmark suite on three open-source models: GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 with 1 trillion parameters. At low concurrency, GPT-OSS and Kimi K2.5 reached about 1,400 tokens per second per user, while DeepSeek R1 exceeded 700 tokens per second per user at concurrency 1. At a low-concurrency point of 100 tokens per second, Kimi K2.5 performed more than 9 times better than the next-best chip, and GPT-OSS delivered close to twice the per-megawatt throughput of GB200's peak point. Accuracy, measured by GSM8k, was on par with Nvidia's chips.
SemiAnalysis noted that all scores were based on single-token prediction without speculative decoding or prefill/decode separation, while the comparison Blackwell results used multi-token prediction. If compared with GB300 with MTP enabled, the peak energy-efficiency advantage would shrink to about 1.5 times. The benchmark numbers were provided by OpenAI, and SemiAnalysis verified the InferenceX runs on-site but did not run the full suite or the AgentX benchmark, which covers long-context multi-turn conversations closer to real production loads. SemiAnalysis also said that comparing Jalapeño with Blackwell was not entirely fair, as the chip is designed to compete with Nvidia's next-generation Vera Rubin, which uses HBM4 and has already started shipping to customers. Even against Rubin, Jalapeño's single-token-per-megawatt output exceeded the multi-token results published by Nvidia and CoreWeave in July, and output per dollar was roughly equal. Rubin's results used speculative decoding, which can reduce per-token cost by 3 to 5 times; Jalapeño has not yet implemented it, so its cost advantage could grow.
The chip has a thermal design power of only 700 watts, compared with 1,200 to 1,400 watts for Blackwell flagship cards. It features a single compute die on TSMC N3P process and an I/O chiplet on N3E, delivering 13.4 PFLOPS of MXFP4 compute per die. It is equipped with six HBM4 memory stacks, providing 216GB capacity and 15.4TB/s bandwidth, the highest among currently shipped or near-shipping accelerators, with Samsung as the likely supplier. The software stack is based on Gluon, a language built on Triton that preserves the SPMD programming model but exposes lower-level abstractions.
AI played a major role in the chip's development, according to SemiAnalysis's report. OpenAI used Codex and GPT-Astra during the design phase, reducing SIMD area by 8% and matrix engine area by 10%. AI-written kernels outperformed human-expert kernels by 1.5 to 1.8 times in some blocks. Iteration speed was also notable: parallel scale expanded from TP8 to a full rack TP32 in eight days, and throughput at certain concurrency levels more than doubled within two weeks. As a demonstration, OpenAI used Codex to port the game "Doom" to the chip, running at 36 FPS. The chip was designed with help from OpenAI's GPT-5.6 Sol model, which itself runs on Nvidia GPUs.
OpenAI said it does not plan to replace its existing chip lineup with Jalapeño and will continue to partner with Nvidia for training. The company also plans to develop second- and third-generation chips. The engineering sample is ready, with mass production ramping up in 2027 and most output coming by the end of that year. The single-rack system, called Vindaloo, holds 128 Jalapeño chips, paired with a Katsu CPU host rack and Chana switch rack; the combined power draw of two racks is about 160 kW. A scale-up domain can consist of up to 16 racks with 2,048 chips. Deployment will be done with neocloud partners, starting with reliability data collection. Nvidia's stock rose 2.19% on the same day, coinciding with the launch of its Jetson Orin Nano 2 robot computer and OpenAI's statement that it would continue using Nvidia chips.
Separately, The Wall Street Journal reported on Aug. 25 that OpenAI's head of data centers, Chris Malone, has left the company. Bloomberg later confirmed the news via an OpenAI spokesperson. Malone, who joined in March 2025 after the Stargate project was announced, had no named successor; his duties were split among several leaders. He is the fourth executive to leave in a recent wave, following chief revenue officer Denise Dresser, chief operating officer Brad Lightcap, and Fidji Simo, a deputy to CEO Altman. TheNextWeb counted seven senior departures or role changes since April. The timing, months before a possible IPO, is unusual.