OpenAI has revealed new benchmark results for its custom-designed Jalapeño chip, showcasing significant performance gains in artificial intelligence inference processing. At the Hot Chips conference, OpenAI presented data indicating that Jalapeño surpasses current state-of-the-art inference processors in both tokens per user and throughput per kilowatt. Richard Ho, OpenAI's head of hardware, described the performance advance as "very, very significant," highlighting the chip's efficiency in serving multiple customers with low latency.
The benchmarks compare Jalapeño against Nvidia's Blackwell system. OpenAI developed Jalapeño in collaboration with Broadcom, utilizing its own AI models to aid in the chip's development. The company plans for Jalapeño to be a multigenerational platform, integrating AI products, models, chips, and memory development. A key design focus for Jalapeño is to reduce delays in the prefill and communication phases of inference, often bottlenecks in processing. This is achieved by minimizing data movement and keeping model state local to optimize compute, memory, and networking for each inference stage.
OpenAI estimates that Jalapeño will see limited deployment at the end of 2026, with more substantial rollout expected in 2027.