OpenAI has shared initial benchmark results for its Jalapeño chip, indicating a significant performance advance over existing state-of-the-art inference processors. Tested on Semianalysis’s InferenceX benchmark, Jalapeño showed improvements in both tokens per user and throughput per kilowatt.
Richard Ho, OpenAI’s head of hardware, stated that the results demonstrate a substantial performance gain, allowing the chip to serve more AI work per unit of power while also providing quicker responses. He noted that Jalapeño is designed for both efficient high-volume customer service and low latency.
The comparison was made against an Nvidia Blackwell system. OpenAI developed Jalapeño in collaboration with Broadcom, utilising its own models in the development process. The company aims for Jalapeño to be a multigenerational platform, integrating AI products, models, chips, and memory.
This full-stack approach allowed OpenAI to address specific friction points in the inference process, particularly minimising delays during the prefill and communication phases, which are often bottlenecks.