OpenAI has unveiled fresh benchmark results for Jalapeño, its new chip system designed to accelerate AI inference at large scale. Presented at the Hot Chips conference, the results suggest the platform delivers higher tokens per user and stronger throughput per kilowatt than leading inference processors in current use.
Richard Ho, OpenAI's head of hardware, said the system marks a major step forward in performance, combining speed with energy efficiency. According to the company, Jalapeño is built to handle more AI workloads per unit of power while keeping response times low, making it suitable for high-volume services that demand fast output.
The comparison was made against an Nvidia Blackwell system, though OpenAI noted the market will continue to evolve before Jalapeño reaches broader deployment. The company expects limited rollout at the end of 2026, followed by a larger expansion in 2027.
First introduced last October, Jalapeño was developed with Broadcom and shaped in part by OpenAI's own models. The company says it is aiming for a multigenerational platform where chips, memory, models, and AI products evolve together.
OpenAI says the design focuses on reducing friction in key inference stages, especially prefill and communication. By keeping model state and KV cache local when needed, the system is intended to cut data movement and improve coordination across compute, memory, and networking. If successful at scale, this approach could help define the next era of efficient AI infrastructure.