OpenAI’s Jalapeño chip is built for fast inference at scale, benchmarks show

OpenAI has unveiled its custom-designed Jalapeño AI inference chip, which aims to improve performance and power efficiency for large-scale AI workloads. The chip is expected to begin small-volume deployment in late 2026, with broader availability in 2027.
Why it matters
This move signals OpenAI's transition toward vertical integration, potentially reducing reliance on Nvidia hardware and optimizing infrastructure for its specific model architectures.
At the Hot Chips conference on Tuesday, OpenAI shared a more detailed look at Jalapeño, including the first batch of benchmark results for the new system. Tested on Semianalysis’s InferenceX benchmark, Jalapeño registered both more tokens per user and more throughput per kilowatt than the currently available state-of-the-art inference processors.
“The bottom line is that the results show a very, very significant performance advance over state of the art,” said Richard Ho, OpenAI’s head of hardware, in a press call. “Jalapeño can serve more AI work per unit of power, while also returning responses more quickly. It’s very efficient to serve a lot of customers, but it can also be very low latency.”
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in