
OpenAI’s Jalapeño chip is built for fast inference at scale,
Tested on SemiAnalysis’ InferenceX benchmark, Jalapeño registered both more tokens per user and more throughput per kilowatt than the currently available state-
Key Takeaways
- →At the Hot Chips conference on Tuesday, OpenAI shared a more detailed look at Jalapeño, including the first batch of benchmark results for the new sys
- →Tested on SemiAnalysis’ InferenceX benchmark, Jalapeño registered both more tokens per user and more throughput per kilowatt than the currently availa
- →“The bottom line is that the results show a very, very significant performance advance over state of the art,” said Richard Ho, OpenAI’s head of hardw
- →“Jalapeño can serve more AI work per unit of power, while also returning responses more quickly
- →It’s very efficient to serve a lot of customers, but it can also be very low latency
Key Takeaways
- At the Hot Chips conference on Tuesday, OpenAI shared a more detailed look at Jalapeño, including the first batch of benchmark results for the new sys
- Tested on SemiAnalysis’ InferenceX benchmark, Jalapeño registered both more tokens per user and more throughput per kilowatt than the currently availa
- “The bottom line is that the results show a very, very significant performance advance over state of the art,” said Richard Ho, OpenAI’s head of hardw
- “Jalapeño can serve more AI work per unit of power, while also returning responses more quickly
- It’s very efficient to serve a lot of customers, but it can also be very low latency
At the Hot Chips conference on Tuesday, OpenAI shared a more detailed look at Jalapeño, including the first batch of benchmark results for the new system. Tested on SemiAnalysis’ InferenceX benchmark, Jalapeño registered both more tokens per user and more throughput per kilowatt than the currently available state-of-the-art inference processors.
“The bottom line is that the results show a very, very significant performance advance over state of the art,” said Richard Ho, OpenAI’s head of hardware, in a press call. “Jalapeño can serve more AI work per unit of power, while also returning responses more quickly. It’s very efficient to serve a lot of customers, but it can also be very low latency.”
Notably, that comparison is against an Nvidia Blackwell system -- but by the time Jalapeño reaches full deployment, the competition may have advanced significantly. Ho estimated that Jalapeño would deploy at the end of 2026 “in very small volumes,” with more significant deployment coming in 2027.
First announced last October, Jalapeño was developed by OpenAI in close collaboration with Broadcom, with OpenAI’s own models assisting in the development process. The company plans to make Jalapeño a multigenerational platform, allowing AI products, models, chips, and memory all developed in concert.
Because of that full-stack approach, OpenAI was able to address specific phases in the inference process that often cause friction during inference processing. In particular, Jalapeño is designed to minimize delays during the prefill and communication phases of processing, which OpenAI says often act as bottlenecks.
“We designed Jalapeño to minimize data movement and communication delays,” the company said in a blog post presenting the results. “This means that model state, including the KV cache used while generating a response, can be explicitly placed and kept local while the system activates the right combination of compute, memory, and networking for each inference phase.”
When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.
Russell Brandom has been covering the tech industry since 2012, with a focus on platform policy and emerging technologies. He previously worked at The Verge and Rest of World, and has written for Wired, The Awl and MIT’s Technology Review. He can be reached at russell.brandom@techcrunch.com or on Signal at 412-401-5489.
Frequently Asked Questions
What is OpenAI’s Jalapeño chip is built for fast?
Tested on SemiAnalysis’ InferenceX benchmark, Jalapeño registered both more tokens per user and more throughput per kilowatt than the currently available state-of-the art.
Frequently Asked Questions
What is OpenAI’s Jalapeño chip is built for fast?
Stay in the loop
Get the latest tech news and AI insights delivered to your inbox. No spam, unsubscribe anytime.
TechVeb Team
Your trusted source for the latest in technology, AI innovations, and digital trends. We bring you in-depth analysis, expert reviews, and comprehensive guides.
Learn more about us →More ai News
What reception awaits Man City at Anfield?
Bus welcomes, banners and flags expected as Man City head to Liverpool on Sunday for their first game since they were found guilty of breaching Premier...
Watch the trailer for ‘The Altruists,’ Netflix’s show about the FTX
A fictionalized Sam Bankman-Fried is coming to your TV screen on November 19.
'Dad dragged me out of bed by my hair to work' - rural abuse victims
The National Rural Crime Network says the true picture of rural abuse is often hidden, with survivors often having more barriers to reach support.