OpenAI’s Jalapeño Chip Isn’t Hot—And That’s A Good Thing

Date:

Share post:

OpenAI presented the first measured results for Jalapeño, its custom inference chip, at the Hot Chips conference on August 25. Against Nvidia’s GB200 and GB300 rack systems, Jalapeño delivered 1.5 to 1.9 times more AI work per watt at peak throughput and 1.7 to 3.6 times lower end-to-end latency across three open-weight models. Those ratios invite a simple reading in which a customer has outbuilt its supplier. The unit of measurement carries more weight, because a watt spent on a processor does not disappear. It leaves as heat.

OpenAI Measured Jalapeño In Watts Rather Than In Chips

The tests ran on InferenceX, a public benchmark from research firm SemiAnalysis. Results were normalized to each accelerator’s published power rating: 700 watts for Jalapeño, 1,200 for the GB200 and 1,400 for the GB300. Those are thermal design power ratings: the heat a cooling system must carry away from each package. A processor performs no mechanical work, so nearly everything it draws leaves as heat. Power draw and heat output describe the same event. OpenAI noted that performance is sometimes reported per chip and argued for a different basis, stating that “the more useful standard is performance per unit of power.” No accelerator at this scale runs cool, and 700 watts remains a substantial heat source. The claim is narrower: by rating, a Jalapeño package sheds half the heat of a GB300, and on the two models tested against that part, between 1.5 and 1.7 times less for every unit of work.

Power Availability, Not Capital, Now Governs AI Deployment

SemiAnalysis, which ran the benchmark alongside OpenAI engineers, reports that OpenAI is limited by data center power rather than budget or floor space. Electricity demand from data centers rose 17% in 2025 while demand from AI-focused facilities surged 50%, according to the International Energy Agency, which also records developers building onsite generation because grid connections arrive too slowly. Nvidia makes the same argument. Jensen Huang told his Taipei keynote in June that for an operator holding a fixed gigawatt, “throughput per watt is revenues.”

SemiAnalysis Called The Blackwell Comparison Incomplete

SemiAnalysis verified runs inside OpenAI’s lab but did not execute the full suite, and the underlying data came from OpenAI. It called the comparison with Blackwell “somewhat incomplete and unfair,” because Jalapeño carries HBM4 memory while the GB200 and GB300 use HBM3E. Nvidia’s Vera Rubin platform also uses HBM4, and against Rubin the two produce almost the same output tokens per dollar, with Rubin’s figures using speculative decoding and Jalapeño’s not. Rubin is shipping while Jalapeño remains at engineering samples. Package ratings also understate facility draw: an ASIC rack of 128 chips draws 130 kilowatts, and the full two-rack system approaches 160. Richard Ho, OpenAI’s head of hardware, estimated on a press call that deployment would begin at the end of 2026 “in very small volumes”.

Jalapeño’s Nine-Month Design Cycle Is The More Transferable Result

OpenAI says its own models carried the team from initial design to tapeout in nine months, a span that sits inside a wider program SemiAnalysis dates at roughly 16 months from first hires to the November 2025 tapeout, with three further months of bring-up on silicon. That compression bears less on this chip’s standing than on the cost of attempting one at all, since custom silicon has long been gated by time and by expertise scarce enough that most companies conclude renting is the rational course. A chip development cycle measured in months rather than years reopens that calculation.

What OpenAI built with that speed matters as much as the speed. An ASIC buys efficiency by surrendering flexibility, which conventionally means it repays its cost where the work holds still, and OpenAI designed against that constraint rather than accepting it. Rather than dividing its silicon into separate pools for prefill and decode—an arrangement that runs efficiently at one traffic mix and strands hardware at every other—it kept a single fungible fleet able to absorb whatever ratio of work arrives. SemiAnalysis described a general inference chip and called the result a surprise, and the team brought three open-weight models outside the original production plan to high performance within two months. The specialization runs to inference as a task rather than to any model or traffic pattern within it, which is what keeps the investment defensible as the work changes.

Source link

LEAVE A REPLY

Please enter your comment!
Please enter your name here

Related articles

Thai Superstar TEN Heads To The U.S. For Showcase Tour This Fall

TEN (Courtesy of ILLIMNT)ILLIMNTThai superstar TEN is coming to the U.S. for his TEN (10) US Showcase Tour...

Should Arsenal Break The Bank To Land Julian Alvarez?

Atletico Madrid's Argentine forward #19 Julian Alvarez is pictured amid boos from home supporters during the Spanish league...

New Jersey 5’s Run Back Kellogg’s Raisin Bran “Fibes” Partnership In NYC, Paving A New Blueprint For Major League Pickleball

MLP 2026 Regular Season Champions New Jersey rebrands once again as the Fibes. From L-R Millie Rane, Will...

ATEEZ, BTS, LE SSERAFIM Named Top Winners

Yunho, Seonghwa, San, Yeosang, Hongjoong, Wooyoung, Jongho and Mingi of ATEEZ attend the '2026 SBS Gayo Daejeon Summer'...