
SAN FRANCISCO, CALIFORNIA - JUNE 02: Open AI CEO Sam Altman speaks during Snowflake Summit 2025 astatine Moscone Center connected June 02, 2025 successful San Francisco, California. Snowflake Summit 2025 runs done June 5th. (Photo by Justin Sullivan/Getty Images)
Getty Images
OpenAI presented the first measured results for Jalapeño, its civilization conclusion chip, at the Hot Chips convention connected August 25. Against Nvidia's GB200 and GB300 rack systems, Jalapeño delivered 1.5 to 1.9 times much AI activity per watt astatine highest throughput and 1.7 to 3.6 times little end-to-end latency crossed 3 open-weight models. Those ratios induce a elemental reference successful which a customer has outbuilt its supplier. The portion of measurement carries much weight, because a watt spent connected a processor does not disappear. It leaves arsenic heat.
OpenAI Measured Jalapeño In Watts Rather Than In Chips
The tests ran connected InferenceX, a nationalist benchmark from investigation patient SemiAnalysis. Results were normalized to each accelerator's published powerfulness rating: 700 watts for Jalapeño, 1,200 for the GB200 and 1,400 for the GB300. Those are thermal creation powerfulness ratings: the power a cooling strategy must transportation distant from each package. A processor performs nary mechanical work, truthful astir everything it draws leaves arsenic heat. Power tie and power output picture the aforesaid event. OpenAI noted that capacity is sometimes reported per spot and based on for a different basis, stating that “the much useful modular is capacity per portion of power.” No accelerator astatine this standard runs cool, and 700 watts remains a important power source. The declare is narrower: by rating, a Jalapeño package sheds half the power of a GB300, and connected the 2 models tested against that part, betwixt 1.5 and 1.7 times little for each portion of work.
Power Availability, Not Capital, Now Governs AI Deployment
SemiAnalysis, which ran the benchmark alongside OpenAI engineers, reports that OpenAI is limited by information halfway powerfulness alternatively than fund aliases level space. Electricity request from information centers roseate 17% successful 2025 while request from AI-focused accommodation surged 50%, according to the International Energy Agency, which besides records developers building onsite procreation because grid connections get excessively slowly. Nvidia makes the aforesaid argument. Jensen Huang told his Taipei keynote successful June that for an usability holding a fixed gigawatt, “throughput per watt is revenues.”
SemiAnalysis Called The Blackwell Comparison Incomplete
SemiAnalysis verified runs wrong OpenAI's laboratory but did not execute the afloat suite, and the underlying information came from OpenAI. It called the comparison pinch Blackwell “somewhat incomplete and unfair,” because Jalapeño carries HBM4 representation while the GB200 and GB300 usage HBM3E. Nvidia's Vera Rubin level besides uses HBM4, and against Rubin the 2 nutrient almost the aforesaid output tokens per dollar, pinch Rubin's figures utilizing speculative decoding and Jalapeño's not. Rubin is shipping while Jalapeño remains astatine engineering samples. Package ratings besides understate installation draw: an ASIC rack of 128 chips draws 130 kilowatts, and the afloat two-rack strategy approaches 160. Richard Ho, OpenAI's caput of hardware, estimated connected a property telephone that deployment would statesman astatine the extremity of 2026 “in very mini volumes”.
Jalapeño's Nine-Month Design Cycle Is The More Transferable Result
OpenAI says its ain models carried the squad from first creation to tapeout successful 9 months, a span that sits wrong a wider programme SemiAnalysis dates astatine astir 16 months from first hires to the November 2025 tapeout, pinch 3 further months of bring-up connected silicon. That compression bears little connected this chip's opinionated than connected the costs of attempting 1 astatine all, since civilization silicon has agelong been gated by clip and by expertise scarce capable that astir companies reason renting is the logical course. A spot improvement rhythm measured successful months alternatively than years reopens that calculation.
What OpenAI built pinch that velocity matters arsenic overmuch arsenic the speed. An ASIC buys ratio by surrendering flexibility, which conventionally intends it repays its costs wherever the activity holds still, and OpenAI designed against that constraint alternatively than accepting it. Rather than dividing its silicon into abstracted pools for prefill and decode—an statement that runs efficiently astatine 1 postulation operation and strands hardware astatine each other—it kept a azygous fungible fleet capable to sorb immoderate ratio of activity arrives. SemiAnalysis described a wide conclusion spot and called the consequence a surprise, and the squad brought 3 open-weight models extracurricular the original accumulation scheme to precocious capacity wrong 2 months. The specialization runs to conclusion arsenic a task alternatively than to immoderate exemplary aliases postulation shape wrong it, which is what keeps the finance defensible arsenic the activity changes.
English (US) ·
Indonesian (ID) ·