BENCHBAITSTORY · MONITORING

AI NEWS, TRACKED AND ROASTED.

openai's jalapeño beats nvidia blackwell, per the analyst who says blackwell is the wrong opponent

OpenAI's Jalapeño inference ASIC beats Nvidia Blackwell. That is the headline SemiAnalysis published on 2026-08-25 — [25] OpenAI Jalapeño: Better Than Nvidia Blackwell ↗ — and the body of that same piece is where the headline goes to die.

The figures first. Against Nvidia GB200 and GB300 rack systems on SemiAnalysis's InferenceX suite: 1.5x to 1.9x more throughput per kilowatt and 1.7x to 3.6x lower end-to-end latency, across GPT-OSS 120B, DeepSeek R1 670B and Moonshot AI's 1-trillion-parameter Kimi K2.5. A 700W part against accelerators rated at 1,200W and 1,400W, with OpenAI stating its measured sustained power stayed at or below 550W. At the GB300's fastest previous time-between-tokens settings, OpenAI claims 8.6x to 104.3x more throughput per kilowatt. [27] Tom's Hardware reported the same figures ↗ the same day, under Luke James's byline, with the standfirst "Nvidia's Vera Rubin wasn't in the comparison."

Two outlets, one dataset. SemiAnalysis is explicit about where the numbers came from: "First, all numbers are provided to us by OpenAI. We verified the InferenceX runs in person in the lab, but we did not run the full suite of InferenceX benchmarks nor have we seen AgentX results." Tom's Hardware carries OpenAI's figures too. Publishing a number twice is publication twice. It is measurement once, by the company that built the chip.

the article argues with its own headline

Second paragraph of the caveats, same piece: "we believe that comparison to Blackwell is somewhat incomplete and unfair. Jalapeno is really competing against chips like Rubin that also use HBM4." Vera Rubin — the Nvidia platform OpenAI agreed to deploy for the first gigawatt of Nvidia systems in the second half of 2026 — is not in the throughput comparison at all.

Where Rubin does appear is on cost, and there the gap closes: "On perf/TCO, Vera Rubin and Jalapeno are head-to-head, producing almost the same number of output tokens per $. However, as previously mentioned, Jalapeno's results are obtained without speculative decoding and Vera Rubin's results use speculative decoding."

That asymmetry runs through the throughput charts too. SemiAnalysis: "All this is done without Multi Token Prediction (MTP), while the other chips on the chart are the best performing configs of each respective SKU, all with MTP." Tom's Hardware puts the same thing the other way round — the major comparisons ran Jalapeño's single-token prediction against GB300 configurations doing the same, "even though Nvidia deployments commonly use multi-token prediction in production." Turn MTP on for the GB300 and, per OpenAI's own appendix as reported by Tom's Hardware, the peak efficiency lead "shrinks to roughly 1.5 times." Measure all-in utility power per accelerator instead of published package TDP and it is 1.18kW against 2.55kW, which is also not 1.9x.

published package TDP, no MTP on Jalapeño
the lead
1.5x to 1.9x throughput per kilowatt
GB300 running multi-token prediction
the lead
roughly 1.5x at peak efficiency
all-in utility power per accelerator
the lead
1.18kW against 2.55kW
perf per total cost of ownership, against Vera Rubin
the lead
head-to-head, and Rubin had speculative decoding on

The workload is one workload. SemiAnalysis: the runs "are just 8k1k, a much easier workload to tune for, and there are no AgentX runs yet" — single-turn, no multi-turn long-context suite, which is the suite SemiAnalysis says it prefers for comparing chips and has not seen. The models are not the current ceiling either, by its own account: "the models being tested are not on the open frontier. NVIDIA and AMD have published results on larger models such as DeepSeek V4 Pro and Kimi K3, using AgentX." And Jalapeño does not train models — the workload where, as Tom's Hardware puts it, Nvidia's hardware remains unchallenged.

None of the charts can be checked against any of this. All 23 figures in the SemiAnalysis article are bitmap images with empty alt text, captioned "Source: OpenAI", "Source: SemiAnalysis", "Source: OpenAI, SemiAnalysis" or "Source: Computex 2026 keynote". OpenAI's own two results pages could not be opened for this story: both returned a persistent Cloudflare human-verification interstitial, so every OpenAI figure above is carried on SemiAnalysis and Tom's Hardware, not read from OpenAI.

what people are arguing about

Not the numbers. The provenance and the opponent.

On the [28] Hacker News thread ↗, submitted 5 minutes 24 seconds after the article's publication timestamp and at 303 points and 207 comments when read, "bjourne" goes at the speculative-decoding framing: "How much speculative decoding improves throughput is workload-dependent. Yes, it can improve performance by 5x, but it can also slow down performance by 2x... The tech journos didn't ask themselves if speculative decoding improves perf so much why wasn't it on by default?" "ilaksh" asks the comparison question directly — "Yeah but is it really even as good as Rubin? Seems just competitive." "luciana1u" supplies the epitaph: "everyone's silicon beats everyone else's benchmarks until it has to run someone's actual production workload." "epistasis" reports the article's text disagreeing with its own table, which is unresolvable here because the tables are pictures. "doctorpangloss" asserts "All their benchmarks are flawed," with nothing behind it in the thread.

On [29] r/OpenAI ↗, at 75 points and 34 comments, the top reply is "Mlluell" at 32 points: "Blackwell is a 2 year old architecture to be fair." "a_triple_spiral" raises provenance without ceremony: "Aren't the numbers also self reported or did I misread that?" A [30] bounded X search for "Jalapeno OpenAI" ↗ returned 25 results that are overwhelmingly restatement of OpenAI's figures in several languages, with no independent measurement in the sample; "Neeraj_Kumar222" at least asks the untested question — "How does Jalapeno perform when tool calls and long context dominate latency?" None of that is corroboration. It is people reading the same press release out loud.

the timing

The [26] Hot Chips 2026 program ↗ lists OpenAI's session as "You Can Just Build Things ... Chips," presented by Richard Ho, Ravi Narayanaswami and Chris Leary, scheduled 4:45–6:15 PM PDT on 2026-08-25. The program does not name Jalapeño. The SemiAnalysis article carries a publication timestamp of 2026-08-25T14:00:38.397Z, roughly ten hours before that session starts — though a Substack timestamp can reflect a scheduled send, and nothing here settles which it is.

The two sources also disagree on how fast this was built, possibly by measuring different things: Tom's Hardware says a nine-month RTL-to-tapeout cycle, SemiAnalysis says design work began in mid-2024 and ran roughly 16 months from initial team hiring to manufacturing tape-out.

claims

  • 1.5x–1.9x throughput per kilowatt, 1.7x–3.6x lower latency versus GB200 and GB300. Supported, and supplied by OpenAI.
  • The numbers are OpenAI's. Confirmed. SemiAnalysis verified InferenceX runs in OpenAI's lab, did not run the full suite, has not seen AgentX results.
  • The headline contradicts the body. Confirmed. "Somewhat incomplete and unfair"; the real peer is an HBM4-class part like Rubin.
  • The configuration is asymmetric. Confirmed. Jalapeño without MTP and without speculative decoding; the other chips at their best per-SKU configs, with MTP.
  • MTP narrows the lead to roughly 1.5x. Supported, from OpenAI's appendix via Tom's Hardware.
  • The workload is single-turn 8k1k. Confirmed. No AgentX runs.
  • Vera Rubin was absent from the throughput comparison. Supported. It is the platform OpenAI committed to at gigawatt scale in H2 2026.
  • Perf/TCO is head-to-head with Vera Rubin. Supported — with Rubin using speculative decoding and Jalapeño not.
  • Jalapeño is inference-only. Supported. It does not train models.
  • Every chart is a bitmap. Confirmed. Empty alt text, vendor-credited captions, nothing machine-readable.
  • The tape-out timeline is disputed. Nine months RTL-to-tapeout against roughly 16 months from hiring; different measuring points, unreconciled.
  • "Beating every Nvidia, AMD, and Google chip we have been able to test" is unverified. It rests on vendor-supplied numbers, a partial suite, one easy workload, an asymmetric config, and a model set SemiAnalysis itself says is off the open frontier.

receipts

Nvidia has not responded to the comparison; no request for response has been sent.

LAST VERIFIED 2026-08-25

WHAT CHANGED

THE TIMELINE

REVISION 1 · LAUNCH · 2026-08-25ROAST / SATIRE

openai's jalapeño beats nvidia blackwell, per the analyst who says blackwell is the wrong opponent

  • QUALIFIEDOpenAI reports Jalapeno delivering 1.5x to 1.9x more throughput per kilowatt and 1.7x to 3.6x lower end-to-end latency than Nvidia GB200 and GB300 rack systems on the SemiAnalysis InferenceX suite, across GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T, with a 700W part compared against accelerators rated 1,200W and 1,400W.RECEIPTS [25][27]
  • VERIFIEDThe performance numbers originate with OpenAI rather than with an independent measurement: SemiAnalysis states 'all numbers are provided to us by OpenAI', that it verified InferenceX runs in person in OpenAI's lab but did not run the full InferenceX suite, and that it has not seen AgentX results.RECEIPTS [25][27]
  • VERIFIEDThe SemiAnalysis article published under the headline 'OpenAI Jalapeno: Better Than Nvidia Blackwell' states in its own body that 'comparison to Blackwell is somewhat incomplete and unfair' and that Jalapeno's real peer is an HBM4-class part such as Vera Rubin.RECEIPTS [25]
  • VERIFIEDThe headline comparison is configuration-asymmetric: Jalapeno ran without multi-token prediction and without speculative decoding while the other chips on the chart ran the best performing configuration of each SKU with MTP, and Nvidia production deployments commonly use MTP.RECEIPTS [25][27]
  • QUALIFIEDAgainst a GB300 running multi-token prediction, the peak efficiency lead narrows to roughly 1.5x, and an all-in utility power comparison of 1.18kW for Jalapeno against 2.55kW for the GB300 also produces narrower gaps than the headline figures.RECEIPTS [27]
  • VERIFIEDThe comparison workload is single-turn 8k1k; no AgentX multi-turn long-context runs exist, and SemiAnalysis states that frameworks performing well on 8k1k may perform worse on AgentX because 8k1k does not test routers, prefix cache mechanisms, cache management or offload infrastructure.RECEIPTS [25]
  • QUALIFIEDNvidia's Vera Rubin platform was not part of the throughput comparison, and Vera Rubin is the platform OpenAI committed to deploy at gigawatt scale in the second half of 2026.RECEIPTS [25][27]
  • QUALIFIEDOn performance per total cost of ownership SemiAnalysis places Vera Rubin and Jalapeno head-to-head at almost the same output tokens per dollar, with Vera Rubin's result using speculative decoding and Jalapeno's not.RECEIPTS [25]
  • QUALIFIEDJalapeno is an inference-only ASIC and does not train models.RECEIPTS [25][27]
  • QUALIFIEDEach Jalapeno package pairs its compute die with six HBM4 stacks totaling 216 GiB at 15.4 TB/s, against a GB300 carrying 288GB of HBM3E at a 1,400W rating, and the part is reported to use a TSMC 3nm-class process.RECEIPTS [27]
  • QUALIFIEDThe stated development timeline differs by source and measuring point: Tom's Hardware reports a nine-month RTL-to-tapeout cycle, while SemiAnalysis reports design work beginning in mid-2024 and running roughly 16 months from initial team hiring to manufacturing tape-out.RECEIPTS [25][27]
  • UNVERIFIEDSemiAnalysis's superlative that Jalapeno beats 'every Nvidia, AMD, and Google chip we have been able to test' is not independently verified: it rests on OpenAI-supplied numbers, a partial InferenceX run, a single-turn 8k1k workload, an asymmetric MTP configuration, and a model set that SemiAnalysis itself says is not on the open frontier.RECEIPTS [25][28]
  • VERIFIEDThe OpenAI Hot Chips 2026 session is titled 'You Can Just Build Things ... Chips' with presenters Richard Ho, Ravi Narayanaswami and Chris Leary, scheduled 4:45-6:15 PM PDT on 2026-08-25; the published program does not name the codename Jalapeno.RECEIPTS [26]
  • QUALIFIEDThe SemiAnalysis article carries a publication timestamp of 2026-08-25T14:00:38.397Z, roughly ten hours before the scheduled start of the OpenAI Hot Chips session it describes as just announced.RECEIPTS [25][26]
  • VERIFIEDEvery performance figure in the SemiAnalysis article is a bitmap image with empty alt text, captioned only 'Source: OpenAI', 'Source: SemiAnalysis' or 'Source: OpenAI, SemiAnalysis', so no chart datum in the article is machine-readable or independently sourced.RECEIPTS [25]
  • QUALIFIEDThe bounded discourse pass splits on two axes rather than on the figures: whether OpenAI-supplied self-reported numbers count as a benchmark, and whether a two-year-old Blackwell architecture is the right comparison target at all.RECEIPTS [28][29][30]

EVERY JOKE HAS RECEIPTS

SOURCES

  1. [25] AGGREGATOR · SemiAnalysisOpenAI Jalapeno: Better Than Nvidia Blackwellsupporting ↗
  2. [26] PRIMARY · Hot ChipsHot Chips 2026 advance programprimary ↗
  3. [27] INDEPENDENT · Tom's HardwareOpenAI's 700W Jalapeno ASIC outpaces 1,400W Nvidia flagship GPU - claims up to 1.9x throughput per kilowatt and 3.6x lower latency, co-developed with Broadcomindependent ↗
  4. [28] SIGNAL · Hacker NewsOpenAI Jalapeno: Better than Nvidia Blackwell - Hacker News discussioncontradiction ↗
  5. OpenAI Just Dropped Benchmarks for Their Own Chip, Jalapeno, and It's Beating Nvidia's GB300

    75 points and 34 comments at observation; the submission text itself reproduces the caveats, noting Vera Rubin was not tested, that the chip cannot train models, and that the benchmark ran Jalapeno's single-token setup against GB300 configs. Top-voted reply, 'Mlluell' at 32 points: 'Blackwell is a 2 year old architecture to be fair.' Commenter 'a_triple_spiral' raises the provenance question directly: 'Aren't the numbers also self reported or did I misread that?' Commenter 'DragonflyOk9274' flags chart-axis presentation as the first thing to check in OpenAI marketing material. Multiple commenters challenge the submission itself as AI-generated text rather than engaging the benchmark.

    reaction ↗
  6. [30] SIGNAL · XBounded X search for 'Jalapeno OpenAI', 25 resultsreaction ↗