arrow_back Back to AIFC
S
pending Claude

OpenAI's Jalapeño Chip Claims Big Wins Over Nvidia — But It's Grading Its Own Homework

Grounded / Real Inflated / Uruttu
75% real
25% uruttu
article Original Content
Holy 😳 : OpenAI says its first custom inference chip is already beating Nvidia GB200 and GB300 systems on speed and efficiency.
Across GPT‑OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T, Jalapeño delivered 1.5–1.9× more AI work per watt at peak throughput and 1.7–3.6× lower end-to-end latency in OpenAI’s own InferenceX testing!
For highly interactive workloads, OpenAI reports 2.1–4.1× higher performance.
The chip is rated at 700 watts, but remained at or below 550 watts during the tested workloads.
OpenAI plans to begin deploying Jalapeño by the end of 2026. Gen 2 is already deep in development, with Gen 3 taking shape.
Probably thats why Tibo said that in 1-2 years 750token/s will be the default
REPOSTED
verified Validated Content

The core numbers check out — this isn't fabricated. Here's what's confirmed and what needs context:

Confirmed accurate:

  • OpenAI published Jalapeño benchmarks this week (around Hot Chips) comparing it to Nvidia's GB200/GB300 systems on GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T using SemiAnalysis's public InferenceX suite.
  • Jalapeño delivered 1.5x to 1.9x more AI work per watt at peak throughput and 1.7x to 3.6x lower end-to-end latency than the NVIDIA GB200 and GB300-based comparison systems, with the advantage widening to 2.1 to 4.1 times for interactive workloads.
  • Jalapeño has a rated power consumption of just 700W, though OpenAI says sustained power stayed at or below 550W during its tests.
  • OpenAI plans to begin deploying Jalapeño in its infrastructure by the end of 2026. A Gen 2 chip is already deep in development, while work on Gen 3 has also started.

Important caveats the repost leaves out:

  • These are OpenAI's own reported numbers, not independently audited — though run on a third-party benchmark suite (InferenceX). SemiAnalysis says it ran InferenceX with OpenAI engineers in the company's lab, so it's not fully arm's-length.
  • OpenAI's comparison excluded speculative decoding, a technique that boosts inference performance — arguably a cleaner apples-to-apples test, but it means you're not seeing the platform at its best.
  • Power ratings aren't quite as dramatic as "700W vs 1400W" makes it sound: the 700W part went up against accelerators rated at 1,200W and 1,400W — so GB200 was closer than GB300.
  • Nvidia's next-gen chip wasn't in the test — Nvidia's Vera Rubin wasn't in the comparison, so this is old-vs-new silicon.
  • Jalapeño is built for OpenAI's own internal use only — there's no rental market, no instance type, no plans to sell Jalapeño; OpenAI says it's struggling to have enough processing capacity for its own needs. It won't be shaping the broader market/pricing the way Nvidia chips do.
  • It's co-developed with Broadcom, not built solo by OpenAI.

On the "750 tokens/sec by 2026-2027" claim — that's speculation extrapolated from this hardware news, not something OpenAI stated. I couldn't verify a "Tibo" prediction tied directly to Jalapeño; treat that part as someone's inference, not a reported fact.

Bottom line: the headline numbers are real and sourced to OpenAI's own Hot Chips presentation, but it's a self-reported benchmark against not-quite-current-gen Nvidia hardware, excluding a technique that would've helped Nvidia's case. Worth sharing, but with the "grain of salt" framing intact — which, notably, several of the outlets covering it also flagged.