OpenAI's Jalapeno chip beat Nvidia's GB300 in first benchmarks - don't expect cheaper AI yet
OpenAI's Jalapeno chip beat Nvidia's GB300 on throughput and latency in its first benchmarks. Here's what that does, and doesn't, mean for your AI costs.
What happened
OpenAI published its first benchmark results for Jalapeño, the custom AI inference chip it built with Broadcom. The tests ran on SemiAnalysis's InferenceX platform, and SemiAnalysis engineers were physically present in OpenAI's labs to watch the runs. Across three open-weight models — GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T — Jalapeño, drawing 700 watts, delivered 1.5 to 1.9 times more throughput per kilowatt than Nvidia's 1,400-watt GB300 system, with latency up to 3.6 times lower. On the kind of low-latency, conversational traffic that makes up most of ChatGPT's real usage, the gap widened to 2.1 to 4.1 times faster response times.
OpenAI's Richard Ho called it "a very, very significant performance advance over state of the art." The comparison was against Nvidia's current top-end system, GB300 — not Vera Rubin, the generation Nvidia has coming next.
What's genuinely new here
This is the first time OpenAI has put real, third-party-witnessed numbers behind a custom chip, rather than just announcing that one is coming. Every major AI lab has been rumored to be building its own silicon for years. This is the first to let someone else watch the benchmark run.
The efficiency angle matters because inference, not training, is now where the money goes. A model gets trained once. Every reply a user gets is a new inference call, and at OpenAI's scale, cutting watts per token is effectively a price cut that hasn't been announced yet. The technical approach is also worth knowing: Jalapeño isn't chasing raw compute. It's built to minimize data movement, keeping the model's working memory close to the silicon doing the work instead of shuttling it across a network during the prefill phase — the part of inference that reads and processes your prompt before generating the first token.
What it means if you're running AI into your operations
Nothing changes in your bill this week, or this year. Jalapeño is aimed at "low-volume production" by the end of 2026, with broader rollout in 2027. No pricing has moved, and OpenAI hasn't promised any.
But the trend line is worth tracking if you're deciding how much AI to build into your operations. Inference cost is the biggest single lever on what's economically worth automating. A task that costs three cents in model calls is worth automating almost automatically; the same task at thirty cents often isn't, unless it also saves real staff time. Every serious efficiency gain in inference — a smaller model, a smarter router, or now custom silicon — eventually shows up as either a lower sticker price or a competitor forced to match one. That's the actual mechanism behind why Sonnet, Gemini Flash, and GPT-5.6 Luna have all gotten cheaper this year while also getting better: the cost to serve a token keeps falling, and someone always breaks ranks and passes the savings on.
The honest caveat
Be skeptical of a chip benchmark published by the company that built the chip. SemiAnalysis's InferenceX is a real, independently run benchmark, and having their engineers on-site for the tests is more credible than a number in a slide deck. But OpenAI chose which models to test and which Nvidia system to compare against, then published the results itself. That's a vendor-run test on a public benchmark — not an independent audit.
Nvidia's answer, Vera Rubin, isn't in this comparison at all, and Nvidia just posted a $96 billion quarter, which tells you the market isn't pricing this as an immediate threat. And "low-volume production in late 2026" is a specific, deliberately modest claim: this chip is not serving your ChatGPT calls today, and it won't be at any real scale for a while.
What to do about it
Don't wait for cheaper inference before automating something that already pays for itself. Hardware efficiency gains take years to reach end-user prices, and workflows that work at today's rates — ticket triage, invoice processing, lead qualification — earn their cost back long before Jalapeño ships in volume. If you're already running AI at high volume and unit economics affect your margin, put this on your radar for 2027 planning. It isn't a reason to change anything you're doing right now.
Want this kind of system in your business? Book a free scoping call.