Two years ago NVIDIA said its Blackwell platform would run trillion-parameter inference at "up to 25x less cost and energy" than Hopper. That figure is in NVIDIA's March 2024 launch material and it travelled a long way without much scrutiny.

The figure NVIDIA leads with now is "up to 10x", in a 2026 blog post built on case studies with inference providers running open-source models.

Most of the commentary I have seen treats the second number as a retreat from the first. It is not, and the reason it is not is the more useful thing to understand — because the same confusion sits inside almost every efficiency claim I get asked to assess.

The footnote is where the number lives

NVIDIA's 25x rested on a specific test. A 1.8-trillion-parameter mixture-of-experts model, measured at fixed latency, with GB200 running at FP4 against H100 running at FP8. The H100 cannot do FP4 at all. The Blackwell side ran inside a single NVLink domain while the H100 comparison was spread across InfiniBand. NVIDIA labelled the result "projected performance subject to change."

None of that is hidden. It is all in the material. It is simply below the number, and almost nobody reads below the number.

Adrian Cockcroft — formerly of AWS, and before that Netflix — worked through the same benchmark. Per GPU at matched precision, he puts the raw silicon gain at 2.5x compute and 2.37x memory bandwidth. For the large-system inference case the 25x was actually describing, his own conclusion is that "people should expect more like 8–10x" — which is close to where NVIDIA's own revised figure landed two years later.

One hardware generation, four different multiplesA horizontal bar chart of four multiples for the same hardware generation, each bar drawn ten pixels per multiple. NVIDIA's March 2024 launch material claimed up to 25 times lower cost and energy than Hopper, measured on a 1.8-trillion-parameter mixture-of-experts benchmark: the longest bar. NVIDIA's 2026 blog post claims up to 10 times, measured with inference providers running open-source models: a bar two fifths as long. Adrian Cockcroft's independent analysis concludes that large-system inference should be expected to reach 8 to 10 times, close to NVIDIA's own revised figure: a bar drawn at the 8 times low end. The same analysis puts the raw per-GPU gain at matched precision at only 2 to 2.5 times, drawn at the 2 times low end: by far the shortest bar. The gap between the last two is the precision change and the interconnect, not the silicon.FOUR NUMBERS, ONE GENERATIONNVIDIA launch claim, Mar 202425xNVIDIA blog claim, 202610xIndependent inference estimate8–10xPer GPU, matched precision2–2.5x

Bars are scaled linearly at 10px per multiple. The two range figures are drawn at the low end of their range — 8x and 2x respectively. Sources: NVIDIA Newsroom, March 2024 · NVIDIA blog, 2026 · Adrian Cockcroft, Blackwell benchmark deep dive, 2024 (both independent figures).

A single number was carrying an architecture change, a precision change and an interconnect change at once, and only the architecture change is something you inherit automatically by buying newer hardware.

The precision change is the part that catches people commercially. Moving from FP8 to FP4 halves the bits, so it moves less data and consumes less memory bandwidth, and it can change model output quality. If your workload tolerates it, that saving is genuinely available to you. If it does not, a portion of the advertised gain is not yours and no amount of new hardware will produce it.

Two numbers, two different questions

Here is why the 10x is not a walkback.

The 25x answered: what is the cost and energy per token on a trillion-parameter mixture-of-experts model, on this rack, at this precision, against the previous generation. The 10x answers something else entirely: what cost reduction did named inference providers actually realise, running open-source models, in production. Different workload, different measurement, different question.

NVIDIA did not retract anything. It published a second, narrower, more operationally grounded figure — which is arguably the more honest of the two, and which happens to sit almost exactly where the independent analysis had already put it two years earlier. The trade press reported both as though they were the same kind of claim about the same thing.

That is the failure mode worth naming, because it is not really NVIDIA's. A benchmark is a controlled measurement of one question. It becomes misleading at the point where someone lifts it out of its conditions and treats it as a general property of the hardware. That lift happens in press coverage, in vendor decks, and — the expensive one — in business cases.

I spent eight years at Amazon, most of it in Corporate IT, and the habit that stuck was refusing to let a vendor's own number stand as the basis for a decision. The Frugal Warrior award I picked up in 2019 was for taking more than $500,000 out of ANZ WAN spend, and none of that came from a vendor's efficiency claim. It came from re-architecting dark fibre and renegotiating the contract underneath it. The number that mattered was the one we measured ourselves.

Meanwhile the inputs got more expensive

While that second figure was being published, the cost of the components underneath it moved sharply the other way.

Micron raised DRAM prices more than 60% quarter on quarter in the quarter ended 28 May 2026, on low-single-digit shipment growth. Counterpoint Research measured memory prices rising 80–90% quarter on quarter across most segments from Q4 2025 into Q1 2026. Micron has said it can fill only 55–60% of core customer demand.

Memory market conditions, 2026Three independent figures on memory market conditions in 2026, shown as separate cards rather than a time series — they share no scale and no common period. First: Micron raised DRAM prices more than 60 percent quarter on quarter in the quarter ended 28 May 2026. Second: Counterpoint Research measured memory prices rising 80 to 90 percent quarter on quarter across most segments into the first quarter of 2026. Third: Micron has said it can fill only 55 to 60 percent of core customer demand.WHAT THE INPUTS DID+60%DRAM PRICE, QoQMICRON, QUARTER ENDED 28 MAY 202680–90%MEMORY PRICES, QoQCOUNTERPOINT RESEARCH, Q1 202655–60%OF CORE DEMAND FILLABLEMICRON, 2026

Three independent figures — each on its own scale, not a time series and not a shared period. Sources: Micron quarterly results, quarter ended 28 May 2026 · Counterpoint Research, Q4 2025 to Q1 2026 · Micron demand-fill commentary, 2026.

The mechanism behind the squeeze is margin, not physics. High-bandwidth memory sells for many times more per unit than conventional DDR5, so manufacturers prioritise it. Because HBM also consumes around three times the wafer capacity of equivalent DDR5 — Micron's own figure, 2026 — prioritising it removes more supply from everything else than the revenue share alone would suggest. That is the crowding-out effect, and it reaches the conventional memory in every server you buy.

I want to be precise about the limit of this, because the tempting version of the argument is wrong. Rising DRAM prices do not mean your EC2 or Azure VM list price is about to rise 60%. Hyperscalers buy on long-term contracts, hold inventory, and set list prices for competitive and contractual reasons. Whether and how any of this reaches your specific bill is not something I can evidence from the outside, and I am not going to pretend otherwise.

What I would not assume is that a cost structure moving this fast underneath your provider leaves your renewal terms untouched.

The one number you control

Put the two halves together and the position is specific.

The efficiency gain you were told to expect from the next hardware generation is smaller than the headline, partly conditional on a precision change your workload may not tolerate, and measured on a workload that is probably not yours. The component costs underneath your provider are rising fast, on named and dated evidence, with a pass-through to your bill that nobody outside your provider can quantify. Both of those sit outside your control.

The delta between what you provisioned and what you measurably use sits inside it. That number does not depend on which GPU generation you are on, what precision you run, or what DRAM costs this quarter. It is available today, from your own account.

In the environments I audit this is not a marginal figure. The pattern repeats: compute sized against cluster averages rather than per-instance P99 utilisation, so the averages hide the spiky workloads and the flat ones get treated the same. Commitment coverage well below baseline spend. Storage growing without lifecycle policies because nobody owns the review. None of it requires a hardware refresh to act on, and none of it requires a view on the memory market to justify.

The reframe I would offer any CTO reading an efficiency claim this year is short. A benchmark tells you what to test. Your own utilisation data tells you what to change.

What I am not claiming

Two things I have deliberately not argued, because I keep seeing them argued badly.

The first is that Blackwell is not a real improvement. It plainly is, NVIDIA published its benchmark conditions and labelled the projection, and the hardware is faster. My argument is about what happens to a number after it leaves the footnote, not about whether the number was honestly produced.

The second is anything about where AI infrastructure spending ends up. I have no edge on that call, and neither does anyone selling you a view on it — which has not stopped a great deal of confident writing in both directions. Component prices are rising and vendor efficiency claims need reading carefully. Those two facts are true whether the wider buildout turns out to be well judged or not, and they are the only two this article depends on.

If you want to know where your own environment sits against the one number here that you actually control, the AWS or Azure Cost & Risk Review measures provisioned capacity against real utilisation and returns a prioritised report. If there is a renewal or commitment decision in front of you and you would rather talk it through first, the contact page will reach me directly.


FAQ

Does a vendor efficiency claim like "25x" ever translate to my bill?

Partially, sometimes, and never at the headline number. A vendor benchmark measures one workload shape on one configuration under conditions the vendor selected. Your bill reflects your workload mix, your utilisation, your commitment coverage and your region. The gap between the two is wide enough that I would not let a vendor figure into a business case. Use it to decide what to test. Do not use it to forecast.

What should I ask a vendor to disclose before I accept a benchmark?

Four things, and the answers are usually available if you ask. What numerical precision was each side of the comparison run at, and can the older hardware even run the newer precision. What was the interconnect topology on each side. Was the figure measured or projected — NVIDIA labelled its own Blackwell number "projected performance subject to change", which is disclosed and easy to miss. And what was the workload, because a 1.8-trillion-parameter mixture-of-experts model at fixed latency tells you very little about your inference pattern. If a vendor will not answer the precision question, that is the answer.

If memory prices are rising, why has my instance list price not changed?

Component cost and cloud list price are not the same line. Hyperscalers buy on long-term contracts, hold inventory, and set list prices for competitive and contractual reasons, so input cost movements surface with a lag and unevenly across service lines. They also do not always surface as a price rise — sometimes it is tighter capacity or a longer lead time on a specific instance family. I cannot tell you the pass-through to your bill, because it is not something I can evidence from the outside. What I would say is that it is a reason to read a renewal more carefully than you did last year.

Should we delay a cloud commitment until memory prices settle?

I would not make a commitment decision on a component price forecast. SK Hynix has said publicly the shortage could run past 2030, and that is one named view from one supplier with an interest in the answer. What is decision-grade is your own coverage ratio and workload stability. If a workload has been stable in production for 12 months or more with no planned architecture change in the next 36, commit. If it is mid-migration, do not. The memory market does not change that logic in either direction.

How do I tell whether a capacity constraint is real or a negotiating position?

Ask for the constraint to be named at the SKU and region level, with a date. A real constraint is specific — a particular instance family in a particular region with a lead time attached — because the person telling you about it is working around it themselves. A negotiating position is general and urgent, and it tends to arrive attached to a commitment term. The second tell is whether the constraint survives a change of scope. If a smaller or differently-shaped ask makes it disappear, it was commercial.