+20% vs −50%. In September, a B200-hour rose 20.5% on Ornn's index ($6.40 to $7.71). That month, OpenAI halved GPT-6 Sol's list price.
Both are true. The gap has to land somewhere.
The math: $/M tokens = $/GPU-hour ÷ millions of tokens per GPU-hour. If the GPU-hour costs 1.2x and the token sells for 0.5x, a provider renting GPUs needs 1.2 ÷ 0.5 = 2.4x more tokens out of every GPU-hour to hold margin.
Where that 2.4x comes from:
- Software on the same chip. Prime Intellect's NVFP4 KV cache fits about 50% more cached tokens per decode GPU.
- New silicon. Cognition reports up to 4.8x the total token throughput on Vera Rubin NVL72 vs GB200 at CoreWeave.
The second one is a building problem. Rubin racks are estimated at 190–230 kW vs 130–140 kW for Grace Blackwell, on 45°C liquid that now cools the networking trays too. About 1.5x the power for up to 4.8x the tokens is roughly 3x the tokens per kW (my estimate, upper bound).
On the facility side:
- A 5 MW hall holds about 21–26 Rubin racks, down from about 35. The limit moves from floor area to busway, CDUs and heat rejection.
- A megawatt capped at 40–60 kW of air can only serve older GPUs, and older GPUs make the more expensive tokens.
At ReadyInfra, when we walk legacy data centers, it is very clear that the ability to cost-effectively retrofit the data hall to accommodate newer GPU can make or break a deal.
My read: falling token prices don't make power cheaper. They make density mandatory. As in No. 01, the scarce thing in a metro is an energized megawatt that can actually take the rack.
If you've retrofit a 2–10 MW hall for liquid, what hit its limit first: power distribution, floor loading or heat rejection?
The conversation on this one is happening on LinkedIn.
Join the discussion →