+20% vs −50%. In September, a B200-hour rose 20.5% on Ornn's index ($6.40 to $7.71). That month, OpenAI halved GPT-6 Sol's list price.

Both are true. The gap has to land somewhere.

The math: $/M tokens = $/GPU-hour ÷ millions of tokens per GPU-hour. If the GPU-hour costs 1.2x and the token sells for 0.5x, a provider renting GPUs needs 1.2 ÷ 0.5 = 2.4x more tokens out of every GPU-hour to hold margin.

Where that 2.4x comes from:

The second one is a building problem. Rubin racks are estimated at 190–230 kW vs 130–140 kW for Grace Blackwell, on 45°C liquid that now cools the networking trays too. About 1.5x the power for up to 4.8x the tokens is roughly 3x the tokens per kW (my estimate, upper bound).

On the facility side:

At ReadyInfra, when we walk legacy data centers, it is very clear that the ability to cost-effectively retrofit the data hall to accommodate newer GPU can make or break a deal.

My read: falling token prices don't make power cheaper. They make density mandatory. As in No. 01, the scarce thing in a metro is an energized megawatt that can actually take the rack.

If you've retrofit a 2–10 MW hall for liquid, what hit its limit first: power distribution, floor loading or heat rejection?

The conversation on this one is happening on LinkedIn.

Join the discussion →