SemiAnalysis puts 1-year H100 leases at $2.35 per GPU-hour, up from $1.70 in October 2025, roughly a 40% jump in five months. My read is simple: this is not a nostalgia trade on Hopper. It is the AI infra market snapping back to its hardest constraint. The companies that locked memory, servers, racks, power, and financing early now get to reprice everyone else.
The useful part of the piece is not the generic “demand boom” line. It is the shape of the move. H100 pricing first bottomed, then ripped higher fast. That breaks the clean analyst story from late 2025 that Blackwell volume would mechanically crush Hopper economics. Spot B200 at $14 per GPU-hour with no availability says something else too: the short-term market has stopped being a real discovery venue. Prices are no longer there to clear excess inventory. They are there to ration access. And if capacity scheduled before August to September 2026 is already spoken for, then “new supply is coming” matters less than “new supply is already allocated.”
I have always thought people tell the GPU shortage story too much like a chip story. In practice it is a delivery and balance-sheet story. H100 still rising four years into the architecture cycle does not mean it regained technical leadership. It means replacement friction is much higher than the market wanted to admit. Moving a large cluster from H100-era deployments to B200 or GB300 is not a board swap. You touch power density, cooling, rack layout, network topology, software tuning, scheduling, failure domains, and customer validation. We already saw this in 2023 and 2024: a new Nvidia platform showing up on slides did not mean rentable capacity showed up at scale. The bottleneck regularly sat in HBM, packaging, server integration, and datacenter power-up. The article does not disclose actual GB300 deployed volume or utilization by SKU, so I do not buy the old model that newer architecture alone should force older pricing down.
The demand side is directionally believable, but I am not fully buying the smoothness of the narrative. Anthropic, Claude Code, multi-agent workflows, and media generation all consume real compute. I believe code agents are especially heavy because they combine long context, tool use, iterative retries, and background execution. But the article leans on a forecast that Claude Code-related commits will exceed 20% of global daily commits by end-2026, and I want more than a forecast there. Commit volume is not the same thing as GPU burn. In many enterprise workflows, the expensive part is not the commit. It is testing, regression runs, UI generation, video assets, repeated debugging loops, and long-lived sessions. Multi-agent systems have the same problem. People count agents because it sounds large. The bill is driven by concurrency, duration, retries, and model mix. None of that is broken out here.
The memory section is where this gets more serious. The piece says LPDDR5 contract prices rose about 4x year over year and DDR5 about 5x. If that is apples-to-apples, this is no longer just “server costs are up.” That is the supply curve shifting left. A lot of AI commentary compresses every shortage into “Nvidia GPU scarcity.” I think that misses where margin is actually being captured. New cluster deployments slip because of ordinary components too: DDR, SSDs, NICs, boards, racks, transformers, power distribution, not just HBM. If OEMs can raise system prices by more than component inflation, they know customers have limited substitutes. If neoclouds can demand larger prepayments and longer terms, credit markets are also accepting scarcity pricing.
One signal here that I do think people underrate is subleasing. Once operators start splitting and re-renting leased clusters, the market has moved from shortage into early financialization. Capacity holders are no longer monetizing GPUs only through products. They are monetizing the reservation itself. I do not think that is healthy, but it is revealing. GPU access starts to behave like a forward contract on power, not ordinary cloud supply. We had shades of this with reserved instances and resale before, but not at this level of bluntness. If that practice spreads, pricing gets stickier and public cloud list prices become even less representative of true marginal supply.
The article is also right that public markets still look colder than private demand. But I would not turn that into a blanket “neoclouds win” story. CoreWeave, Nebius, IREN and similar names trade with skepticism not only because investors fear terminal value decay. They also have heavy capital structures, refinancing risk, customer concentration risk, electricity exposure, and execution risk. Higher rental yields lift ROIC today. The question is whether that yield is built on durable scarcity or on a temporary delivery squeeze. If GB300 ramps faster in the second half of 2026, or if inference efficiency improves again through more aggressive quantization, distillation, caching, routing, or better workload shaping, the operators who expanded hardest at peak pricing will feel the pressure first. We already saw parts of the stack cut unit token costs materially in 2024 with KV cache gains, speculative decoding, and routing tricks. Demand does not move in a straight line forever.
So my conclusion is narrow but important. The H100 price spike is real, but it does not prove old GPUs are permanently repriced upward. It proves AI infrastructure is still governed by supply-chain timing and financing structure more than the headline model race. The risky part of the story is the suggestion that rental prices now only go one way. I do not buy that. If even one of these reverses — memory loosens, server delivery speeds up, or inference efficiency jumps — 1-year lease pricing can soften before the hype does. This looks more like a delivery squeeze than a permanent reset.