Skip to content

Blog / Artificial Intelligence

The GPU shortage, explained without the mythology

XXXFuel Editors

5 min read

Artificial Intelligence

It was never only NVIDIA. HBM, CoWoS packaging, and power are why a purchase order can clear and the cluster can still be late.

The GPU shortage, explained without the mythology
Flagship GPU street price index (2022=100)

Unit: index

2022
100
2023
164
2024
148
2025
132
2026
121

Shortage moved up the stack: power, not just silicon.

Hardware · Supply chain · 2026

People say “GPU shortage” when they mean a stack of shortages that happen to show up as a missing H100 or B200. The silicon is one constraint. High-bandwidth memory is another. CoWoS-style packaging is a third. Then you need a building that can feed the rack.

It was never only NVIDIA

NVIDIA designs the accelerator most buyers still write into a purchase order. TSMC and the OSATs assemble it. SK Hynix, Micron, and Samsung try to feed it HBM. A utility tries to plug it in. Any one of those can make the cluster late after the PO has cleared.

Through 2026, public earnings-call commentary from the memory vendors still described HBM — including HBM3E — as allocated. TSMC’s own language on AI frontend and backend capacity has been “tight,” with CoWoS oversubscribed. Sell-side notes (treat as estimates, not a TSMC 10-K) have put 2026 CoWoS demand far above 2024 output, with NVIDIA taking on the order of three-fifths of that packaging capacity and the top three customers the rest of the air in the room.

Lead times quoted in the channel for H100/H200-class parts have sat in the three-to-four-quarter range when they are available at all. That is not a law of physics. It is what a constrained backend looks like when everyone learned to hoard.

Unused reservation was cheaper than a waitlist. That lesson from 2023–24 did not expire in one easing cycle.

Power is the shortage you can see from the highway

A purchase order can clear and the cluster can still be late because the substation is not there. Transformers, interconnect queues, and liquid cooling are why hyperscalers talk about campuses, not about GPUs, on earnings calls. If you are a government that wants to “fix the GPU shortage” this year, the lever you can actually pull is permitting and power, not a new 2 nm fab that will not yield for you in 2026.

What a startup should do instead of mythology

  • Rent. A closet of accelerators is a working-capital problem dressed up as strategy.
  • Design so you can swap vendors. If your stack is one SKU from one cloud, you are on their waitlist, not yours.
  • Measure tokens and latency, not “we have H100s.” The shortage is an excuse to overbuy.
  • If you need fine-tuning, budget the extra memory; HBM is the quiet tax.

What this is not

It is not “NVIDIA cannot make chips.” It is not “AI demand is fake.” Both slogans were useful to someone in 2024. In 2026 the boring version is: packaging and memory are still the binding constraints on the parts everyone wants, power is the binding constraint on turning those parts on, and hoarding amplifies both. When a vendor says supply is “easing,” ask which layer — wafers, HBM, CoWoS, or megawatts.

Sources

TSMC quarterly commentary on AI backend tightness (2025–26). Memory-vendor HBM allocation language from 2025–26 earnings. Sell-side CoWoS share estimates (NVIDIA ~60%) — estimates, flagged as such. Channel lead-time reports for Hopper/Blackwell-class parts, 2026. IEA Energy and AI for the power-system half of the same story.

Blackwell did not repeal scarcity

A new generation of accelerators changes the shape of the shortage. It does not delete it. More FLOMs per watt still need HBM on the package and a rack that can reject the heat. Liquid cooling is why a 2026 campus tour looks like a plant room, not like a 2019 air aisle. If your facilities team has never specced a CDU, you are not buying GPUs. You are buying a construction project.

Export controls are a fourth queue

US rules on which parts can go where created a parallel market: downgraded SKUs, more memory-bandwidth gymnastics, and a lot of procurement fiction. That is not the same as CoWoS being full. It can make CoWoS look fuller in the regions that are allowed to buy the good parts. Do not mix the two in one sentence.

Neoclouds, TPUs, and the “just use someone else” slide

Renting from a specialist cloud is how most startups should meet 2026. Google’s TPU inventory, AMD’s MI300-class parts, and a row of neoclouds exist because NVIDIA’s allocation is a political object inside every hyperscaler. Switching is a software problem. Teams that coupled themselves to one CUDA kernel and one SKU discovered that in 2024 and are still paying for it. Teams that kept a portable serving stack treat the shortage as a price, not a personality crisis.

Ask a vendor who says “supply is easing” to name the layer: wafers, HBM, CoWoS, substations, or export licenses. If they cannot name it, they are reading a headline to you.

How to brief a board without folklore

Three slides: (1) which layer is tight for the SKU you actually buy — HBM, packaging, or power; (2) lead time in quarters, not slogans; (3) the fallback vendor and the software cost of switching. If the AI lead cannot fill those, they are repeating a podcast. Capacity will ease on some layers in 2027. Buildings will not. Plan the interconnect first if you are the one signing a campus.

Used and off-lease Hopper still exists. It is not a strategy for a training run. It is a strategy for inference you should have been distilling anyway. Distill, then rent. Buying a trophy rack because a competitor tweeted a cluster photo is how this shortage ate working capital in 2024.

More from the desk