Why cloud capacity is suddenly tight: the AI data center power gap, explained
TechScripts Nepal · Kathmandu · September 24, 2026 · 3 min read
If cloud capacity has felt tighter or pricier than expected this year, it's not your imagination. Analysts and reporting throughout 2026 point to a real, structural gap between how much AI and cloud compute demand exists and how much data center capacity is physically available to serve it.
The numbers behind the squeeze
Morgan Stanley estimates that U.S. data centers will need roughly 68 gigawatts of additional power between 2026 and 2028. Of that, only about 15 gigawatts is tied to projects already under construction, and another 15 gigawatts could come from available or already-contracted grid capacity. That leaves a gap of roughly 38 gigawatts — demand with no clear supply path yet.
It's not just a paper estimate. Capacity bottlenecks inside major cloud platforms have already forced real, paying customers to look elsewhere: reporting indicates Microsoft's Azure was unable to meet compute demand for some customers, with at least one large customer signing with a competing provider as a result, and GitHub reportedly having to route some of its AI agent traffic to AWS because Azure didn't have the capacity to handle it internally.
The bottleneck isn't what you'd expect
For the last few years, the assumed constraint on AI infrastructure was chip supply — GPUs specifically. That's shifted. The binding constraint now is physical: transformers, switchgear, batteries, and the utility power availability to run it all. High-voltage transformer lead times have stretched to three to five years or longer in 2026. You can't shorten that by writing a bigger check to a chip vendor — it's a hardware manufacturing and grid-interconnection problem, and those move on industrial timelines, not software timelines.
Of roughly 16 gigawatts of new data center capacity announced for 2026 across more than 140 US projects, only about 5 gigawatts is actually under construction. The rest is stuck at the "announced" stage, waiting on the same physical bottlenecks.
What this actually means if you're planning cloud spend
You don't need to be running a hyperscaler to feel this. A few practical implications:
- "The cloud is infinite" is no longer a safe planning assumption. Capacity-constrained regions and instance types are a real possibility, not a hypothetical. If a workload has a hard deadline, confirm capacity rather than assuming it's there when you need it.
- Reserved capacity and long-term commitments carry more weight than they used to. In a genuinely constrained market, providers prioritize customers with existing commitments over new, unplanned demand.
- Multi-cloud isn't just a resilience talking point anymore — it's a practical hedge. If one provider hits a capacity wall in a region you depend on, having a second option evaluated in advance saves you from scrambling during an actual shortage.
- Costs are likely to reflect scarcity, not just usage. Budget conversations that assumed steadily falling cloud unit costs may need revisiting if capacity stays this tight into 2027 and beyond.
None of this means AI or cloud adoption plans need to stop — it means the planning has to account for a resource that, for the first time in a while, is genuinely finite in the near term.