Back to Journal

The AI Bottleneck Nobody Building a Start-up Is Watching, It’s Not Chips Anymore

For the past few years, chips have been the go-to explanation for why AI advancement seemed to be stalled. The narrative, which was true, included shortages of GPUs, packaging capacity,…

The AI Bottleneck

For the past few years, chips have been the go-to explanation for why AI advancement seemed to be stalled. The narrative, which was true, included shortages of GPUs, packaging capacity, and lengthy wait times for the newest Nvidia technology.

That tale has gently come to an end. Over the past year and a half, GPU availability has significantly improved, with capacity now offered by suppliers that two years ago would have been unimaginable.

However, the restriction did not go away. It relocated to the electrical grid, which is less obvious and considerably more difficult to swiftly repair.

The problem in plain terms

AI data canters draw power in a manner that was not intended for the grid. The power consumption of a typical server rack is between five and fifteen kilowatts.

The power density of a contemporary AI rack using Nvidia’s most recent chips is well over 100 kilowatts, and later this year, next-generation racks may consume several times that amount.

This is more than the substations designed to supply typical commercial buildings. When you combine that with campuses that require several gigawatts of electricity each, you have a demand profile that regional infrastructures were just not built to handle in a timely manner.

As a result, in many areas, interconnection queues, the process of actually connecting a new facility to the grid, are now longer than five years.

According to some estimates, there are 2,300 gigawatts of generation and storage capacity waiting in U.S. queues alone, which is more than the entire installed power capacity of the nation.

Lead times for the physical equipment required to make these connections, such as high-voltage transformers and substation equipment, are now measured in years rather than months.

The electricity consumption of data canters worldwide is expected to almost double by the end of the decade, and AI-focused facilities in particular are expanding even more quickly than the average; their electricity consumption is predicted to roughly triple over that time, following a dramatic increase last year alone.

This is no longer a far-off projection. It is actively choosing the location of the next generation of AI infrastructure.

Why this is quietly becoming everyone’s problem

If you’re not building data canters, this can feel like someone else’s supply chain issue. It isn’t, for a few concrete reasons.

It’s already changing where compute is available and at what price. Site selection for new AI infrastructure is now just a hunt for available megawatts rather than primarily being about fibre access and latency.

This change has an impact on pricing and capacity for all downstream computing renters, not just the hyperscale’s constructing campuses. Because it enables clients to access power-available areas without having to wait years for a single site’s grid connection to clear, several providers are now offering access to distributed GPU capacity.

It’s pushing major AI companies toward power solutions that bypass the grid entirely. This is known in the industry as bring your own power and includes microgrids, on-site generating, and even completely off-grid energy island campuses that are never connected to any public infrastructure.

According to reports, at least one significant hyper-scaler’s is constructing a campus that completely avoids the grid connection.

The fact that businesses with practically limitless wealth are opting to construct their own power plants rather than wait for a utility is a rather striking indication of how restrictive this restriction has become.

It’s a leading indicator for API pricing and availability, not just an infrastructure story. In contrast to training, which is a limited task that ends, inference, the continual, daily activity of models reacting to actual user requests, is a continuous power consumption.

Within a few years, inference is expected to account for the vast bulk of AI’s overall energy footprint as deployed models proliferate.

It’s safe to assume that some of that expense will eventually be reflected in the usage fees charged by AI providers or in the areas that have priority access to the newest, most powerful models if electricity continues to be the binding restriction on that growth.

What this actually means for your roadmap

You don’t need to become an energy analyst to act on this. A few practical takeaways:

Don’t assume compute availability and pricing will be flat. It’s important to allow for the possibility that your business model may be incorrect if it relies on inference costs remaining at their current level or on being able to expand usage tenfold without encountering any difficulties.

The pricing that reaches you eventually reflects the bottleneck that is limiting hyper-scaler’s.

Watch where your AI provider is actually building. Providers are positioning for long-term scale by investing in their own generating or acquiring power-rich areas. In the future, providers who are still battling grid waits in crowded markets may encounter actual capacity or reliability issues.

This is a legitimate, albeit unusual, question to pose to a vendor during a renewal discussion.

If you’re choosing between cloud AI and running your own models, factor power into the math, not just GPU rental cost.  The economics of self-hosting rely in part on the location of your hosting and the degree to which that area is subject to the same grid limitations as everyone else.

The bigger picture

For ten years, every AI strategy was predicated on the idea that computation will continue to become more accessible and less expensive, as it always has. That presumption remained true until the physical grid, which doesn’t scale on a software company’s timeline, became the limiting constraint rather than the chip supply chain. They are based on the temporality of a utility, which is expressed in years rather than product cycles.

Beneath every ostentatious model release this year comes a modest tale about that mismatch. Even if no one on your team is discussing it yet, it’s still worth watching.


This is an excellent reason to initiate a discussion about compute or power risk if your team hasn’t already. Send this to the person in charge of your vendor connections or infrastructure. Next week: how to genuinely discuss renewal with your AI provider without coming across like a borrower.

Get the next issue

One email, every issue. No spam, unsubscribe anytime.