This year, Google’s TPU Ironwood has been the subject of headlines like “[Hyperscaler] Unveils Chip That Could Dethrone Nvidia.” every few weeks. Amazon’s Trainium 3. Microsoft’s Maia 200. Meta’s MTIA, with a roadmap for four new generations in two years. Every one of them is presented as a David vs. Goliath scenario. To date, none of them have negatively impacted Nvidia’s status as the headlines suggest, and knowing why is more helpful than the headlines themselves.
The numbers that actually matter
By 2026, Nvidia will control around 80% of the AI accelerator market, and data center revenue will account for almost all of the company’s overall revenue. That hasn’t changed much even if four of the biggest tech firms in the world, each with billions of dollars and some of the world’s top chip designers, are concurrently producing rival processors. If bespoke silicon were genuinely going head-to-head with Nvidia and losing, that would be one story. The more realistic explanation is that it’s not really fighting the same struggle.

Custom ASICs are actually expanding quickly, much faster than the GPU market as a whole, but they are primarily focused on inference workloads, which now account for about two-thirds of all AI compute expenditures, as opposed to the training workloads that initially established Nvidia’s dominance. The majority of the headlines are the headlines, the headlines, and the headlines, the headlines, skip.
Training versus inference, the distinction that actually explains everything
Training a frontier model is a bounded, extremely demanding job, a relatively small number of runs, each needing massive, flexible, bleeding-edge compute, run by a handful of labs with the expertise to push hardware to its absolute limit. This is Nvidia’s home turf, built on fifteen-plus years of CUDA software investment that every AI lab’s tooling, libraries, and institutional knowledge are built around.
Switching away from that ecosystem isn’t just a hardware decision, it means rewriting software stacks, retraining engineering teams, and giving up a decade of tooling maturity. That’s a switching cost custom silicon still hasn’t fully cracked, and multiple analysts tracking this space specifically flag it as the real moat, more than raw chip performance.
Inference, the ongoing, repetitive work of a deployed model actually answering real user requests, millions of times a day, is a completely different problem. It’s predictable, repetitive, and runs at a scale where a chip purpose-built for one specific, well-understood workload can be meaningfully cheaper to operate than a general-purpose GPU, even if it’s less flexible. That’s exactly the gap custom ASICs are built for, and it’s why Google can credibly claim its TPUs cut total cost of ownership by double-digit percentages against comparable Nvidia setups for the specific, narrow inference workloads they’re designed around.
So when Amazon says Trainium 3 processes a majority of its own Bedrock token throughput, or Google says its TPUs power the bulk of Gemini’s training and inference, that’s real and significant, but it’s happening almost entirely inside each company’s own infrastructure, for its own workloads, not out in the open merchant market competing directly for the customers Nvidia actually sells to.
The part almost no coverage mentions, everyone’s fighting over the same factory
Here’s the structural detail that undercuts the “chip war” framing more than any other, and it’s the single most useful thing to understand if you’re trying to read this space clearly: nearly every advanced AI chip on the planet , Nvidia’s, Google’s, Amazon’s, Microsoft’s, is manufactured by the same company, TSMC, on the same handful of advanced process nodes. Google, Amazon, Microsoft, and Nvidia aren’t just competing for AI customers.
They’re all competing for limited manufacturing slots at the same fabricator, which is a fundamentally different and much more cooperative-looking competitive structure than “chip war” suggests. A genuinely useful reframe: this isn’t Nvidia versus the hyperscalers. It’s everyone, Nvidia included, competing for a scarce resource one level further down the supply chain.
There’s a second cooperative wrinkle worth naming: the same hyper-scalers building chips to compete with Nvidia remain, simultaneously, some of Nvidia’s single largest customers, Microsoft and Amazon especially are building custom silicon for narrow, internal, cost-sensitive workloads while continuing to buy enormous volumes of Nvidia GPUs for everything else. That’s not really a war. It’s diversification, and it’s the kind of thing any sophisticated enterprise buyer does with a critical, concentrated supplier regardless of industry, you don’t bet your entire cost structure on one vendor if you can help it, especially when roughly 40% of that vendor’s total revenue comes from exactly the four companies now building alternatives.

What a more honest headline would say
If the coverage matched the underlying data more closely, the story wouldn’t be “Nvidia is under threat.” It would read something closer to: hyper-scalers are successfully reducing their own internal inference costs with purpose-built chips, while remaining structurally dependent on Nvidia for training and for anything sold externally to other companies that need flexible, CUDA-compatible infrastructure.
That’s a real, meaningful shift, inference cost reduction at hyper-scaler scale is genuinely significant for anyone renting compute from these providers, but it’s a narrower and more specific claim than “challenging Nvidia’s dominance,” and most projections for Nvidia’s market share even five years out still put it well above 50%, not displaced.
What this actually means if you’re building or buying AI infrastructure
If you’re renting inference capacity at real volume, it’s worth specifically asking your cloud provider whether a custom-silicon option exists for your workload. The cost advantages on well-matched inference workloads are real and increasingly well-documented, Google and Amazon both publish specific total-cost-of-ownership claims against comparable Nvidia setups. For a steady-state, well-understood inference workload specifically, it’s worth benchmarking against the alternative rather than defaulting to GPU pricing out of habit.
If your workload needs training flexibility, broad software compatibility, or you’re not locked into a single cloud provider, Nvidia’s ecosystem moat is still the safer default, and the data backs that up. The custom silicon options are largely cloud-locked, Google’s TPUs run on Google Cloud, Amazon’s Trainium on AWS, which is a real constraint if portability across providers matters to your business.
Don’t read “hyper-scaler ships new chip” headlines as a signal that Nvidia pricing or availability is about to shift dramatically. The market share data doesn’t support that read yet, and the structural relationship, hyper-scalers simultaneously competing with and depending on Nvidia, while all competing for the same TSMC manufacturing capacity, is a lot more stable and interdependent than the “chip war” framing implies.
The bigger picture
This is a genuinely interesting story about how large companies manage dependency on a critical, concentrated supplier, diversify where it’s cheap and workload-specific to do so, stay deeply dependent where switching costs are prohibitive. It’s a less interesting story, so far, as a tale of Nvidia’s dominance actually cracking. Worth knowing the difference before the next “chip war” headline shapes how you think about your own infrastructure costs.
If your team is deciding between GPU and custom-silicon inference pricing and nobody’s actually run the comparison yet, forward this along as the nudge to benchmark it properly. Next week: a closer look at what OpenAI’s own custom chip effort with Broadcom signals about where even the model labs think this is heading.
