Four Rival AI Models Failed Together. A Week Later, Nobody Has Said Why.
ChatGPT, Claude, Gemini and Grok all went down inside the same window on September 3. Our deal records suggest the four competitors have less independent plumbing than the word competitor implies.
On September 3, starting around 15:30 UTC, ChatGPT, Claude, Gemini and Grok all returned errors at once. Login failures and elevated error rates ran for roughly one to two hours before service came back.
A week later there is still no confirmed common root cause. Anthropic pointed at infrastructure. xAI pointed at a failure in its Memphis compute center. Four companies that compete on everything produced the same failure inside the same window and explained it separately.
The obvious read is coincidence, and it might be. Here is the less obvious read.
**Competitors, one substrate**
We track 644 deals. Nvidia is a named counterparty in 46 of them, 39 times as the investor, seller or provider and 7 times as the recipient. No other company in the tracker comes close.
That number understates it, because the interesting deals are the ones where Nvidia is not a party at all and is still in the room. Anthropic is paying Lambda an estimated $35 billion over six years for 350MW in Texas. Lambda is Nvidia-backed. The data center is leased by Nvidia. Anthropic separately committed $45 billion over six years to Nscale for 460MW of Nvidia Vera Rubin capacity. Lambda raised $1 billion in short-dated debt arranged by JPMorgan specifically to buy Nvidia GPUs for a dedicated Microsoft deployment. AWS is deploying 2 million additional Nvidia GPUs.
Then, on the same day as the outage, Nvidia agreed to buy Hugging Face for about $12.93 billion, the distribution layer for more than three million models.
None of that caused a specific outage. It does mean that when you draw the dependency graph for four rival assistants, the lines converge faster than the org charts do.
**Why this is a hard failure to diagnose**
Concentration of this kind does not produce a clean single point of failure that someone can name in a status page. It produces correlated risk. Shared silicon generations, shared firmware and driver stacks, shared colocation partners, shared power interconnects, shared networking gear. Any of those can degrade in several places at once without any one operator seeing more than their own slice.
That is also why the honest answer from each company is the one they gave. Each of them told the truth about their own layer. Nobody owns the layer underneath all four.
**The counter-evidence, which is real**
The industry is not uniformly consolidating. MBZUAI in Abu Dhabi just released six fully open models up to 375B parameters with weights, code, data and methodology, which is a genuine second source. Saudi Arabia's Humain is building on AMD, and EuroHPC is funding an AMD-based machine in Finland. If you want a substrate that is not Nvidia, it exists, and there is more of it than there was a year ago. Whether anyone with production traffic is actually running on it is a different question.
**What to do with this**
Go and check what your AI failover actually falls over to. Most teams that added a backup provider after the July ChatGPT outage now route from one frontier model to another. On September 3 both of those were down. A real second path means a different silicon vendor and a different data center operator, not just a different logo in the API call.
Then watch for one thing over the next month: a published root cause analysis from any of the four. If a hyperscale-adjacent incident hits four independent providers and none of them can say why in public, that absence is the finding. It tells you the failure lived somewhere none of them can see, and that is the layer worth asking your vendor about.