The Three AI Outages Overlapped for 85 Minutes
Four official status pages logged the Claude, Grok, and ChatGPT incidents of 3 September 2026 starting 2 hours 21 minutes apart, but all three had unresolved incidents running at once for 85 minutes. Anthropic and OpenAI's own incident histories show that kind of overlap on 4 separate days in the past 31.
This piece has a Chinese companion on aire: 三家 AI 紅燈重疊的 85 分鐘 (in Chinese)
Four incidents, logged on four official status pages, starting 2 hours and 21 minutes apart: Claude’s Sonnet 5 alert at 20:37 Taipei time, Claude’s broader multi-model incident at 21:26, Grok’s outage four minutes later at 21:30, and ChatGPT/Codex at 22:58, all on 3 September 2026.
Line them up on one clock and the idea that all three went down at once turns out to be compressed. The window where Claude, Grok, and ChatGPT all had an unresolved incident running at the same time was 85 minutes. Widen the same clock to the past 31 days, and two of those vendors, Anthropic and OpenAI, have logged unresolved incidents overlapping before, on four separate days, not one.
Opening accounts at three AI vendors is not the same as holding three insurance policies.
The Four Outages Started 2 Hours and 21 Minutes Apart
| Service | Incident name (official) | Start | End | Official duration |
|---|---|---|---|---|
| Claude | Elevated errors for Claude Sonnet 5 | Sep 3, 20:37 | Sep 3, 20:56 | ~19 min |
| Claude | Elevated errors for multiple models | Sep 3, 21:26 | Sep 4, 00:23 | ~2h 57m |
| Grok (Web) | Models outage | Sep 3, 21:30 | Sep 4, 01:07 | 3h 37m |
| Grok in X | Models outage | Sep 3, 21:30 | Sep 4, 01:05 | 3h 35m |
| ChatGPT/Codex | Elevated errors across ChatGPT and Codex | Sep 3, 22:58 | Sep 4, 00:55 | ~1h 57m |
All times are Taipei (UTC+8). Claude’s and OpenAI’s rows come from each vendor’s own statuspage incident data; the two Grok rows come from status.x.ai’s incident pages, which already display in GMT+8. Five rows, four incidents: Grok’s outage is logged twice, once for the web app and once for the X app, both starting at 21:30.
The start times run 20:37, 21:26, 21:30, 22:58, 2 hours and 21 minutes apart, not simultaneous. The window where Claude, Grok, and ChatGPT all had an unresolved incident running at the same time starts at the latest of the three start times, 22:58, and ends at the earliest of the three end times, 00:23: 85 minutes. Using the impact-ended time Anthropic wrote into its own update instead of the incident’s formal close, 00:16 rather than 00:23, the same overlap comes out to 78 minutes.
Claude logged two incidents that night. The first affected only Claude Sonnet 5 and lasted 19 minutes. The second was broader: a 21:50 update listed the affected models verbatim: “An exhaustive list of affected models: Mythos/Fable 5.1, Mythos/Fable 5, Opus 5, Opus 4.8, Opus 4.6.” The affected components included claude.ai, the Claude API, Claude Code, and Claude Cowork, with impact marked major. By 23:25, Anthropic wrote: “The only affected models right now are Opus 4.8 and Opus 5. The rest of the models have recovered to baseline error rate.” The closing update said: “Impact has ended as of 9:16 PT / 16:16 UTC.” That put the end of impact at 00:16 Taipei time on 4 September. The rest of the models had recovered to baseline before 23:25; Opus 4.8 and Opus 5 took at least 51 minutes longer.
Grok’s two incidents ran the longest of the night, 3 hours 37 minutes and 3 hours 35 minutes by the page’s own count. The incident page’s language was short: “Grok is experiencing issues. We are working on restoring service as quickly as possible.” It closed with: “We have resolved the situation, and traffic is healthy again.”
OpenAI logged its start time at 22:58. Its incident note added: “Some Codex remote control users may need to pair their mobile device again following this incident.”
Engadget’s own 3 September report timed the same night in Pacific hours: Grok “was down for three and a half hours, beginning at 6:30AM PT, according to a company status page”; Claude “also had problems Thursday, starting around 6:30AM PT”, citing the same status page whose incident name appears in the table above; and ChatGPT users “began experiencing problems with the AI service beginning around 7:30AM PT, according to reports on downdetector.com. The issue was resolved by 9:55AM PT, OpenAI said, though it didn’t elaborate on a cause.” Converted to Taipei time, the closing figures line up with the status-page numbers above. The starting figures need one qualification each: Engadget’s “around 6:30AM PT” for Claude rounds up from the status page’s own 21:26 (6:26AM PT), and the 7:30AM PT figure for ChatGPT is a Downdetector report time, not an entry on OpenAI’s status page. This piece uses the status pages, not user reports, as the record.
Two Vendors Overlapped on Four Days in the Past 31
Pull Anthropic’s and OpenAI’s own incident histories, their public statuspage incident data, each vendor’s full set of start and resolved times, take the union within each vendor, then the intersection across the two, and a second pattern shows up.
The two feeds’ shared coverage window runs from 20:43 on 4 August to 17:22 on 4 September, about 30.9 days. Claude logged 27 incidents in that window; OpenAI logged 25. The share of time with any unresolved incident running was 5.7% for Claude and 11.0% for OpenAI.
That intersection comes out as six stretches, each one a period when Claude and OpenAI both had an unresolved incident running, adding up to 6.0 hours across four separate days:
- 13 August, 22:33–00:08 (95 min)
- 1 September, 01:23–01:52 (29 min)
- 1 September, 03:21–04:28 (67 min)
- 2 September, 00:02–00:22 (20 min)
- 2 September, 01:05–02:07 (62 min)
- 3 September, 22:58–00:23 (85 min)
The 85 minutes on 3 September is the second-longest of the six; the longest is the 95 minutes on 13 August. Two vendors overlapping has happened on four separate days in the past month. What made 3 September different is that a third vendor joined in, at a time a lot of people happened to be watching.
Two limits belong next to those numbers. First, each statuspage feed only returns the most recent 50 or 25 incidents, so the window is capped at 31 days; this ratio cannot be extrapolated to a full year. In the public sources I could find, no one has published this particular cross-vendor overlap calculation — which is a statement about what has been published, not about how much data exists. The 31-day window reflects how the data was pulled here, not a limit on the underlying record. Second, an unresolved incident is not the same as an outage; most of the incidents above affected only some models or some regions, and a red status page does not mean every request failed.
No One Has Explained What Happened in Memphis
After Grok recovered, SpaceXAI posted an apology on X. Engadget’s 3 September report embeds the post in full:
We are sorry for the issues you may have experienced with Grok following an outage at our Memphis compute center this morning. We’d also like to apologize to our impacted compute partners. All systems have now been restored and are functioning nominally.
Two phrases in that post carry the whole section: Memphis compute center, and impacted compute partners. Elon Musk followed up separately, saying the company was “taking corrective action to ensure this does not happen again.” SpaceXAI did not disclose what caused the issue in Memphis, and as of this writing, still hasn’t.
Separately, Anthropic announced on 6 May 2026 that it had “signed an agreement with SpaceX to use all of the compute capacity at their Colossus 1 data center. This gives us access to more than 300 megawatts of new capacity (over 220,000 NVIDIA GPUs) within the month. This additional capacity will directly improve capacity for Claude Pro and Claude Max subscribers.” TechCrunch’s 20 May report, citing SpaceX’s SEC S-1 filing, said Anthropic would pay $1.25 billion per month through May 2029, with either side able to terminate the contract on 90 days’ notice, and described the deal as securing “the entire output of the Colossus 1 data center near Memphis, Tennessee”. Anthropic’s own announcement never mentions Memphis by name; that detail comes from TechCrunch’s and CNBC’s reporting, not from Anthropic itself.
Claude’s second incident started at 21:26; Grok’s started four minutes later, at 21:30. Both times come from each vendor’s own status page. No source connects the four minutes.
Engadget itself was careful about how far to take this juxtaposition: the timing of SpaceXAI’s outage, it wrote, “roughly lines up with issues experienced by other AI platforms Thursday morning”, not more than that.
Laid on the same timeline, these facts raise a question that no one has answered yet:
- SpaceXAI said “our Memphis compute center”, without naming which one. Memphis has more than one: the original Colossus supercomputer runs out of a converted Electrolux factory in the city’s Boxtown district, and in March 2025 xAI bought a second, 1-million-square-foot site in Whitehaven, on Tulane Road, for $80 million, according to property records. No source names Colossus 1 as the one that failed.
- The phrase “compute partners” is plural, and unnamed. Engadget described them as unnamed and said they were apparently affected by the issues. No source, first-hand or otherwise, names them.
- Anthropic’s status page, from its first update to its close, only ever describes which models had elevated error rates and which had recovered. It gives no cause and names no third party. Engadget’s own reporting notes: “Anthropic and OpenAI didn’t immediately respond to requests for comment.”
- Engadget put the two incidents side by side and, in the same piece, said as much itself:
It’s also not entirely clear which of SpaceXAI’s “compute partners” may have been affected.
The company hasn’t said if the unspecified issue was related to SpaceXAI’s data center. The two companies signed a deal earlier this year for the Claude maker to lease compute from Musk’s AI firm.
One detail is easy to miss in Anthropic’s own announcement. It says: “We train and run Claude on a range of AI hardware—AWS Trainium, Google TPUs, and NVIDIA GPUs—and continue to explore opportunities to bring additional capacity online.” The same announcement lists several other compute deals alongside the SpaceX agreement, up to 5 gigawatts with Amazon, 5 gigawatts with Google and Broadcom coming online in 2027, a strategic partnership with Microsoft and NVIDIA that includes $30 billion of Azure capacity, and a $50 billion infrastructure investment with Fluidstack. Colossus 1 is one contract among several, not the whole of where Claude runs.
Until one of the parties involved says more, this line stops at a question.
Google’s Records Show No Gemini Incident on 3 September
What happened to Gemini that night depends on who you ask.
Three of Google’s own status channels show nothing for the window. The Google AI Studio and Gemini API status page’s incident history runs back to December 2024, with no entry for 3 September; the most recent entry is 15 July 2026, an AI Studio Build issue. Google Cloud’s incident feed logged zero incidents in the window, though that feed only holds six entries in total, too short to serve as a full history. Google Workspace’s status dashboard shows “No incidents” for the period 28 August through 4 September.
Third-party monitors and user reports tell a different story. StatusGator logged two Gemini anomalies on 3 September, flagged on its own listing as “Never officially acknowledged by Google”. Downdetector reports, relayed through media coverage, began climbing from around 8:44am Eastern time (20:44 Taipei) that day. Those are third-party accounts, not entries in an official record.
Google’s own status page explains the gap. A billing-tier note on the page reads: “Free-tier requests use sheddable capacity, while billed-tier requests are protected by critical priority.” Further down: “Individual customer availability may vary depending on billing status and surface used: free tier, billed tier, as well as the chosen model and API features in use.”
Free-tier capacity is built to be shed under pressure. Dropping it is a design choice, not an incident, so it does not appear in an incident log. A user can feel Gemini go down and Google’s own status page can stay green at the same time.
The accurate way to put it: Google’s own records do not log anything for that evening. Whether users were affected is a question the status page was never built to answer.
On 18 November 2025, Cloudflare’s Name Appeared on Another Company’s Status Page
A genuine single point of failure has a specific shape, and 18 November 2025 is the reference case. Cloudflare’s own post-incident report says its network “began experiencing significant failures to deliver core network traffic” at 11:20 UTC. The report’s own timeline table marks a different, later moment as the formal start of customer impact, 11:28 UTC: “Deployment reaches customer environments, first errors observed on customer HTTP traffic.” Both numbers are Cloudflare’s own; they describe two different moments, when the network first began failing, and when the change reached customers and produced the first customer-facing errors. By the same timeline, “Main impact resolved” at 14:30 UTC and “All services resolved” at 17:06 UTC. The cause, per Cloudflare: a change to a database permissions system that caused a “feature file” used by its Bot Management system to double in size and exceed a runtime limit.
The mark that failure left is on someone else’s page. Google’s AI Studio status history still carries an entry from that day: “Apps in Build may not load due to global Cloudflare outage. Mitigations are underway.”, detected 18 November at 20:35 and resolved at 23:09. That page renders its timestamps in the reader’s own time zone; read from Taipei (UTC+8), those two moments are 12:35 and 15:09 UTC, both inside Cloudflare’s own 11:28–17:06 UTC impact window. When one infrastructure provider goes down, its name tends to turn up on the status pages of the companies that depend on it. That is what happened here.
Nothing like that shows up for 3 September 2026. None of the four status pages behind this piece points to a shared third party. Checked directly against Cloudflare’s own incident feed, there was no global incident in the 21:00–01:30 Taipei window spanning 3 to 4 September; that day’s feed shows only regional items, a Hong Kong 5xx issue from 17:21 to 17:36 and a Seattle 522 issue among them. The mark that a shared failure leaves on other companies’ status pages is absent this time.
The Shared Third-Party Theory Has Changed Once
Within two days, the community’s guess about a shared cause changed once, from Cloudflare to Microsoft Azure, with some versions naming Azure’s East US region specifically.
Azure’s own Post Incident Review page has no entry for 3 September. Sorted by most recent, the newest PIR on the page is dated 23 July 2026, “Post Incident Review (PIR) – Network connectivity – Issues accessing resources in West US”; nothing has been added since. But the same page states its own scope: “From June 1, 2022, this includes PIRs for broad issues as described in our documentation.” Regional or subscription-specific incidents are handled through individual Service Health notifications and do not appear on this page at all. An empty PIR page is not the same as a clean night, and the same gap shows up on the other side of this piece: Google explains its own gap with a sheddable-capacity policy; Microsoft explains this one with a broad-issues threshold for what gets a public PIR. Both are limits the companies wrote themselves, not evidence that nothing happened.
One fact runs the other way. SpaceXAI’s own apology names the failure as coming from “our Memphis compute center”, its own facility, not a cloud vendor. That does not square with a theory that all three incidents shared a cloud-provider cause.
The theory is not baseless, either. Anthropic’s own May announcement lists a “strategic partnership with Microsoft and NVIDIA that includes $30 billion of Azure capacity” among its compute deals. Anthropic and OpenAI having some presence on Azure is verifiable. That this particular outage came from that overlap is not; no source makes that connection.
The versions circulating disagree with each other. One outlet’s headline reads “Gemini Survived When ChatGPT, Claude, and Grok Collapsed: Azure Is at Fault”, arguing Gemini’s survival came from running on Google Cloud rather than Azure. Another account blames an Azure East US regional failure, and counts Gemini among the vendors affected, the opposite of the first version’s claim. Neither cites a statement from Microsoft. No Microsoft statement on this incident exists in the public record. And none of the four official status pages behind this piece mentions Azure, or any other shared third party.
A real single point of failure leaves the same name on other companies’ status pages, that is what 18 November 2025 shows, down to the entry still sitting on Google AI Studio’s own history. Neither candidate for 3 September, Cloudflare or Azure, has left that mark.
Three Accounts Do Not Solve Two Problems
Opening three accounts takes minutes. Two problems after that don’t.
The first is whether the work can actually move. A companion piece on aire, How long does it take to switch from Claude to Codex? (in Chinese), breaks the friction into three layers: credentials, memory, and tooling. That piece’s own reading is that credentials are usually the easiest of the three, and that memory and tooling are where a switch actually gets stuck. A second account is not a usable backup unless the work can move across all three layers. Otherwise it is a spare login page.
The second is knowing when to switch. Status pages don’t answer that question. Google’s own sheddable-capacity language already says as much: dropped free-tier requests don’t get logged as an incident. Downdetector’s Gemini reports were climbing from 20:44 Taipei that evening; Google’s three official channels have no matching entry to this day. In Claude’s 3 September updates, some models had recovered to baseline while Opus 4.8 and Opus 5 remained affected for at least 51 minutes longer. A status page records what the provider has marked as an incident; it does not say whether the request in front of you will succeed.
The 31-day numbers sit there regardless: 5.7% for Claude, 11.0% for OpenAI, two vendors overlapping on four separate days. They can’t answer how many accounts is enough, because that isn’t what they measure. The same companion piece recalculated Claude’s incident count from official history and found 326 incidents between January and July 2026, more than double the third-party figures that have been circulating; that recalculation belongs to that piece, not this one, and it’s worth carrying over here as a reminder that availability numbers don’t hold up to being copied without checking.
The question worth actually calculating is a different one: which pieces of work cannot tolerate an 85-minute window in which all three vendors have unresolved incidents at once? Whatever the answer, that’s where the backup belongs.