AI users found fewer places to turn Thursday morning.
ChatGPT, Claude, and Grok suffered overlapping outages on Sept. 3, while Gemini users also reported problems. For businesses increasingly relying on multiple AI providers as a fallback strategy, the disruption exposed an uncomfortable possibility: switching models does little to help when several services fail at once.
The confirmed incidents involving ChatGPT, Claude, and Grok overlapped for more than an hour, turning what could have been isolated vendor outages into a broader test of AI resilience.
Failures spread across chatbots and coding tools
Anthropic and SpaceXAI began reporting problems with Claude and Grok within minutes of each other shortly after 9:25 a.m. ET, according to their incident records. OpenAI followed at 10:43 a.m. All three providers had restored service by 1:07 p.m.
Problems reached workplace tools alongside consumer AI chatbots. Anthropic listed Claude’s API and Claude Code among the affected products. Cowork also encountered errors, and OpenAI’s outage affected Codex.
SpaceXAI’s incident record lists a three-hour, 37-minute outage for Grok, the group's longest confirmed incident. Gemini users also reported failed responses and slow performance, Gulf News reported, although Google had not confirmed a widespread consumer outage at the time.
Vendor explanations leave a common cause unresolved
Official responses from the major AI companies differed in detail. An OpenAI spokesperson told USA TODAY that “a routing error” made ChatGPT and Codex unavailable for some users. Anthropic also attributed Claude’s outage to infrastructure issues, according to the report.
SpaceXAI gave the most specific explanation. “We are sorry for the issues you may have experienced with Grok following an outage at our Memphis compute center this morning,” the company wrote on X. SpaceXAI also apologized to its “impacted compute partners.”
No vendor has linked the incidents, and their timing alone does not establish a common cause.
What eWeek found: AI failover plans have a weak spot
After aligning the vendors’ published incident timelines, eWeek found that ChatGPT, Claude, and Grok were simultaneously affected for 93 minutes. Gemini was excluded from the calculation because available reports did not provide sufficiently consistent start and recovery times.
The overlap matters because switching providers is one of the simplest ways organizations try to build resilience into AI workflows. But applications tied to a single vendor’s API need preconfigured failover logic to switch automatically, while coding tools and intelligent automation may remain unavailable until engineers reroute them or the original service returns.
Teams developing AI resilience plans can test those plans before an outage by temporarily removing their primary model from a workflow. Employees should be able to complete the same task through the fallback without requesting new access, rebuilding prompts or context, or waiting for engineers to reconfigure integrations.
A failed handoff can expose gaps in authentication, API compatibility, context transfer, permissions, or employee training before those problems surface during a real outage.
One AI outage can often be routed around. Several at once reveal whether a backup plan actually exists in the workflow or merely on a vendor list.
Related reading: Find out how Astra and SafeMind take different approaches to enterprise cyber defense and where those strategies may overlap.


