Most Companies Have No Clue How Expensive a Cloud Outage Really Is
Tyler
Co-Founder & CEO

A year ago, "the cloud is down" was still a sentence that made headlines. A server room caught fire, a region went dark, an engineer fat-fingered a config push, and for a few hours half the internet felt it. It was rare enough to be a story.
It is not rare anymore. It is Tuesday.
Over the past 90 days, I have watched the causes of cloud disruption shift under our feet, watched AI platforms rack up outage numbers that would have been unthinkable a year ago, and watched the SLA credit math that's supposed to make enterprises whole fall further behind the actual cost of downtime. None of this happened quietly. It is all sitting in plain sight, in earnings calls, in incident postmortems, in research reports nobody in the boardroom is reading closely enough.
So let's read them closely.
The Nature of the Problem Is Changing, Not Just the Frequency
For most of the last decade, "cloud outage" meant something physical: a failed drive, a downed power feed, a severed fiber line. Power still holds that title: it remains the single leading cause of major outages today. But InfoWorld's recent analysis of the Uptime Institute's seventh Annual Outage Analysis makes a point that deserves more attention than it's getting: the fastest-growing share of impactful outages no longer comes from hardware at all. IT and networking issues now account for 23% of impactful outages, and the share caused by human error, specifically staff failing to follow established procedures, rose 10 percentage points year over year. Today's major outages are increasingly shaped by complexity itself: control-plane failures, cascading configuration errors, and automated systems making bad decisions faster than any human could catch them.
That distinction matters enormously for anyone trying to manage downtime risk. Physical infrastructure failures are the kind of thing you can diversify against: multi-region, multi-AZ, redundant power. Complexity and process failures are different. They propagate through the exact automation and interdependency that make modern cloud platforms powerful in the first place, and they don't show up on a backup generator's maintenance schedule. The more sophisticated the system, the more ways it has to fail in a way nobody designed for. The Uptime Institute puts a price on that shift: 54% of organizations said their most recent significant outage cost more than $100,000, and 20% said it cost more than $1 million.
TechRadar's reporting backs this up with a number that should concern anyone running a board-level risk conversation: a small handful of hyperscalers now underpin the infrastructure for a huge share of the EU's digital economy, and enterprise contingency planning has not kept pace with how concentrated that dependency has become. When three providers control roughly two-thirds of the global cloud market, "the cloud" isn't a redundancy strategy. It's a few companies' control planes, and you are downstream of all of them at once.
What This Actually Costs, and Why AI Is Now Part of the Problem
Before we get to the SLA math, it is worth being honest about the size of the number. A June 2026 report from Splunk, now part of Cisco, conducted with Oxford Economics across 2,000 executives at Global 2000 companies, put the annual cost of unplanned downtime at $600 billion, up roughly 50% in just two years.
What downtime costs in 2026
- $600 billion a year in unplanned downtime, up roughly 50% in two years
- ~$15,000 per minute of downtime, with the average company absorbing close to $300 million a year before anyone formally calls it a crisis
- A major incident knocks an average of 3.4% off a company's stock price
- Ransomware payouts have nearly tripled to $40 million; regulatory fines now average $51 million
Here is the part that should reframe how you think about your AI spend. Companies have poured money into AI specifically to prevent this, a median of $24.5 million a year on AI systems meant to catch problems before customers do. And yet AI is increasingly the thing that breaks. Half of the organizations surveyed reported downtime tied to incorrect AI automation or model drift, and nearly a third traced outages to bugs introduced by embedding AI into production systems. Splunk calls it the reliability paradox: the harder companies lean on AI to eliminate operational risk, the more they manufacture a newer, less predictable version of it.
What makes this category so dangerous is that it often fails silently. As RTInsights described it in June 2026, a distributed AI node that loses connectivity may not crash at all. It keeps serving predictions on an increasingly stale model state, outputs that look normal while the damage compounds beneath them. A logistics system routing on five-minute-old traffic data is not slow, it is wrong. A maintenance model reading a sensor feed that stopped updating twenty minutes ago is not delayed, it is dangerous. Traditional outages turn a dashboard red; these do not, which is exactly why they run longer and cost more before anyone acts. RTInsights pegs the going rate of IT downtime at more than $33,000 per minute, and rising.
Two structural gaps make it worse. The first is visibility: only 38% of technology executives in the Splunk survey said they could consistently identify the root cause of a downtime incident, despite heavy spending on monitoring. The second is shadow AI. Two-thirds of organizations report employees using unapproved AI tools to write code and make decisions, with no central record of what data those tools touch or how their outputs move into production. You cannot file an SLA claim, or even reconstruct an outage, for a failure you never saw.
Then AI Made It Dramatically Worse
If outages were merely changing in character, that would be enough to justify a hard look at your downtime exposure. But the last twelve months added a second variable that compounds the first: the AI platforms now wired into daily enterprise workflows are failing at a rate that has no precedent.
Ookla's analysis of 471 days of U.S. Downdetector data, 3.7 million user-reported problems across ChatGPT, Claude, Gemini, and Microsoft Copilot, found that the number of high-disruption days (days where a platform's outage reports spiked to more than 10x its own median) rose from 6 in Q1 2025, to 16 in Q4 2025, to 51 in Q1 2026: roughly a 750% increase over the year. The analysis, by Ookla researcher Luke Kehoe, was independently covered by IEEE ComSoc's Technology Blog, Mobile World Live, and Telecoms.com.
Here is where you have to read the data carefully, because the headline metric can invert the real picture. On that spike measure, Claude accounted for 39 of the 51 high-disruption days, Gemini 7, Copilot 3, and ChatGPT just 2. Read quickly, that makes OpenAI look like the most reliable name in the group. It is not, and the gap between those two readings is the whole point. Downdetector only registers a high-disruption day when a platform's reports spike far above its own routine volume, so a service that is busy and frequently troubled every single day rarely trips the wire, while a fast-growing platform looks catastrophic on a single bad afternoon.
When we checked the vendors' own status pages instead, the ranking inverted. OpenAI's public status history, counted entry by entry, logs 88 separate incidents across 51 distinct days between March 18 and June 16, 2026, more than any platform on the chart above. Google's Gemini API and AI Studio history tells a similar story beneath a flat headline: almost the same incident count in Q1 2026 as in Q1 2025, 19 versus 19, but the severity mix shifted hard, from zero major incidents in the 2025 window to nine of nineteen in 2026.
The lesson is not that one platform is uniquely bad. It is that no single dashboard, Downdetector or a vendor's own status page, tells you the truth on its own. The entire category of AI infrastructure your teams now depend on for code generation, customer support, and internal tooling has become measurably less reliable in the same window it became measurably more embedded in how work gets done.
That embedding is exactly what turned a reliability problem into an SLA problem. TechTimes reported that GitHub's AI agent infrastructure crisis forced Microsoft to reroute load to AWS just to keep services running, a quiet admission that the boundaries between providers are blurrier than the contracts assume, and that an outage on "your" platform can now originate in someone else's infrastructure entirely. Days later, Copilot suffered its second major outage in eleven days, a multi-hour authentication failure between Copilot and Microsoft Graph, which exposed something most enterprise buyers had missed entirely: unlike Exchange Online, SharePoint, or Teams, Copilot carries no financially backed uptime guarantee under standard Microsoft 365 SLAs at all.
This is the part most procurement and FinOps teams have not caught up to: the SLA you negotiated for compute and storage was never written with "the AI agent your engineers rely on all day is now down 51 days a quarter" in mind.
The SLA Math Was Never Built for This
Here's where it gets expensive, and where the right response actually matters. Cloud provider SLA credits are structured as a percentage of your monthly service fee, scaled to how badly your uptime missed the target: typically something like a 10% credit below 99.99% availability, climbing to 25% below 99.0%, and 100% only in the most extreme failures below 95.0%. That structure made rough sense when outages were rare, short, and proportional to what you paid for the affected service.
It makes much less sense now, and the gap shows up first in what companies actually recover. A detailed SoftwareSeni analysis of the October 20, 2025 AWS US-East-1 outage, which generated more than 8 million Downdetector reports and affected over 1,000 companies, found that the SLA credits enterprises actually received covered, on average, only about 8% of the real business loss the outage caused. TechRadar's reporting puts a range on just how large that loss can get: a separate 15-hour AWS outage in fall 2025 produced estimated material damages of $38 million to $581 million, against an SLA credit structure capped at a fraction of the monthly bill. Part of that 8% gap is structural and no amount of process will close it. A meaningful part of it isn't: it's a tracking and claims-filing gap, and that part is entirely fixable.
Start with what's genuinely fixed. SLA credits are capped by what you pay your cloud provider each month, while your business loss is capped by nothing: it scales with revenue, customer commitments, and the size of what you've built on top of the infrastructure that went down. A small SaaS company and a large enterprise can experience the exact same multi-hour outage and walk away with credits in the same narrow band, even though one of them absorbed a loss two orders of magnitude larger than the other. No amount of internal process changes that ceiling. What it does change is whether you actually collect the full amount the contract already owes you, which is the part most companies leave on the table.
The November 2025 Cloudflare outage put a number on what that looks like in practice: an estimated $250 million in broader economic impact, with individual companies like Shopify reporting roughly $4 million in direct losses and over $170 million in downstream effects across their merchant ecosystem. None of that gets made whole by a standard SLA credit, and even the credit itself only gets paid if someone inside the organization maps the outage timeline to the SLA terms and files the claim before the window closes. If you already have that process in place, the right tooling, and the discipline to run it across every provider you use, this is a manageable problem: you're capturing what you're owed, and the math above is simply the ceiling you're working within. If you don't have that in place, the growing frequency and cost of these outages is exactly the kind of risk that should drive a buy decision. Building and maintaining that capability yourself is a real option, but for most teams it's a worse use of budget than buying a system purpose-built to do it.
Concentration Risk Is Compounding the Problem
Layer one more thing on top of this: the infrastructure underneath "the cloud" is more concentrated, and more fragile in specific physical locations, than most enterprise risk models account for. CryptoBriefing reported that a fire at a third-party Delhi colocation facility on June 9 forced an emergency power shutdown, knocking out networking equipment serving Google Cloud's asia-south2 region and causing latency spikes and packet loss for users across Delhi, Chennai, Mumbai, and the surrounding metro areas, a reminder that for all the abstraction of "the cloud," it still runs on physical buildings that can burn, flood, or lose power, and that a single facility going dark can take a meaningful slice of a region's enterprise traffic with it.
Combine that physical concentration with the market concentration TechRadar flagged (three providers controlling roughly two-thirds of global cloud infrastructure), and you get a risk profile where a fire in one data center, a bad config push in one control plane, or an AI agent crisis in one provider's stack can cascade across thousands of companies that have no direct relationship to whatever actually failed.
Regulators are moving faster than risk committees
The EU's Digital Operational Resilience Act (DORA) now requires financial entities to maintain documented registers of cloud provider performance and enforce SLA accountability. As of November 2025, AWS, Azure, and Google Cloud are formally designated as critical ICT infrastructure under direct EU supervision. That gap between what regulators require and what most enterprises track is worth closing before it closes on its own terms.
What You Should Actually Be Doing Right Now
- Stop treating SLA credits as something your cloud provider proactively gives you. They don't. The credit is yours contractually, but claiming it requires documented proof of an outage window mapped against the SLA terms, and providers have no incentive to make that process easy. If no one owns this internally, the default outcome is zero recovery.
- Build (or buy) continuous SLA monitoring across every provider you use, not just your primary one. The GitHub/Microsoft/AWS rerouting episode is the clearest evidence yet that outages don't respect the boundary between "your" cloud provider and the providers behind them. Monitoring one contract while your workloads quietly depend on three is monitoring the wrong thing.
- Re-read your AI vendor SLAs with the last 90 days of outage data next to them. If your Copilot, Claude, or Gemini-dependent workflows aren't covered by an SLA that reflects 51 high-disruption days a quarter as the new normal, you're carrying risk your contract doesn't acknowledge exists.
- Quantify your actual downtime exposure, not just your contractual entitlement. The 8% coverage gap is an average. Your number could be worse. Map your last few major provider incidents against your real business impact (lost revenue, SLA breaches with your own customers, support costs) and compare that to what you actually claimed and received.
- Treat third-party and physical concentration risk as a board-level question, not an IT footnote. A data center fire in another country can hit your uptime. Ask your providers directly about facility-level redundancy and concentration, and document the answer the way DORA already requires regulated entities to.
Things To Think About
The story used to be that cloud outages were rare, physical, and roughly proportional to what you paid when they happened. None of that is true anymore. Outages are increasingly caused by complexity and automation rather than hardware. AI platforms, now load-bearing infrastructure for daily work, are failing at a rate that grew 750% in a year. And the SLA credit system meant to make you whole when any of this happens was built for a world where outages were the exception, not the operating condition.
The gap between what you're owed and what you actually recover isn't a rounding error. For most enterprises, it's the difference between an outage being a bad day and an outage being a six- or seven-figure unrecovered loss. Closing that gap requires the same thing it always has: someone (or something) that watches every SLA, every provider, every incident, continuously, and turns "we were down" into a filed and paid claim before the window closes.
That's the problem Next Signal exists to solve. Cloud outages are inevitable. Accountability isn't. Not unless you build it.
Frequently Asked Questions
How much do cloud and AI outages cost businesses in 2026?
Unplanned downtime costs businesses an estimated $600 billion a year, up roughly 50% in two years, according to a Splunk and Oxford Economics report. That is about $15,000 per minute overall, and IT downtime specifically runs more than $33,000 per minute per RTInsights.
Do SLA credits cover the cost of a cloud outage?
No. In the October 20, 2025 AWS US-East-1 outage, SLA credits covered on average only about 8% of the real business loss, per a SoftwareSeni analysis. Credits are capped by your monthly cloud spend, not by the loss you incur, so the recovery gap widens as a company grows.
Is ChatGPT more reliable than Claude or Gemini?
It depends entirely on the metric. Downdetector's spike measure flagged ChatGPT only twice in Q1 2026 because that metric compares a platform to its own baseline. On its own status page, OpenAI logged 88 incidents across 51 days in a comparable window, more than any platform. No single dashboard tells the full reliability story.
Does Microsoft Copilot have a financially backed SLA?
No. Unlike Exchange Online, SharePoint, or Teams, Microsoft 365 Copilot carries no financially backed uptime guarantee under standard Microsoft 365 SLAs, which left enterprises with no contractual protection during Copilot's June 2026 outages.
What is DORA and which cloud providers does it cover?
The Digital Operational Resilience Act (DORA) is an EU regulation on operational resilience for financial entities. As of November 2025, AWS, Azure, and Google Cloud are formally designated as critical ICT third-party providers under direct EU supervision.
Sources
Industry data and reporting cited in this article:
- InfoWorld: The Causes of Cloud Outages Are Changing
- TechRadar: What Cloud Outages Tell Us About Putting Your Eggs in One Basket
- Splunk / Oxford Economics: The Hidden Costs of Downtime (via Fast Company)
- RTInsights: The Hidden Cost of AI Infrastructure Downtime
- TechTimes: GitHub's AI Agent Crisis Forces Microsoft to Tap AWS as Outages Break Enterprise SLAs
- TechTimes: Microsoft Copilot Fails Twice in June, Enterprise IT Has No SLA Protection for AI Downtime
- CryptoBriefing: Google Cloud Outage in India Triggered by Third-Party Data Center Fire
- IEEE ComSoc Technology Blog: Ookla, AI Platform Reliability Decreases as Outages Surge
- Mobile World Live: Ookla Finds AI Platform Outages Surge as Adoption Grows
- Telecoms.com: AI App Disruption Is on the Up
- Google Cloud / Gemini API Service Health
- OpenAI Status History
- SoftwareSeni: Calculating the True Cost of Cloud Outages and Downtime
- Cloudflare: 18 November 2025 Outage
- Digital Operational Resilience Act (DORA), Regulation (EU) 2022/2554