One air conditioner, 28 hours, $25 million

by Rick Snide | Oct 6, 2026 | IT Services

In May, a single air conditioner problem in a Virginia data center took down Coinbase's trading engine, stopped FanDuel from paying out bets during an NBA playoff game, and knocked a humanitarian aid platform offline — all at the same time, for the better part of a day and a half. 

What actually happened 

Here's what actually happened. Amazon Web Services runs a big chunk of the internet out of a region called us-east-1, in Virginia. One data-center hall inside that region had a cooling failure — multiple units went down, temperatures climbed past safe limits, and the automatic safety response kicked in: shut the servers off before they cook themselves. Reasonable move for the hardware. Bad day for everyone renting space on it. 

That one hall's outage didn't stay contained. It cascaded into a dozen other AWS services — load balancers, container orchestration, databases, search, machine learning tools, the works — for roughly 28 hours region-wide. By Coinbase's own account, they halted trading for close to six hours and didn't call it fully resolved for about 20. FanDuel couldn't process cash-outs mid-game. CME Group's institutional trading tool started throwing errors. And this wasn't AWS's first time with this exact problem — it's the third major us-east-1 failure since 2021. 

Why this matters 

One estimate put the total damage across affected businesses at around $25 million. AWS's own service credits for an outage like this typically run 10-30% of your monthly bill. Do that math for a second: if you're paying AWS $10,000 a month and an outage costs your business $50,000 in lost revenue, your maximum compensation is somewhere around $1,000-$3,000. That gap is yours to eat. 

I'm not bringing this up to pile on AWS — every major cloud provider has had a version of this — but because it's a good, expensive reminder of something a lot of businesses assume without ever checking: "the cloud" isn't one indestructible thing. It's someone else's data center, with someone else's air conditioning, and when that fails, your stuff goes down with it, on their timeline, not yours. 

What to do about it 

So the actual question worth asking isn't "should we trust the cloud" — you should, mostly, it's still more reliable than most businesses could build themselves. The question is: do you know what happens to your business if your cloud provider has a bad day? Not a philosophical question — a practical one. What's actually down? For how long can you survive it? Is there a backup path, even a manual one, for the two or three systems you genuinely cannot live without for a day? 

Most businesses I talk to have never actually answered that question, because it's uncomfortable and the cloud has mostly just... worked. Until the day it doesn't. 

The bottom line 

Here's my ask: pick the one system you absolutely cannot lose for a day — payroll, your CRM, whatever runs your ordering — and find out this week what your actual plan is if it goes down. If the honest answer is "I assume it won't," that's worth fixing before it's the thing that does. 

— Rick