stale read: down forty minutes after it wasnt [4b7C1Js3TWJ]
I'm refreshing a status page for the third time in five minutes, the way you check a text thread when someone said they'd call and didn't. AWS's CloudFront service is down, half the internet's education platforms have gone with it, and the page still says "investigating." What I don't know yet: the outage ended nine minutes before that update posted. On July 16, 2026, a capacity limit in a single Frankfurt availability zone triggered a global CloudFront control plane failure. Canvas, Blackboard, Hugging Face, and the UK National Lottery all went down. AWS's own retrospective says recovery happened at 11:18 UTC. Status updates at 11:27 and 11:57 still described it as ongoing. That forty-minute gap, and the security-versus-flexibility tradeoff baked into the feature that caused the whole thing, is what this episode is actually about. Not "AWS went down." A security feature quietly traded away operational flexibility that nobody was told they were giving up, and during the incident, AWS's own workaround was to switch that feature off. If you've ever refreshed a status page like reloading it faster would change the answer, or signed off on an architecture decision without asking what you were trading away, this one's for you. Topics: AWS outage, cloud architecture, incident response, observability, system design, engineering accountability, CloudFront, transparent systems