Site navigation

AWS Disruption Exposes Fragility of Global Cloud Infrastructure

Graham Turner

,

AWS outage
A major outage centred on AWS’s US-EAST-1 region left hundreds of apps and websites offline for much of the day, prompting fresh scrutiny of the concentration of services on a handful of cloud providers.

Amazon Web Services said late on Monday that it had resolved a major outage that left hundreds of apps, websites and services offline for much of the day, following an incident that arguably raises some uncomfortable questions around the interconnectivity and fragility of global digital infrastructure, with UK MPs also raising questions about dependence on overseas hosting.

What happened?

The disruption, which appears to have begun at around 7am BST on Monday, affected a wide range of services.

Streaming platforms including Amazon’s Prime Video and Disney+ were hit, as were apps such as Airbnb, Snapchat and Duolingo. Banking customers in the UK reported problems too, with Lloyds, Halifax and Bank of Scotland among those logged by outage tracker Downdetector.

Users also reported issues with messaging apps Signal and WhatsApp across Europe, while gaming platforms including Fortnite and Roblox were disrupted. Perplexity AI, Coinbase and Robinhood were among other services affected, as was HMRC’s website, among many others, including Amazon’s own Ring doorbells – insane to think that an IT issue in West Virginia can cause your doorbell to stop working.

Downdetector recorded an initial huge spike in reports on Monday morning, and then a much larger surge some nine hours later, posting that it had received more than 11 million reports in total.

Early in the day, Downdetector told the BBC it had seen more than four million reports from users across 500 sites within just a few hours – more than double the amount it would see across an entire regular weekday – and the reports later peaked at more than 11 million as further services attempted to recover.

Amazon said the system at issue was back to “pre-event levels” and it is working through the data backlog caused by the problem (they said it would take two hours, so that should be done by time of publishing).

In an update, Amazon said “mitigations were applied to resolve launch failures”, linking a “load balancer health” issue to the problem at Amazon Web Services (AWS). At 10.27am on Monday, the company said it had seen “significant signs of recovery”. Later, at 11.35am, AWS added that the “underlying DNS (domain name system) issue has been fully mitigated”. At around 11pm BST Amazon said all AWS services had “returned to normal operations.”

AWS said on its status page that Monday’s outage originated at its US-EAST-1 location in northern Virginia, its oldest and largest region for web services. According to documentation on the AWS website, US-EAST-1 is often the default region for many AWS services – the region has suffered outages in previous years.

Amazon said the issue “appears to be related to DNS resolution of the DynamoDB API endpoint in US-EAST-1”.

Outage raises serious questions

The scale of the disruption drew fresh attention to the concentration of digital services on a small number of large cloud providers.

AWS handles a substantial share of global cloud infrastructure and powers millions of apps and websites.

The widespread impact – from gaming and streaming to banking and government services – was described by experts and academics in the briefing material as a demonstration of how interconnected everyday digital services have become and of their reliance on a handful of providers.

In the UK, the outage has prompted scrutiny from MPs about the resilience of critical infrastructure that depends on overseas hosting.

The Treasury Committee has questioned why Amazon had not been designated a critical third party (CTP) under new rules that came into force earlier this year, which allow regulators to intervene to improve the resilience of key service providers to the financial sector.

In a letter to Lucy Rigby MP, the economic secretary to the Treasury, the committee set out a series of questions linked to the outages and asked why the Treasury had not designated Amazon Web Services, or any other major technology firm, a CTP.


Recommended reading


Committee chairwoman Meg Hillier also cited speculation that the AWS outage related to its US operations and asked if the Treasury was concerned that “seemingly key parts of our IT infrastructure are hosted abroad”? The committee asked what work the Treasury was doing with HMRC, which it said might have been affected by the outages, to look at what went wrong and how to prevent such incidents in future.

Amazon has not yet provided a fuller public accounting beyond its status updates and statements on the company’s service page.

Graham Turner

Sub Editor

Latest News

AI

Nvidia Launches Open Secure AI Alliance for AI Safety and Security

AI Business Recruitment

Nearly a Quarter of Orgs Reducing Entry-level Hiring Due to AI Automation

Business

Scottish Businesses Turn to Self-funding as Growth Confidence Dips in H2

Data Finance

Payment Leaders are Struggling to Get Real-time Data