Site navigation

Google Report Sheds Light on UK Data Centre Heatwave Outage

Michael Behr

,

Google data centre heatwave
The record high temperatures were too much for the data centre.

With a data centre knocked out in the UK heatwave, Google has provided insights on what caused cloud services to fail across parts of Europe.

The company released its incident report, detailing the cause of the shutdown. According to this, one of its data centre’s cooling systems, including redundant ones, failed, with Google noticing the issue around 2:20pm. This forced the company to shut down its servers.

However, the report did not mention why the cooling systems failed.

Google’s report added that, despite efforts by its engineers to fix the problem during the day, the servers were ultimately switched off at 6PM to prevent more damage being done to the machines.

“We powered down this part of the zone to prevent an even longer outage or damage to machines. This caused a partial failure of capacity in that zone, leading to instance terminations, service degradation, and networking issues for a subset of customers,” the report noted.

The cooling systems was repaired at around 10pm when temperatures had dipped slightly. Services were restored to operation by 11:28 on the 20th of July.

On July 19th, the UK was hit by a major heatwave, with parts of the country recording temperature above 40°C.

The heat put strain on data centres across the country, which require cool temperatures to operate. This caused several, belonging to Google and Oracle, to shut down.

While these temperatures are manageable for data centres, the affected ones were not designed to withstand the unprecedented UK heat.

This is turn led to service outages for customers as the companies were forced to take their centres offline to protect the infrastructure.

This made services unavailable across Google’s Europe West 2 zone, including all virtual machines running on Google’s Compute Engine – approximately 35% of all VMs in the area.

The outage occurred despite Google designing its systems so that regional services can survive the failure of a single zone.


Recommended


Google blamed the data centre outage on it inadvertently modifying traffic routing to avoid all three zones in the Europe West 2 region, rather than just the impacted Europe West 2-a zone. This also stopped the company from accessing data backups.

In order to prevent a similar outage occurring, the company said that it would carefully re-test its failover automation to ensure stronger resilience in its failover protocols during large-scale events.

In addition, it said it will investigate and develop more advanced methods to progressively decrease the thermal load within a single data centre, reducing the probability that a full shutdown is required.

Google noted that a Full Incident Report is being prepared that will provide additional details on the cause of the outage.


Get the latest news from DIGIT direct to your inbox

Our newsletter covers the latest technology and IT news from Scotland and beyond, as well as in-depth features and exclusive interviews with leading figures and rising stars.

To subscribe, click here.

Michael Behr

Senior Staff Writer

Latest News

AI

Nvidia Launches Open Secure AI Alliance for AI Safety and Security

AI Business Recruitment

Nearly a Quarter of Orgs Reducing Entry-level Hiring Due to AI Automation

Business

Scottish Businesses Turn to Self-funding as Growth Confidence Dips in H2

Data Finance

Payment Leaders are Struggling to Get Real-time Data