Site navigation

GitHub Admits “Deep Architectural Work” Needed to Address Outages

Tom Quinn

,

GitHub outages
The platform has added new safeguards, improved monitoring and begun isolating critical components to prevent future cascades.

Dev platform Github has admitted that “deep architectural work” is needed to resolve the root causes behind a series of recent outages that left almost all of its services critically impaired.

In a post-mortem blog, GitHub’s Chief Technology Officer, Vladimir Fedorov, said that the major incidents across February and March were a result of “extremely rapid usage growth” across the platform, which had exposed scaling limitations in parts of its architecture.  

Specifically, GitHub’s instability was put down to spiking load, with the platform hammered by more traffic than it could handle, and the system unable to block or filter misbehaving clients from spamming it with excessive requests.

In early February, for instance, two popular apps shipped updates that began generating more than ten times their usual API traffic, a problem that surfaced only as users slowly adopted the new versions.

Fedorov also said that after deploying a new model and “trying to get it to customers as quickly as possible”, the platform had cut a critical cache refresh window from 12 to 2 hours, with subsequent issues initially going unnoticed due to lower weekend traffic and limited monitoring.

On February 9th, the platform saw a major incident when peak traffic, widespread client‑app updates and a new model release combined to overwhelm a key database cluster, with legacy architecture storing user settings, authentication and management strained.

Compounding those issues, Fedorov said that “architectural coupling” within the system had allowed localised issues to quickly spread to other critical services, which snowballed into larger outages.

Those outages continued into March, when request failures hit 40% for github.com, around 43% for the GitHub API, and 21% for GitHub Copilot requests. Another incident, days later, saw 95% of workflow runs fail to start within five minutes for GitHub Actions, with an average delay of 30 minutes, while 10% of workflow runs failed with an infrastructure error.

Problems peaked on March 19th, when the average error rate for GitHub’s Copilot Coding Agent service hit 99%, caused by an authentication failure inside the system, which stopped the service from being able to connect to the database it depends on. 

According to GitHub’s SVP of engineering, Jakub Oleksy, the platform needs “deep architectural work”, and has begun that by introducing a killswitch that lets engineers instantly disable the caching system, implementing automated monitoring for credential lifecycle events, and isolating the cache on a dedicated host, limiting any future issues to the services that depend on it.


Recommended reading


Fedorov added that, as a near-term priority, the platform had expedited capacity planning and aimed to complete a full audit of critical data and compute infrastructure to address the scale of growth, with plans to migrate infrastructure to Azure to “enable both vertical scaling within regions and horizontal scaling across regions”.

Engineers are also working to isolate key dependencies so that systems like GitHub Actions and Git will not be impacted by shared infrastructure issues, and protect downstream components during spikes, what the CTO called “breaking apart the monolith.”

“We know GitHub is critical digital infrastructure, and we are taking urgent action to ensure our platform is available when and where you need it,” concluded Fedorov.

Tom Quinn

Staff Writer, DIGIT

Latest News

AI

Nvidia Launches Open Secure AI Alliance for AI Safety and Security

AI Business Recruitment

Nearly a Quarter of Orgs Reducing Entry-level Hiring Due to AI Automation

Business

Scottish Businesses Turn to Self-funding as Growth Confidence Dips in H2

Data Finance

Payment Leaders are Struggling to Get Real-time Data