After two major outages at the tail end of the year, internet infrastructure giant Cloudflare is pushing hard to build renewed resilience in hopes it can prevent similar blackouts from happening again.
Dubbed “Code Orange: Fail Small”, the operation will see Cloudflare prioritise a slow roll-out of iterative improvements to its network that will increase resilience to errors or mistakes that lead to widespread outage.
In a blog post, the company’s Chief Technical Officer, Dane Knecht, wrote that both of Cloudflare’s outages followed a similar pattern, happening moments after the firm deployed a configuration change in its data centres in hundreds of cities around the world.
The first outage, which happened in November and took down services like YouTube, ChatGPT and Sage, was sparked by an automatic update to Cloudflare’s Bot Management classifier, while December’s issue was caused by changes to a security tool meant to protect customers from a recently discovered vulnerability in the open source React framework.
These “deeply embarrassing” episodes exposed what the Cloudflare exec described as a serious gap in how the company deploys configuration changes, with the system prioritising speed over caution – until now considered an advantage that allowed clients to make changes that would distribute globally in seconds.
Outlining the firm’s plans, Knecht said the Fail Small operation will be organised across three main areas that address those issues that triggered the global incidents experienced in the last two months.
The most important of these workstreams, according to Knecht, is the introduction of controlled rollouts for configuration changes, bringing them in line with Cloudflare’s software‑release policy, which requires teams to define success metrics, outline a rollout plan, and document what happens if a deployment fails.
At the moment, the company’s Quicksilver software component sees any new DNS record or security rule reach 90% of servers on the network within seconds, a speed which allowed for ‘breaking change’ to propagate across Cloudflare’s systems before being tested.
By the end of the Code Orange mission, the progress of configuration updates will be more carefully monitored, with a deployment toolkit used to begin automatic rollbacks if something goes wrong.
“We expect this to allow us to quickly catch the kinds of issues that occurred in these past two incidents long before they become widespread problems,” said Knecht.
To make sure of that, the company is also in the process of reviewing interface contracts between every critical product and service across its network to test how it handles failure at any point.
By assuming the worst, that failure will occur between each interface, Cloudflare said that widespread outages like those caused by changes to its Bot Management service could have been avoided.
Knecht admitted that in that incident, there were at least two key interfaces where, “if we had assumed failure was going to happen, we could have handled it gracefully to the point that it was unlikely any customer would have been impacted.”
Recommended reading
- Cloudflare Scrambles to Recover From Major Outages
- Who Gets to Train the Internet? Interview With Cloudflare’s VP of Product
- What’s the Latest With the AWS Outage?
The third area of improvement will focus on Cloudflare’s internal “break glass” procedures, which let pre-approved staff members temporarily elevate their access to handle urgent, high‑severity issues.
The company hopes that doing away with “circular dependences” and allowing smoother access to high-level tools might solve emergencies faster, but that must be balanced with keeping customer data safe and preventing unauthorised access.
By the end of Q1’26, Cloudflare said it plans to have made significant headway, but that the company doesn’t see this work as something to be completed. Rather, they form a sea change in how it will approach product and engineering work from now on.
“Some of these goals will be evergreen. We will always need to better handle circular dependencies as we launch new software, and our break glass procedures will need to update to reflect how our security technology changes over time,” said Knecht.
“We failed our users and the Internet as a whole in these past two incidents. We have work to do to make it right.”





