Site navigation

What Can Businesses Learn From Google’s Post-Mortem Culture?

Graham Turner

,

Google post-mortem
At DIGIT’s Cloud First Summit, Google Engineering Manager Joydip Basu detailed how to successfully engender constructive post-mortem practices.

Failures and outages can be an everyday occurrence in a tech business (or any, for that matter). What’s important is how they’re dealt with.

Joydip Basu, an engineering manager at Google, is a huge proponent of his company’s philosophy that the fallout from an incident should be dealt with constructively and free from bad faith finger-pointing.

As he puts it: “How do we learn from failures? How can you, too?”

What are post-mortems?

At it’s most basic, a post-mortem is a written record of an incident – a customer can’t do something or there’s been a security breach, for example.

“You’re always going to have outages,” says Basu. What can make it a learning experience is properly documenting the event, with a detailed summary of the root cause, followed by effective and preventative actions to reduce likelihood or impact of a recurrence.

Before you even begin writing one, however, Basu is insistent that a post-mortem should be blameless. “This is core to Google’s SRE values,” as he says.

He adds: “Post-mortems are a great tool to help you understand complex systems. Investing in how you do them is time well spent as we’re never going to reach a point where we don’t have to do them as tech keeps changing along with the processes and use-cases in which they’re deployed.”

How should we write post-mortems?

The ethos behind this is a natural continuation of and reaffirmation of being objective, constructive and blameless – essentially, “fixing systems and processes, not people,” according to Basu.

Basu expands on why this is so critical: “Psychological safety enables interpersonal risk, allowing people to ask questions and admit failures. You want people to contribute to post-mortems freely, if there’s a culture of blame, people will be hesitant to share.”

Conversely a post-mortem shouldn’t be an opportunity to boast.

On this, Basu says: “Learn, but don’t celebrate heroism” – in this context, it means that post-mortems “aren’t a vehicle for sharing tales of glory, it’s not a tool for setting an example or creating role models”.

When should you write a post-mortem?

This is when you must consider your impact criteria – at what point does this event warrant chronicling: how many users were affected; how much revenue was lost, the potential impact of the event – these are all usable metrics and from there you can rank the severity of the vulnerability or fault.

“Even with near misses – i.e. the criteria would have been met if it wasn’t for luck – you should write a post-mortem,” according to Basu.

Finally, if an outage or incident presents an interesting learning opportunity, such as if there are interested parties that could benefit from reading the post-mortem, then you should “do a quick write up with relaxed with reviews and action items,” says Basu.

Who should write a post-mortem?

According to Basu, writing a post-mortem should be a “collaborative effort” which uses “real time/asynchronous collaboration tools whenever possible”. However, he notes that one person should be the owner of the post-mortem and all other collaborators should be considered as contributors.

Who will read your post-mortem?

Engineers, managers and others in the team whom the incident directly affects should be the first stop as they “only require a little background as they are already likely familiar with the context,” says Basu.

For these people, the main interest in the post-mortem will be the root cause analysis, as well as a balanced action item plan.


Recommended


Secondly, Basu says that directors, architects and affected teams should read as many post-mortems as possible. With these, “a detailed background is required as these are people who will not be very familiar with the context,” according to Basu.

What incident highlights/lowlights should be captured?

Root cause, trigger and impact – these are the critical aspects to a good post-mortem. Basu goes on to detail an example that incapsulates each of these elements.

  • Canary metrics didn’t detect a bug in a previously unused feature (root cause)
  • Feature was enabled by accident and rolled out (trigger)
  • Product ordering was unavailable for four hours, resulting in revenue loss  (impact)

Following this, Basu lays out an action item plan that covers who these issues would be addressed

  • Implement canary analysis to detect failures before rollout to all workers
  • Prevent accidental enabling of features by additional rollout checklist item

Basu closes out by extolling the virtues of reflecting on, “the lessons learned, what went well/poorly? Where did you get lucky?”

No business with a tech focus is going to run smoothly 100% of time – being constructive and judicious in how chronicle and action an incident is what can make or break a company.


Fintech Summit 2022

Our next conference is the Fintech Summit, held live and in-person at Edinburgh’s Dynamic Earth on 15th September.

Now in its ninth year, the Fintech Summit is Scotland’s largest annual gathering for financial technology professionals, providing an opportunity to reconnect with industry peers, build new relationships and explore the latest developments across the sector.

To secure your free place at the Summit, please visit: www.fintech-summit.co.uk

Graham Turner

Sub Editor

Latest News

AI

Nvidia Launches Open Secure AI Alliance for AI Safety and Security

AI Business Recruitment

Nearly a Quarter of Orgs Reducing Entry-level Hiring Due to AI Automation

Business

Scottish Businesses Turn to Self-funding as Growth Confidence Dips in H2

Data Finance

Payment Leaders are Struggling to Get Real-time Data