Blameless incident reviews
DefaultWhen something breaks in production, we learn from it calmly and in writing. The post-incident review explains what happened and changes the system so the same thing is less likely next time.
Blameless by design
Section titled “Blameless by design”People do not come to work to cause outages. A review that hunts for someone to pin it on just teaches everyone to hide problems, so ours look for the conditions that let the incident happen, not the person who happened to be nearest.
It helps to hold the Retrospective Prime Directive in mind:
Regardless of what we discover, we understand and truly believe that everyone did the best job they could, given what they knew at the time, their skills and abilities, the resources available, and the situation at hand.
What a review covers
Section titled “What a review covers”- A plain timeline: what happened, when, and how we noticed.
- The impact, described in terms a client would recognise.
- The contributing factors, plural. Real incidents rarely have a single cause, so we keep asking why (the Five Whys are usually enough) until we reach a real root cause rather than a convenient place to stop.
- Where we got lucky. The things that went right by chance rather than by design, like the one person who knew the system happening to be online, are the gaps that will not always be filled. They belong in the write-up as much as what went wrong.
- A short list of actions with owners, tracked like any other work.
Make it count
Section titled “Make it count”Share the write-up rather than filing it away. The value is the change that follows, and the fact that other teams learn from it too. If nothing changes as a result, the review was just theatre.