post-mortem
Installation
SKILL.md
Post-Mortem (Blameless Incident Review)
Overview
A post-mortem is a structured, blameless review held after an incident, outage, regression, missed launch, or failed experiment. The goal is not to assign fault but to learn how the system (people, process, code, and organization) produced the outcome, and to commit to durable changes that reduce the chance of recurrence.
This skill operationalizes the Google SRE blameless post-mortem template, the Etsy "morgue" tradition, John Allspaw's "How Complex Systems Fail" reading, Charles Perrow's Normal Accident Theory, and Sidney Dekker's Field Guide to Understanding "Human Error". Where the companion discovery/pre-mortem/ skill imagines failure before it happens, post-mortem learns from failure that already did.
Core Capabilities
- Severity classification — Sev 0-4 + near-miss thresholds, with post-mortem requirement and SLA per level
- Blameless facilitation — five ground rules, the Dekker New View, hindsight/counterfactual removal, the Allspaw test
- 10-section authoring — header → summary → impact → timeline → what went well/wrong → factors → RCA → actions → lessons
- Root-cause analysis — 5 Whys for clear chains, Causal Tree for multi-factor Sev 0/1/2; root-cause vs contributing-factor
- Action-item discipline — single owner, real ticket, testable acceptance criteria, recurring completion audit
- Distribution & archive — audience-appropriate formats; searchable, service-tagged archive