observability-and-reliability
Installation
SKILL.md
Observability and reliability
Monitoring tells you a thing you predicted is happening. Observability lets you ask a question you did not anticipate. Production failures are mostly the unanticipated kind.
Instrument for questions you have not thought of yet
Emit structured events with enough context to slice afterwards — request identifiers, user or tenant, version, dependency, outcome, duration. Free-text logs are unsearchable at volume and become expensive noise.
Propagate a correlation identifier across every hop. Without it, a distributed system is a set of independent stories and reconstructing one request is manual archaeology.
Measure what the user experiences at the percentile they experience it. A p50 latency graph is mostly a graph of the people who were not affected.