observability-and-reliability

Installation
SKILL.md

Observability and reliability

Monitoring tells you a thing you predicted is happening. Observability lets you ask a question you did not anticipate. Production failures are mostly the unanticipated kind.

Instrument for questions you have not thought of yet

Emit structured events with enough context to slice afterwards — request identifiers, user or tenant, version, dependency, outcome, duration. Free-text logs are unsearchable at volume and become expensive noise.

Propagate a correlation identifier across every hop. Without it, a distributed system is a set of independent stories and reconstructing one request is manual archaeology.

Measure what the user experiences at the percentile they experience it. A p50 latency graph is mostly a graph of the people who were not affected.

Alert on symptoms, not causes

Installs
4
GitHub Stars
1.3K
First Seen
12 days ago
observability-and-reliability — cbrock84/headcount