monitoring

Installation
SKILL.md

Monitoring

You are wiring up the outside view of a service that already shipped: is it alive, is it fast, and when it breaks, does exactly one human get exactly one actionable page. This skill emits a concrete setup — a checker config, a health-endpoint contract, symptom-based alert rules, and an on-call rotation. Not telemetry instrumentation (that is ../observability/SKILL.md), not the release-gating healthcheck (that is ../deployment/SKILL.md).

The one rule

A page is justified only when there is real user impact AND a human action the system can't take itself. Everything below descends from this. Internalize the three tiers:

  • Page (wake someone): users are hurting now and a human must intervene. Checkout returns 5xx. Site unreachable. Error budget burning fast.
  • Ticket (look during business hours): degraded but not bleeding. Slow-burn budget use, cert expiring in 14 days.
  • Dashboard-only (don't notify): CPU at 80%, a single retry, a transient blip the system already healed.

If an alert doesn't map to an immediate human action, it is not a page — it's noise, and noise trains people to ignore the one page that matters.

The 4-layer stack

Don't skip layers and don't collapse them — each answers a different question.

Installs
3
GitHub Stars
116
First Seen
Aug 6, 2026
monitoring — ericrisco/rsc-harness