monitoring
Monitoring
You are wiring up the outside view of a service that already shipped: is it alive, is it fast, and when it breaks, does exactly one human get exactly one actionable page. This skill emits a concrete setup — a checker config, a health-endpoint contract, symptom-based alert rules, and an on-call rotation. Not telemetry instrumentation (that is ../observability/SKILL.md), not the release-gating healthcheck (that is ../deployment/SKILL.md).
The one rule
A page is justified only when there is real user impact AND a human action the system can't take itself. Everything below descends from this. Internalize the three tiers:
- Page (wake someone): users are hurting now and a human must intervene. Checkout returns 5xx. Site unreachable. Error budget burning fast.
- Ticket (look during business hours): degraded but not bleeding. Slow-burn budget use, cert expiring in 14 days.
- Dashboard-only (don't notify): CPU at 80%, a single retry, a transient blip the system already healed.
If an alert doesn't map to an immediate human action, it is not a page — it's noise, and noise trains people to ignore the one page that matters.
The 4-layer stack
Don't skip layers and don't collapse them — each answers a different question.