qa-observability

Installation
SKILL.md

QA Observability

Use telemetry as a QA signal and a debugging substrate. Treat logs, metrics, traces, and profiles as evidence for test outcomes, release readiness, and production regressions.

Core references live in data/sources.json. Prefer primary docs and re-check volatile external facts before recommending versions, pricing, or vendor features.

Quick Start (Default)

If key context is missing, ask for: critical user journeys, service/dependency inventory, environments (local/staging/prod), current telemetry stack, and current SLO/SLA commitments.

  1. Establish the minimum bar: correlation IDs, structured logs, traces, and golden metrics (latency, traffic, errors, saturation).
  2. Verify propagation: confirm traceparent and your request ID flow across boundaries end-to-end.
  3. Make failures diagnosable: every integration or E2E failure should capture a trace link or trace ID plus correlated logs, and critical degraded paths should expose structured error metadata such as rate-limit codes, retry hints, and state-transition markers.
  4. Define SLIs/SLOs and an error budget policy; wire multi-window burn-rate alerts.
  5. Produce artifacts: a readiness checklist, an SLO definition, and alert rules using assets/checklists/template-observability-readiness-checklist.md, assets/monitoring/slo/slo-definition.yaml, and assets/monitoring/slo/prometheus-alert-rules.yaml.

Default QA stance

Installs
162
GitHub Stars
79
First Seen
Jan 23, 2026
qa-observability — vasilyu1983/ai-agents-public