qa-observability
Installation
SKILL.md
QA Observability
Use telemetry as a QA signal and a debugging substrate. Treat logs, metrics, traces, and profiles as evidence for test outcomes, release readiness, and production regressions.
Core references live in data/sources.json. Prefer primary docs and re-check volatile external facts before recommending versions, pricing, or vendor features.
Quick Start (Default)
If key context is missing, ask for: critical user journeys, service/dependency inventory, environments (local/staging/prod), current telemetry stack, and current SLO/SLA commitments.
- Establish the minimum bar: correlation IDs, structured logs, traces, and golden metrics (latency, traffic, errors, saturation).
- Verify the telemetry transport with one known request or fault: application emission, context propagation, collector acceptance, and backend query. Stop the claim at the last observed stage; static SDK or collector configuration proves configuration only.
- Make failures diagnosable: every integration or E2E failure should capture a trace link or trace ID plus correlated logs, and critical degraded paths should expose structured error metadata such as rate-limit codes, retry hints, and state-transition markers.
- When dashboards, SLOs, or alerting are in scope, define the relevant SLI/SLO and error-budget policy, verify that the dashboard selects the known telemetry, and test alert evaluation and routing with a synthetic condition and a controlled non-production receiver. Do not notify real responders without explicit authorization. Record backend-query, dashboard-selection, rule-evaluation, route, and delivery results independently; basic instrumentation can stop after backend-query proof.
- Produce the artifacts that match the task: a readiness checklist, SLO definition, or alert rules using
assets/checklists/template-observability-readiness-checklist.md,assets/monitoring/slo/slo-definition.yaml, andassets/monitoring/slo/prometheus-alert-rules.yaml. Preserve environment, service revision, trace/request ID, query window, sampling decision, and the result of each stage exercised.