devops-sre-practices

Installation
SKILL.md

SRE Practices

Purpose

Implement Site Reliability Engineering practices: define SLIs/SLOs aligned with business goals, manage error budgets with burn rate alerts, systematically reduce toil, conduct blameless incident analysis, build production readiness reviews, and mature reliability culture across the organization.

Agent Protocol

Trigger

Exact user phrases: "SRE", "site reliability", "SLI", "SLO", "error budget", "error budget policy", "burn rate", "toil", "toil reduction", "reliability engineering", "postmortem", "incident analysis", "5 whys", "production readiness review", "PRR", "reliability dashboard", "multi-window", "multi-burn-rate", "reliability maturity", "SLO monitoring", "service level objective", "service level indicator".

Input Context

  • Current monitoring and alerting stack (Prometheus, Datadog, Grafana)
  • Existing incident response process
  • Team size and on-call rotation structure
  • Current service-level objectives (if any)
  • Known reliability pain points and past incidents
  • Business context (revenue-critical services vs internal tools)
Installs
10
GitHub Stars
21
First Seen
May 30, 2026
devops-sre-practices — j4flmao/agent-skills