site-reliability-engineering

Installation
SKILL.md

Site Reliability Engineering

A comprehensive methodology for designing, operating, and improving reliable production systems. Rooted in Google SRE principles and extended with modern practices for incident command, observability engineering, error budget governance, and operational excellence.

When to Load This Skill

Trigger What It Means
"Design reliability into this system" SLO/SLI framework, error budget policy, resilience architecture
"Run an incident postmortem" Blameless postmortem with timeline, 5 Whys, action tracking
"Improve our on-call" Rotation design, alert tuning, toil reduction, escalation policy
"Build observability" The Four Golden Signals, dashboard design, alert rule patterns
"Do a reliability review" Architecture review against SRE principles, risk assessment
"I need an incident commander" Incident command framework, role cards, communication templates
"Automate this operational task" Toil assessment, automation decision tree, runbook pattern

Loading Order

Installs
1
GitHub Stars
167
First Seen
13 days ago
site-reliability-engineering — magnus919/hermes-profiles