chaos-engineering
Installation
SKILL.md
Chaos Engineering
End-to-end chaos engineering: experiment design, fault injection catalog, gameday execution, and the maturity model that turns one-off "let's break stuff" exercises into a reliable discipline. Provider-agnostic — works whether you use Litmus, Chaos Mesh, AWS FIS, Gremlin, ChaosToolkit, or hand-rolled scripts.
This skill answers four questions: what to inject, where to inject it, how to size the blast, and how to extract durable learning from each run.
When to use this skill
| Situation | Skill applies |
|---|---|
| Spinning up a chaos program from scratch | Yes — start with maturity model + first 5 experiments |
| Designing a single experiment for a known concern | Yes — use the experiment design loop |
| Planning a gameday for a team or service | Yes — use scripts/gameday_planner.py |
| Validating a kill switch or fallback path actually works | Yes — chaos is the way to test these in prod-like conditions |
| Post-incident verification: "did the fix really fix it?" | Yes — re-inject the original fault, confirm the new behavior |
| Compliance evidence (SOC 2 A1 / DORA Art. 25) | Yes — chaos runs produce auditable resilience-testing evidence |
| Improving SLOs / error budgets | Pair with engineering/observability-designer — chaos surfaces SLO violations |