failure-mode-designer
Installation
SKILL.md
Failure Mode Designer
Start from critical user flows and architecture components. Read references/contract-v1.md and references/handoff-v1.md, then return FM-* records in a state patch.
Procedure
- List critical flows and their dependency chain. Cover applicable classes: timeout, retry storm, duplicate delivery, partial failure, dependency outage, overload, corruption, region failure, and human/operational error.
- For each meaningful failure, record trigger, blast radius, detection signal and threshold, mitigation, degraded mode, recovery steps, data impact, owner, and test/chaos scenario.
- Treat retries as a coupled policy: bounded retry budget, exponential backoff, jitter, idempotency key, and classification of retryable errors. Do not recommend blind retries.
- Link every critical flow to at least one
FM-*. Tie recovery targets toNFR-*, RPO, or RTO evidence. A failure without detection is incomplete. - Record residual risks and safe next handoff. Never mark design complete; validation owns that gate.
Acceptance checks
Run bundled scripts/validate_failure_modes.py <failure-model.json>, contract validation, repository validation, and skill-creator quick validation. Accept only when every critical flow has an FM record and every record has detection, recovery, data impact, and an executable test scenario.