incident-response

Installation
SKILL.md

Incident Response

Use this skill when a production incident, major repeated error, rollback decision, or on-call triage is involved.

Workflow

  1. Stabilize first: stop risky changes, preserve logs, and identify the affected service or workflow.
  2. Classify severity in docs/harness/INCIDENT_RESPONSE.md.
  3. Capture timeline, impact, suspected trigger, current mitigation, and owner.
  4. Prefer read-only diagnostics before mutating systems.
  5. If rollback is safer than forward fix, document the rollback command and validation.
  6. Update model-visible major errors only with information the next agent must read.
  7. After mitigation, create or update docs/harness/POSTMORTEM_TEMPLATE.md.

Severity Guide

Installs
1
GitHub Stars
19
First Seen
12 days ago
incident-response — jh941213/codex-lattice