dead-letter
๐ฎ dead-letter โ Failed Task Triage & Capture
Capture a failed or blocked task before context clears. Categorizes failure modes into a 9-part taxonomy, distinguishes transient from permanent failures, verifies actual state before any retry, preserves partial work and orphaned resources, and generates either a bounded autonomous retry prompt or an escalation decision point. Recovery follows a strict ordered sequence โ classify โ decide โ precondition โ fix โ verify-against-baseline โ close โ tracked step-by-step in the record so a retry resumes exactly where recovery stopped. At close-out, deterministic failures can be converted into a repro test pack: the capture that documents a failure becomes the regression test that prevents its recurrence. The status sweep clusters open records by root cause โ repeated failures across different tasks surface as one systemic finding with one fix, not a storm of individual retries.
When to Use
- An agent returns "I can't proceed", "blocked on", "error:", or a partial result.
- After a long-running subagent batch or background command fails.
- Before reporting an unresolvable failure to the user or orchestrator.
- When an in-progress task must be abandoned and resumable later.
- When sweeping open dead-letter records (
dead-letter status) โ past-SLA or high-retry records need triage. - When a bug or blocker is worth handing to Dev/QA as a reproduction case (
repro pack). - On a scheduled sweep (
nightly failure triage) โ open records get clustered by root cause so systemic issues surface once, not per-task.