distributed-failure-analyzer
When to Use
You are diagnosing or preventing a failure in a distributed system and the root cause is not immediately obvious. The symptom could be: a timeout that might mean the remote node is dead, or might mean the network is congested, or might mean the node is alive but paused. A write that appeared to succeed but whose data is now missing. A leader that is still writing after another leader was elected. A lock that was held correctly but still allowed two writers simultaneously.
This skill imposes a diagnostic framework: every unexplained distributed system failure traces to one of three root fault categories — network faults, clock unreliability, or process pauses — and each category has a bounded set of mechanisms and well-understood mitigations. The skill maps symptoms to categories, categories to mechanisms, and mechanisms to concrete fixes.
Use it reactively (incident post-mortem, production debugging) or proactively (design review, codebase audit for timing anti-patterns).
Cross-references:
replication-failure-analyzer— for failures specific to replication lag, failover, and quorum behaviorconsistency-model-selector— for selecting isolation and consistency guarantees that prevent a class of failures at the application layer
Context and Input Gathering
Before analysis, collect the following. Ask the user for any that are missing.