skills/skills.volces.com/distributed-failure-analyzer

distributed-failure-analyzer

Installation
SKILL.md

When to Use

You are diagnosing or preventing a failure in a distributed system and the root cause is not immediately obvious. The symptom could be: a timeout that might mean the remote node is dead, or might mean the network is congested, or might mean the node is alive but paused. A write that appeared to succeed but whose data is now missing. A leader that is still writing after another leader was elected. A lock that was held correctly but still allowed two writers simultaneously.

This skill imposes a diagnostic framework: every unexplained distributed system failure traces to one of three root fault categories — network faults, clock unreliability, or process pauses — and each category has a bounded set of mechanisms and well-understood mitigations. The skill maps symptoms to categories, categories to mechanisms, and mechanisms to concrete fixes.

Use it reactively (incident post-mortem, production debugging) or proactively (design review, codebase audit for timing anti-patterns).

Cross-references:

  • replication-failure-analyzer — for failures specific to replication lag, failover, and quorum behavior
  • consistency-model-selector — for selecting isolation and consistency guarantees that prevent a class of failures at the application layer

Context and Input Gathering

Before analysis, collect the following. Ask the user for any that are missing.

Installs
2
First Seen
Apr 24, 2026
distributed-failure-analyzer from skills.volces.com