databricks-cluster-forensics
Installation
SKILL.md
Databricks Cluster Forensics
The operational SRE spine of the pack — what a Databricks engineer reaches for at 2 AM when the compute layer is broken or unexplained. It correlates a cluster's live event stream across API surfaces to name the failure with its actual error code and its version-specific mitigation, not "network problem, try again".
Overview
Six real compute-layer failures live in this skill; each has a deterministic detector and an on-demand reference: