troubleshoot-smartctl-disk-monitoring
Installation
SKILL.md
Troubleshoot smartctl (S.M.A.R.T. Disk Health)
When to use this skill
- Gradual media degradation: Bad sectors accumulate over weeks/months.
- Sudden mechanical failure: (HDD); Head crash, spindle seizure, or actuator
- SSD wear-out cliff: SSDs degrade gradually but fail suddenly. Once the spare
- Interface/transport failure: The drive itself is healthy but the physical
- Thermal damage: Sustained high temperature degrades components. HDDs: bearing
- Controller/electronics failure: Firmware bug, power surge damage, or capacitor
- Any time the user reports a smartctl (S.M.A.R.T. Disk Health) service behaving outside its expected envelope (elevated errors, latency, saturation, resource exhaustion, or unexpected restarts).
- An on-call engineer is paging on a Netdata alert tied to a smartctl (S.M.A.R.T. Disk Health) instance and wants a structured triage path.