troubleshoot-smartctl-disk-monitoring

Installation
SKILL.md

Troubleshoot smartctl (S.M.A.R.T. Disk Health)

When to use this skill

  • Gradual media degradation: Bad sectors accumulate over weeks/months.
  • Sudden mechanical failure: (HDD); Head crash, spindle seizure, or actuator
  • SSD wear-out cliff: SSDs degrade gradually but fail suddenly. Once the spare
  • Interface/transport failure: The drive itself is healthy but the physical
  • Thermal damage: Sustained high temperature degrades components. HDDs: bearing
  • Controller/electronics failure: Firmware bug, power surge damage, or capacitor
  • Any time the user reports a smartctl (S.M.A.R.T. Disk Health) service behaving outside its expected envelope (elevated errors, latency, saturation, resource exhaustion, or unexpected restarts).
  • An on-call engineer is paging on a Netdata alert tied to a smartctl (S.M.A.R.T. Disk Health) instance and wants a structured triage path.

Key facts

Installs
11
Repository
netdata/skills
GitHub Stars
2
First Seen
Jun 1, 2026
troubleshoot-smartctl-disk-monitoring — netdata/skills