gke-ai-troubleshooting-handle-disruption-gpu-tpu
Installation
SKILL.md
Handle Disruption on GPUs and TPUs Troubleshooting
🔍 Diagnostic Workflow
Step 0: Context Acquisition
- Mandatory: When a user asks to debug or investigate an actual workload
disruption, node crash, or unexpected restart without providing complete
cluster details, you MUST immediately halt and request all missing mandatory
parameters (
project_id,location,cluster_name,timestamp) BEFORE delivering theories or general diagnostic commands. Only skip context acquisition if the user explicitly requests a generic reusable runbook or provides a complete static telemetry/log dump for offline analysis. - Optional:
node_name,workload_name,workload_namespace,nodepool_name.