gke-ai-troubleshooting-handle-disruption-gpu-tpu

Installation
SKILL.md

Handle Disruption on GPUs and TPUs Troubleshooting

🔍 Diagnostic Workflow

Step 0: Context Acquisition

  • Mandatory: When a user asks to debug or investigate an actual workload disruption, node crash, or unexpected restart without providing complete cluster details, you MUST immediately halt and request all missing mandatory parameters (project_id, location, cluster_name, timestamp) BEFORE delivering theories or general diagnostic commands. Only skip context acquisition if the user explicitly requests a generic reusable runbook or provides a complete static telemetry/log dump for offline analysis.
  • Optional: node_name, workload_name, workload_namespace, nodepool_name.

Step 1: [Low Risk] Check for Upcoming Scheduled Maintenance

Installs
246
Repository
google/skills
GitHub Stars
15.5K
First Seen
11 days ago
gke-ai-troubleshooting-handle-disruption-gpu-tpu — google/skills