skills/smithery.ai/groq-incident-runbook

groq-incident-runbook

Installation
SKILL.md

Groq Incident Runbook

Overview

Rapid incident response procedures for Groq-related outages.

Prerequisites

  • Access to Groq dashboard and status page
  • kubectl access to production cluster
  • Prometheus/Grafana access
  • Communication channels (Slack, PagerDuty)

Severity Levels

Level Definition Response Time Examples
P1 Complete outage < 15 min Groq API unreachable
P2 Degraded service < 1 hour High latency, partial failures
P3 Minor impact < 4 hours Webhook delays, non-critical errors
P4 No user impact Next business day Monitoring gaps
Installs
1
First Seen
Mar 21, 2026
groq-incident-runbook from smithery.ai