guidewire-observability-and-incident-response
Installation
SKILL.md
Guidewire Observability and Incident Response
Overview
Run a production Guidewire Cloud API integration with the dashboards, alerts, and runbooks an on-call engineer can act on at 3am. This skill consolidates the operational layer: what to measure, what to alert on, how to triage the top five incident classes, and how to close the loop with a post-incident review that prevents recurrence rather than performing root-cause theater.
Five operational failures this skill prevents:
- Vanity dashboards — graphs of total request count, no per-endpoint p99, no SLI burn rate; on-call sees green during a real outage because the right thing was never measured.
- Alert fatigue — every transient
5xxpages someone; in three weeks the team mutes the channel; a real incident two weeks later goes unnoticed for an hour. - No triage tree — on-call wakes up to "401 spike", does not know whether to rotate a secret, restart the integration, or call GCC; loses 20 minutes Googling.
- Skipped post-incident review — the same root cause produces three incidents in a quarter because no one wrote down the action items from the first one.
- Common-errors table living in nine Slack threads — operators cannot find the recovery for an error they have seen before, ask the same question in #ops, the answer takes 45 minutes.