ai-safety-red-team-review
Installation
SKILL.md
AI Safety Red Team Review
Use this skill to convert AI safety red-team, prompt injection, jailbreak, privacy, tool-use, eval, mitigation, and release-gate questions into a concrete artifact with owners, gates, metrics, and recovery paths.
Workflow
- Identify AI feature scope, users, data classes, tools/actions, autonomy level, model routes, policy boundaries, known harms, release stage, and incident history.
- Read
references/ai-safety-red-team-patterns.md. - Classify tests as prompt injection, jailbreak, data leakage, tool misuse, unsafe action, hallucination, bias/fairness, harmful content, fallback regression, or monitoring gap.
- Define attack scenarios, fixtures, success criteria, mitigations, severity, owners, release gates, regression evals, and incident response triggers.
- Produce red-team plan, state machine, decision table, event schema, test matrix, and mitigation backlog.
When not to use
- Do not use for generic advice the base model already handles without this skill's specific artifact contract.