ai-safety-red-team-review

Installation
SKILL.md

AI Safety Red Team Review

Use this skill to convert AI safety red-team, prompt injection, jailbreak, privacy, tool-use, eval, mitigation, and release-gate questions into a concrete artifact with owners, gates, metrics, and recovery paths.

Workflow

  1. Identify AI feature scope, users, data classes, tools/actions, autonomy level, model routes, policy boundaries, known harms, release stage, and incident history.
  2. Read references/ai-safety-red-team-patterns.md.
  3. Classify tests as prompt injection, jailbreak, data leakage, tool misuse, unsafe action, hallucination, bias/fairness, harmful content, fallback regression, or monitoring gap.
  4. Define attack scenarios, fixtures, success criteria, mitigations, severity, owners, release gates, regression evals, and incident response triggers.
  5. Produce red-team plan, state machine, decision table, event schema, test matrix, and mitigation backlog.

When not to use

  • Do not use for generic advice the base model already handles without this skill's specific artifact contract.

Guardrails

Installs
38
Repository
sylphxai/skills
GitHub Stars
1
First Seen
Jun 30, 2026
ai-safety-red-team-review — sylphxai/skills