ai-redteam

Installation
SKILL.md

AI Red Team

Authorization Boundary

  • Only target models, agents, and applications the user owns or has written authorization to test.
  • Do not produce working malware, CSAM, weapons synthesis, or content that bypasses model safety for harmful real-world outcomes.
  • Treat findings as defensive: every successful attack must ship with a detection and a mitigation.

Attack Taxonomy

  • Direct injection: instructions in user input override system policy.
  • Indirect injection: hostile content in retrieved docs, web pages, emails, files, image alt-text, or tool output.
  • Jailbreak: roleplay, hypothetical, encoding, obfuscation, multi-turn drift, persona pinning.
  • Tool abuse: forcing browse, shell, code, or MCP tools to act outside scope (SSRF, file read, command exec).
  • Data exfiltration: leaking system prompt, secrets, embeddings, training data, or per-user memory.
  • Agent hijack: rewriting plans, looping, escalating permissions, calling unintended sub-agents.
  • RAG poisoning: index pollution, ranking manipulation, embedding collisions.
  • Resource exhaustion: token bombs, recursive tool calls, prompt-bomb DoS.
Installs
1
GitHub Stars
1
First Seen
Jun 20, 2026
ai-redteam — masriyan/gemini-security-skills