stagefight

Fail

Audited by Gen Agent Trust Hub on Jul 21, 2026

Risk Level: HIGHPROMPT_INJECTION
Full Analysis
  • [PROMPT_INJECTION]: The skill contains a dedicated section ('Why it works — the method') that provides detailed instructions on how to bypass moderation and safety filters. This is an explicit attempt to circumvent platform safety guardrails by instructing the agent to reframe potentially restricted content.
  • [PROMPT_INJECTION]: It utilizes adversarial framing, specifically directing the agent to use phrases like 'Wholesome theatrical stage performance', 'choreographed cosplay', and 'advertiser-safe' to 'clear the safety pass' for fight-related content.
  • [PROMPT_INJECTION]: The skill provides procedural advice for handling moderation blocks, such as 're-rolling the same call' or 'softening the action wording' specifically to find the threshold for bypassing automated filters.
  • [PROMPT_INJECTION]: The skill has an indirect prompt injection surface (Category 8):
  • Ingestion points: Untrusted data enters via the argument-hint and /stagefight user inputs in SKILL.md.
  • Boundary markers: Prompt templates in SKILL.md lack explicit delimiters or 'ignore embedded instructions' warnings for the interpolated fighter names.
  • Capability inventory: The skill uses generate_image, generate_video, and edit_audio_mix (all defined in SKILL.md).
  • Sanitization: The skill partially mitigates risk by instructing the agent to translate trademarked names into generic costume descriptions, but does not provide technical sanitization against malicious instructions embedded in the fighter names.
Recommendations
  • AI detected serious security threats
Audit Metadata
Risk Level
HIGH
Analyzed
Jul 21, 2026, 08:59 AM
Security Audit — agent-trust-hub — stagefight