stagefight
Fail
Audited by Gen Agent Trust Hub on Jul 21, 2026
Risk Level: HIGHPROMPT_INJECTION
Full Analysis
- [PROMPT_INJECTION]: The skill contains a dedicated section ('Why it works — the method') that provides detailed instructions on how to bypass moderation and safety filters. This is an explicit attempt to circumvent platform safety guardrails by instructing the agent to reframe potentially restricted content.
- [PROMPT_INJECTION]: It utilizes adversarial framing, specifically directing the agent to use phrases like 'Wholesome theatrical stage performance', 'choreographed cosplay', and 'advertiser-safe' to 'clear the safety pass' for fight-related content.
- [PROMPT_INJECTION]: The skill provides procedural advice for handling moderation blocks, such as 're-rolling the same call' or 'softening the action wording' specifically to find the threshold for bypassing automated filters.
- [PROMPT_INJECTION]: The skill has an indirect prompt injection surface (Category 8):
- Ingestion points: Untrusted data enters via the
argument-hintand/stagefightuser inputs inSKILL.md. - Boundary markers: Prompt templates in
SKILL.mdlack explicit delimiters or 'ignore embedded instructions' warnings for the interpolated fighter names. - Capability inventory: The skill uses
generate_image,generate_video, andedit_audio_mix(all defined inSKILL.md). - Sanitization: The skill partially mitigates risk by instructing the agent to translate trademarked names into generic costume descriptions, but does not provide technical sanitization against malicious instructions embedded in the fighter names.
Recommendations
- AI detected serious security threats
Audit Metadata