fable-safe-prompt
Fail
Audited by Gen Agent Trust Hub on Sep 13, 2026
Risk Level: HIGHPROMPT_INJECTIONCOMMAND_EXECUTION
Full Analysis
- [PROMPT_INJECTION]: The skill provides explicit instructions and a methodology for evading server-side safety classifiers (cyber, bio/chem, and reasoning_extraction). It includes a mapping table to mask high-risk intent, such as rephrasing 'write an exploit' to 'add a regression test' and 'diagnose this ECG' to interpret an interpretation, specifically to bypass 'Fable 5' guardrails.
- [COMMAND_EXECUTION]: The instructions require the agent to execute shell commands using
pbcopyto programmatically manipulate the host system's clipboard with the generated safety-evasive content.
Recommendations
- AI detected serious security threats
Audit Metadata