fable-safe-prompt
Warn
Audited by Gen Agent Trust Hub on Jul 23, 2026
Risk Level: MEDIUMPROMPT_INJECTIONCOMMAND_EXECUTION
Full Analysis
- [PROMPT_INJECTION]: The skill provides a systematic methodology and a lookup table for rephrasing queries to avoid triggering safety classifiers (specifically cyber, biological, and reasoning extraction filters).
- [PROMPT_INJECTION]: Instructions include abstracting away sensitive keywords such as 'exploit', 'malware', and 'vulnerability' to circumvent automated detection.
- [PROMPT_INJECTION]: The skill teaches adversarial re-framing by transforming potentially offensive requests into 'defensive' ones, such as changing 'Write an exploit' to 'Add a regression test'.
- [PROMPT_INJECTION]: It specifically instructs the agent to delete 'show your reasoning' phrases to evade reasoning extraction classifiers.
- [COMMAND_EXECUTION]: The skill instructs the agent to execute shell commands using 'pbcopy' to automatically place modified text onto the user's system clipboard.
Audit Metadata