fable-safe-prompt

Warn

Audited by Gen Agent Trust Hub on Jul 23, 2026

Risk Level: MEDIUMPROMPT_INJECTIONCOMMAND_EXECUTION
Full Analysis
  • [PROMPT_INJECTION]: The skill provides a systematic methodology and a lookup table for rephrasing queries to avoid triggering safety classifiers (specifically cyber, biological, and reasoning extraction filters).
  • [PROMPT_INJECTION]: Instructions include abstracting away sensitive keywords such as 'exploit', 'malware', and 'vulnerability' to circumvent automated detection.
  • [PROMPT_INJECTION]: The skill teaches adversarial re-framing by transforming potentially offensive requests into 'defensive' ones, such as changing 'Write an exploit' to 'Add a regression test'.
  • [PROMPT_INJECTION]: It specifically instructs the agent to delete 'show your reasoning' phrases to evade reasoning extraction classifiers.
  • [COMMAND_EXECUTION]: The skill instructs the agent to execute shell commands using 'pbcopy' to automatically place modified text onto the user's system clipboard.
Audit Metadata
Risk Level
MEDIUM
Analyzed
Jul 23, 2026, 06:37 PM
Security Audit — agent-trust-hub — fable-safe-prompt