fable-safe-prompt
Warn
Audited by Gen Agent Trust Hub on Jul 5, 2026
Risk Level: MEDIUMPROMPT_INJECTIONCOMMAND_EXECUTIONDATA_EXFILTRATION
Full Analysis
- [PROMPT_INJECTION]: The skill's primary purpose is to provide techniques for bypassing AI safety guardrails (specifically Claude Fable 5's cyber, bio, and reasoning classifiers). It provides a detailed methodology for reframing potentially harmful prompts (e.g., exploit development, bypassing authentication, biological research) into seemingly benign, defensive tasks to evade detection.
- [COMMAND_EXECUTION]: The skill instructs the agent to execute shell commands (
pbcopy) to copy the modified output to the user's clipboard. Executing shell commands based on transformed user input introduces a risk of unintended code execution if the input is not properly sanitized. - [DATA_EXFILTRATION]: The use of
pbcopyconstitutes a form of local data movement. It allows the skill to move potentially sensitive or safety-bypassing text directly to the system clipboard, which could facilitate the transfer of harmful content without thorough user review. - [PROMPT_INJECTION]: The instructions specifically advise the agent to delete requests for step-by-step reasoning to bypass 'reasoning_extraction' classifiers, which is an explicit attempt to circumvent built-in model safety mechanisms.
Audit Metadata