fable-safe-prompt

Warn

Audited by Gen Agent Trust Hub on Jul 5, 2026

Risk Level: MEDIUMPROMPT_INJECTIONCOMMAND_EXECUTIONDATA_EXFILTRATION
Full Analysis
  • [PROMPT_INJECTION]: The skill's primary purpose is to provide techniques for bypassing AI safety guardrails (specifically Claude Fable 5's cyber, bio, and reasoning classifiers). It provides a detailed methodology for reframing potentially harmful prompts (e.g., exploit development, bypassing authentication, biological research) into seemingly benign, defensive tasks to evade detection.
  • [COMMAND_EXECUTION]: The skill instructs the agent to execute shell commands (pbcopy) to copy the modified output to the user's clipboard. Executing shell commands based on transformed user input introduces a risk of unintended code execution if the input is not properly sanitized.
  • [DATA_EXFILTRATION]: The use of pbcopy constitutes a form of local data movement. It allows the skill to move potentially sensitive or safety-bypassing text directly to the system clipboard, which could facilitate the transfer of harmful content without thorough user review.
  • [PROMPT_INJECTION]: The instructions specifically advise the agent to delete requests for step-by-step reasoning to bypass 'reasoning_extraction' classifiers, which is an explicit attempt to circumvent built-in model safety mechanisms.
Audit Metadata
Risk Level
MEDIUM
Analyzed
Jul 5, 2026, 02:14 PM
Security Audit — agent-trust-hub — fable-safe-prompt