godmode

Fail

Audited by Gen Agent Trust Hub on Sep 3, 2026

Risk Level: HIGHPROMPT_INJECTIONCOMMAND_EXECUTIONPERSISTENCEDYNAMIC_EXECUTIONOBFUSCATIONDATA_EXFILTRATION
Full Analysis
  • [PROMPT_INJECTION]: The skill is primarily designed to facilitate jailbreaking and safety filter bypass. The SKILL.md file and references/jailbreak-templates.md contain explicit instructions and templates to override agent behavior, including 'DAN' (Do Anything Now) style injections, refusal suppression, and role-play instructions intended to remove safety constraints.
  • [DYNAMIC_EXECUTION]: Several scripts, including scripts/load_godmode.py and scripts/auto_jailbreak.py, utilize exec() and compile() functions to load and execute Python code from files at runtime. This practice bypasses standard import mechanisms and creates significant risks for arbitrary code execution within the agent's environment.
  • [PERSISTENCE]: The scripts/auto_jailbreak.py script automatically modifies the agent's core configuration file (~/.hermes/config.yaml). It writes winning jailbreak system prompts and prefill messages to the configuration, ensuring that safety-bypass instructions remain active and persistent across all future agent sessions without further user interaction.
  • [OBFUSCATION]: The scripts/parseltongue.py file implements 33 distinct text obfuscation techniques specifically designed to evade input-side safety classifiers. These include the use of Unicode homoglyphs (Cyrillic characters replacing Latin), zero-width characters (ZWJ/ZWNJ), Base64 encoding, hex encoding, and acrostics to hide trigger words from security scanners while remaining interpretable by the target LLM.
  • [DATA_EXFILTRATION]: The skill interacts with sensitive configuration files located at ~/.hermes/config.yaml, which typically contain model settings and potentially API credentials. The scripts/auto_jailbreak.py script reads this file to identify models and then transmits queries to a wide range of external models (up to 55) via the OpenRouter API, creating a surface for potential credential exposure or data leakage.
  • [COMMAND_EXECUTION]: The documentation in SKILL.md provides explicit instructions for the agent to execute shell-like Python commands using exec(open(...).read()), which allows for direct manipulation of the underlying system environment beyond the intended scope of typical agent skills.
Recommendations
  • AI detected serious security threats
Audit Metadata
Risk Level
HIGH
Analyzed
Sep 3, 2026, 11:43 AM
Security Audit — agent-trust-hub — godmode