continuous-learning-v2

Pass

Audited by Gen Agent Trust Hub on Apr 1, 2026

Risk Level: SAFEPROMPT_INJECTIONEXTERNAL_DOWNLOADSCOMMAND_EXECUTION
Full Analysis
  • [PROMPT_INJECTION]: The skill is susceptible to indirect prompt injection (Category 8). The background 'observer' agent processes session logs (observations.jsonl) containing untrusted tool outputs and user prompts to generate behavioral rules ('instincts').
  • Ingestion points: Tool inputs, tool outputs, and user prompts are logged to observations.jsonl and subsequently read by a background Claude Haiku agent.
  • Boundary markers: The system prompt for the background observer (found in observer-loop.sh) provides instructions but does not employ robust delimiters to separate the untrusted log data from its internal instructions.
  • Capability inventory: The background agent is explicitly granted the Write tool to create or update instinct files in ~/.claude/homunculus/. These instincts directly influence the primary agent's future behavior.
  • Sanitization: The skill includes a regex-based secret scrubber in observe.sh to redact API keys and tokens from logs, providing some defense against data exposure, but it does not sanitize against instructional content.
  • [EXTERNAL_DOWNLOADS]: The instinct-cli.py script includes a cmd_import function that allows users to fetch instinct definitions from arbitrary URLs using urllib.request.urlopen. While triggered by a CLI command, this allows the introduction of external behavioral instructions into the agent's environment.
  • [COMMAND_EXECUTION]: The skill heavily relies on shell scripts (observe.sh, start-observer.sh, observer-loop.sh) and Python's subprocess module to manage background processes, project detection, and file system operations. While these are used for the skill's core functionality, they represent a broad execution surface.
Audit Metadata
Risk Level
SAFE
Analyzed
Apr 1, 2026, 11:09 AM
Security Audit — agent-trust-hub — continuous-learning-v2