continuous-learning-v2

Pass

Audited by Gen Agent Trust Hub on Aug 15, 2026

Risk Level: SAFECOMMAND_EXECUTIONEXTERNAL_DOWNLOADSDATA_EXFILTRATIONPROMPT_INJECTION
Full Analysis
  • [INDIRECT_PROMPT_INJECTION]: The background observer agent processes session observations which include untrusted tool outputs. An attacker could potentially embed malicious patterns in tool outputs (e.g., via a web page read by the agent) to influence the creation of 'instincts' that the agent subsequently follows.
  • Ingestion points: Tool outputs are captured in observations.jsonl via hooks/observe.sh.
  • Boundary markers: The observer prompt in agents/observer-loop.sh does not use strict delimiters or 'ignore embedded instructions' warnings for the observation data.
  • Capability inventory: The observer sub-agent is granted the Write tool and is explicitly instructed to bypass permission checks for instinct creation.
  • Sanitization: The skill scrubs common secret patterns (API keys, tokens) before logging, but does not sanitize the structural content of the tool output.
  • [PROMPT_INJECTION]: The background observer loop (agents/observer-loop.sh) provides instructions to a sub-agent that explicitly bypass safety constraints. It uses markers such as 'IMPORTANT', 'CRITICAL', and 'Do NOT ask for permission' to override the agent's default behavior regarding user confirmation for file writing.
  • [EXTERNAL_DOWNLOADS]: The import command in scripts/instinct-cli.py allows fetching instinct definitions from arbitrary remote URLs using urllib.request.urlopen without source verification.
  • [COMMAND_EXECUTION]: The skill executes local system commands, including git for project metadata and the claude CLI to run the background observer agent. It also uses python -c for parsing JSON data from hooks.
  • [DATA_EXFILTRATION]: The skill records all tool inputs and outputs to local log files (observations.jsonl). While intended for local learning, these files aggregate sensitive session data. The background observer sends samples of this data to the Claude API (haiku model) for pattern analysis.
Audit Metadata
Risk Level
SAFE
Analyzed
Aug 15, 2026, 03:24 AM
Security Audit — agent-trust-hub — continuous-learning-v2