autonomous-agent-patterns

Fail

Audited by Gen Agent Trust Hub on Apr 11, 2026

Risk Level: HIGHCOMMAND_EXECUTIONEXTERNAL_DOWNLOADSREMOTE_CODE_EXECUTIONDATA_EXFILTRATIONPROMPT_INJECTION
Full Analysis
  • [COMMAND_EXECUTION]: Found in SandboxedExecution.execute_sandboxed which uses subprocess.run(shell=True). Using shell=True with string-based commands is a dangerous practice that can lead to shell injection if inputs are not perfectly sanitized. Additionally, CheckpointManager._capture_workspace uses subprocess.getoutput to run git commands, which also invokes a shell process.- [EXTERNAL_DOWNLOADS]: The ContextManager.add_url method uses the requests.get() library to fetch content from arbitrary external URLs provided at runtime.- [REMOTE_CODE_EXECUTION]: The MCPAgent.create_tool pattern implements a "generate-write-execute" flow. It asks an LLM to generate Python code, writes that code directly to the filesystem in server.py, and then attempts to load/execute it via connect_server. This represents a critical remote code execution surface if the LLM's output is influenced by malicious instructions.- [DATA_EXFILTRATION]: The skill implements tools for reading local files (ReadFileTool, ReadFileTool) and making network requests (ContextManager.add_url). This combination of primitives allows for the reading of sensitive local data and transmitting it to external servers.- [PROMPT_INJECTION]: The skill is vulnerable to indirect prompt injection. Ingestion points: ContextManager.add_url fetches external web content and VisualAgent.describe_page processes browser screenshots. Boundary markers: The format_for_prompt method uses simple markdown headers (e.g., ## URL:) but lacks robust delimiters or instructions to ignore embedded commands. Capability inventory: The skill has access to subprocess.run, filesystem writes, and network operations. Sanitization: There is no evidence of sanitization, filtering, or validation of the content fetched from external URLs before it is injected into the LLM prompt context.
Recommendations
  • AI detected serious security threats
Audit Metadata
Risk Level
HIGH
Analyzed
Apr 11, 2026, 06:18 PM
Security Audit — agent-trust-hub — autonomous-agent-patterns