code-runner

Pass

Audited by Gen Agent Trust Hub on Aug 26, 2026

Risk Level: SAFECOMMAND_EXECUTIONREMOTE_CODE_EXECUTIONPROMPT_INJECTION
Full Analysis
  • [COMMAND_EXECUTION]: The skill provides a run_command tool that allows model-controlled shell command execution within a disposable git worktree. Evidence found in src/code_runner/tools.py and tool_use.py where subprocess.run is invoked with bash -lc. Mitigations include a denylist of protected files (e.g., .git, .env) and regex-based blocking of destructive command patterns like git reset --hard and rm -rf.
  • [REMOTE_CODE_EXECUTION]: The skill evaluates dynamically generated or provided code. The apply.py script uses the Python compile() function to validate the syntax of LLM-generated code before committing it to disk. Additionally, the 'Definition of Done' (DoD) mechanism in src/code_runner/dod.py executes arbitrary commands to verify task success.
  • [PROMPT_INJECTION]: The skill's architecture is vulnerable to indirect prompt injection via the processing of untrusted data. Ingestion points include the read_file and search_code tools which feed file content and search results back into the agent context. Although system prompts provide operational rules, the skill lacks rigid structural boundaries or sanitization logic to prevent the model from obeying instructions hidden within the source code or test outputs it processes.
Audit Metadata
Risk Level
SAFE
Analyzed
Aug 26, 2026, 06:01 PM
Security Audit — agent-trust-hub — code-runner