code-runner
Pass
Audited by Gen Agent Trust Hub on Aug 26, 2026
Risk Level: SAFECOMMAND_EXECUTIONREMOTE_CODE_EXECUTIONPROMPT_INJECTION
Full Analysis
- [COMMAND_EXECUTION]: The skill provides a
run_commandtool that allows model-controlled shell command execution within a disposable git worktree. Evidence found insrc/code_runner/tools.pyandtool_use.pywheresubprocess.runis invoked withbash -lc. Mitigations include a denylist of protected files (e.g.,.git,.env) and regex-based blocking of destructive command patterns likegit reset --hardandrm -rf. - [REMOTE_CODE_EXECUTION]: The skill evaluates dynamically generated or provided code. The
apply.pyscript uses the Pythoncompile()function to validate the syntax of LLM-generated code before committing it to disk. Additionally, the 'Definition of Done' (DoD) mechanism insrc/code_runner/dod.pyexecutes arbitrary commands to verify task success. - [PROMPT_INJECTION]: The skill's architecture is vulnerable to indirect prompt injection via the processing of untrusted data. Ingestion points include the
read_fileandsearch_codetools which feed file content and search results back into the agent context. Although system prompts provide operational rules, the skill lacks rigid structural boundaries or sanitization logic to prevent the model from obeying instructions hidden within the source code or test outputs it processes.
Audit Metadata