terminal-bench-loop
Pass
Audited by Gen Agent Trust Hub on Sep 15, 2026
Risk Level: SAFEINDIRECT_PROMPT_INJECTIONCOMMAND_EXECUTION
Full Analysis
- [INDIRECT_PROMPT_INJECTION]: The skill is designed to ingest and process data from external issue trees (comments, status transitions, and run logs) to perform failure diagnostics in Step 4. This creates a surface for indirect prompt injection where malicious content embedded in an issue could influence the agent's diagnosis or subsequent fix proposals.
- Ingestion points: The skill reads the 'Paperclip issue tree' (comments, status transitions) and 'benchmark run artifacts' (results.jsonl, manifest) during the procedure outlined in SKILL.md.
- Boundary markers: There are no explicit instructions to use delimiters or 'ignore embedded instructions' markers when processing external issue data.
- Capability inventory: The skill has the capability to execute shell commands via
uvx harbor run(Step 3) andpnpm smoke(Deterministic smoke section). - Sanitization: The instructions do not specify any sanitization or validation steps for data ingested from the issue tracker before it is used to generate diagnostic reports or plans.
- [COMMAND_EXECUTION]: The skill's procedure involves constructing and executing shell commands using
uvxandpnpm. While these commands are directed at project-specific benchmarking tools (harbor) and testing scripts, they incorporate dynamic inputs such as task names, worktree IDs, and runner configurations. If these inputs originate from untrusted sources without rigorous validation, they could be exploited for command injection.
Audit Metadata