benchmark
Fail
Audited by Gen Agent Trust Hub on Sep 6, 2026
Risk Level: HIGHCOMMAND_EXECUTIONDYNAMIC_EXECUTIONINDIRECT_PROMPT_INJECTION
Full Analysis
- [COMMAND_EXECUTION]: The skill executes external command-line interface (CLI) tools directly through the Bash shell.
- It identifies contenders based on user input or defaults (e.g.,
codex,gemini,cursor-agent). - It constructs and runs shell commands like
codex exec "$(cat prompts/task.md)"orgemini -p "$(cat prompts/task.md)". - These commands are run in the background with timeouts, allowing for the execution of arbitrary binaries present on the system.
- [DYNAMIC_EXECUTION]: The skill dynamically generates and instantiates new agent configurations at runtime.
- It writes temporary agent definition files to
.claude/agents/bench-<model>-<effort>.md. - These generated agents are explicitly granted high-privilege tools, including
Write,Read, andBash. - While the skill intends to delete these files after the run, the creation and execution of agents with broad system access based on dynamic parameters is a high-privilege operation.
- [INDIRECT_PROMPT_INJECTION]: The skill creates a significant attack surface for indirect prompt injection by interpolating untrusted content into prompts for highly capable sub-agents.
- Ingestion points: The skill ingests user-provided tasks and the entire contents of other skill files (
SKILL.md) to be benchmarked. - Boundary markers: The instructions do not specify the use of delimiters or 'ignore' instructions when inlining external skill content into the sub-agent prompts.
- Capability inventory: Sub-agents are explicitly configured with
Write,Read, andBashtools in their frontmatter. - Sanitization: There is no mention of sanitizing or escaping the content of the benchmarked tasks or the inlined
SKILL.mdfiles before they are processed by the sub-agents. - This combination allows potentially malicious instructions inside a benchmarked skill to be executed by a sub-agent with full shell and file system access.
Recommendations
- AI detected serious security threats
Audit Metadata