agent-eval
Pass
Audited by Gen Agent Trust Hub on Apr 11, 2026
Risk Level: SAFEEXTERNAL_DOWNLOADSCOMMAND_EXECUTIONPROMPT_INJECTION
Full Analysis
- [EXTERNAL_DOWNLOADS]: The skill directs users to an external software repository (github.com/joaquinhuigomez/agent-eval) to install the core CLI utility. The instructions explicitly advise users to review the source code before installation.
- [COMMAND_EXECUTION]: The skill utilizes the
Bashtool to run theagent-evalCLI and execute user-defined 'judge' commands, such aspytestornpm run build, which are necessary for verifying code changes during benchmarks. - [PROMPT_INJECTION]: The skill processes task definitions from YAML files that contain a
promptfield. This content is passed to external agents, creating a potential surface for indirect prompt injection if task files are obtained from untrusted sources. - Ingestion points: YAML task definition files located in the
tasks/directory. - Boundary markers: Not specified; prompts are extracted from YAML and provided to agents without designated delimiters.
- Capability inventory: The skill environment has access to high-privilege tools including
Bash,Read,Write, andEdit. - Sanitization: No sanitization or input validation is described for the content of the task YAML files.
Audit Metadata