agent-eval
Pass
Audited by Gen Agent Trust Hub on Sep 1, 2026
Risk Level: SAFEEXTERNAL_DOWNLOADSCOMMAND_EXECUTIONINDIRECT_PROMPT_INJECTION
Full Analysis
- [EXTERNAL_DOWNLOADS]: The documentation references an external software repository (github.com/joaquinhuigomez/agent-eval) for installing the agent-eval CLI tool.
- [COMMAND_EXECUTION]: The tool is designed to execute shell commands (e.g., pytest, npm run build) defined in YAML task files to verify agent performance.
- [INDIRECT_PROMPT_INJECTION]: The skill processes untrusted data in the form of YAML task definitions which include natural language prompts and shell commands.
- Ingestion points: YAML files located in the tasks/ directory.
- Boundary markers: No specific delimiters or warnings to ignore instructions within the task definitions are mentioned.
- Capability inventory: The skill uses the Bash tool to execute judge commands defined in the tasks.
- Sanitization: No evidence of sanitization or validation of the commands or prompts within the YAML files.
Audit Metadata