agent-eval

Pass

Audited by Gen Agent Trust Hub on Sep 1, 2026

Risk Level: SAFEEXTERNAL_DOWNLOADSCOMMAND_EXECUTIONINDIRECT_PROMPT_INJECTION
Full Analysis
  • [EXTERNAL_DOWNLOADS]: The documentation references an external software repository (github.com/joaquinhuigomez/agent-eval) for installing the agent-eval CLI tool.
  • [COMMAND_EXECUTION]: The tool is designed to execute shell commands (e.g., pytest, npm run build) defined in YAML task files to verify agent performance.
  • [INDIRECT_PROMPT_INJECTION]: The skill processes untrusted data in the form of YAML task definitions which include natural language prompts and shell commands.
  • Ingestion points: YAML files located in the tasks/ directory.
  • Boundary markers: No specific delimiters or warnings to ignore instructions within the task definitions are mentioned.
  • Capability inventory: The skill uses the Bash tool to execute judge commands defined in the tasks.
  • Sanitization: No evidence of sanitization or validation of the commands or prompts within the YAML files.
Audit Metadata
Risk Level
SAFE
Analyzed
Sep 1, 2026, 02:38 AM
Security Audit — agent-trust-hub — agent-eval