agent-eval

Pass

Audited by Gen Agent Trust Hub on Apr 11, 2026

Risk Level: SAFEEXTERNAL_DOWNLOADSCOMMAND_EXECUTIONPROMPT_INJECTION
Full Analysis
  • [EXTERNAL_DOWNLOADS]: The skill directs users to an external software repository (github.com/joaquinhuigomez/agent-eval) to install the core CLI utility. The instructions explicitly advise users to review the source code before installation.
  • [COMMAND_EXECUTION]: The skill utilizes the Bash tool to run the agent-eval CLI and execute user-defined 'judge' commands, such as pytest or npm run build, which are necessary for verifying code changes during benchmarks.
  • [PROMPT_INJECTION]: The skill processes task definitions from YAML files that contain a prompt field. This content is passed to external agents, creating a potential surface for indirect prompt injection if task files are obtained from untrusted sources.
  • Ingestion points: YAML task definition files located in the tasks/ directory.
  • Boundary markers: Not specified; prompts are extracted from YAML and provided to agents without designated delimiters.
  • Capability inventory: The skill environment has access to high-privilege tools including Bash, Read, Write, and Edit.
  • Sanitization: No sanitization or input validation is described for the content of the task YAML files.
Audit Metadata
Risk Level
SAFE
Analyzed
Apr 11, 2026, 03:42 AM
Security Audit — agent-trust-hub — agent-eval