skills/sehoon787/my-claude/agent-eval/Gen Agent Trust Hub

agent-eval

Pass

Audited by Gen Agent Trust Hub on Mar 24, 2026

Risk Level: SAFE
Full Analysis
  • [SAFE]: No malicious instructions, obfuscation, or unauthorized data access patterns were detected in the skill content.\n- [COMMAND_EXECUTION]: The skill demonstrates the execution of shell commands for benchmarking purposes, including running the 'agent-eval' CLI and executing 'judge' commands such as pytest. This is consistent with the skill's stated utility.\n- [EXTERNAL_DOWNLOADS]: The documentation points to an external GitHub repository (github.com/joaquinhuigomez/agent-eval) for the tool's source code. This is an informative reference for installation and review.\n- [PROMPT_INJECTION]: The skill defines a system for processing YAML task files that contain shell commands. While this creates an indirect prompt injection surface, it is a core feature of the benchmarking tool. \n
  • Ingestion points: YAML files in the tasks/ directory. \n
  • Boundary markers: None. \n
  • Capability inventory: Bash tool for executing commands defined in YAML. \n
  • Sanitization: None; commands are executed as defined in the task files.
Audit Metadata
Risk Level
SAFE
Analyzed
Mar 24, 2026, 07:43 AM
Security Audit — agent-trust-hub — agent-eval