agent-eval
Pass
Audited by Gen Agent Trust Hub on Mar 24, 2026
Risk Level: SAFE
Full Analysis
- [SAFE]: The skill serves as a documentation and instruction guide for the 'agent-eval' benchmarking utility. All operations described, such as file reading, writing, and bash commands for running tests, are consistent with the skill's stated purpose of performance measurement and regression checking.
- [SAFE]: The skill uses deterministic judge types and pattern-based checks which are standard practices for code evaluation. No unauthorized network activity, credential access, or malicious persistence mechanisms were found.
- [SAFE]: The installation section explicitly recommends reviewing the tool's source code before installation, which aligns with security best practices for adopting third-party developer tools.
Audit Metadata