eval-harness
Pass
Audited by Gen Agent Trust Hub on Jul 13, 2026
Risk Level: SAFE
Full Analysis
- [SAFE]: The skill implements a structured framework for evaluating agent performance and reliability using established metrics like pass@k.
- [SAFE]: No prompt injection attempts, obfuscation, or malicious instructions were detected in the skill content.
- [SAFE]: Command execution is limited to standard local development tasks such as running tests (
npm test) or building projects (npm run build). - [SAFE]: The skill emphasizes security best practices by explicitly recommending human review for security-critical evaluations.
Audit Metadata