eval-harness
Pass
Audited by Gen Agent Trust Hub on Apr 1, 2026
Risk Level: SAFECOMMAND_EXECUTION
Full Analysis
- [SAFE]: The skill provides structured guidance and templates for implementing evaluation-driven development, focusing on benchmarking agent reliability.
- [COMMAND_EXECUTION]: Instructs the agent to execute standard build and test commands (e.g., npm test, npm run build) to determine the success of code-based tasks.
- [SAFE]: The framework manages all evaluation data and logs locally within the project structure, avoiding any external data transfer or unsafe dependencies.
Audit Metadata