eval-harness

Pass

Audited by Gen Agent Trust Hub on Apr 1, 2026

Risk Level: SAFECOMMAND_EXECUTION
Full Analysis
  • [SAFE]: The skill provides structured guidance and templates for implementing evaluation-driven development, focusing on benchmarking agent reliability.
  • [COMMAND_EXECUTION]: Instructs the agent to execute standard build and test commands (e.g., npm test, npm run build) to determine the success of code-based tasks.
  • [SAFE]: The framework manages all evaluation data and logs locally within the project structure, avoiding any external data transfer or unsafe dependencies.
Audit Metadata
Risk Level
SAFE
Analyzed
Apr 1, 2026, 11:08 AM
Security Audit — agent-trust-hub — eval-harness