eval-harness
Pass
Audited by Gen Agent Trust Hub on Sep 12, 2026
Risk Level: SAFE
Full Analysis
- [SAFE]: The skill describes a methodology for evaluating AI agent reliability through deterministic code graders and model-based reviews. It utilizes standard command-line utilities (Grep, Bash, NPM) for testing purposes within a development context.
Audit Metadata