eval-harness

Pass

Audited by Gen Agent Trust Hub on Sep 12, 2026

Risk Level: SAFE
Full Analysis
  • [SAFE]: The skill describes a methodology for evaluating AI agent reliability through deterministic code graders and model-based reviews. It utilizes standard command-line utilities (Grep, Bash, NPM) for testing purposes within a development context.
Audit Metadata
Risk Level
SAFE
Analyzed
Sep 12, 2026, 03:41 PM
Security Audit — agent-trust-hub — eval-harness