eval-harness

Pass

Audited by Gen Agent Trust Hub on Aug 15, 2026

Risk Level: SAFE
Full Analysis
  • [SAFE]: The skill provides instructional templates for defining and running performance evaluations (evals) for AI agents.
  • [SAFE]: Command execution examples are standard development utilities (grep, npm test, npm run build) used for deterministic grading of code changes.
  • [SAFE]: All storage and retrieval operations target local project directories (e.g., .claude/evals/), following established software engineering patterns.
  • [SAFE]: No evidence of prompt injection, obfuscation, or unauthorized data exfiltration was found in the provided documentation or metadata.
Audit Metadata
Risk Level
SAFE
Analyzed
Aug 15, 2026, 03:23 AM
Security Audit — agent-trust-hub — eval-harness