eval-harness
Pass
Audited by Gen Agent Trust Hub on Aug 15, 2026
Risk Level: SAFE
Full Analysis
- [SAFE]: The skill provides instructional templates for defining and running performance evaluations (evals) for AI agents.
- [SAFE]: Command execution examples are standard development utilities (grep, npm test, npm run build) used for deterministic grading of code changes.
- [SAFE]: All storage and retrieval operations target local project directories (e.g., .claude/evals/), following established software engineering patterns.
- [SAFE]: No evidence of prompt injection, obfuscation, or unauthorized data exfiltration was found in the provided documentation or metadata.
Audit Metadata