agent-evals
Pass
Audited by Gen Agent Trust Hub on Sep 10, 2026
Risk Level: SAFE
Full Analysis
- [SAFE]: The skill is entirely composed of informational and technical documentation regarding agent performance evaluation. It provides guidance on creating evaluation rubrics, implementing checker nodes for trajectory validation, and using standardized telemetry formats.
- [EXTERNAL_DOWNLOADS]: The skill references technical documentation from Anthropic regarding harness design for long-running applications (SKILL.md §5a). This reference is provided for architectural context and is documented neutrally.
- [INDIRECT_PROMPT_INJECTION]: While the skill discusses processing production data and user transcripts, it provides extensive mitigation strategies, including the use of deterministic checker nodes and narrow pass/fail rubrics to validate agent behavior. No implementation of this surface is present in the skill itself.
Audit Metadata