anthropic-evaluations

Pass

Audited by Gen Agent Trust Hub on Apr 28, 2026

Risk Level: SAFEEXTERNAL_DOWNLOADSNO_CODE
Full Analysis
  • [SAFE]: The skill content is purely documentation and templates intended to assist users in designing agent evaluations. No malicious patterns such as prompt injection, credential harvesting, or unauthorized data access were found.
  • [EXTERNAL_DOWNLOADS]: The skill references several external, well-known frameworks (such as Harbor, Promptfoo, Braintrust, and Langfuse) and research benchmarks (such as SWE-bench, Terminal-Bench, and OSWorld). These references are provided for informational purposes and do not involve automated execution or installation of untrusted code.
  • [NO_CODE]: The skill consists entirely of Markdown documentation and YAML templates. It does not include any executable scripts, binaries, or active code components, which inherently limits its security risk profile.
Audit Metadata
Risk Level
SAFE
Analyzed
Apr 28, 2026, 08:49 AM
Security Audit — agent-trust-hub — anthropic-evaluations