ai-evals

Pass

Audited by Gen Agent Trust Hub on Aug 12, 2026

Risk Level: SAFE
Full Analysis
  • [SAFE]: The skill is a documentation-only resource providing architectural and methodological guidance for designing trustworthy AI evaluation programs. It contains no executable scripts, binaries, or automated tasks that could perform unauthorized actions.
  • [EXTERNAL_DOWNLOADS]: The documentation references and provides integration snippets for several industry-standard evaluation frameworks and tools, including Hugging Face's lighteval, OpenAI's promptfoo, UK AISI's inspect-ai, and AWS Bedrock's evaluation services. These tools are sourced from well-known and trusted organizations.
  • [DATA_EXFILTRATION]: The skill emphasizes data privacy by instructing users to de-identify PII (Personally Identifiable Information) and sensitive data from production logs and support tickets before they are utilized in evaluation datasets.
  • [PROMPT_INJECTION]: The skill includes extensive guidance on identifying and mitigating security risks, with dedicated sections for safety red-teaming that cover jailbreak families, indirect prompt injection through tool outputs, and adversarial robustness testing.
Audit Metadata
Risk Level
SAFE
Analyzed
Aug 12, 2026, 09:09 PM
Security Audit — agent-trust-hub — ai-evals