ai-evaluation

Pass

Audited by Gen Agent Trust Hub on May 5, 2026

Risk Level: SAFEPROMPT_INJECTION
Full Analysis
  • [PROMPT_INJECTION]: The skill documents an evaluation pipeline that processes untrusted data, creating a surface for indirect prompt injection.
  • Ingestion points: The example script in SKILL.md interpolates case.input and actual (untrusted model output) into the judge prompt.
  • Boundary markers: The prompt template lacks delimiters or specific instructions to isolate untrusted content from the instructions.
  • Capability inventory: Frontmatter enables the Bash tool, and the script uses the anthropic library for network-based LLM calls.
  • Sanitization: The provided implementation does not validate or sanitize inputs before interpolation.
Audit Metadata
Risk Level
SAFE
Analyzed
May 5, 2026, 06:21 PM
Security Audit — agent-trust-hub — ai-evaluation