ai-evaluation
Pass
Audited by Gen Agent Trust Hub on May 5, 2026
Risk Level: SAFEPROMPT_INJECTION
Full Analysis
- [PROMPT_INJECTION]: The skill documents an evaluation pipeline that processes untrusted data, creating a surface for indirect prompt injection.
- Ingestion points: The example script in SKILL.md interpolates
case.inputandactual(untrusted model output) into the judge prompt. - Boundary markers: The prompt template lacks delimiters or specific instructions to isolate untrusted content from the instructions.
- Capability inventory: Frontmatter enables the
Bashtool, and the script uses theanthropiclibrary for network-based LLM calls. - Sanitization: The provided implementation does not validate or sanitize inputs before interpolation.
Audit Metadata