evaluating-llms-harness
Warn
Audited by Snyk on May 19, 2026
Risk Level: MEDIUM
Full Analysis
MEDIUM W011: Third-party content exposure detected (indirect prompt injection risk).
- Third-party content exposure detected (high risk: 0.90). The skill's workflows and docs (e.g., references/api-evaluation.md and references/custom-tasks.md in SKILL.md) explicitly ingest and evaluate responses from public API models (OpenAI, Anthropic, local/OpenAI-compatible endpoints) and arbitrary HuggingFace/public datasets or custom API endpoints, meaning the agent will fetch and interpret untrusted, third-party/user-generated content as part of its evaluation pipeline.
Issues (1)
W011
MEDIUMThird-party content exposure detected (indirect prompt injection risk).
Audit Metadata