evaluating-llms-harness

Warn

Audited by Snyk on May 19, 2026

Risk Level: MEDIUM
Full Analysis

MEDIUM W011: Third-party content exposure detected (indirect prompt injection risk).

  • Third-party content exposure detected (high risk: 0.90). The skill's workflows and docs (e.g., references/api-evaluation.md and references/custom-tasks.md in SKILL.md) explicitly ingest and evaluate responses from public API models (OpenAI, Anthropic, local/OpenAI-compatible endpoints) and arbitrary HuggingFace/public datasets or custom API endpoints, meaning the agent will fetch and interpret untrusted, third-party/user-generated content as part of its evaluation pipeline.

Issues (1)

W011
MEDIUM

Third-party content exposure detected (indirect prompt injection risk).

Audit Metadata
Risk Level
MEDIUM
Analyzed
May 19, 2026, 09:58 AM
Issues
1
Security Audit — snyk — evaluating-llms-harness