ai-data-curation-pretraining

Warn

Audited by Snyk on Aug 12, 2026

Risk Level: MEDIUM
Full Analysis

MEDIUM W011: Third-party content exposure detected (indirect prompt injection risk).

  • Third-party content exposure detected (medium risk: 0.30). The skill’s required runtime workflow ingests and LLM-reads extracted free text from CommonCrawl WARC records during the extraction→filtering stages (e.g., via datatrove HTMLExtractor/trafilatura), so outsider-authored web content can be processed without selecting a specific item first.

Issues (1)

W011
MEDIUM

Third-party content exposure detected (indirect prompt injection risk).

Audit Metadata
Risk Level
MEDIUM
Analyzed
Aug 12, 2026, 09:09 PM
Issues
1
Security Audit — snyk — ai-data-curation-pretraining