ai-data-curation-pretraining
Warn
Audited by Snyk on Aug 12, 2026
Risk Level: MEDIUM
Full Analysis
MEDIUM W011: Third-party content exposure detected (indirect prompt injection risk).
- Third-party content exposure detected (medium risk: 0.30). The skill’s required runtime workflow ingests and LLM-reads extracted free text from CommonCrawl WARC records during the extraction→filtering stages (e.g., via datatrove HTMLExtractor/trafilatura), so outsider-authored web content can be processed without selecting a specific item first.
Issues (1)
W011
MEDIUMThird-party content exposure detected (indirect prompt injection risk).
Audit Metadata