ai-data-curation-pretraining
Pass
Audited by Gen Agent Trust Hub on Sep 23, 2026
Risk Level: SAFENO_CODEEXTERNAL_DOWNLOADSINDIRECT_PROMPT_INJECTION
Full Analysis
- [SAFE]: The skill serves as a documentation-only reference for data curation tasks. It contains no executable code, scripts, or command triggers, which significantly reduces the potential attack surface.\n- [EXTERNAL_DOWNLOADS]: The skill references a wide array of research papers, datasets, and tools hosted by trusted entities like Hugging Face, EleutherAI, AllenAI, and the European Commission. These links are provided for informational purposes and correspond to official, well-known domains.\n- [INDIRECT_PROMPT_INJECTION]: The skill focuses on the processing of untrusted web-scale data (CommonCrawl) and addresses the inherent surface vulnerability of training on external content. \n
- Ingestion points: Documents the use of
WARCReaderandHTMLExtractorto ingest web data (references/web-curation-pipeline.md).\n - Boundary markers: Recommends explicit heuristic rules (Gopher/C4) and safety classifiers to isolate high-quality text.\n
- Capability inventory: No code-execution capabilities or scripts are provided within the skill folder.\n
- Sanitization: Provides detailed guidance on PII scrubbing (regex and classifiers) and n-gram decontamination to ensure dataset integrity.
Audit Metadata