ai-data-curation-pretraining

Pass

Audited by Gen Agent Trust Hub on Aug 12, 2026

Risk Level: SAFE
Full Analysis
  • [EXTERNAL_DOWNLOADS]: The skill references a wide range of external resources, including official documentation from the European Commission, research papers on arXiv and Nature, and GitHub repositories from trusted organizations such as HuggingFace, AllenAI, NVIDIA, and EleutherAI. These references are provided for educational and research purposes and target well-known, reputable services.
  • [COMMAND_EXECUTION]: Includes instructional examples for common data engineering tasks, such as using the AWS CLI for accessing the public CommonCrawl S3 bucket and Python snippets for processing data with standard libraries like datatrove and fasttext. These examples are benign and strictly follow the skill's purpose as a functional reference.
  • [SAFE]: The skill explicitly promotes and provides instructions for security-best-practices, including PII (Personally Identifiable Information) scrubbing using regex and classifiers, and rigorous decontamination protocols to ensure that benchmark evaluation data does not leak into training corpora.
Audit Metadata
Risk Level
SAFE
Analyzed
Aug 12, 2026, 09:09 PM
Security Audit — agent-trust-hub — ai-data-curation-pretraining