training-data
Pass
Audited by Gen Agent Trust Hub on Aug 6, 2026
Risk Level: SAFE
Full Analysis
- [SAFE]: The skill provides instructions and code snippets for preparing training datasets, which is its stated purpose. It correctly identifies the requirements for various TRL trainers and provides best practices for chat templates and EOS tokens.
- [SAFE]: Dependencies such as 'transformers', 'datasets', and 'distilabel' are standard libraries from well-known and trusted organizations in the machine learning ecosystem. The skill suggests pinning versions, which is a security best practice.
- [SAFE]: Data operations, including loading JSONL files and pushing to the Hugging Face Hub ('push_to_hub'), follow best practices for dataset management and use legitimate API patterns. The skill explicitly advises using private repositories and recording provenance.
- [SAFE]: The skill includes comprehensive guardrails against common errors in dataset preparation, such as template mismatches and lack of decontamination, which improves the integrity of the fine-tuning process.
Audit Metadata