fine-tuning-with-trl
Pass
Audited by Gen Agent Trust Hub on Oct 1, 2026
Risk Level: SAFE
Full Analysis
- [SAFE]: The skill consists of legitimate technical documentation for machine learning fine-tuning. It utilizes well-known and trusted libraries including
trl,transformers,peft, anddatasets. All code examples follow standard industry practices for RLHF (Reinforcement Learning from Human Feedback) pipelines. - [INDIRECT_PROMPT_INJECTION]: The skill describes workflows that ingest external datasets (e.g., from the Hugging Face Hub) to train models. While processing untrusted data is a theoretical vector for indirect prompt injection, it is the primary and intended purpose of this machine learning skill. The instructions rely on established library trainers (SFTTrainer, DPOTrainer) which are standard for these tasks.
Audit Metadata