simpo-training
Pass
Audited by Gen Agent Trust Hub on Oct 1, 2026
Risk Level: SAFEEXTERNAL_DOWNLOADSCOMMAND_EXECUTIONREMOTE_CODE_EXECUTIONINDIRECT_PROMPT_INJECTION
Full Analysis
- [REMOTE_CODE_EXECUTION]: Clones the Hugging Face
alignment-handbookrepository and executes its training scripts (e.g.,scripts/run_simpo.py) to perform the alignment optimization. - [EXTERNAL_DOWNLOADS]: Fetches configuration files and datasets from trusted infrastructure, including Hugging Face's official repositories, Anthropic's preference datasets, and NVIDIA's research data.
- [COMMAND_EXECUTION]: Provides instructions to execute shell commands for Conda environment creation, dependency installation via pip (including Flash Attention 2), and launching distributed training jobs using the
accelerate launchutility. - [INDIRECT_PROMPT_INJECTION]:
- Ingestion points: The skill processes external training data from sources like UltraFeedback and HH-RLHF as defined in
SKILL.mdandreferences/datasets.md. - Boundary markers: Data is processed using standard dataset loaders without specific adversarial instruction delimiters defined in the provided snippets.
- Capability inventory: Includes the capability to write training checkpoints to local storage (
SKILL.md), execute local and cloned Python scripts, and download external resources via git and pip. - Sanitization: Employs data quality filters for length and diversity (LSH-based deduplication) in
references/datasets.md, though it does not implement specific sanitization for adversarial prompt content, which is typical for the intended training use case.
Audit Metadata