simpo-training

Pass

Audited by Gen Agent Trust Hub on Oct 1, 2026

Risk Level: SAFEEXTERNAL_DOWNLOADSCOMMAND_EXECUTIONREMOTE_CODE_EXECUTIONINDIRECT_PROMPT_INJECTION
Full Analysis
  • [REMOTE_CODE_EXECUTION]: Clones the Hugging Face alignment-handbook repository and executes its training scripts (e.g., scripts/run_simpo.py) to perform the alignment optimization.
  • [EXTERNAL_DOWNLOADS]: Fetches configuration files and datasets from trusted infrastructure, including Hugging Face's official repositories, Anthropic's preference datasets, and NVIDIA's research data.
  • [COMMAND_EXECUTION]: Provides instructions to execute shell commands for Conda environment creation, dependency installation via pip (including Flash Attention 2), and launching distributed training jobs using the accelerate launch utility.
  • [INDIRECT_PROMPT_INJECTION]:
  • Ingestion points: The skill processes external training data from sources like UltraFeedback and HH-RLHF as defined in SKILL.md and references/datasets.md.
  • Boundary markers: Data is processed using standard dataset loaders without specific adversarial instruction delimiters defined in the provided snippets.
  • Capability inventory: Includes the capability to write training checkpoints to local storage (SKILL.md), execute local and cloned Python scripts, and download external resources via git and pip.
  • Sanitization: Employs data quality filters for length and diversity (LSH-based deduplication) in references/datasets.md, though it does not implement specific sanitization for adversarial prompt content, which is typical for the intended training use case.
Audit Metadata
Risk Level
SAFE
Analyzed
Oct 1, 2026, 07:50 AM
Security Audit — agent-trust-hub — simpo-training