train-persona

Pass

Audited by Gen Agent Trust Hub on Mar 17, 2026

Risk Level: SAFEEXTERNAL_DOWNLOADSCOMMAND_EXECUTIONPROMPT_INJECTION
Full Analysis
  • [EXTERNAL_DOWNLOADS]: The skill downloads model weights and configurations from Hugging Face (e.g., deepseek-ai/DeepSeek-R1-Distill-Qwen-7B). This is a well-known service for machine learning assets and is considered a safe source for these dependencies.
  • [COMMAND_EXECUTION]: The run.sh script uses uv run to execute local Python scripts such as train_persona.py and upgrade_traces.py. This is the intended mechanism for managing the training and evaluation workflow.
  • [PROMPT_INJECTION]: The skill processes external training data which could potentially contain indirect prompt injections targeting the persona being trained.
  • Ingestion points: Training examples are loaded from JSONL files via the --train-file and --input CLI arguments.
  • Boundary markers: The logic does not currently implement specific delimiters or instructions to ignore instructions embedded within the training data.
  • Capability inventory: The skill performs model training via the trl and transformers libraries and includes placeholders for LLM-based critique.
  • Sanitization: No content sanitization or structured validation for the training text was identified in the provided scripts.
Audit Metadata
Risk Level
SAFE
Analyzed
Mar 17, 2026, 06:35 AM
Security Audit — agent-trust-hub — train-persona