fine-tuning-expert

Pass

Audited by Gen Agent Trust Hub on Aug 26, 2026

Risk Level: SAFEPROMPT_INJECTIONCOMMAND_EXECUTION
Full Analysis
  • [INDIRECT_PROMPT_INJECTION]: The skill facilitates the ingestion of training data from external files and remote sources, which serves as an entry point for untrusted data that could influence model training.
  • Ingestion points: The script references/dataset-preparation.md (via the load_custom_dataset function) and the primary SKILL.md (via load_dataset) process data in JSON, JSONL, and Parquet formats.
  • Capability inventory: The skill allows for local file system writes (checkpoints), network requests to established AI services (Hugging Face Hub, OpenAI), and the execution of external binaries for model conversion.
  • Boundary markers: The instructions do not define specific delimiters or warnings to isolate training data content from the model's instructions during the training process.
  • Sanitization: While a quality_filter is provided to exclude common AI refusal phrases, it lacks robust sanitization to detect or neutralize malicious instructions embedded in the datasets.
  • [COMMAND_EXECUTION]: The model deployment utilities in references/deployment-optimization.md utilize subprocess.run to call external Python scripts and binaries, such as the llama-quantize tool from the llama.cpp project. These commands are executed as part of a legitimate workflow for model quantization and format conversion.
Audit Metadata
Risk Level
SAFE
Analyzed
Aug 26, 2026, 05:57 PM
Security Audit — agent-trust-hub — fine-tuning-expert