fine-tuning-expert
Pass
Audited by Gen Agent Trust Hub on Aug 26, 2026
Risk Level: SAFEPROMPT_INJECTIONCOMMAND_EXECUTION
Full Analysis
- [INDIRECT_PROMPT_INJECTION]: The skill facilitates the ingestion of training data from external files and remote sources, which serves as an entry point for untrusted data that could influence model training.
- Ingestion points: The script
references/dataset-preparation.md(via theload_custom_datasetfunction) and the primarySKILL.md(viaload_dataset) process data in JSON, JSONL, and Parquet formats. - Capability inventory: The skill allows for local file system writes (checkpoints), network requests to established AI services (Hugging Face Hub, OpenAI), and the execution of external binaries for model conversion.
- Boundary markers: The instructions do not define specific delimiters or warnings to isolate training data content from the model's instructions during the training process.
- Sanitization: While a
quality_filteris provided to exclude common AI refusal phrases, it lacks robust sanitization to detect or neutralize malicious instructions embedded in the datasets. - [COMMAND_EXECUTION]: The model deployment utilities in
references/deployment-optimization.mdutilizesubprocess.runto call external Python scripts and binaries, such as thellama-quantizetool from thellama.cppproject. These commands are executed as part of a legitimate workflow for model quantization and format conversion.
Audit Metadata