create-gpt

Warn

Audited by Gen Agent Trust Hub on Mar 17, 2026

Risk Level: MEDIUMREMOTE_CODE_EXECUTIONCOMMAND_EXECUTIONEXTERNAL_DOWNLOADSPROMPT_INJECTION
Full Analysis
  • [REMOTE_CODE_EXECUTION]: Multiple scripts including evaluate.py, hp_search.py, infer.py, train_grpo.py, and train_sft.py load models using the Hugging Face Transformers library with the trust_remote_code=True parameter. This setting allows the execution of custom code provided within the model's remote repository.
  • [REMOTE_CODE_EXECUTION]: The scripts/rewards/task_reward.py file contains functionality to dynamically load and execute arbitrary Python code from a file via importlib and exec_module. This is used to load pluggable reward functions for GRPO training, but serves as a direct vector for code execution if an untrusted file path is provided.
  • [COMMAND_EXECUTION]: The skill uses subprocess.run throughout its orchestration scripts (iterative_train.py, teacher_student_loop.py, export_gguf.py) to execute various Python scripts and CLI utilities. While these primarily target local files, they expand the attack surface by allowing the agent to launch numerous sub-processes with varying arguments.
  • [EXTERNAL_DOWNLOADS]: The skill is configured to download pre-trained model weights from Hugging Face and various Python dependencies from PyPI. While these sources are well-known, the automated fetching of large binary files and libraries remains a point of external data ingestion.
  • [PROMPT_INJECTION]: The skill is vulnerable to indirect prompt injection in several locations:
  • Ingestion points: infer.py, router.py, and teacher.py take user-provided input and incorporate it into prompts for language models.
  • Boundary markers: The skill uses simple text markers (e.g., "Input:") which are insufficient to prevent a model from obeying instructions embedded within the data.
  • Capability inventory: The skill possesses the ability to execute shell commands, write files to disk, and make external network calls via the scillm utility.
  • Sanitization: No sanitization or robust escaping is performed on the user-provided data before it is interpolated into model prompts.
Audit Metadata
Risk Level
MEDIUM
Analyzed
Mar 17, 2026, 06:34 AM
Security Audit — agent-trust-hub — create-gpt