create-classifier

Warn

Audited by Gen Agent Trust Hub on Mar 17, 2026

Risk Level: MEDIUMCOMMAND_EXECUTIONDATA_EXFILTRATIONREMOTE_CODE_EXECUTIONPROMPT_INJECTIONEXTERNAL_DOWNLOADS
Full Analysis
  • [COMMAND_EXECUTION]: Multiple scripts, including assess_task.py and model_selector.py, construct shell commands using string interpolation and execute them via subprocess.run with bash -lc. This pattern is susceptible to command injection if input parameters like task names or model backbones are not strictly sanitized.
  • [DATA_EXFILTRATION]: The scripts/mine_transcripts.py utility is designed to access and parse sensitive user data, including Claude conversation logs (~/.claude/projects) and Codex history (~/.codex/history.jsonl), to extract training data from previous interactions.
  • [REMOTE_CODE_EXECUTION]: The use of torch.load() in evaluate.py and iterative_train.py to load model checkpoints presents a risk of arbitrary code execution during deserialization if a malicious model file is used.
  • [PROMPT_INJECTION]: Data augmentation utilities like hf_augment.py and collect_bridge_data.py utilize the scillm tool with hardcoded prompt templates. This creates an attack surface for indirect prompt injection where malicious content in the training data could influence the language model's output during the augmentation process.
  • [EXTERNAL_DOWNLOADS]: The skill frequently interacts with Hugging Face using the datasets and huggingface_hub libraries to download external datasets and model weights for model initialization and data enrichment.
Audit Metadata
Risk Level
MEDIUM
Analyzed
Mar 17, 2026, 06:34 AM
Security Audit — agent-trust-hub — create-classifier