train-convo-steering
Pass
Audited by Gen Agent Trust Hub on Mar 17, 2026
Risk Level: SAFEPROMPT_INJECTIONCOMMAND_EXECUTIONEXTERNAL_DOWNLOADS
Full Analysis
- [PROMPT_INJECTION]: The skill constructs prompts for an external LLM judge by directly interpolating user text into the prompt string in train_convo_steering/judge_deepseek.py, creating a surface for prompt injection attacks.
- [PROMPT_INJECTION]: The skill is susceptible to indirect prompt injection as it processes untrusted conversation logs during its nightly analysis phase. Ingestion points: user text from _out/live_logs.jsonl. Boundary markers: absent in the judge prompt template. Capability inventory: influences agent response presets and interacts with shared memory. Sanitization: no validation or escaping of the user text is present before LLM processing.
- [COMMAND_EXECUTION]: Integration with the project's memory and taxonomy components is achieved through dynamic module loading using importlib.util in train_convo_steering/cli.py and memory_integration.py. This pattern of loading code from computed paths represents a risk if the skill's environment is compromised.
- [EXTERNAL_DOWNLOADS]: The skill performs network operations to reach an external API endpoint (DeepSeek via Chutes.ai) during the nightly processing routine to analyze interaction data and update learned priors.
Audit Metadata