train-convo-steering

Pass

Audited by Gen Agent Trust Hub on Mar 17, 2026

Risk Level: SAFEPROMPT_INJECTIONCOMMAND_EXECUTIONEXTERNAL_DOWNLOADS
Full Analysis
  • [PROMPT_INJECTION]: The skill constructs prompts for an external LLM judge by directly interpolating user text into the prompt string in train_convo_steering/judge_deepseek.py, creating a surface for prompt injection attacks.
  • [PROMPT_INJECTION]: The skill is susceptible to indirect prompt injection as it processes untrusted conversation logs during its nightly analysis phase. Ingestion points: user text from _out/live_logs.jsonl. Boundary markers: absent in the judge prompt template. Capability inventory: influences agent response presets and interacts with shared memory. Sanitization: no validation or escaping of the user text is present before LLM processing.
  • [COMMAND_EXECUTION]: Integration with the project's memory and taxonomy components is achieved through dynamic module loading using importlib.util in train_convo_steering/cli.py and memory_integration.py. This pattern of loading code from computed paths represents a risk if the skill's environment is compromised.
  • [EXTERNAL_DOWNLOADS]: The skill performs network operations to reach an external API endpoint (DeepSeek via Chutes.ai) during the nightly processing routine to analyze interaction data and update learned priors.
Audit Metadata
Risk Level
SAFE
Analyzed
Mar 17, 2026, 06:36 AM
Security Audit — agent-trust-hub — train-convo-steering