auto-review-loop-llm

Pass

Audited by Gen Agent Trust Hub on Sep 15, 2026

Risk Level: SAFEDATA_EXFILTRATIONINDIRECT_PROMPT_INJECTIONCOMMAND_EXECUTION
Full Analysis
  • [DATA_EXFILTRATION]: The skill transmits the LLM_API_KEY and "Full research context" (claims, methods, results) to an external endpoint defined by ${LLM_BASE_URL} via curl or an MCP tool. While the documentation suggests well-known and trusted providers, the mechanism allows the configuration of any arbitrary URL, which could lead to credential and data exfiltration if misconfigured to point to a malicious server.
  • [INDIRECT_PROMPT_INJECTION]: The skill ingests untrusted data from external LLM responses and uses it to drive the "Phase C: Implement Fixes" stage.
  • Ingestion points: External LLM API responses received via the llm-chat MCP tool or curl (Phase A/B).
  • Boundary markers: None identified in the prompt templates or parsing logic to prevent the LLM from injecting malicious instructions into the suggested fixes.
  • Capability inventory: The skill is granted broad permissions including Bash(*), Write, and Edit to modify the local project environment.
  • Sanitization: No validation or sanitization of the LLM-suggested fixes is described before they are implemented in the code or documentation.
  • [COMMAND_EXECUTION]: The skill uses Bash(*) to handle large file writes (cat << 'EOF' > file). This operation is performed autonomously without seeking user confirmation, which increases the impact of a compromised LLM response that might attempt to write malicious scripts or configuration files.
Audit Metadata
Risk Level
SAFE
Analyzed
Sep 15, 2026, 02:08 PM
Security Audit — agent-trust-hub — auto-review-loop-llm