adversarial-review
Fail
Audited by Gen Agent Trust Hub on Mar 29, 2026
Risk Level: HIGHCOMMAND_EXECUTIONPROMPT_INJECTIONDATA_EXFILTRATION
Full Analysis
- [COMMAND_EXECUTION]: The skill constructs shell commands that interpolate untrusted data, such as code diffs and user intents, directly into command arguments for external tools. Evidence: Step 3 in SKILL.md instructs the agent to execute codex exec and claude commands where the prompt variable contains external content. Risk: Malicious diffs or user input containing shell metacharacters like backticks or semicolons could result in arbitrary command execution on the host system.
- [PROMPT_INJECTION]: The skill is susceptible to indirect prompt injection by processing external data (diffs and plans) and passing it to models without sanitization or delimiters. Ingestion points: Code diffs, referenced plans, and user messages identified in SKILL.md Step 2. Boundary markers: None are defined in the reviewer prompt template in Step 3. Capability inventory: Shell command execution via external CLI tools, file system enumeration (ls), and temporary directory creation. Sanitization: No escaping or validation is performed on the untrusted data before interpolation into the prompt.
- [DATA_EXFILTRATION]: The skill reads project source code and transmits it to external model providers via CLI tools for analysis. Evidence: The review process described in SKILL.md Step 3 explicitly includes project diffs in the data sent to the external service commands. Context: This behavior is central to the skill's purpose and targets established service providers.
Recommendations
- AI detected serious security threats
Audit Metadata