adversarial-review

Warn

Audited by Gen Agent Trust Hub on Aug 26, 2026

Risk Level: MEDIUMCOMMAND_EXECUTIONINDIRECT_PROMPT_INJECTIONDYNAMIC_EXECUTION
Full Analysis
  • [COMMAND_EXECUTION]: The skill instructs the agent to construct and execute shell commands (codex exec and claude -p) by interpolating a prompt variable. This variable contains data from untrusted sources, including recent code diffs and user messages. If the agent fails to properly escape this input before shell interpolation, it allows for arbitrary command execution on the host environment.
  • [INDIRECT_PROMPT_INJECTION]: The skill ingests untrusted data that could influence the agent's behavior.
  • Ingestion points: Recent code diffs, referenced plans, and user messages identified in SKILL.md (Step 2).
  • Boundary markers: There are no explicit boundary markers or instructions to isolate the diff content from the rest of the prompt, allowing embedded instructions in the code to potentially manipulate the reviewer model's output.
  • Capability inventory: The skill uses shell execution capabilities (codex exec, claude) and performs file operations in a temporary directory.
  • Sanitization: There is no evidence of sanitization or filtering for the data interpolated into the prompts or the output received from the secondary models.
  • [DYNAMIC_EXECUTION]: The skill dynamically generates shell commands based on the detected model identity to interact with external tools.
Audit Metadata
Risk Level
MEDIUM
Analyzed
Aug 26, 2026, 11:03 PM
Security Audit — agent-trust-hub — adversarial-review