adversarial-legal-review-en

Pass

Audited by Gen Agent Trust Hub on Aug 25, 2026

Risk Level: SAFEPROMPT_INJECTION
Full Analysis
  • [INDIRECT_PROMPT_INJECTION]: The skill is designed to ingest and analyze external, high-stakes documents (such as legal opinions or memos), which serve as untrusted data inputs. This creates a surface for indirect prompt injection, where an attacker could embed instructions within the text to manipulate the AI's adversarial debate logic or force a biased verdict.\n
  • Ingestion points: The skill processes user-supplied documents through the Read tool as described in the SKILL.md workflow.\n
  • Boundary markers: While the skill uses isolated tiers to reduce bias, it lacks explicit technical delimiters (e.g., specific XML tagging) to clearly separate untrusted document content from the agent's internal instructions.\n
  • Capability inventory: The skill utilizes file Read and Write capabilities to perform its analysis and generate reports.\n
  • Sanitization: There are no defined mechanisms for sanitizing or escaping the content of the documents before they are analyzed by the LLM.
Audit Metadata
Risk Level
SAFE
Analyzed
Aug 25, 2026, 11:09 AM
Security Audit — agent-trust-hub — adversarial-legal-review-en