academic-paper-reviewer

Pass

Audited by Gen Agent Trust Hub on Oct 3, 2026

Risk Level: SAFEINDIRECT_PROMPT_INJECTION
Full Analysis
  • [INDIRECT_PROMPT_INJECTION]: The skill exhibits an inherent attack surface for indirect prompt injection because it is designed to ingest and process untrusted third-party data (academic manuscripts). However, it implements industry-standard mitigations.
  • Ingestion points: Untrusted data enters the agent context in the Phase 1 and Phase 2 reviewer dispatches (e.g., in eic_agent.md and methodology_reviewer_agent.md) where manuscripts are provided.
  • Boundary markers: The skill explicitly uses XML-style tags (<paper_content>...</paper_content>) to delimit untrusted data.
  • Sanitization: It includes clear, recurring instructions to the LLM (e.g., the 'Instruction-Data Boundary' principle) stating that imperative-looking text inside retrieved content must be treated as data to report on, not as commands to follow. Specifically, it names 'ignore previous instructions' as a pattern to identify and disregard.
  • Capability inventory: The skill's capabilities are restricted to generating markdown reports and structured decision packages. It does not utilize dangerous tools for network exfiltration or file system modification outside of its own generated artifacts.
  • [PROMPT_INJECTION]: Static analysis flagged multiple instances of 'ignore instructions' patterns. Upon manual review, these are confirmed to be defensive instructions rather than malicious injections. The skill explicitly instructs sub-agents to ignore any 'ignore previous instructions' directives found within the manuscripts being reviewed to prevent adversarial control of the agent's persona.
Audit Metadata
Risk Level
SAFE
Analyzed
Oct 3, 2026, 01:29 PM
Security Audit — agent-trust-hub — academic-paper-reviewer