deep-research

Pass

Audited by Gen Agent Trust Hub on Oct 2, 2026

Risk Level: SAFEINDIRECT_PROMPT_INJECTIONDATA_EXFILTRATIONCOMMAND_EXECUTIONPROMPT_INJECTION
Full Analysis
  • [INDIRECT_PROMPT_INJECTION]: The skill is designed to ingest and process substantial amounts of untrusted third-party content (web pages, fetched PDFs, manuscripts) across its 13-agent pipeline. This creates a significant surface for indirect prompt injection attacks where malicious instructions could be hidden in academic papers or web content.
  • Ingestion points: Content is ingested in agents/bibliography_agent.md, agents/source_verification_agent.md, and agents/synthesis_agent.md from external databases and PDFs.
  • Boundary markers: The skill implements a robust instruction-data-boundary marker across almost all agent files, explicitly instructing the AI to treat retrieved content as data and not commands.
  • Capability inventory: The pipeline possesses significant capabilities, including network access to academic APIs, local file writing, and execution of shell utilities like pdftotext (in agents/timeline_extraction_agent.md).
  • Sanitization: The skill includes explicit instructions to detect and ignore imperative-looking text aimed at the agent within retrieved data.
  • [DATA_EXFILTRATION]: The skill contains an optional ARS_CROSS_MODEL feature (detailed in agents/devils_advocate_agent.md and agents/research_architect_agent.md) that allows the agent to send research blueprints and reviewed materials to external AI models for blind critiques or disagreement checks. While the skill requires explicit user consent and identifies the external provider/model before sending, it constitutes an intentional path for data to leave the local environment.
  • [COMMAND_EXECUTION]: The skill uses the pdftotext command-line utility in agents/timeline_extraction_agent.md to scan the first page of local PDFs for publication dates. This is a legitimate functional requirement but involves subprocess execution.
  • [PROMPT_INJECTION]: The SKILL.md routing core defines a [direct-mode] 'escape hatch' token. If a user message begins with this prefix, the system strips it and skips certain intent-clarification and routing steps to execute commands directly. While intended for power-user control, this is a mechanism that overrides standard behavioral logic.
Audit Metadata
Risk Level
SAFE
Analyzed
Oct 2, 2026, 01:10 PM
Security Audit — agent-trust-hub — deep-research