problem-doc-model-selector

Pass

Audited by Gen Agent Trust Hub on Sep 5, 2026

Risk Level: SAFEINDIRECT_PROMPT_INJECTIONCOMMAND_EXECUTION
Full Analysis
  • [INDIRECT_PROMPT_INJECTION]: The skill processes untrusted user-supplied documents (PDF, Word, TXT) from the problem_files/ directory and extracts their content to generate structured reports and JSON manifests. These outputs are subsequently used to guide the behavior of the agent in later stages of the paper-writing workflow. This creates a surface for indirect prompt injection where a malicious document could contain hidden instructions designed to influence the agent's actions.
  • Ingestion points: The scripts/analyze_problem.py script reads all files in the problem_files/ directory using Path.rglob('*').
  • Boundary markers: Extracted text is placed into Markdown templates and JSON files. While these provide structural separation, the skill lacks explicit instructions or delimiters that warn the agent to ignore potentially adversarial directives embedded within the processed text.
  • Capability inventory: The skill facilitates the execution of local workflow scripts (workflow_guard.py, update_workflow_memory.py) and performs extensive file system writes to paper_output/.
  • Sanitization: The normalize_text function in scripts/analyze_problem.py uses regular expressions to clean whitespace and newlines, but it does not sanitize or filter for potential prompt injection patterns.
  • [COMMAND_EXECUTION]: The SKILL.md instructions explicitly direct the agent to execute multiple shell commands to run Python scripts located in the .claude/skills/ directory (e.g., workflow_guard.py). These scripts manage the state and integrity of the overall workflow, representing a dependency on local executable code within hidden project directories.
Audit Metadata
Risk Level
SAFE
Analyzed
Sep 5, 2026, 02:58 PM