inno-idea-eval

Pass

Audited by Gen Agent Trust Hub on Sep 21, 2026

Risk Level: SAFECOMMAND_EXECUTIONINDIRECT_PROMPT_INJECTION
Full Analysis
  • [COMMAND_EXECUTION]: The skill instructs the agent to execute a local Python script (search_ai_papers.py) to perform novelty verification against academic databases including arXiv, Semantic Scholar, and OpenAlex. This relies on the existence of a shared research skill in the local environment.
  • [INDIRECT_PROMPT_INJECTION]: The skill processes untrusted content from research papers, external codebases, and user-provided ideas, and interpolates this content into prompts for the evaluation agent.
  • Ingestion points: Data is read from Ideation/references/papers/ (LaTeX files), Experiment/code_references/ (source code and READMEs), and selected_idea.txt.
  • Boundary markers: Evaluation templates use Markdown headers to separate sections, but they do not employ explicit boundary delimiters or instructions for the agent to disregard instructions embedded in the ingested data.
  • Capability inventory: The skill has the ability to read and write files within the project workspace and execute shell commands via a local script.
  • Sanitization: There is no evidence of sanitization, filtering, or escaping of the ingested content before it is provided to the Eval Agent in the evidence_block.
Audit Metadata
Risk Level
SAFE
Analyzed
Sep 21, 2026, 05:59 PM