inno-idea-eval
Pass
Audited by Gen Agent Trust Hub on Sep 21, 2026
Risk Level: SAFECOMMAND_EXECUTIONINDIRECT_PROMPT_INJECTION
Full Analysis
- [COMMAND_EXECUTION]: The skill instructs the agent to execute a local Python script (
search_ai_papers.py) to perform novelty verification against academic databases including arXiv, Semantic Scholar, and OpenAlex. This relies on the existence of a shared research skill in the local environment. - [INDIRECT_PROMPT_INJECTION]: The skill processes untrusted content from research papers, external codebases, and user-provided ideas, and interpolates this content into prompts for the evaluation agent.
- Ingestion points: Data is read from
Ideation/references/papers/(LaTeX files),Experiment/code_references/(source code and READMEs), andselected_idea.txt. - Boundary markers: Evaluation templates use Markdown headers to separate sections, but they do not employ explicit boundary delimiters or instructions for the agent to disregard instructions embedded in the ingested data.
- Capability inventory: The skill has the ability to read and write files within the project workspace and execute shell commands via a local script.
- Sanitization: There is no evidence of sanitization, filtering, or escaping of the ingested content before it is provided to the Eval Agent in the
evidence_block.
Audit Metadata