auto-research
Pass
Audited by Gen Agent Trust Hub on Sep 26, 2026
Risk Level: SAFEINDIRECT_PROMPT_INJECTIONCOMMAND_EXECUTION
Full Analysis
- [INDIRECT_PROMPT_INJECTION]: The skill defines a research loop that explicitly relies on ingesting untrusted data from multiple external sources.
- Ingestion points: The workflow in Step 1 and Step 4 requires the agent to read "user feedback," "project instructions," "research records," and "actual failure traces."
- Boundary markers: While the skill suggests documenting a "research contract" to define success, it provides no instructions for sanitizing or delimiting raw input from traces or feedback to prevent them from being interpreted as commands.
- Capability inventory: The agent is instructed to "edit candidates" (Step 2), "run the evaluation command" (Step 3), and execute focused tasks to "test one hypothesis" (Step 5).
- Sanitization: There are no requirements for escaping or validating the content of failure logs or feedback before they are used to generate new hypotheses or trial commands.
- [COMMAND_EXECUTION]: The core of the skill involves managing and executing shell commands for benchmarking and evaluation.
- The agent is tasked with recording evaluation commands, turnaround times, and resource limits.
- This functional capability is a primary feature of the skill, but when combined with the data ingestion surface, it presents a vector where malicious instructions in a processed log could influence the construction of the evaluation commands.
Audit Metadata