research-implement-feature
Pass
Audited by Gen Agent Trust Hub on Sep 17, 2026
Risk Level: SAFECOMMAND_EXECUTIONEXTERNAL_DOWNLOADSINDIRECT_PROMPT_INJECTIONDYNAMIC_EXECUTION
Full Analysis
- [COMMAND_EXECUTION]: The skill is designed to execute arbitrary shell commands via the Bash tool to verify implementation rungs and run the final artifact.
- Evidence: Phase 2 and Phase 3 instructions require running "acceptance checks" and "success commands" (e.g.,
python scripts/run.py --smoke) to confirm the code functions as expected. - [EXTERNAL_DOWNLOADS]: The skill supports cloning external repositories from a URL provided in the user arguments.
- Evidence: The
BASE_REPOparameter in the YAML frontmatter and the Phase 0 instructions allow the agent to clone a repository to use as a starting point for implementation. - [INDIRECT_PROMPT_INJECTION]: The skill has a vulnerability surface for indirect prompt injection as it ingests untrusted external data which is then used to guide code generation and shell execution.
- Ingestion points: Phase 0 reads content from external sources such as paper PDFs, fetched READMEs, or issue threads as part of the research phase (SKILL.md).
- Boundary markers: The skill attempts to mitigate this by referencing
shared-references/injection-hygiene.mdand instructing the agent to treat fetched content as data, not instructions. - Capability inventory: The skill possesses extensive capabilities including
Bash(*),Write, andEdit, which could be abused if instructions inside the external data are accidentally followed. - Sanitization: The process relies on the agent's adherence to the hygiene guidelines to prevent data from redirecting the build or execution flow.
- [DYNAMIC_EXECUTION]: The skill inherently relies on generating code and then executing it immediately to verify progress.
- Evidence: The core "Feature Ladder" logic (Phase 1-3) involves writing new Python or script files and then executing them via the shell to satisfy acceptance gates.
Audit Metadata