research-implement-feature

Pass

Audited by Gen Agent Trust Hub on Sep 17, 2026

Risk Level: SAFECOMMAND_EXECUTIONEXTERNAL_DOWNLOADSINDIRECT_PROMPT_INJECTIONDYNAMIC_EXECUTION
Full Analysis
  • [COMMAND_EXECUTION]: The skill is designed to execute arbitrary shell commands via the Bash tool to verify implementation rungs and run the final artifact.
  • Evidence: Phase 2 and Phase 3 instructions require running "acceptance checks" and "success commands" (e.g., python scripts/run.py --smoke) to confirm the code functions as expected.
  • [EXTERNAL_DOWNLOADS]: The skill supports cloning external repositories from a URL provided in the user arguments.
  • Evidence: The BASE_REPO parameter in the YAML frontmatter and the Phase 0 instructions allow the agent to clone a repository to use as a starting point for implementation.
  • [INDIRECT_PROMPT_INJECTION]: The skill has a vulnerability surface for indirect prompt injection as it ingests untrusted external data which is then used to guide code generation and shell execution.
  • Ingestion points: Phase 0 reads content from external sources such as paper PDFs, fetched READMEs, or issue threads as part of the research phase (SKILL.md).
  • Boundary markers: The skill attempts to mitigate this by referencing shared-references/injection-hygiene.md and instructing the agent to treat fetched content as data, not instructions.
  • Capability inventory: The skill possesses extensive capabilities including Bash(*), Write, and Edit, which could be abused if instructions inside the external data are accidentally followed.
  • Sanitization: The process relies on the agent's adherence to the hygiene guidelines to prevent data from redirecting the build or execution flow.
  • [DYNAMIC_EXECUTION]: The skill inherently relies on generating code and then executing it immediately to verify progress.
  • Evidence: The core "Feature Ladder" logic (Phase 1-3) involves writing new Python or script files and then executing them via the shell to satisfy acceptance gates.
Audit Metadata
Risk Level
SAFE
Analyzed
Sep 17, 2026, 02:24 AM
Security Audit — agent-trust-hub — research-implement-feature