diagnosing-bugs
Pass
Audited by Gen Agent Trust Hub on Sep 25, 2026
Risk Level: SAFECOMMAND_EXECUTIONDYNAMIC_EXECUTIONINDIRECT_PROMPT_INJECTION
Full Analysis
- [COMMAND_EXECUTION]: The skill's core workflow requires the agent to execute various shell commands, including test runners, curl for HTTP requests, and custom diagnostic scripts. This capability is necessary for the feedback loop approach to debugging.
- [DYNAMIC_EXECUTION]: The agent is instructed to create and run several types of dynamic scripts and harnesses, such as Throwaway harnesses, Bisection harnesses, and scripts based on the hitl-loop.template.sh template. This involves generating executable code at runtime to verify hypotheses and reproduce bugs.
- [INDIRECT_PROMPT_INJECTION]: The skill has a significant attack surface for indirect prompt injection as it is designed to ingest and process untrusted external data. (1) Ingestion points: The agent reads potentially attacker-controlled content from files like HAR files, log dumps, network traces, and event payloads (e.g., in SKILL.md under 'Ways to construct one'). (2) Boundary markers: The instructions do not specify the use of delimiters or ignore warnings for the data being processed. (3) Capability inventory: The agent possesses extensive capabilities, including shell command execution, file system access, and network operations via curl. (4) Sanitization: There are explicit instructions in the Redact section of SKILL.md to replace secrets with and to avoid showing credentials in outputs.
Audit Metadata