eval-skills
Pass
Audited by Gen Agent Trust Hub on Sep 26, 2026
Risk Level: SAFEINDIRECT_PROMPT_INJECTIONCOMMAND_EXECUTION
Full Analysis
- [INDIRECT_PROMPT_INJECTION]: The skill processes untrusted external code and data, creating a surface for indirect prompt injection attacks. \n
- Ingestion points: The skill ingests a 'Target skill' (external SKILL.md file) and 'Golden cases' (untrusted user inputs such as screenshots, prompts, or files) in the 'Inputs you need' section and Step 1. \n
- Boundary markers: Step 2 mandates 'blind' runs where the runner agent receives only the case input and is isolated from the main conversation context to prevent instructions from leaking between environments. \n
- Capability inventory: The skill spawns subagents, writes to a 'throwaway sandbox dir' on the filesystem, and executes shell commands for environment maintenance. \n
- Sanitization: The skill employs best-effort sandboxing and a mandatory post-run verification step where the agent must 'sweep the live checkout (git status) and clean anything the run leaked' to maintain host integrity.
- [COMMAND_EXECUTION]: The skill utilizes shell commands to manage and verify the state of the local development environment. \n
- Evidence: Step 2 includes instructions to 'sweep the live checkout (git status) and clean anything the run leaked' to ensure environment isolation and remediation of side effects.
Audit Metadata