verification-before-completion
Pass
Audited by Gen Agent Trust Hub on Sep 4, 2026
Risk Level: SAFEINDIRECT_PROMPT_INJECTIONPROMPT_INJECTION
Full Analysis
- [INDIRECT_PROMPT_INJECTION]: The skill establishes a workflow that requires the agent to process and trust external data sources, creating a vulnerability surface.
- Ingestion points: Output from external commands such as test runners, linters, and build systems (SKILL.md).
- Boundary markers: Absent. The instructions do not provide delimiters or warnings to treat external output as untrusted or to ignore embedded instructions within that output.
- Capability inventory: The skill is used in context with subprocess execution, file system access, and version control operations (Commits/PRs).
- Sanitization: Absent. There are no instructions to validate or filter the content of the command output before the agent uses it to verify success.
- Risk: A malicious file being tested or linted could produce output that contains instructions the agent might follow, leading to unauthorized actions during the verification phase.
- [PROMPT_INJECTION]: The skill employs authoritative and coercive language common in prompt injection attacks to override the agent's default behavior.
- Evidence: The text uses terms like "The Iron Law", "non-negotiable", and "Violating the letter of this rule is violating the spirit of this rule."
- Behavioral Steering: It includes a direct threat ("If you lie, you'll be replaced") to enforce compliance, which mimics adversarial pressure tactics used to bypass standard operational guidelines.
Audit Metadata