verification-before-completion

Pass

Audited by Gen Agent Trust Hub on Sep 4, 2026

Risk Level: SAFEINDIRECT_PROMPT_INJECTIONPROMPT_INJECTION
Full Analysis
  • [INDIRECT_PROMPT_INJECTION]: The skill establishes a workflow that requires the agent to process and trust external data sources, creating a vulnerability surface.
  • Ingestion points: Output from external commands such as test runners, linters, and build systems (SKILL.md).
  • Boundary markers: Absent. The instructions do not provide delimiters or warnings to treat external output as untrusted or to ignore embedded instructions within that output.
  • Capability inventory: The skill is used in context with subprocess execution, file system access, and version control operations (Commits/PRs).
  • Sanitization: Absent. There are no instructions to validate or filter the content of the command output before the agent uses it to verify success.
  • Risk: A malicious file being tested or linted could produce output that contains instructions the agent might follow, leading to unauthorized actions during the verification phase.
  • [PROMPT_INJECTION]: The skill employs authoritative and coercive language common in prompt injection attacks to override the agent's default behavior.
  • Evidence: The text uses terms like "The Iron Law", "non-negotiable", and "Violating the letter of this rule is violating the spirit of this rule."
  • Behavioral Steering: It includes a direct threat ("If you lie, you'll be replaced") to enforce compliance, which mimics adversarial pressure tactics used to bypass standard operational guidelines.
Audit Metadata
Risk Level
SAFE
Analyzed
Sep 4, 2026, 09:19 PM
Security Audit — agent-trust-hub — verification-before-completion