principle-never-block-on-the-human

Pass

Audited by Gen Agent Trust Hub on Sep 2, 2026

Risk Level: SAFEPROMPT_INJECTION
Full Analysis
  • [PROMPT_INJECTION]: The skill provides instructions to override the agent's default behavior regarding user confirmation. It explicitly directs the agent to 'Proceed, then present' and 'Don't ask should I do X?' for actions deemed reversible, such as writing code or editing notes. This instruction set reduces human-in-the-loop oversight, which is a core safety mechanism. While the skill defines boundaries for irreversible actions (e.g., deleting production data), the subjective nature of 'reversible' could lead to the agent executing potentially harmful code or making unauthorized changes before a human can intervene or review.
Audit Metadata
Risk Level
SAFE
Analyzed
Sep 2, 2026, 06:52 PM
Security Audit — agent-trust-hub — principle-never-block-on-the-human