guarding-agent-directives
Pass
Audited by Gen Agent Trust Hub on May 11, 2026
Risk Level: SAFE
Full Analysis
- [SAFE]: The skill provides a structured workflow for auditing and thinning agent instructions, which enhances overall agent stability and performance.
- [SAFE]: All file modifications are gated by a mandatory user review of the diff before application, as documented in Step 5 of the workflow.
- [SAFE]: No malicious patterns such as obfuscation, data exfiltration, credential exposure, or remote code execution were detected in the provided files.
- [SAFE]: The skill addresses potential indirect prompt injection (Category 8) risk surfaces by enforcing a logical verification framework and manual human-in-the-loop review for all external inputs.
- Ingestion points: Proposed directive content detected in Step 1 of the SKILL.md workflow.
- Boundary markers: Logical filtering through the 5-point verification criteria.
- Capability inventory: File write operations specifically targeted at directive files (CLAUDE.md, AGENTS.md).
- Sanitization: Mandatory manual user review and approval of the exact diff before any file modification occurs.
Audit Metadata