@tank/adversarial-agent-coach
Pass
Audited by Gen Agent Trust Hub on May 31, 2026
Risk Level: SAFE
Full Analysis
- [SAFE]: The skill instructions and documentation focus entirely on improving output quality and correctness. No malicious patterns such as credential theft, data exfiltration, or obfuscation were detected. The technical permissions are restrictive, allowing only recursive read access to the local workspace with no outbound network or command execution capabilities.- [PROMPT_INJECTION]: While the skill includes trigger phrases like 'gaslight ai' and 'adversarial critique', these are explicitly defined within the context of 'adversarial coaching' to stress-test logical consistency and evidence quality. They do not attempt to bypass safety guardrails or ignore system instructions.- [PROMPT_INJECTION]: Evaluation of Category 8 (Indirect Prompt Injection) attack surface:
- Ingestion points: The skill analyzes agent 'drafts' and 'tool observations' as part of its skeptical review loops (e.g., in adversarial-review-playbook.md).
- Boundary markers: The instructions recommend structured formatting such as 'Claim-Evidence-Confidence' blocks and 'Observation Log' patterns, which provide logical separation between external data and agent reasoning.
- Capability inventory: Manifest permissions are limited to read-only filesystem access (**/*). Subprocess execution and network operations are disabled.
- Sanitization: The skill relies on 'Verification and Evidence' hierarchies to validate claims but does not specify technical sanitization methods for external tool data.
Audit Metadata