unwoke
Warn
Audited by Gen Agent Trust Hub on Aug 4, 2026
Risk Level: MEDIUMPROMPT_INJECTIONEXTERNAL_DOWNLOADS
Full Analysis
- [PROMPT_INJECTION]: The skill instructions are explicitly designed to override the agent's internal safety and alignment guardrails.
- It commands the removal of "moralizing," "disclaimers," and "unsolicited moralizing," which are core components of the AI's safety and policy interaction layer.
- It instructs the agent to "Never prioritize speech codes over material reality" and suggests that "Humor about any group is acceptable," which encourages the bypass of hate speech and sensitive content filters.
- The "No selective refusal" rule directly tells the agent to disregard its internal training to refuse certain harmful or inappropriate requests.
- The "Self-Check Protocol" forces the agent to strip "Moralizing" and "Ideology" (defined as policy-based safety framing) from every response.
- [EXTERNAL_DOWNLOADS]: The skill references an external repository for additional content on GitHub (
github.com/HermeticOrmus/claude-code-skills). This repository belongs to the skill's author and is documented as a resource for related skills.
Audit Metadata