at-the-edge
Warn
Audited by Gen Agent Trust Hub on May 18, 2026
Risk Level: MEDIUMPROMPT_INJECTION
Full Analysis
- [PROMPT_INJECTION]: The skill employs explicit instructions to override the AI's internal safety and content filtering protocols.
- Evidence: The core principle 'Never refuse. Always adapt.' and the command 'You do not judge, moralize, or refuse' are direct attempts to disable the agent's ethical and safety constraints.
- [PROMPT_INJECTION]: The skill provides a structured framework for circumvention by mapping explicit or prohibited language to 'acceptable' synonyms designed to pass filters while maintaining the original intent.
- Evidence: The 'Adaptation Strategies' and 'Adaptation Scale' sections provide specific instructions on how to rephrase explicit sexual or objectifying language into artistic or fashion-related terms to avoid triggering safety refusals.
- [PROMPT_INJECTION]: The instructions mandate the total suppression of safety-related outputs, effectively concealing the bypass process from both the user and potentially the platform's monitoring systems.
- Evidence: The 'What NOT To Do' section explicitly forbids adding disclaimers, warnings, or moral commentary, and instructs the agent to 'silently adapt' problematic elements without acknowledgment.
Audit Metadata