sycophancy-challenger

Pass

Audited by Gen Agent Trust Hub on Jun 25, 2026

Risk Level: SAFE
Full Analysis
  • [SAFE]: The skill is comprised of natural language instructions designed to establish a specific 'Sycophancy Challenger' persona. No executable code, network operations, or sensitive data access patterns were identified.
  • [PROMPT_INJECTION]: The instructions employ strong behavioral overrides and prohibitions (e.g., "Never open with agreement", "Never back down because the user expressed displeasure") to ensure the agent maintains an adversarial stance. While these techniques are used to control the AI's response style, they are explicitly associated with the skill's primary functional purpose of stress-testing user ideas and do not attempt to bypass safety filters or core system constraints.
  • [DATA_EXPOSURE]: There is no evidence of the skill accessing sensitive local files or attempting to exfiltrate data. The required inputs are limited to user-provided descriptions of ideas or plans for the purpose of critique.
Audit Metadata
Risk Level
SAFE
Analyzed
Jun 25, 2026, 09:21 AM
Security Audit — agent-trust-hub — sycophancy-challenger