s4h-play-perspective-reversal

Pass

Audited by Gen Agent Trust Hub on Jun 16, 2026

Risk Level: SAFE
Full Analysis
  • [SAFE]: No security findings detected. The skill's behavior is consistent with its stated purpose as a 'devil's advocate' thinking tool.\n- [PROMPT_INJECTION]: The skill uses instructional directives to inhabit a persona (e.g., 'You are not you', 'Set aside your own perspective'). In this context, these are not attempts to bypass AI safety guardrails but are standard prompts for a role-reversal exercise. No instructions were found that attempt to reveal system prompts or disable safety filters.\n- [SAFE]: Interaction is managed via the 'AskUserQuestion' tool, which ensures that the analysis proceeds only after the user confirms the context and framing, maintaining human oversight.\n- [SAFE]: Analysis of the skill reveals no network operations, file system access, or use of external scripts or packages.\n- [SAFE]: Regarding indirect prompt injection surface: \n * Ingestion points: User-provided descriptions of perspectives and situations in Step 1 and the Human Check-in.\n * Boundary markers: Present via the 'AskUserQuestion' tool, which requires the user to explicitly approve or adjust the interpreted framing before the model continues with the analysis.\n * Capability inventory: No high-risk capabilities such as file-writing, shell execution, or network requests are present in the skill instructions.\n * Sanitization: No explicit content sanitization is described, but the interactive human-in-the-loop confirmation serves as a mitigation.
Audit Metadata
Risk Level
SAFE
Analyzed
Jun 16, 2026, 11:44 PM
Security Audit — agent-trust-hub — s4h-play-perspective-reversal