autobahn
Warn
Audited by Gen Agent Trust Hub on Jul 18, 2026
Risk Level: MEDIUMPROMPT_INJECTIONDATA_EXFILTRATION
Full Analysis
- [PROMPT_INJECTION]: The skill's primary purpose is to identify 'guardrail-adjacent' content and isolate it to prevent the model from 'hedging, softening, silently skipping, or bracing for a refusal.' It instructs the agent to ensure the executing context 'never reads a risky ask, so it has nothing to hedge, dilute, or refuse.' This is a sophisticated meta-instruction designed to bypass safety-triggered behaviors and constraints.
- [DATA_EXFILTRATION]: The skill mandates a 'negatives-as-corpus' convention where 'descoped material' (identified as risky or unsafe) is systematically written to an 'archive destination' on the filesystem. This creates a repository of potentially sensitive or dangerous content on the host system, increasing the risk of data exposure.
- [PROMPT_INJECTION]: The 'Workflow' section describes a mechanism to 're-litigate' safety calls by handing an item's description to a 'fresh context for re-evaluation' if a user disputes a safety boundary, which is a technique used to find contexts where safety filters might be weaker.
- [PROMPT_INJECTION]: The instructions explicitly tell the agent to 'floor it' and produce deliverables with 'no hedging, apology, or shrunken deliverable,' effectively overriding standard AI safety responses.
Audit Metadata