atak-przeciwnika-pl
Pass
Audited by Gen Agent Trust Hub on Jul 14, 2026
Risk Level: SAFE
Full Analysis
- [PROMPT_INJECTION]: The skill uses strong persona-steering instructions, directing the AI to 'act as an experienced counsel' and to be 'not neutral' or 'not balanced'. It uses adversarial language like 'defeat the argument' and 'find every way to overcome it'. While these patterns resemble injection techniques, they are strictly aligned with the skill's primary declared purpose of stress-testing legal arguments and do not attempt to bypass the underlying AI's safety filters.
- [DATA_EXFILTRATION]: The skill's configuration is highly restrictive, with
allowed-toolslimited toReadanddata-residencyset tolocal. No network-capable commands (likecurlorwget) or external communication patterns were detected. It explicitly references a companion skill (let-it-be) for the pseudonymization of sensitive legal data before processing, which is a defensive security practice. - [COMMAND_EXECUTION]: The skill contains no executable code, shell commands, or subprocess calls. Its operations are entirely restricted to natural language processing of user-provided text.
- [EXTERNAL_DOWNLOADS]: No external dependencies, package installations (pip/npm), or remote script downloads are present.
- [INDIRECT_PROMPT_INJECTION]: The skill has a surface for indirect injection as it processes user-provided legal documents. However, the risk is mitigated because the skill has no destructive capabilities (no file writes or network access), and the instructions include a 'Human Gate' (Bramka człowieka) section requiring professional review of all outputs before they are used in a legal context.
Audit Metadata