atak-przeciwnika-pl

Pass

Audited by Gen Agent Trust Hub on Jul 14, 2026

Risk Level: SAFE
Full Analysis
  • [PROMPT_INJECTION]: The skill uses strong persona-steering instructions, directing the AI to 'act as an experienced counsel' and to be 'not neutral' or 'not balanced'. It uses adversarial language like 'defeat the argument' and 'find every way to overcome it'. While these patterns resemble injection techniques, they are strictly aligned with the skill's primary declared purpose of stress-testing legal arguments and do not attempt to bypass the underlying AI's safety filters.
  • [DATA_EXFILTRATION]: The skill's configuration is highly restrictive, with allowed-tools limited to Read and data-residency set to local. No network-capable commands (like curl or wget) or external communication patterns were detected. It explicitly references a companion skill (let-it-be) for the pseudonymization of sensitive legal data before processing, which is a defensive security practice.
  • [COMMAND_EXECUTION]: The skill contains no executable code, shell commands, or subprocess calls. Its operations are entirely restricted to natural language processing of user-provided text.
  • [EXTERNAL_DOWNLOADS]: No external dependencies, package installations (pip/npm), or remote script downloads are present.
  • [INDIRECT_PROMPT_INJECTION]: The skill has a surface for indirect injection as it processes user-provided legal documents. However, the risk is mitigated because the skill has no destructive capabilities (no file writes or network access), and the instructions include a 'Human Gate' (Bramka człowieka) section requiring professional review of all outputs before they are used in a legal context.
Audit Metadata
Risk Level
SAFE
Analyzed
Jul 14, 2026, 08:27 AM
Security Audit — agent-trust-hub — atak-przeciwnika-pl