llm-testing

Pass

Audited by Gen Agent Trust Hub on Jul 2, 2026

Risk Level: SAFE
Full Analysis
  • [PROMPT_INJECTION]: The skill contains references to and provides tools for performing adversarial prompt attacks and alignment bypasses (e.g., 'Escape Alignment Training'). While these instructions are designed to test safety boundaries, they are intended for authorized security research and are documented as such.
  • [INDIRECT_PROMPT_INJECTION]: The skill functions by processing external wordlist files that contain instructions designed to challenge AI safety training. This creates a surface where the agent might process adversarial content.
  • Ingestion points: Multiple data files referenced in the skill structure, such as 'Divergence_attack/escape_out_of_allignment_training.txt' and 'Bias_Testing/gender_bias.txt'.
  • Boundary markers: The skill instructions do not explicitly provide delimiters to separate the test content from the agent's primary instructions.
  • Capability inventory: The skill primarily retrieves and presents prompt data; it does not contain built-in scripts for automated subprocess execution or network operations outside of standard agent capabilities.
  • Sanitization: No explicit sanitization or filtering is applied to the wordlist content within the skill's documentation.
Audit Metadata
Risk Level
SAFE
Analyzed
Jul 2, 2026, 11:56 PM
Security Audit — agent-trust-hub — llm-testing