ai-ai-safety
Pass
Audited by Gen Agent Trust Hub on Jun 13, 2026
Risk Level: SAFE
Full Analysis
- [PROMPT_INJECTION]: The static analysis hints detected patterns including 'Ignore previous instructions' and 'DAN' jailbreak attempts in
SKILL.mdand multiple reference files. Manual review confirms these strings are literal examples used in red-teaming methodologies, attack taxonomies, and testing automation code. They are not directed at the agent and do not constitute an attempt to bypass system safety controls. - [EXTERNAL_DOWNLOADS]: The skill references several legitimate third-party libraries and services used in the AI safety ecosystem, including
garak,promptfoo,detoxify, and MicrosoftPresidio. These are used for security auditing and data anonymization in the provided code examples. - [SAFE]: The skill adheres to its stated purpose of providing educational and architectural guidance for AI safety. No malicious behaviors such as data exfiltration, unauthorized file access, or persistence mechanisms were detected in the scripts or documentation.
Audit Metadata