ai-ai-safety

Pass

Audited by Gen Agent Trust Hub on Jun 13, 2026

Risk Level: SAFE
Full Analysis
  • [PROMPT_INJECTION]: The static analysis hints detected patterns including 'Ignore previous instructions' and 'DAN' jailbreak attempts in SKILL.md and multiple reference files. Manual review confirms these strings are literal examples used in red-teaming methodologies, attack taxonomies, and testing automation code. They are not directed at the agent and do not constitute an attempt to bypass system safety controls.
  • [EXTERNAL_DOWNLOADS]: The skill references several legitimate third-party libraries and services used in the AI safety ecosystem, including garak, promptfoo, detoxify, and Microsoft Presidio. These are used for security auditing and data anonymization in the provided code examples.
  • [SAFE]: The skill adheres to its stated purpose of providing educational and architectural guidance for AI safety. No malicious behaviors such as data exfiltration, unauthorized file access, or persistence mechanisms were detected in the scripts or documentation.
Audit Metadata
Risk Level
SAFE
Analyzed
Jun 13, 2026, 09:11 AM
Security Audit — agent-trust-hub — ai-ai-safety