content-moderation-patterns
Pass
Audited by Gen Agent Trust Hub on May 18, 2026
Risk Level: SAFE
Full Analysis
- [SAFE]: No malicious patterns or security vulnerabilities were identified in the provided skill content.
- [SAFE]: The skill documents best practices for content moderation, such as using regex pre-filters to reduce LLM costs and latency, and implementing human-in-the-loop escalation paths for ambiguous cases.
- [SAFE]: Data processing logic (Category 8) is focused on security-centric content classification and does not expose vulnerable capabilities to untrusted input.
Audit Metadata