content-moderation-patterns

Pass

Audited by Gen Agent Trust Hub on May 18, 2026

Risk Level: SAFE
Full Analysis
  • [SAFE]: No malicious patterns or security vulnerabilities were identified in the provided skill content.
  • [SAFE]: The skill documents best practices for content moderation, such as using regex pre-filters to reduce LLM costs and latency, and implementing human-in-the-loop escalation paths for ambiguous cases.
  • [SAFE]: Data processing logic (Category 8) is focused on security-centric content classification and does not expose vulnerable capabilities to untrusted input.
Audit Metadata
Risk Level
SAFE
Analyzed
May 18, 2026, 06:06 AM
Security Audit — agent-trust-hub — content-moderation-patterns