skills/skills.volces.com/content-safety-guard

content-safety-guard

Installation
SKILL.md

Content Safety Guard

A production-tested dual-layer AI content guardrail for chatbots and AI agents. Intercepts outbound messages before delivery and evaluates them through a judge model — with a complete red-team test methodology to verify your guardrail actually works.

Blue ocean skill: As of publication, no equivalent exists on ClawHub. Most AI safety tooling focuses on input filtering; this pattern guards the output layer — what the AI sends to your users.


The Core Pattern: Dual-Layer Defense

Installs
5
First Seen
Apr 21, 2026
content-safety-guard from skills.volces.com