guardrails-safety

Pass

Audited by Gen Agent Trust Hub on Jul 8, 2026

Risk Level: SAFE
Full Analysis
  • [PROMPT_INJECTION]: The skill contains numerous regex patterns and attack examples used for defensive classification and training. For example, references/input-guardrail-patterns.md includes heuristics for detecting 'DAN' jailbreaks and instruction overrides. These are functional components of the guardrail system.
  • [COMMAND_EXECUTION]: A SandboxExecutor implementation is provided in references/logic-guardrail-patterns.md that uses subprocess.run to execute commands in an isolated environment. This is presented as a security control for high-risk tool calls.
  • [SAFE]: Multiple reference files (e.g., references/architecture-patterns.md, references/deployment-pipelines.md) contain large volumes of repetitive technical text. While this appears to be auto-generated filler content, it does not contain executable or harmful logic.
Audit Metadata
Risk Level
SAFE
Analyzed
Jul 8, 2026, 12:24 PM
Security Audit — agent-trust-hub — guardrails-safety