safety-guard

Pass

Audited by Gen Agent Trust Hub on Sep 12, 2026

Risk Level: SAFE
Full Analysis
  • [INDIRECT_PROMPT_INJECTION]: The skill is designed to process and intercept tool inputs such as shell commands and file paths. While this constitutes an attack surface for indirect prompt injection, the skill's logic is explicitly defensive.
  • Ingestion points: Intercepts strings passed to the Bash tool and file paths passed to Write, Edit, and MultiEdit (SKILL.md).
  • Boundary markers: Implementation relies on 'PreToolUse' hooks to evaluate inputs before tool execution.
  • Capability inventory: The skill manages interactions with Bash, Write, Edit, and MultiEdit tools.
  • Sanitization: Inputs are compared against a 'Watched patterns' list of destructive commands and directory allow-lists to enforce constraints.
  • [DATA_EXPOSURE]: The skill documentation mentions logging blocked actions to ~/.claude/safety-guard.log. This is a standard practice for diagnostic logging in local development environments and does not indicate unauthorized data exfiltration.
Audit Metadata
Risk Level
SAFE
Analyzed
Sep 12, 2026, 03:40 PM
Security Audit — agent-trust-hub — safety-guard