output-sanitizer

Pass

Audited by Gen Agent Trust Hub on Sep 15, 2026

Risk Level: SAFEINDIRECT_PROMPT_INJECTION
Full Analysis
  • [INDIRECT_PROMPT_INJECTION]: The skill processes untrusted agent output to detect and redact sensitive data. This ingestion of external data creates a potential surface for indirect prompt injection, where malicious instructions embedded in the data could attempt to bypass redaction logic.
  • Ingestion points: The skill processes agent response text within the 'Sanitization Protocol' defined in SKILL.md.
  • Boundary markers: None identified; the skill treats input text as content to be scanned without explicit delimiters for instructions vs. data.
  • Capability inventory: The skill has file-read permissions but no network or shell access (SKILL.md), which limits the impact of potential injection attacks.
  • Sanitization: The skill focuses on redacting secrets and PII rather than sanitizing against instruction-based attacks within the input stream.
Audit Metadata
Risk Level
SAFE
Analyzed
Sep 15, 2026, 02:44 AM
Security Audit — agent-trust-hub — output-sanitizer