reply-classification

Pass

Audited by Gen Agent Trust Hub on Apr 10, 2026

Risk Level: SAFE
Full Analysis
  • [PROMPT_INJECTION]: The skill documents a workflow for processing untrusted data (inbound email replies), which constitutes an indirect prompt injection surface. The documentation explicitly identifies this risk and provides appropriate mitigation strategies, such as human escalation triggers and sanitization requirements.
  • Ingestion points: Email subject and body content (SKILL.md).
  • Boundary markers: The skill recommends separating instructions from email content using structured prompts and delimiters.
  • Capability inventory: The skill is strictly instructional and does not include scripts or tools with executable capabilities.
  • Sanitization: The instructions emphasize the need for content sanitization and human oversight for sensitive or low-confidence intents.
Audit Metadata
Risk Level
SAFE
Analyzed
Apr 10, 2026, 07:38 PM
Security Audit — agent-trust-hub — reply-classification