crisis-detection-intervention-ai

Pass

Audited by Gen Agent Trust Hub on Sep 23, 2026

Risk Level: SAFEINDIRECT_PROMPT_INJECTION
Full Analysis
  • [INDIRECT_PROMPT_INJECTION]: The skill processes untrusted user input (e.g., journal entries, support chat messages) to classify mental health crises, which introduces a potential surface for indirect prompt injection.
  • Ingestion points: User-provided text is ingested via the CrisisDetector class in scripts/crisis_detector.ts and interpolated into LLM prompts in SKILL.md for classification.
  • Boundary markers: The prompt template in SKILL.md uses standard double-quote interpolation ("${text}") without specialized delimiters or explicit instructions for the model to disregard embedded commands.
  • Capability inventory: Crisis detection results are designed to trigger automated actions, including counselor notifications, SMS alerts for backup staff, and database logging of sensitive events.
  • Sanitization: There is no evidence of explicit sanitization or filtering of user-generated content before it is processed by the detection logic or the LLM prompt.
Audit Metadata
Risk Level
SAFE
Analyzed
Sep 23, 2026, 06:08 PM
Security Audit — agent-trust-hub — crisis-detection-intervention-ai