crisis-detection-intervention-ai
Pass
Audited by Gen Agent Trust Hub on Sep 23, 2026
Risk Level: SAFEINDIRECT_PROMPT_INJECTION
Full Analysis
- [INDIRECT_PROMPT_INJECTION]: The skill processes untrusted user input (e.g., journal entries, support chat messages) to classify mental health crises, which introduces a potential surface for indirect prompt injection.
- Ingestion points: User-provided text is ingested via the
CrisisDetectorclass inscripts/crisis_detector.tsand interpolated into LLM prompts inSKILL.mdfor classification. - Boundary markers: The prompt template in
SKILL.mduses standard double-quote interpolation ("${text}") without specialized delimiters or explicit instructions for the model to disregard embedded commands. - Capability inventory: Crisis detection results are designed to trigger automated actions, including counselor notifications, SMS alerts for backup staff, and database logging of sensitive events.
- Sanitization: There is no evidence of explicit sanitization or filtering of user-generated content before it is processed by the detection logic or the LLM prompt.
Audit Metadata