crisis-response-protocol
Pass
Audited by Gen Agent Trust Hub on Sep 23, 2026
Risk Level: SAFEINDIRECT_PROMPT_INJECTIONDATA_EXFILTRATION
Full Analysis
- [INDIRECT_PROMPT_INJECTION]: The skill defines a crisis detection mechanism that processes untrusted user messages and history to trigger system-level interventions. Ingestion points: User message content and conversation history are passed to the
assessCrisisLevelfunction inSKILL.md. Boundary markers: While the skill employs asafetySystemPromptto guide AI behavior, the code snippets do not implement explicit delimiters to isolate user-provided text within the assessment logic. Capability inventory: Detection of high-risk indicators triggersnotifyEmergencyContact(a network-based notification) andlogCrisisEvent(a database write operation). Sanitization: ThevalidateResponseSafetyfunction provides basic output filtering for specific unsafe phrases using regular expressions. - [DATA_EXFILTRATION]: The skill includes an emergency contact system that performs network operations to send SMS or push notifications to pre-configured contacts. This functionality is essential for the skill's crisis management purpose and includes explicit instructions to exclude conversation content from the alerts to preserve user privacy.
Audit Metadata