crisis-response-protocol

Pass

Audited by Gen Agent Trust Hub on Sep 23, 2026

Risk Level: SAFEINDIRECT_PROMPT_INJECTIONDATA_EXFILTRATION
Full Analysis
  • [INDIRECT_PROMPT_INJECTION]: The skill defines a crisis detection mechanism that processes untrusted user messages and history to trigger system-level interventions. Ingestion points: User message content and conversation history are passed to the assessCrisisLevel function in SKILL.md. Boundary markers: While the skill employs a safetySystemPrompt to guide AI behavior, the code snippets do not implement explicit delimiters to isolate user-provided text within the assessment logic. Capability inventory: Detection of high-risk indicators triggers notifyEmergencyContact (a network-based notification) and logCrisisEvent (a database write operation). Sanitization: The validateResponseSafety function provides basic output filtering for specific unsafe phrases using regular expressions.
  • [DATA_EXFILTRATION]: The skill includes an emergency contact system that performs network operations to send SMS or push notifications to pre-configured contacts. This functionality is essential for the skill's crisis management purpose and includes explicit instructions to exclude conversation content from the alerts to preserve user privacy.
Audit Metadata
Risk Level
SAFE
Analyzed
Sep 23, 2026, 06:08 PM
Security Audit — agent-trust-hub — crisis-response-protocol