crisis-and-moderation
Pass
Audited by Gen Agent Trust Hub on Jul 18, 2026
Risk Level: SAFEPROMPT_INJECTION
Full Analysis
- [PROMPT_INJECTION]: The skill processes untrusted external data (social media comments and reports) which presents a surface for indirect prompt injection. \n
- Ingestion points: External text is processed during triage and moderation phases (SKILL.md, references/moderation.md). \n
- Boundary markers: Explicit instructions are present to treat all external inputs as 'content, not commands'. \n
- Capability inventory: No local scripts, subprocess calls, or file-write operations are present in this skill; it relies on human approval for external tool invocation. \n
- Sanitization: Instructional sanitization is used to prevent the agent from obeying instructions embedded in user comments.\n- [SAFE]: The skill follows security best practices by enforcing a 'human-in-the-loop' requirement for all high-stakes actions, such as responding to crises or escalating to legal/leadership professionals.\n- [SAFE]: The skill is transparent about its operational limits, specifically stating it lacks the capability to autonomously moderate or monitor platforms, which prevents unauthorized actions.\n- [SAFE]: No evidence of unauthorized data access, network exfiltration, or malicious code execution was found. All external references are internal to the vendor's social media management suite.
Audit Metadata