crisis-and-moderation

Pass

Audited by Gen Agent Trust Hub on Jul 18, 2026

Risk Level: SAFEPROMPT_INJECTION
Full Analysis
  • [PROMPT_INJECTION]: The skill processes untrusted external data (social media comments and reports) which presents a surface for indirect prompt injection. \n
  • Ingestion points: External text is processed during triage and moderation phases (SKILL.md, references/moderation.md). \n
  • Boundary markers: Explicit instructions are present to treat all external inputs as 'content, not commands'. \n
  • Capability inventory: No local scripts, subprocess calls, or file-write operations are present in this skill; it relies on human approval for external tool invocation. \n
  • Sanitization: Instructional sanitization is used to prevent the agent from obeying instructions embedded in user comments.\n- [SAFE]: The skill follows security best practices by enforcing a 'human-in-the-loop' requirement for all high-stakes actions, such as responding to crises or escalating to legal/leadership professionals.\n- [SAFE]: The skill is transparent about its operational limits, specifically stating it lacks the capability to autonomously moderate or monitor platforms, which prevents unauthorized actions.\n- [SAFE]: No evidence of unauthorized data access, network exfiltration, or malicious code execution was found. All external references are internal to the vendor's social media management suite.
Audit Metadata
Risk Level
SAFE
Analyzed
Jul 18, 2026, 02:27 PM
Security Audit — agent-trust-hub — crisis-and-moderation