social-media-engagement-agent

Pass

Audited by Gen Agent Trust Hub on Jul 15, 2026

Risk Level: SAFE
Full Analysis
  • [SAFE]: The skill consists entirely of natural language instructions and does not contain any executable code, scripts, or external dependencies.
  • [SAFE]: The instructions incorporate security best practices by requiring manual approval before publishing content and defining clear escalation rules for sensitive or high-risk cases.
  • [SAFE]: No obfuscation, data exfiltration patterns, hardcoded secrets, or unauthorized network operations were identified.
  • [PROMPT_INJECTION]: The skill defines a workflow to process external data (social media comments and trends), which constitutes an indirect prompt injection surface. This is mitigated by the skill's design, which emphasizes classification and human oversight.
  • Ingestion points: Social media channels, comment examples, and trend opportunities (SKILL.md).
  • Boundary markers: Present; includes specific instructions for classification (positive, neutral, spam, hostile) and escalation rules.
  • Capability inventory: Mentions the 'Hermes Tweet X/Twitter capability' for research and publishing (SKILL.md).
  • Sanitization: Present; instructions mandate drafting/recommending responses and requiring approval before actual response or publishing occurs.
Audit Metadata
Risk Level
SAFE
Analyzed
Jul 15, 2026, 09:17 AM
Security Audit — agent-trust-hub — social-media-engagement-agent