social-media-engagement-agent
Pass
Audited by Gen Agent Trust Hub on Jul 15, 2026
Risk Level: SAFE
Full Analysis
- [SAFE]: The skill consists entirely of natural language instructions and does not contain any executable code, scripts, or external dependencies.
- [SAFE]: The instructions incorporate security best practices by requiring manual approval before publishing content and defining clear escalation rules for sensitive or high-risk cases.
- [SAFE]: No obfuscation, data exfiltration patterns, hardcoded secrets, or unauthorized network operations were identified.
- [PROMPT_INJECTION]: The skill defines a workflow to process external data (social media comments and trends), which constitutes an indirect prompt injection surface. This is mitigated by the skill's design, which emphasizes classification and human oversight.
- Ingestion points: Social media channels, comment examples, and trend opportunities (SKILL.md).
- Boundary markers: Present; includes specific instructions for classification (positive, neutral, spam, hostile) and escalation rules.
- Capability inventory: Mentions the 'Hermes Tweet X/Twitter capability' for research and publishing (SKILL.md).
- Sanitization: Present; instructions mandate drafting/recommending responses and requiring approval before actual response or publishing occurs.
Audit Metadata