user-generated-content
Pass
Audited by Gen Agent Trust Hub on Sep 18, 2026
Risk Level: SAFEINDIRECT_PROMPT_INJECTION
Full Analysis
- [INDIRECT_PROMPT_INJECTION]: The skill defines workflows for the agent to monitor and curate content from Xiaohongshu users via hashtags, mentions, and direct messages. This creates a vulnerability surface where the agent could inadvertently process malicious instructions embedded within the user-generated text.
- Ingestion points: SKILL.md Step 2 describes reviewing branded hashtags (e.g., #MyGlowUp), DM tags, and community mentions.
- Boundary markers: The instructions do not specify the use of delimiters or 'ignore' instructions to isolate untrusted user content from the agent's core logic.
- Capability inventory: The skill enables the agent to select content, communicate with users via DMs for permissions, and repost content to main feeds and stories.
- Sanitization: The framework lacks explicit steps for sanitizing or validating ingested user text to prevent the execution of hidden instructions.
Audit Metadata