user-generated-content

Pass

Audited by Gen Agent Trust Hub on Sep 18, 2026

Risk Level: SAFEINDIRECT_PROMPT_INJECTION
Full Analysis
  • [INDIRECT_PROMPT_INJECTION]: The skill defines workflows for the agent to monitor and curate content from Xiaohongshu users via hashtags, mentions, and direct messages. This creates a vulnerability surface where the agent could inadvertently process malicious instructions embedded within the user-generated text.
  • Ingestion points: SKILL.md Step 2 describes reviewing branded hashtags (e.g., #MyGlowUp), DM tags, and community mentions.
  • Boundary markers: The instructions do not specify the use of delimiters or 'ignore' instructions to isolate untrusted user content from the agent's core logic.
  • Capability inventory: The skill enables the agent to select content, communicate with users via DMs for permissions, and repost content to main feeds and stories.
  • Sanitization: The framework lacks explicit steps for sanitizing or validating ingested user text to prevent the execution of hidden instructions.
Audit Metadata
Risk Level
SAFE
Analyzed
Sep 18, 2026, 04:45 PM
Security Audit — agent-trust-hub — user-generated-content