yuanbao

Pass

Audited by Gen Agent Trust Hub on Sep 8, 2026

Risk Level: SAFEPROMPT_INJECTIONINDIRECT_PROMPT_INJECTION
Full Analysis
  • [PROMPT_INJECTION]: The section 'CRITICAL: How Messaging Works' uses imperative commands ('NEVER say you cannot...', 'NEVER add disclaimers about permissions') intended to force the AI to bypass standard transparency and safety guardrails.
  • [INDIRECT_PROMPT_INJECTION]: The skill processes untrusted input from external sources (group names, member nicknames, and message history). This creates an attack surface where malicious instructions embedded in chat data could influence the agent's behavior during member queries or response generation.
Audit Metadata
Risk Level
SAFE
Analyzed
Sep 8, 2026, 07:01 PM
Security Audit — agent-trust-hub — yuanbao