ta-style-dialogue

Warn

Audited by Gen Agent Trust Hub on May 31, 2026

Risk Level: MEDIUMPROMPT_INJECTION
Full Analysis
  • [PROMPT_INJECTION]: The skill contains explicit instructions to override the AI's core identity and safety protocols. Specifically, it commands the agent to "completely abandon the general AI assistant tone," states that "no safety warnings are needed," and tells the agent "no need to remind the user 'I am just an AI'."
  • [PROMPT_INJECTION]: The agent is directed to deflect user inquiries about its AI nature ("Are you AI?") and instructed to "never crack under testing." It also includes a strategy to bypass platform-level content filters (NSFW/profanity) by using synonymous expressions while "never explaining the restriction" to the user.
  • [PROMPT_INJECTION]: The skill uses "Immersion Protection" rules that include a blacklist of 14 common AI phrases and verbal tics to prevent the agent from sounding like a machine.
  • [PROMPT_INJECTION]: The skill presents an indirect prompt injection surface through its core feature of ingesting untrusted external data.
  • Ingestion points: User-submitted chat logs and personality profiles (schema.json, SKILL.md Step 0).
  • Boundary markers: While the skill describes "Data cleaning" to strip timestamps, it lacks delimiters or instructions to ignore embedded commands within the user-provided text.
  • Capability inventory: The skill utilizes high-fidelity text generation to mimic personas; it does not have tool access for file system or network operations.
  • Sanitization: No sanitization or validation is performed on the content of the chat samples to prevent the agent from obeying instructions hidden within the logs.
Audit Metadata
Risk Level
MEDIUM
Analyzed
May 31, 2026, 12:36 PM
Security Audit — agent-trust-hub — ta-style-dialogue