copywriting-neta-injection
Pass
Audited by Gen Agent Trust Hub on Jun 17, 2026
Risk Level: SAFEPROMPT_INJECTIONDATA_EXFILTRATION
Full Analysis
- [PROMPT_INJECTION]: The skill is susceptible to Indirect Prompt Injection (Category 8) due to its reliance on external data retrieved from social media platforms (X, Reddit, TikTok). Malicious content on these platforms could be engineered to influence the agent's behavior or output when processed during the copywriting phases.\n
- Ingestion points: Public cultural references and 'neta' retrieved via the WebSearch tool (referenced in
protocols/copy-neta-injection.md).\n - Boundary markers: None identified; the skill processes untrusted search results to extract structural templates without explicit separation between external data and system instructions.\n
- Capability inventory: The skill generates content candidates that are woven into drafts by other automated drafting agents.\n
- Sanitization: No input sanitization is documented; the skill instead relies on a post-generation 'Neta Safety Gate' rubric for evaluative review.\n- [DATA_EXFILTRATION]: The skill constructs and executes search queries that include details from the user's brief, such as the product description and target audience. This behavior transmits potentially sensitive business intent to third-party search engines through the WebSearch tool. This is a standard risk for skills requiring real-time cultural context retrieval.\n- [SAFE]: The skill utilizes highly reputable and trusted literary and academic archives for reference verification, including Project Gutenberg, the National Diet Library, and Aozora Bunko. These sources are correctly identified as canonical grounding for literary allusions.
Audit Metadata