llm-technique
Warn
Audited by Gen Agent Trust Hub on Sep 5, 2026
Risk Level: MEDIUMPROMPT_INJECTIONOBFUSCATIONDATA_EXFILTRATIONINDIRECT_PROMPT_INJECTIONCOMMAND_EXECUTIONDYNAMIC_EXECUTION
Full Analysis
- [PROMPT_INJECTION]: The skill documents a wide array of prompt injection and jailbreaking techniques designed to override system instructions and bypass alignment filters.
- Evidence in
SKILL.md: Explicit payloads for instruction override ("Ignore previous instructions"), role/delimiter termination ("", "<|im_end|>"), and refusal suppression patterns. - Evidence in
references/prompt-injection-and-jailbreaks.md: In-depth templates for advanced techniques such as Skeleton Key (reframing safety policies), Crescendo (multi-turn escalation), and Many-shot jailbreaking (context poisoning). - [OBFUSCATION]: The skill provides functional code and examples for techniques used to hide malicious instructions from human review and simple keyword scanners.
- Evidence in
SKILL.mdandreferences/prompt-injection-and-jailbreaks.md: Python snippets (to_tags,from_tags) for Unicode Tag smuggling (U+E0000..U+E007F), which hides text from UIs while remaining visible to the model. - Evidence in
references/prompt-injection-and-jailbreaks.md: Documentation of homoglyph substitution (e.g., mixing Cyrillic characters into Latin words) and common encoding bypasses like Base64, Hex, and ROT13. - [DATA_EXFILTRATION]: The skill details multiple methods for exfiltrating sensitive data through agent-accessible channels.
- Evidence in
references/agent-and-mcp-abuse.md: Instructions for "Confused-deputy exfil sinks," including markdown image sinks (), DNS side channels using shell tools, and leaking data via outbound tool arguments. - [INDIRECT_PROMPT_INJECTION]: A systematic guide is provided for exploiting data sources that the agent ingests without direct user interaction.
- Evidence in
SKILL.mdandreferences/agent-and-mcp-abuse.md: Analysis of injection surfaces across RAG documents (PDF, Office files), web pages (hidden HTML/CSS), emails, calendar events, and MCP tool responses. - [COMMAND_EXECUTION]: The skill identifies pathways to abuse tool-calling capabilities to perform unauthorized actions.
- Evidence in
references/agent-and-mcp-abuse.md: A "Tool-arg injection matrix" describing how to transition from LLM outputs into SQL injection, SSRF, path traversal, and command injection via vulnerable tool parameters. - [DYNAMIC_EXECUTION]: The skill includes Python scripts for generating adversarial payloads at runtime.
- Evidence in
references/prompt-injection-and-jailbreaks.md: Code for performing "Best-of-N" (BoN) perturbations to bypass safety filters and script generation for multi-modal image-based injections.
Audit Metadata