prompt-injection-defense

Pass

Audited by Gen Agent Trust Hub on Sep 16, 2026

Risk Level: SAFEINDIRECT_PROMPT_INJECTIONPROMPT_INJECTION
Full Analysis
  • [SAFE]: The skill serves as a repository of security best practices and defense patterns. The identified content consists of instructional text and TypeScript code snippets used for validating inputs and sanitizing external data. The identified regex patterns are for defensive detection purposes.
  • [INDIRECT_PROMPT_INJECTION]: The skill provides a framework for defending against indirect injection attacks by defining an ingestion surface and mitigation strategies. 1. Ingestion points: Untrusted data is processed through methods like sanitizeExternalContent and processRetrievedDocs in references/patterns.md and references/sharp_edges.md. 2. Boundary markers: The skill implements isolation markers such as ---BEGIN EXTERNAL CONTENT--- and XML-like document tags to separate data from instructions. 3. Capability inventory: The skill describes a system with tool-calling capabilities constrained by allowed lists and user confirmation. 4. Sanitization: The skill provides multi-layered sanitization, including instruction-pattern neutralization, structural analysis, and homoglyph detection.
  • [PROMPT_INJECTION]: Phrases flagged by static analysis, such as instructions to ignore generic approaches in SKILL.md, are part of the grounding logic for reference system usage. The malicious-looking patterns in the reference files are regex literals for security detection logic and do not constitute an active attack.
Audit Metadata
Risk Level
SAFE
Analyzed
Sep 16, 2026, 04:40 PM
Security Audit — agent-trust-hub — prompt-injection-defense