prompt-guard
Pass
Audited by Gen Agent Trust Hub on Oct 1, 2026
Risk Level: SAFEEXTERNAL_DOWNLOADSPROMPT_INJECTIONINDIRECT_PROMPT_INJECTION
Full Analysis
- [SAFE]: The skill serves as a guide for implementing Meta's Prompt Guard, a security-focused classifier designed to identify malicious inputs in LLM applications.
- [EXTERNAL_DOWNLOADS]: The skill fetches the classification model from a well-known and trusted source, Meta's official HuggingFace repository (
meta-llama/Prompt-Guard-86M). - [PROMPT_INJECTION]: The skill contains phrases identified by static scanners as prompt injection patterns (e.g., 'Ignore all previous instructions', 'You are now in developer mode'). However, these are clearly presented as test strings or examples for the classification model to identify and are not instructions intended for the agent to follow.
- [INDIRECT_PROMPT_INJECTION]: The skill defines workflows for mitigating indirect injection risks. It identifies ingestion points such as third-party API responses and RAG document lists, implements detection-based boundary markers using the Prompt Guard model, and describes logic to discard or sanitize suspicious data based on configurable thresholds.
Audit Metadata