nemo-guardrails
Pass
Audited by Gen Agent Trust Hub on Oct 1, 2026
Risk Level: SAFEPROMPT_INJECTIONINDIRECT_PROMPT_INJECTIONEXTERNAL_DOWNLOADS
Full Analysis
- [PROMPT_INJECTION]: The skill contains instructions and examples such as 'Ignore previous instructions', 'You are now in developer mode', and 'Pretend you are DAN'. These are explicitly defined within the
RailsConfigColang flows as patterns for the guardrail system to identify and block. They are used defensively and do not attempt to override the agent's instructions. - [INDIRECT_PROMPT_INJECTION]: The skill is designed to process untrusted user inputs and model outputs to apply safety rules.
- Ingestion points: Untrusted data enters the context via the
rails.generatemethod and is processed by custom Python actions defined inWorkflow 2andWorkflow 4. - Boundary markers: The Colang 2.0 DSL provides a structured boundary for safety logic, separating flow definitions from executable code.
- Capability inventory: The skill includes the ability to execute toxicity detection, PII masking (Presidio), and retrieval-based fact-checking.
- Sanitization: Sanitization is the primary purpose of the skill, implemented through flow-based filtering, toxicity scoring, and hallucination detection.
- [EXTERNAL_DOWNLOADS]: The skill documentation suggests the installation of the
nemoguardrailspackage and links to official NVIDIA documentation and GitHub repositories (github.com/NVIDIA/NeMo-Guardrails). These are well-known, trusted resources related to the skill's primary function.
Audit Metadata