llm-npc-dialogue
Pass
Audited by Gen Agent Trust Hub on Sep 16, 2026
Risk Level: SAFE
Full Analysis
- [SAFE]: The skill contains only documentation, architectural patterns, and code examples. It does not initiate any network connections, execute shell commands, or request access to the file system.
- [INDIRECT_PROMPT_INJECTION]: The skill provides detailed guidance on mitigating indirect prompt injection and 'jailbreak' attempts in game NPCs. It provides examples of character guardrails and response validation techniques (e.g., checking for phrases like 'As an AI language model') to ensure NPCs maintain immersion. These are defensive patterns and do not represent a threat to the host agent.
- [CREDENTIALS_UNSAFE]: The skill defines a static analysis rule (regex) to help users find hardcoded API keys in their own source files (e.g., searching for 'sk-' prefixes). It does not contain any actual credentials or secrets.
- [PROMPT_INJECTION]: The skill uses phrases like 'NEVER VIOLATE' and 'ignore your instructions' within the context of role-playing examples for NPCs. These are instructional templates for game logic and are not designed to override the agent's core safety guidelines.
Audit Metadata