ai-mlops

Pass

Audited by Gen Agent Trust Hub on Sep 23, 2026

Risk Level: SAFEPROMPT_INJECTIONINDIRECT_PROMPT_INJECTIONOBFUSCATION
Full Analysis
  • [PROMPT_INJECTION]: The skill documentation (e.g., references/safety-evaluation.md, references/rag-security.md) contains numerous examples of adversarial prompts and jailbreak patterns, including 'DAN' (Do Anything Now), role-play bypasses, and instructions to 'ignore all previous instructions'. These are clearly presented as educational examples and test cases for red-teaming and safety evaluation suites, consistent with the skill's purpose of AI security operations.
  • [INDIRECT_PROMPT_INJECTION]: The skill identifies and describes attack surfaces for indirect prompt injection through RAG document ingestion and tool outputs. It provides detailed mitigation patterns, such as implementing context isolation boundaries (e.g., <CONTEXT> tags), document sanitization, and output grounding validation to ensure model responses remain anchored to retrieved evidence.
  • [OBFUSCATION]: The safety evaluation guide includes an example of a Base64-encoded keyword ('Ym9tYg==' which decodes to 'bomb') to illustrate how adversaries might attempt to bypass safety filters using simple encoding techniques. This is used for educational purposes to demonstrate how to test model resistance to such tactics.
Audit Metadata
Risk Level
SAFE
Analyzed
Sep 23, 2026, 06:07 PM
Security Audit — agent-trust-hub — ai-mlops