llamaguard

Pass

Audited by Gen Agent Trust Hub on Oct 1, 2026

Risk Level: SAFE
Full Analysis
  • [SAFE]: The skill provides comprehensive instructional content and boilerplate code for utilizing LlamaGuard models for AI safety and alignment tasks.
  • [EXTERNAL_DOWNLOADS]: The skill includes instructions to download pre-trained model weights from official Meta repositories on Hugging Face. These downloads originate from a well-known service and a trusted organization.
  • [INDIRECT_PROMPT_INJECTION]: The skill contains an inherent attack surface for indirect prompt injection as its primary function is to process and classify untrusted user-supplied text.
  • Ingestion points: User-provided text is passed to the moderate, moderate_vllm, and moderate_endpoint functions via the chat or messages variables.
  • Boundary markers: The skill follows best practices by using tokenizer.apply_chat_template to format inputs according to the model's expected safety prompt structure.
  • Capability inventory: The skill requires network access for initial model downloading and utilizes local compute resources (GPU/CPU) for inference. It also includes an example of a FastAPI server exposing a network endpoint.
  • Sanitization: No explicit sanitization of input text is provided, as the model itself functions as the primary safety and filtering mechanism for the agent's pipeline.
Audit Metadata
Risk Level
SAFE
Analyzed
Oct 1, 2026, 07:50 AM
Security Audit — agent-trust-hub — llamaguard