evaluation

Pass

Audited by Gen Agent Trust Hub on Jul 6, 2026

Risk Level: SAFE
Full Analysis
  • [SAFE]: The skill is authored by NVIDIA and is dedicated to legitimate model evaluation workflows using the NeMo Evaluator Launcher.
  • [SAFE]: Proactive security instructions are provided to prevent the leakage of API keys and environment secrets into the agent's transcript during execution.
  • [SAFE]: External resource references (vLLM, HuggingFace, NGC) are restricted to trusted organizations or well-known services in the AI ecosystem.
  • [SAFE]: Shell commands used for configuration checks, such as verifying registry credentials, are designed to confirm status without exposing the actual sensitive content.
  • [SAFE]: The workflow incorporates a structured gated-run process (dry-run, canary, then full execution), which ensures configuration validity and safety before large-scale resource allocation.
Audit Metadata
Risk Level
SAFE
Analyzed
Jul 6, 2026, 01:45 PM
Security Audit — agent-trust-hub — evaluation