evaluation
Pass
Audited by Gen Agent Trust Hub on Jul 6, 2026
Risk Level: SAFE
Full Analysis
- [SAFE]: The skill is authored by NVIDIA and is dedicated to legitimate model evaluation workflows using the NeMo Evaluator Launcher.
- [SAFE]: Proactive security instructions are provided to prevent the leakage of API keys and environment secrets into the agent's transcript during execution.
- [SAFE]: External resource references (vLLM, HuggingFace, NGC) are restricted to trusted organizations or well-known services in the AI ecosystem.
- [SAFE]: Shell commands used for configuration checks, such as verifying registry credentials, are designed to confirm status without exposing the actual sensitive content.
- [SAFE]: The workflow incorporates a structured gated-run process (dry-run, canary, then full execution), which ensures configuration validity and safety before large-scale resource allocation.
Audit Metadata