model-evaluator

Pass

Audited by Gen Agent Trust Hub on Jul 31, 2026

Risk Level: SAFE
Full Analysis
  • [SAFE]: The skill provides rigorous guidelines for model evaluation, focusing on technical and operational aspects like metric selection (NDCG, MAP, Precision/Recall), dataset integrity (leakage, contamination), and human evaluation rubrics.
  • [SAFE]: No external network operations, remote code execution patterns, or suspicious script generations were detected. All instructions are purely natural language guidance.
  • [SAFE]: The skill mentions common evaluation tools like lm-evaluation-harness, promptfoo, and DeepEval as reference anchors for users, which is appropriate for a developer-centric skill.
  • [SAFE]: No obfuscation, prompt injection, or persistence mechanisms are present. The content is transparent and aligns with the stated purpose of model quality assessment.
Audit Metadata
Risk Level
SAFE
Analyzed
Jul 31, 2026, 11:46 AM
Security Audit — agent-trust-hub — model-evaluator