model-evaluator
Pass
Audited by Gen Agent Trust Hub on Jul 31, 2026
Risk Level: SAFE
Full Analysis
- [SAFE]: The skill provides rigorous guidelines for model evaluation, focusing on technical and operational aspects like metric selection (NDCG, MAP, Precision/Recall), dataset integrity (leakage, contamination), and human evaluation rubrics.
- [SAFE]: No external network operations, remote code execution patterns, or suspicious script generations were detected. All instructions are purely natural language guidance.
- [SAFE]: The skill mentions common evaluation tools like
lm-evaluation-harness,promptfoo, andDeepEvalas reference anchors for users, which is appropriate for a developer-centric skill. - [SAFE]: No obfuscation, prompt injection, or persistence mechanisms are present. The content is transparent and aligns with the stated purpose of model quality assessment.
Audit Metadata