model-evaluator

Installation
SKILL.md

Model Evaluator

Measure model quality honestly, surface failure modes early, and turn scores into decisions.

What this skill owns

Use this skill to:

  • design evaluation plans for classical ML, ranking, recommender, forecasting, and LLM systems
  • choose metrics, baselines, holdouts, and gating criteria
  • audit datasets, leakage risk, label quality, and benchmark validity
  • analyze errors, cohort slices, calibration, fairness, robustness, and drift
  • interpret online experiments, human review results, and operational quality signals
  • produce go/no-go recommendations, scope restrictions, or follow-up eval roadmaps

What this skill does not own

Installs
1
Repository
heymoezy/porter
First Seen
Jul 31, 2026
model-evaluator — heymoezy/porter