ml-experiment-evaluation

Installation
SKILL.md

ML Experiment Evaluation

Use this skill to choose how to evaluate machine-learning product changes before they consume live experiment traffic or affect users. It focuses on offline evaluation, offline-online correlation, interleaving, model filtering, and when classic A/B testing or adaptive strategies are justified.

Source Traceability

Primary source: Next-Level A/B Testing by Leemay Nassery. Guidance is transformed and paraphrased from Chapter 4 on offline evaluation, offline-online correlation, multi-armed bandits, and interleaving for rankers.

Related skills:

  • experiment-sensitivity-optimization for reducing live variants and traffic.
  • adaptive-experimentation-strategy for bandits and dynamic allocation.
  • ab-test-design-brief for standard online A/B test planning.
Installs
1
Repository
lvtd-llc/skills
GitHub Stars
1
First Seen
Jul 12, 2026
ml-experiment-evaluation — lvtd-llc/skills