experiment-results-interpreter
Installation
SKILL.md
experiment-results-interpreter
Overview
The ab-test-validity-checklist skill confirms that an experiment
was run correctly - clean SRM, honest peeking discipline, pre-declared
OEC. This skill covers the next question: given a valid experiment,
what does the result actually mean, and is it safe to ship?
The two most common failure modes at this stage, per Kohavi, Tang, Xu Trustworthy Online Controlled Experiments (Cambridge Univ. Press, 2020, ISBN 9781108724265), are:
- Shipping a result that is statistically significant but not practically meaningful.
- Shipping a result that will not persist because it reflects novelty, primacy, or interaction artefacts rather than genuine long-term user value.