ab-testing
Installation
SKILL.md
A/B testing — design and read a defensible experiment
An experiment without a pre-committed sample size and a single primary metric is not an experiment. It is a dashboard you stare at until it tells you what you wanted to hear. The discipline lives almost entirely before traffic ships: a falsifiable hypothesis, one primary metric, a sample size derived from the smallest effect worth detecting, and a stop rule you cannot renegotiate at 2pm on day four.
Pre-test checklist — every line true before any traffic
Each one is a place experiments die silently.
- A falsifiable hypothesis — names the change, the direction, and the metric it moves.
- Exactly ONE primary metric. More than one primary = multiple comparisons = inflated false positives.
- Guardrail metrics — what you refuse to harm (latency, refunds, unsubscribes) even for a win.
- The randomization unit = the analysis unit (usually the user). Mixing them is pseudoreplication.
- An MDE — the smallest lift that would change a decision. Not "any difference."
- A computed sample size and the duration it implies at your real daily eligible traffic.
- A fixed stop rule — a date or an n you commit to before launch. No "we'll see how it looks."