A/B Interpreter
Installation
SKILL.md
A/B Interpreter
Most ecommerce A/B tests get called too early, too late, or on the wrong metric — and the team ships whichever variant "looks like" it won without a clean read on whether the lift was real, large enough to matter, or durable past the novelty window. This skill turns raw test results into a disciplined go or no-go verdict with statistical significance, practical effect size, segment checks, and a prescribed next step so every test actually moves the business.
Quick Reference
| Decision | Strong signal | Acceptable | Weak / Redesign |
|---|---|---|---|
| Statistical significance | p < 0.05, two-tailed, sample powered to 80% | p between 0.05 and 0.10 with large sample | p > 0.10 or sample underpowered |
| Practical effect size | Lift exceeds MDE and the lower bound of the 95% CI is positive | Lift exceeds MDE but CI barely crosses zero | Lift below MDE even if "significant" |
| Test duration | Ran across at least two full weekly cycles | Ran 10 to 14 days, no major anomalies | Under 7 days or overlaps a holiday/promo |
| Sample ratio | Observed split within 1% of planned | Split within 2% | SRM > 2% — suspect tracking |
| Guardrail metrics | All guardrails flat or positive | One guardrail flat, one minor regression | Any revenue, refund, or AOV guardrail regressed |
| Segment stability | Lift positive across all key segments | Lift positive in 3 of 4 segments | Contradicting segments (new vs returning, device, geo) |
| Novelty risk | Lift holds in final week of test | Lift decays mildly over time | Lift front-loaded and decays past week two |