A/B tests and the cost of peeking
Your new ranking model shows +1.2% click-through after three days of an A/B test, with p = 0.03. Do you ship it?
A p-value means what it says only under the plan it was computed for: a sample size fixed in advance, one look, one metric, and a correct randomization. Check a running test every day and stop the first time p dips below 0.05, and "p < 0.05" happens far more often than 5% of the time when nothing changed. This lesson builds the p-value from scratch, then shows exactly how that breaks.
Run tests where nothing changed
An A/A test shows both groups the same system. Any difference you see is noise. Before reading p-values, look at the noise.
About 5%. That's what the 0.05 threshold is: the fraction of no-difference experiments that look at least this extreme by chance. The p-value is the probability, if nothing changed, of a difference as large as the one observed. Run them:
z = observed difference ÷ its standard error. The shaded tails beyond ±1.96 hold 5% of a standard normal: those tests would be called "significant".
How much does a click rate wobble?
Each user either clicks or doesn't: a Bernoulli variable with variance \(p(1-p)\). The average of \(n\) independent ones has variance \(p(1-p)/n\). The difference of two independent averages adds their variances:
SE falls like \(1/\sqrt n\). Halving it takes 4× the data, and detecting an effect half as large takes 4× the traffic. That square law is the whole economics of experimentation.
Can this test even see the effect?
Power is the probability of detecting a real effect of a given size. To reach it, the effect must stand about \(z_{1-\alpha/2} + z_{\text{power}}\) standard errors clear of zero, which is 1.96 + 0.84 = 2.8 for α = 0.05 and 80% power:
\(7.85 \times (0.0475 + 0.0484) / 0.001^2 \approx 753{,}000\). Small relative lifts on small rates need enormous samples. An underpowered test mostly returns "not significant", and when it does come back significant it overstates the effect (the winner's curse). Try your own numbers:
Ship or not
Eight experiments, 2,000 users per arm per day, a 5% base click rate. Some variants are genuinely +10% better; others are identical to control. Watch the results day by day and decide whenever you like: ship B, or keep A. The planned length is 14 days, and the test was sized for that.
New cases, no hints
A chi-square test of the split gives \(\chi^2 = 32.4\), p ≈ \(10^{-8}\). That's a sample ratio mismatch. Something (a crash, a redirect, bot filtering, logging) removes users from one arm, and probably not at random, so every metric comparison is suspect. Check SRM before reading any metric.
The independent unit is the user, not the impression. Heavy users contribute many correlated impressions, so the naive SE is too small and p-values come out too small. Use the delta method for a ratio metric, or a bootstrap that resamples users. (The third option is also true of a ratio metric, but it's a question of which quantity you want, not an error in the test.)
\(1 - 0.95^{20} \approx 0.64\). Pick one primary decision metric in advance, treat the rest as guardrails or exploration, and correct for multiplicity (Bonferroni, or Benjamini–Hochberg for false discovery rate) when you must test many.
Short tests overstate novelty (and understate learning effects, the opposite pattern). Plot the effect by exposure day or by user cohort, run long enough to see it flatten, and keep a long-term holdback to measure the durable effect.
Write the test yourself
Standard library only: math.erfc gives normal tail probabilities, and statistics.NormalDist().inv_cdf gives quantiles. ⌘/Ctrl+Enter runs.
Exercise 1: two-proportion z-test
Exercise 2: sample size and a sample-ratio check
Answer out loud, then check
- Sequential designs: group-sequential tests with alpha spending (e.g. O'Brien–Fleming boundaries that are strict early), or always-valid p-values and confidence sequences (mSPRT).
- Or pre-commit a fixed horizon from a power calculation and decide only at the end. Dashboards can show data without showing "significance".
- Bonferroni across planned looks is valid but conservative.
- CUPED: regress out a pre-experiment covariate (the same metric in the prior period), \(Y' = Y - \theta(X - \bar X)\) with \(\theta = \mathrm{cov}(X,Y)/\mathrm{var}(X)\). The variance falls by the factor \(1-\rho^2\), with no bias because X is pre-treatment.
- Also: more sensitive metrics (proxies), trimming or capping heavy-tailed metrics, stratification, and triggering analysis only on users actually exposed to the change.
- Offline labels come from logs produced by the old ranker: position and exposure bias, and items it never showed have no labels.
- The metric mismatch: NDCG on logged clicks vs the business outcome (dwell, retention). Candidate generation limits what any ranker can surface.
- Train/serve skew in features, latency changes, a novelty baseline, or an underpowered test. Diagnose with counterfactual or interleaving evaluation, slice analysis, and a check that serving features match training.
- When treated and control users interact: marketplaces sharing inventory or budget, social feeds, ride-hailing supply. Treatment leaks into control (interference, a SUTVA violation).
- Use cluster randomization (by region or social graph cluster), switchback designs (alternate by time block), or budget-split or two-sided designs in marketplaces.
- The price is fewer independent units, so larger variance. Plan power accordingly.
- Agree in advance on an overall evaluation criterion and guardrail thresholds: latency, crashes, revenue, and long-term proxies such as retention.
- Quantify the trade-off with confidence intervals, not point estimates, and check that the regression is real (SRM, novelty, segment concentration).
- Options: ship to a subset, iterate on the regression, or keep a holdback to measure the long-term effect. Write down the reasoning.
Sources: Kohavi, Tang & Xu, Trustworthy Online Controlled Experiments (2020); Deng et al., Improving the Sensitivity of Online Controlled Experiments by Utilizing Pre-Experiment Data (WSDM 2013); Johari et al., Always Valid Inference (2015).