← Atelier
Lesson 07 · Experiments · ~30 min + code

A/B tests and the cost of peeking

Your new ranking model shows +1.2% click-through after three days of an A/B test, with p = 0.03. Do you ship it?

Your call:

A p-value means what it says only under the plan it was computed for: a sample size fixed in advance, one look, one metric, and a correct randomization. Check a running test every day and stop the first time p dips below 0.05, and "p < 0.05" happens far more often than 5% of the time when nothing changed. This lesson builds the p-value from scratch, then shows exactly how that breaks.

1 · What p means

Run tests where nothing changed

An A/A test shows both groups the same system. Any difference you see is noise. Before reading p-values, look at the noise.

You run 1,000 A/A tests, each with 10,000 users per arm and a 5% click rate. In how many does the two-sided test say p < 0.05?

About 5%. That's what the 0.05 threshold is: the fraction of no-difference experiments that look at least this extreme by chance. The p-value is the probability, if nothing changed, of a difference as large as the one observed. Run them:

z-scores of simulated A/A tests

z = observed difference ÷ its standard error. The shaded tails beyond ±1.96 hold 5% of a standard normal: those tests would be called "significant".

2 · The standard error

How much does a click rate wobble?

Each user either clicks or doesn't: a Bernoulli variable with variance \(p(1-p)\). The average of \(n\) independent ones has variance \(p(1-p)/n\). The difference of two independent averages adds their variances:

\[ \begin{aligned} \mathrm{SE}(\hat p_B - \hat p_A) &= \sqrt{\frac{p(1-p)}{n_A} + \frac{p(1-p)}{n_B}} \\ z &= \frac{\hat p_B - \hat p_A}{\mathrm{SE}} \end{aligned} \]
To halve the standard error, how many more users do you need?

SE falls like \(1/\sqrt n\). Halving it takes 4× the data, and detecting an effect half as large takes 4× the traffic. That square law is the whole economics of experimentation.

3 · Power and sample size

Can this test even see the effect?

Power is the probability of detecting a real effect of a given size. To reach it, the effect must stand about \(z_{1-\alpha/2} + z_{\text{power}}\) standard errors clear of zero, which is 1.96 + 0.84 = 2.8 for α = 0.05 and 80% power:

\[ n \approx \frac{(z_{1-\alpha/2} + z_{\text{power}})^2\,\big(p_A(1-p_A) + p_B(1-p_B)\big)}{(p_B - p_A)^2} \]
Base click rate 5%. You care about a +2% relative lift (5.0% → 5.1%). Roughly how many users per arm for 80% power?

\(7.85 \times (0.0475 + 0.0484) / 0.001^2 \approx 753{,}000\). Small relative lifts on small rates need enormous samples. An underpowered test mostly returns "not significant", and when it does come back significant it overstates the effect (the winner's curse). Try your own numbers:

Sample size calculator (two-sided α, two arms of equal size)
4 · Play it

Ship or not

Eight experiments, 2,000 users per arm per day, a 5% base click rate. Some variants are genuinely +10% better; others are identical to control. Watch the results day by day and decide whenever you like: ship B, or keep A. The planned length is 14 days, and the test was sized for that.

5 · Other ways tests lie

New cases, no hints

You split traffic 50/50 and get 50,900 users in A and 49,100 in B. The metric difference is significant. What first?

A chi-square test of the split gives \(\chi^2 = 32.4\), p ≈ \(10^{-8}\). That's a sample ratio mismatch. Something (a crash, a redirect, bot filtering, logging) removes users from one arm, and probably not at random, so every metric comparison is suspect. Check SRM before reading any metric.

You randomize by user, but compute click-through per impression and use the standard error formula from section 2 with n = impressions. What goes wrong?

The independent unit is the user, not the impression. Heavy users contribute many correlated impressions, so the naive SE is too small and p-values come out too small. Use the delta method for a ratio metric, or a bootstrap that resamples users. (The third option is also true of a ratio metric, but it's a question of which quantity you want, not an error in the test.)

You track 20 independent metrics in an A/A test. What's the chance at least one shows p < 0.05?

\(1 - 0.95^{20} \approx 0.64\). Pick one primary decision metric in advance, treat the rest as guardrails or exploration, and correct for multiplicity (Bonferroni, or Benjamini–Hochberg for false discovery rate) when you must test many.

A new recommendation surface lifts engagement 8% in week one, 3% in week two, and 1% in week three. Most likely?

Short tests overstate novelty (and understate learning effects, the opposite pattern). Plot the effect by exposure day or by user cohort, run long enough to see it flatten, and keep a long-term holdback to measure the durable effect.

6 · Implement, unaided

Write the test yourself

Standard library only: math.erfc gives normal tail probabilities, and statistics.NormalDist().inv_cdf gives quantiles. ⌘/Ctrl+Enter runs.

Exercise 1: two-proportion z-test

Exercise 2: sample size and a sample-ratio check

7 · Staff-level follow-ups

Answer out loud, then check

How do you let teams monitor experiments continuously without inflating false positives?
  • Sequential designs: group-sequential tests with alpha spending (e.g. O'Brien–Fleming boundaries that are strict early), or always-valid p-values and confidence sequences (mSPRT).
  • Or pre-commit a fixed horizon from a power calculation and decide only at the end. Dashboards can show data without showing "significance".
  • Bonferroni across planned looks is valid but conservative.
How would you reduce the variance of an experiment without more traffic?
  • CUPED: regress out a pre-experiment covariate (the same metric in the prior period), \(Y' = Y - \theta(X - \bar X)\) with \(\theta = \mathrm{cov}(X,Y)/\mathrm{var}(X)\). The variance falls by the factor \(1-\rho^2\), with no bias because X is pre-treatment.
  • Also: more sensitive metrics (proxies), trimming or capping heavy-tailed metrics, stratification, and triggering analysis only on users actually exposed to the change.
The new ranker wins offline (NDCG +3%) but is flat online. Why might that happen?
  • Offline labels come from logs produced by the old ranker: position and exposure bias, and items it never showed have no labels.
  • The metric mismatch: NDCG on logged clicks vs the business outcome (dwell, retention). Candidate generation limits what any ranker can surface.
  • Train/serve skew in features, latency changes, a novelty baseline, or an underpowered test. Diagnose with counterfactual or interleaving evaluation, slice analysis, and a check that serving features match training.
When is user-level randomization invalid, and what do you do instead?
  • When treated and control users interact: marketplaces sharing inventory or budget, social feeds, ride-hailing supply. Treatment leaks into control (interference, a SUTVA violation).
  • Use cluster randomization (by region or social graph cluster), switchback designs (alternate by time block), or budget-split or two-sided designs in marketplaces.
  • The price is fewer independent units, so larger variance. Plan power accordingly.
How do you make a ship decision when the primary metric wins but a guardrail regresses?
  • Agree in advance on an overall evaluation criterion and guardrail thresholds: latency, crashes, revenue, and long-term proxies such as retention.
  • Quantify the trade-off with confidence intervals, not point estimates, and check that the regression is real (SRM, novelty, segment concentration).
  • Options: ship to a subset, iterate on the regression, or keep a holdback to measure the long-term effect. Write down the reasoning.

Sources: Kohavi, Tang & Xu, Trustworthy Online Controlled Experiments (2020); Deng et al., Improving the Sensitivity of Online Controlled Experiments by Utilizing Pre-Experiment Data (WSDM 2013); Johari et al., Always Valid Inference (2015).