← Atelier
Lesson 11 · Evaluation · ~30 min + code

Leakage and the leaderboard that lies

Your churn model scores AUC 0.97 in validation. In production, the next month, it scores 0.62. Nothing is broken in the code. What happened?

A validation score is a prediction about the future. It's only as honest as the conditions it was measured under: the same information the model will have, on data it hasn't seen, in the way it will be used. Every leak breaks one of those conditions.

1 · What validation promises

Three sets, three jobs

The training set fits parameters. The validation set chooses among models: features, hyperparameters, early stopping. The test set estimates how the chosen model will do on new data, once.

You try 200 configurations and report the best one's score on the test set. Is that score a fair estimate of production performance?

Each score is truth plus noise. The maximum of 200 is biased upward by the luckiest noise: you've fit the test set through selection, without training on it. A large test set shrinks the bias but doesn't remove it. Choose on validation, confirm on a test set you touch once, and treat a public leaderboard you submit to repeatedly as a validation set, not a test.

2 · Prediction time

What did the model know, and when?

Every prediction happens at a moment \(t\). Features may only use information recorded before \(t\). The label describes what happens after \(t\).

A nightly warehouse job computes "lifetime sessions per user" over all history, and you join it onto historical training rows. What's wrong?

For a row dated in March, "lifetime" computed tonight includes April through September. Users who stayed accumulated more sessions after March, so the feature partly encodes the label. The fix is a point-in-time (as-of) join: each training row gets the feature's value as it was at that row's timestamp. Feature stores exist largely to make this the default.

3 · Play it

Validate, then deploy

Twelve months of subscription data. Pick features and a validation split, validate, then deploy to month 12, which is the future. Features come from the tables named. A colleague's first attempt is preselected. Goal: production AUC ≥ 0.70 with validation within 0.03 of it, a model that's good and honestly measured.

Features
Validation split

A logistic regression is trained in your browser each time. The data is simulated: 1,500 users joining over 12 months, churn driven by a hidden engagement trait, with drift over time.

4 · The usual leaks

A short taxonomy

LeakWhat happensTell-tale signFix
Target leakageA feature is a consequence of the label (closed-account flag, refund issued, diagnosis code)One feature alone is nearly perfectCheck each feature's availability time; drop post-outcome fields
Temporal leakageRandom splits on time-ordered data, or features computed with future dataRandom-split score ≫ time-split scoreSplit by time; point-in-time features
Group leakageThe same user, patient, session or near-duplicate item is in both train and testGreat on known users, poor on new onesSplit by group; dedupe
Preprocessing leakageScalers, PCA, target encoders or feature selection fit on all data before splittingSmall, persistent optimismFit every step inside each training fold (pipelines)
Selection leakageTuning or reporting on the test set, many leaderboard submissionsTest score drops on a fresh holdoutTouch the test set once; nested CV
5 · New cases, no hints

Apply it somewhere else

A chest X-ray classifier scores 0.95 AUC. Many patients have several X-rays, and the split was by image. What's likely?

Group leakage. Images from the same patient share identifying features, so the test set contains near-copies of training data. Split by patient, and report performance on new patients (and new hospitals, if that's the deployment setting).

You forecast daily demand. You use 5-fold cross-validation with random folds. What goes wrong?

Use rolling-origin (forward-chaining) validation: train on days up to \(t\), validate on the next window, move forward. Add a gap equal to the feature and label horizon so overlapping windows don't leak.

You select the 20 most correlated features on the full dataset, then cross-validate a model using them. With pure-noise features and labels, what AUC does CV report?

Selecting on all rows picks the features that correlate with the labels by chance, including in the validation folds, so CV rewards that chance. Selection has to happen inside each training fold. This is the classic case in Hastie, Tibshirani & Friedman (§7.10.2).

6 · Implement, unaided

Write the honest versions

Plain editor, numpy, hidden tests. ⌘/Ctrl+Enter runs.

Exercise 1: AUC from ranks, with ties

Exercise 2: a point-in-time feature

7 · Staff-level follow-ups

Answer out loud, then check

How do you detect leakage before it reaches production?
  • Too-good-to-be-true checks: single-feature AUC near 1, one feature dominating importance, results far above published or prior baselines.
  • Audit every feature's source table and timestamp against the prediction time. Compare random-split vs time-split vs group-split scores.
  • Shadow deployment: score live traffic before launch and compare against the offline estimate. Adversarial validation can also reveal train/test distribution differences.
How would you design validation for a recommender that must serve both returning and new users?
  • A temporal split (train up to T, evaluate after T) with features computed as of the prediction time.
  • Report separately for returning users and for users unseen in training (group split), and slice by item age for cold-start items.
  • Evaluate against the candidates the system would actually produce (lesson 9), and check that offline gains predict online ones over time.
Bias vs variance: how do you tell which one you have, and what do you do?
  • Plot learning curves: train and validation loss vs data size or training time.
  • High bias: both are high and close together. Add capacity or features, reduce regularization, train longer.
  • High variance: a large train-validation gap. Add data or augmentation, regularize (weight decay, dropout, early stopping), simplify.
  • If validation behaves oddly (better than train, unstable across seeds), suspect leakage or a split problem before tuning.
Your model degrades a month after launch. How do you tell leakage from drift?
  • Leakage: it was never as good as measured. Re-evaluate the offline setup with a proper time split and point-in-time features. If the offline score drops to match production, it was leakage.
  • Drift: it was good and the world changed. Monitor feature distributions (PSI/KS), prediction distribution, calibration, and label-delayed performance by cohort.
  • Responses: fix the pipeline for leakage; retrain cadence, recency weighting, or robustness for drift.

Sources: Kaufman et al., Leakage in Data Mining: Formulation, Detection, and Avoidance (TKDD 2012); Hastie, Tibshirani & Friedman, The Elements of Statistical Learning, §7.10; Roelofs et al., A Meta-Analysis of Overfitting in Machine Learning (NeurIPS 2019).