Leakage and the leaderboard that lies
Your churn model scores AUC 0.97 in validation. In production, the next month, it scores 0.62. Nothing is broken in the code. What happened?
A validation score is a prediction about the future. It's only as honest as the conditions it was measured under: the same information the model will have, on data it hasn't seen, in the way it will be used. Every leak breaks one of those conditions.
Three sets, three jobs
The training set fits parameters. The validation set chooses among models: features, hyperparameters, early stopping. The test set estimates how the chosen model will do on new data, once.
Each score is truth plus noise. The maximum of 200 is biased upward by the luckiest noise: you've fit the test set through selection, without training on it. A large test set shrinks the bias but doesn't remove it. Choose on validation, confirm on a test set you touch once, and treat a public leaderboard you submit to repeatedly as a validation set, not a test.
What did the model know, and when?
Every prediction happens at a moment \(t\). Features may only use information recorded before \(t\). The label describes what happens after \(t\).
For a row dated in March, "lifetime" computed tonight includes April through September. Users who stayed accumulated more sessions after March, so the feature partly encodes the label. The fix is a point-in-time (as-of) join: each training row gets the feature's value as it was at that row's timestamp. Feature stores exist largely to make this the default.
Validate, then deploy
Twelve months of subscription data. Pick features and a validation split, validate, then deploy to month 12, which is the future. Features come from the tables named. A colleague's first attempt is preselected. Goal: production AUC ≥ 0.70 with validation within 0.03 of it, a model that's good and honestly measured.
A logistic regression is trained in your browser each time. The data is simulated: 1,500 users joining over 12 months, churn driven by a hidden engagement trait, with drift over time.
A short taxonomy
| Leak | What happens | Tell-tale sign | Fix |
|---|---|---|---|
| Target leakage | A feature is a consequence of the label (closed-account flag, refund issued, diagnosis code) | One feature alone is nearly perfect | Check each feature's availability time; drop post-outcome fields |
| Temporal leakage | Random splits on time-ordered data, or features computed with future data | Random-split score ≫ time-split score | Split by time; point-in-time features |
| Group leakage | The same user, patient, session or near-duplicate item is in both train and test | Great on known users, poor on new ones | Split by group; dedupe |
| Preprocessing leakage | Scalers, PCA, target encoders or feature selection fit on all data before splitting | Small, persistent optimism | Fit every step inside each training fold (pipelines) |
| Selection leakage | Tuning or reporting on the test set, many leaderboard submissions | Test score drops on a fresh holdout | Touch the test set once; nested CV |
Apply it somewhere else
Group leakage. Images from the same patient share identifying features, so the test set contains near-copies of training data. Split by patient, and report performance on new patients (and new hospitals, if that's the deployment setting).
Use rolling-origin (forward-chaining) validation: train on days up to \(t\), validate on the next window, move forward. Add a gap equal to the feature and label horizon so overlapping windows don't leak.
Selecting on all rows picks the features that correlate with the labels by chance, including in the validation folds, so CV rewards that chance. Selection has to happen inside each training fold. This is the classic case in Hastie, Tibshirani & Friedman (§7.10.2).
Write the honest versions
Plain editor, numpy, hidden tests. ⌘/Ctrl+Enter runs.
Exercise 1: AUC from ranks, with ties
Exercise 2: a point-in-time feature
Answer out loud, then check
- Too-good-to-be-true checks: single-feature AUC near 1, one feature dominating importance, results far above published or prior baselines.
- Audit every feature's source table and timestamp against the prediction time. Compare random-split vs time-split vs group-split scores.
- Shadow deployment: score live traffic before launch and compare against the offline estimate. Adversarial validation can also reveal train/test distribution differences.
- A temporal split (train up to T, evaluate after T) with features computed as of the prediction time.
- Report separately for returning users and for users unseen in training (group split), and slice by item age for cold-start items.
- Evaluate against the candidates the system would actually produce (lesson 9), and check that offline gains predict online ones over time.
- Plot learning curves: train and validation loss vs data size or training time.
- High bias: both are high and close together. Add capacity or features, reduce regularization, train longer.
- High variance: a large train-validation gap. Add data or augmentation, regularize (weight decay, dropout, early stopping), simplify.
- If validation behaves oddly (better than train, unstable across seeds), suspect leakage or a split problem before tuning.
- Leakage: it was never as good as measured. Re-evaluate the offline setup with a proper time split and point-in-time features. If the offline score drops to match production, it was leakage.
- Drift: it was good and the world changed. Monitor feature distributions (PSI/KS), prediction distribution, calibration, and label-delayed performance by cohort.
- Responses: fix the pipeline for leakage; retrain cadence, recency weighting, or robustness for drift.
Sources: Kaufman et al., Leakage in Data Mining: Formulation, Detection, and Avoidance (TKDD 2012); Hastie, Tibshirani & Friedman, The Elements of Statistical Learning, §7.10; Roelofs et al., A Meta-Analysis of Overfitting in Machine Learning (NeurIPS 2019).