Why leakage misleads
A model should learn from past examples and perform on cases it has not seen. Leakage breaks that boundary: it may expose the answer, include a future event or let the evaluation set influence preprocessing. The metric rises while real-world performance stays poor.
Suppose you predict whether a customer will miss a payment next month. A collections-contact flag recorded after the missed payment gives away the outcome. That feature may be predictive in a dataset and unavailable at the moment a real prediction is needed.
Four patterns to look for
These checks catch many problems before trusting a score:
- Target leakage: a field directly encodes the outcome or a downstream consequence, such as a cancellation code when predicting cancellations.
- Time leakage: features include events after the prediction timestamp, or a random split uses future records to predict the past.
- Preprocessing leakage: scaling, imputation, feature selection or text vocabulary is learned on all rows before splitting.
- Group leakage: related records from one person, device, patient or company appear in both train and test, making test cases less independent than future deployment cases.
A leakage-resistant workflow
Write down the real prediction setting first. It determines the right split and what counts as a valid feature.
Define prediction time and target
For each row, say when a prediction is made, what future period the label covers and what action it supports. This sets a clear cutoff for allowed information.
Audit feature availability
Ask when each field is recorded. Remove columns derived from the target, post-outcome workflows or future events. Check timestamps and data dictionaries, not only column names.
Match the split to deployment
For future predictions, train on earlier periods and test on a later one. For grouped observations, keep groups together. Use random splitting only when examples are independent and deployment will resemble that sample.
Split before fitting transformations
Fit imputers, scalers, encoders and feature selectors on training data only. Apply the fitted transformations to validation and test sets without refitting.
Keep preprocessing inside a pipeline
A pipeline applies fit steps inside each cross-validation fold and reduces accidental information flow. Evaluate the whole pipeline as the model.
Investigate suspicious scores
Compare with a simple baseline, inspect unusually powerful features, test time- or group-aware splits and check duplicates or shared identifiers. Document the split so another analyst can reproduce it.
A surprisingly high score deserves a feature audit
Very high performance can be real, but inspect feature availability dates and near-perfect predictors before sharing it. Compare a realistic holdout with cross-validation and be clear about the population each score represents.
Leakage can enter through unsupervised transformations too. A transformation that learns statistics from the full dataset can use test information without ever seeing the target. A scikit-learn Pipeline helps keep fit operations inside training folds; the test set is for evaluation, not model decisions.
Treat the split as part of the model
A credible metric needs a credible simulation of prediction time. Define what is knowable, split accordingly, fit preprocessing on training records only and review suspiciously strong features.

