Skip to main content
When a model trains, it repeatedly makes predictions on the training examples, computes a loss, and updates its parameters (e.g., via gradient descent). Improving performance on the training set, however, does not guarantee good performance on new, unseen data. The real goal of supervised learning is to generalize — to learn patterns that apply beyond the examples the model saw during training. That is why we split data into training, validation, and test sets.
Splitting data into training, validation, and test sets helps you measure and improve generalization. Use the validation set for development and hyperparameter tuning; reserve the test set for a final, unbiased evaluation.

What is overfitting?

Overfitting happens when a model learns the training data too closely — including noise or idiosyncrasies — instead of learning underlying patterns that generalize. For instance, a model predicting house prices might memorize particular past sales records rather than learning the true drivers of price (size, location, number of bedrooms, etc.). Such a model will likely perform well on the training set but poorly on new houses. Analogy: memorizing practice exam answers only helps if the real exam repeats the same questions. If the exam asks new questions that evaluate the same concepts, memorization fails.

Train / Validation / Test: Roles and workflows

Use separate sets with clear roles: Common workflow:
  1. Train candidate models on the training set.
  2. Evaluate them on the validation set to choose hyperparameters and guard against overfitting.
  3. After selecting a final model, evaluate exactly once on the test set to estimate real-world performance.

Typical split ratios and strategies

There is no one-size-fits-all split. Here are common choices: If data is scarce, k-fold cross-validation (train/validation repeated across k partitions) helps produce more reliable performance estimates. See scikit-learn’s model_selection docs for practical implementations: https://scikit-learn.org/stable/modules/cross_validation.html

Watch out for data leakage (important)

Data leakage occurs when information from the validation or test sets unintentionally influences the training process. This leads to overly optimistic evaluation metrics because the model has in effect “seen” part of the holdout data. Examples of leakage:
  • A feature that directly or indirectly contains the target value.
  • The same entity appears in both training and test sets (e.g., repeated measurements of the same house).
  • Preprocessing that is fit on the full dataset (scaling, imputation, feature selection) before splitting.
Data leakage invalidates evaluation. Always split your dataset before computing statistics (mean, std) for scaling or performing feature selection. Treat the test set as unseen and never peek at it while developing models.

Detecting and addressing overfitting

Typical signs:
  • Training performance keeps improving while validation performance stagnates or worsens.
  • A very complex model achieves near-perfect training scores but much worse validation scores.
Ways to reduce overfitting:
  • Use simpler models or add regularization (L1/L2, dropout).
  • Reduce model capacity (e.g., limit decision tree depth).
  • Increase training data or apply data augmentation.
  • Use early stopping based on validation performance.
  • Apply cross-validation to obtain more reliable hyperparameter estimates.

Short summary

  • Train on the training set.
  • Tune and monitor on the validation set.
  • Evaluate finally on the test set.
  • Guard against data leakage by keeping strict separation and by fitting preprocessing inside cross-validation folds.
  • Use cross-validation for small datasets to get stable estimates.

Exercise preview

In the exercise that follows, the dataset is already split into training, validation, and test sets. A decision tree is trained without restrictions and will likely memorize the training examples, scoring near-perfectly on training but poorly on validation. The remedy in this exercise is to limit the tree depth (reduce capacity). Later, you’ll be asked to inspect a seemingly high-scoring dataset to find where leakage occurred.
Further reading and references:

Watch Video

Practice Lab