> ## Documentation Index
> Fetch the complete documentation index at: https://notes.kodekloud.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Overfitting and the TrainValidationTest Split

> Explains overfitting, the train validation test split, data leakage risks, and techniques like cross validation, regularization, and early stopping to improve model generalization.

When a model trains, it repeatedly makes predictions on the training examples, computes a loss, and updates its parameters (e.g., via gradient descent). Improving performance on the training set, however, does not guarantee good performance on new, unseen data. The real goal of supervised learning is to generalize — to learn patterns that apply beyond the examples the model saw during training. That is why we split data into training, validation, and test sets.

<Callout icon="lightbulb" color="#1CB2FE">
  Splitting data into training, validation, and test sets helps you measure and improve generalization. Use the validation set for development and hyperparameter tuning; reserve the test set for a final, unbiased evaluation.
</Callout>

## What is overfitting?

Overfitting happens when a model learns the training data too closely — including noise or idiosyncrasies — instead of learning underlying patterns that generalize. For instance, a model predicting house prices might memorize particular past sales records rather than learning the true drivers of price (size, location, number of bedrooms, etc.). Such a model will likely perform well on the training set but poorly on new houses.

Analogy: memorizing practice exam answers only helps if the real exam repeats the same questions. If the exam asks new questions that evaluate the same concepts, memorization fails.

## Train / Validation / Test: Roles and workflows

Use separate sets with clear roles:

| Split | Primary use | When to use |
| - | - | - |
| Training set | Fit model parameters (weights, tree splits, etc.) | Used iteratively during training |
| Validation set | Tune hyperparameters, early stopping, and monitor generalization | Used repeatedly during development |
| Test set | Final, unbiased performance estimate | Used only once (or rarely) after model selection |

Common workflow:

1. Train candidate models on the training set.
2. Evaluate them on the validation set to choose hyperparameters and guard against overfitting.
3. After selecting a final model, evaluate exactly once on the test set to estimate real-world performance.

## Typical split ratios and strategies

There is no one-size-fits-all split. Here are common choices:

| Strategy | Example split | When to use |
| - | - | - |
| Standard | `70% / 15% / 15%` | Good general-purpose starting point |
| More training | `80% / 10% / 10%` | When you want more data to train with |
| Cross-validation | `k-fold CV` (e.g., 5- or 10-fold) | When the dataset is small; gives robust estimates |

If data is scarce, k-fold cross-validation (train/validation repeated across k partitions) helps produce more reliable performance estimates. See scikit-learn's model\_selection docs for practical implementations: [https://scikit-learn.org/stable/modules/cross\_validation.html](https://scikit-learn.org/stable/modules/cross_validation.html)

## Watch out for data leakage (important)

Data leakage occurs when information from the validation or test sets unintentionally influences the training process. This leads to overly optimistic evaluation metrics because the model has in effect “seen” part of the holdout data.

Examples of leakage:

* A feature that directly or indirectly contains the target value.
* The same entity appears in both training and test sets (e.g., repeated measurements of the same house).
* Preprocessing that is fit on the full dataset (scaling, imputation, feature selection) before splitting.

<Callout icon="warning" color="#FF6B6B">
  Data leakage invalidates evaluation. Always split your dataset before computing statistics (mean, std) for scaling or performing feature selection. Treat the test set as unseen and never peek at it while developing models.
</Callout>

## Detecting and addressing overfitting

Typical signs:

* Training performance keeps improving while validation performance stagnates or worsens.
* A very complex model achieves near-perfect training scores but much worse validation scores.

Ways to reduce overfitting:

* Use simpler models or add regularization (L1/L2, dropout).
* Reduce model capacity (e.g., limit decision tree depth).
* Increase training data or apply data augmentation.
* Use early stopping based on validation performance.
* Apply cross-validation to obtain more reliable hyperparameter estimates.

## Short summary

* Train on the training set.
* Tune and monitor on the validation set.
* Evaluate finally on the test set.
* Guard against data leakage by keeping strict separation and by fitting preprocessing inside cross-validation folds.
* Use cross-validation for small datasets to get stable estimates.

## Exercise preview

In the exercise that follows, the dataset is already split into training, validation, and test sets. A decision tree is trained without restrictions and will likely memorize the training examples, scoring near-perfectly on training but poorly on validation. The remedy in this exercise is to limit the tree depth (reduce capacity). Later, you'll be asked to inspect a seemingly high-scoring dataset to find where leakage occurred.

```bash theme={null}
Welcome to the KodeKloud Hands-On lab

   ___  ____  ____  ____  _   _  _   _  _   _  ____  _   _  ____ 
  / _ \|  _ \|  _ \|  _ \| | | || \ | || \ | ||  _ \| \ | |/ ___|
 | | | | | | | | | | | | | | | ||  \| ||  \| || | | |  \| | |  _ 
 | |_| | |_| | |_| | |_| | |_| || |\  || |\  || |_| | |\  | |_| |
  \___/|____/|____/|____/ \___/ |_| \_||_| \_||____/|_| \_|\____|

All rights reserved

controlplane ~ on ☁ (us-east-1) → python3 split.py && wc -l train.csv val.csv test.csv
211 train.csv
46 val.csv
46 test.csv

controlplane ~ on ☁ (us-east-1) → python3 train_tree.py --max-depth none
train MAE: $0
val   MAE: $48,210

controlplane ~ on ☁ (us-east-1) → python3 train_tree.py --max-depth 4
train MAE: $22,400
val   MAE: $31,050
```

Further reading and references:

* scikit-learn cross-validation and model selection: [https://scikit-learn.org/stable/modules/cross\_validation.html](https://scikit-learn.org/stable/modules/cross_validation.html)
* Practical guide to overfitting and regularization: [https://developers.google.com/machine-learning/crash-course/regularization-for-simplicity](https://developers.google.com/machine-learning/crash-course/regularization-for-simplicity)
* Guide on preventing data leakage: [https://www.kdnuggets.com/2019/09/data-leakage-machine-learning.html](https://www.kdnuggets.com/2019/09/data-leakage-machine-learning.html)

<CardGroup>
  <Card title="Watch Video" icon="video" cta="Learn more" href="https://learn.kodekloud.com/user/courses/machine-learning-fundamentals/module/7af5e691-bb9b-4904-9a7f-c20fcb32a4ea/lesson/82b090b4-ee00-4cf6-959b-62d5ae2da207" />

  <Card title="Practice Lab" icon="flask-conical" cta="Learn more" href="https://learn.kodekloud.com/user/courses/machine-learning-fundamentals/module/7af5e691-bb9b-4904-9a7f-c20fcb32a4ea/lesson/0103818c-e5f7-4473-9eef-20d8dfcf86bc" />
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.