> ## Documentation Index
> Fetch the complete documentation index at: https://notes.kodekloud.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Understanding RandomForest Regressor

> Overview of Random Forest regression, explaining regression tasks, ensemble decision trees, operation, hyperparameters, use cases, and practical tips for training and evaluation

In supervised machine learning, regression refers to predicting continuous numerical outcomes (for example, house prices or temperatures) rather than discrete labels. Regression models estimate a function that maps input features to a real-valued target.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/MGkgrGfKHDtoCnUb/images/Kubeflow/KServe-and-Katib/Understanding-RandomForest-Regressor/regression-house-price-size-scatter-line.jpg?fit=max&auto=format&n=MGkgrGfKHDtoCnUb&q=85&s=e97256909444b8f892558658b27457db" alt="A presentation slide titled &#x22;What is Regression?&#x22; stating that regression predicts numerical values. On the right is a scatter plot of house price vs. size with gray data points and a blue best-fit line showing the predicted price." width="1920" height="1080" data-path="images/Kubeflow/KServe-and-Katib/Understanding-RandomForest-Regressor/regression-house-price-size-scatter-line.jpg" />
</Frame>

Common regression problems include:

* Predicting house prices from property features
* Estimating salaries from experience and skills
* Forecasting temperature or energy consumption
* Projecting future sales and revenue

Random Forest is an ensemble learning technique that builds many decision trees and combines their outputs to produce a single prediction. For regression tasks, predictions from individual trees are averaged, which typically yields better accuracy and stability than a single decision tree.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/MGkgrGfKHDtoCnUb/images/Kubeflow/KServe-and-Katib/Understanding-RandomForest-Regressor/random-forest-vs-decision-tree-overfitting.jpg?fit=max&auto=format&n=MGkgrGfKHDtoCnUb&q=85&s=f78cd204fb3ec856b3a7f2f30305ec77" alt="An infographic titled &#x22;What is Random Forest?&#x22; comparing a single decision tree (which memorizes training data and overfits, showing a 520k prediction) with a random forest that averages multiple trees for a more stable 310k prediction. It also notes that random forests reduce overfitting and work for both classification and regression." width="1920" height="1080" data-path="images/Kubeflow/KServe-and-Katib/Understanding-RandomForest-Regressor/random-forest-vs-decision-tree-overfitting.jpg" />
</Frame>

## How Random Forest works (high level)

* Bootstrap sampling (bagging): Train each tree on a random sample of the training set drawn with replacement. This produces diverse training sets and reduces variance.
* Feature subsampling: At every split, each tree evaluates only a random subset of features. This lowers correlation between trees and improves ensemble performance. A common heuristic is to try `sqrt(p)` features for classification and about `p/3` for regression, where `p` is the number of features (actual defaults vary by library).
* Independent tree growth: Trees are grown independently and are often grown deep to capture complex relationships.
* Aggregation: For regression, the final prediction is the mean (or sometimes the median) of all tree predictions. For classification, majority voting decides the label.

## Why Random Forests are effective

* Reduced overfitting: Averaging many uncorrelated trees reduces variance compared to a single tree.
* Robustness to noise and outliers: Individual noisy observations have a limited effect on the averaged prediction.
* Captures non-linear relationships: Trees model complex interactions and non-linearities without heavy feature engineering.
* Handles high-dimensional data: Random forests perform well with many features and can implicitly rank feature importance.

## Common use cases

* Finance: risk scoring and asset-price prediction
* Healthcare: patient outcome or biomarker value prediction
* Retail and recommendations: demand forecasting and customer lifetime value
* Energy and weather: load forecasting and temperature prediction

## Training and prediction summary

1. Create many decision trees; each tree uses a bootstrap sample of the training set.
2. When splitting a node, consider a random subset of features to find the best split.
3. For a new observation, run it down every tree to obtain individual predictions.
4. Combine tree outputs — average them for regression — to produce the final prediction.

<Callout icon="lightbulb" color="#1CB2FE">
  Practical tip: libraries like scikit-learn expose key hyperparameters for RandomForestRegressor such as `n_estimators` (number of trees), `max_depth` (maximum tree depth), and `max_features` (number of features considered at each split). Tuning these parameters helps control bias–variance trade-offs and improves generalization. See the scikit-learn docs for details: [https://scikit-learn.org/stable/modules/generated/sklearn.ensemble.RandomForestRegressor.html](https://scikit-learn.org/stable/modules/generated/sklearn.ensemble.RandomForestRegressor.html)
</Callout>

## Key hyperparameters and their purpose

| Parameter | Purpose | Typical default / example |
| - | - | - |
| `n_estimators` | Number of trees in the forest. More trees often improve performance but increase compute. | `100` (common starting point) |
| `max_depth` | Maximum depth of each tree; controls complexity and overfitting. | `None` (grow until leaves are pure) |
| `max_features` | Number of features to consider at each split. Reduces correlation between trees. | `p/3` (regression) or `sqrt(p)` (classification) |
| `min_samples_split` | Minimum number of samples required to split an internal node. | `2` |
| `min_samples_leaf` | Minimum number of samples required to be at a leaf node. | `1` |
| `bootstrap` | Whether bootstrap samples are used when building trees (bagging). | `True` |

## Practical considerations

* Feature scaling is not required for tree-based methods, but categorical features should be encoded appropriately.
* Use out-of-bag (OOB) error (if `bootstrap=True`) as a quick validation metric without a separate holdout.
* For very large datasets, consider limiting `max_depth` or increasing `min_samples_leaf` to speed up training.
* Use feature importance from the trained forest to guide feature selection or interpretability, but be aware of biases (e.g., toward variables with more categories or continuous variables with many split points).

## Links and references

* [scikit-learn RandomForestRegressor documentation](https://scikit-learn.org/stable/modules/generated/sklearn.ensemble.RandomForestRegressor.html)
* [Kaggle — Random Forest Guide](https://www.kaggle.com/learn/overview)
* [Ensemble methods overview (Wikipedia)](https://en.wikipedia.org/wiki/Ensemble_learning)

<CardGroup>
  <Card title="Watch Video" icon="video" cta="Learn more" href="https://learn.kodekloud.com/user/courses/kubeflow/module/d9b1b119-0c6f-494b-b063-8eccd99dbff7/lesson/323d4a34-7341-47eb-8033-f6b9e6cf4bf2" />
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.