Skip to main content
In supervised machine learning, regression refers to predicting continuous numerical outcomes (for example, house prices or temperatures) rather than discrete labels. Regression models estimate a function that maps input features to a real-valued target.
A presentation slide titled "What is Regression?" stating that regression predicts numerical values. On the right is a scatter plot of house price vs. size with gray data points and a blue best-fit line showing the predicted price.
Common regression problems include:
  • Predicting house prices from property features
  • Estimating salaries from experience and skills
  • Forecasting temperature or energy consumption
  • Projecting future sales and revenue
Random Forest is an ensemble learning technique that builds many decision trees and combines their outputs to produce a single prediction. For regression tasks, predictions from individual trees are averaged, which typically yields better accuracy and stability than a single decision tree.
An infographic titled "What is Random Forest?" comparing a single decision tree (which memorizes training data and overfits, showing a 520k prediction) with a random forest that averages multiple trees for a more stable 310k prediction. It also notes that random forests reduce overfitting and work for both classification and regression.

How Random Forest works (high level)

  • Bootstrap sampling (bagging): Train each tree on a random sample of the training set drawn with replacement. This produces diverse training sets and reduces variance.
  • Feature subsampling: At every split, each tree evaluates only a random subset of features. This lowers correlation between trees and improves ensemble performance. A common heuristic is to try sqrt(p) features for classification and about p/3 for regression, where p is the number of features (actual defaults vary by library).
  • Independent tree growth: Trees are grown independently and are often grown deep to capture complex relationships.
  • Aggregation: For regression, the final prediction is the mean (or sometimes the median) of all tree predictions. For classification, majority voting decides the label.

Why Random Forests are effective

  • Reduced overfitting: Averaging many uncorrelated trees reduces variance compared to a single tree.
  • Robustness to noise and outliers: Individual noisy observations have a limited effect on the averaged prediction.
  • Captures non-linear relationships: Trees model complex interactions and non-linearities without heavy feature engineering.
  • Handles high-dimensional data: Random forests perform well with many features and can implicitly rank feature importance.

Common use cases

  • Finance: risk scoring and asset-price prediction
  • Healthcare: patient outcome or biomarker value prediction
  • Retail and recommendations: demand forecasting and customer lifetime value
  • Energy and weather: load forecasting and temperature prediction

Training and prediction summary

  1. Create many decision trees; each tree uses a bootstrap sample of the training set.
  2. When splitting a node, consider a random subset of features to find the best split.
  3. For a new observation, run it down every tree to obtain individual predictions.
  4. Combine tree outputs — average them for regression — to produce the final prediction.
Practical tip: libraries like scikit-learn expose key hyperparameters for RandomForestRegressor such as n_estimators (number of trees), max_depth (maximum tree depth), and max_features (number of features considered at each split). Tuning these parameters helps control bias–variance trade-offs and improves generalization. See the scikit-learn docs for details: https://scikit-learn.org/stable/modules/generated/sklearn.ensemble.RandomForestRegressor.html

Key hyperparameters and their purpose

Practical considerations

  • Feature scaling is not required for tree-based methods, but categorical features should be encoded appropriately.
  • Use out-of-bag (OOB) error (if bootstrap=True) as a quick validation metric without a separate holdout.
  • For very large datasets, consider limiting max_depth or increasing min_samples_leaf to speed up training.
  • Use feature importance from the trained forest to guide feature selection or interpretability, but be aware of biases (e.g., toward variables with more categories or continuous variables with many split points).

Watch Video