Skip to main content
At this point we know a model makes predictions using parameters, and training is the process of adjusting those parameters. To guide that process we need a numeric measure that tells the model how good or bad a particular set of parameters is — this measure is provided by a loss function. Continuing the house-price example: if a model predicts 700,000butthehomeactuallysoldfor700,000 but the home actually sold for 800,000, we need a precise way to quantify that mistake so the model can improve. A loss function converts prediction errors into a single scalar value that training algorithms can minimize.

What is a loss function?

A loss function assigns a numerical penalty to a model’s prediction based on how far it deviates from the true value. Lower loss means better predictions. Different tasks and model behaviors call for different loss functions (regression vs classification, sensitivity to outliers, etc.).

Error and why we transform it

The simplest measure of prediction error for a single example is: error = y_hat - y where y_hat is the prediction and y is the true value. Example:
  • error = 700000 - 800000 = -100000 — the model underpredicted by 100,000.
Raw error has a drawback: positive and negative errors can cancel when averaged across many examples. For instance, an underprediction of 100,000 and an overprediction of 100,000 average to zero, which would misleadingly indicate perfect performance. To avoid cancellation we transform the error. A common choice is squaring the error: L = (y_hat - y) ** 2 Squaring has two desirable effects:
  • Removes the sign so errors of opposite sign don’t cancel.
  • Penalizes larger errors more heavily (e.g., a 100,000 error is punished much more than a 5,000 error).

Mean Squared Error (MSE)

When training on many examples we use the average loss across the dataset. For regression tasks, a widely used loss is Mean Squared Error (MSE): MSE = (1/N) * sum((y_hat_i - y_i) ** 2 for i in 1..N) MSE provides a single scalar feedback signal summarizing how wrong the model is on average. During training, a decreasing MSE generally indicates the model is improving. Example properties:
  • MSE has units of the target squared (e.g., dollars²).
  • The square root of MSE is RMSE (root mean squared error), which returns to the original units.
  • If you need less sensitivity to outliers, consider MAE (mean absolute error).

Python example (NumPy)

This concise example computes per-example errors, squared errors, and MSE.
Note how the raw errors average to zero ((-100000 + 100000) / 2 == 0). If you used plain average error you would mistakenly conclude perfect predictions. MSE avoids that by squaring errors before averaging.

Common loss functions (when to use what)

Loss functions convert prediction errors into a numeric signal that an optimizer can minimize. Choose a loss based on the task (regression vs classification) and the behavior you want during training (e.g., sensitivity to outliers).
MSE squares the target units (e.g., dollars²). If interpretability in original units matters, use RMSE. If you want less sensitivity to extreme errors, consider MAE.

From loss to learning

Computing loss is only half the story. Optimization algorithms such as gradient descent (and its variants) use the loss to compute parameter updates that reduce it. Tracking the chosen loss during training is the primary way to measure model progress.

Watch Video