What is a loss function?
A loss function assigns a numerical penalty to a model’s prediction based on how far it deviates from the true value. Lower loss means better predictions. Different tasks and model behaviors call for different loss functions (regression vs classification, sensitivity to outliers, etc.).Error and why we transform it
The simplest measure of prediction error for a single example is:error = y_hat - y
where y_hat is the prediction and y is the true value.
Example:
error = 700000 - 800000 = -100000— the model underpredicted by 100,000.
L = (y_hat - y) ** 2
Squaring has two desirable effects:
- Removes the sign so errors of opposite sign don’t cancel.
- Penalizes larger errors more heavily (e.g., a 100,000 error is punished much more than a 5,000 error).
Mean Squared Error (MSE)
When training on many examples we use the average loss across the dataset. For regression tasks, a widely used loss is Mean Squared Error (MSE):MSE = (1/N) * sum((y_hat_i - y_i) ** 2 for i in 1..N)
MSE provides a single scalar feedback signal summarizing how wrong the model is on average. During training, a decreasing MSE generally indicates the model is improving.
Example properties:
- MSE has units of the target squared (e.g., dollars²).
- The square root of MSE is RMSE (root mean squared error), which returns to the original units.
- If you need less sensitivity to outliers, consider MAE (mean absolute error).
Python example (NumPy)
This concise example computes per-example errors, squared errors, and MSE.(-100000 + 100000) / 2 == 0). If you used plain average error you would mistakenly conclude perfect predictions. MSE avoids that by squaring errors before averaging.
Common loss functions (when to use what)
Loss functions convert prediction errors into a numeric signal that an optimizer can minimize. Choose a loss based on the task (regression vs classification) and the behavior you want during training (e.g., sensitivity to outliers).
MSE squares the target units (e.g., dollars²). If interpretability in original units matters, use RMSE. If you want less sensitivity to extreme errors, consider MAE.
From loss to learning
Computing loss is only half the story. Optimization algorithms such as gradient descent (and its variants) use the loss to compute parameter updates that reduce it. Tracking the chosen loss during training is the primary way to measure model progress.Links and references
- Scikit-learn: mean_squared_error
- Khan Academy: Mean squared error
- Gradient descent overview: https://en.wikipedia.org/wiki/Gradient_descent