> ## Documentation Index
> Fetch the complete documentation index at: https://notes.kodekloud.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Loss Functions and Mean Squared Error

> Explains loss functions for machine learning, focusing on mean squared error, its properties, examples, alternatives like MAE and RMSE, and guidance on choosing losses

At this point we know a model makes predictions using parameters, and training is the process of adjusting those parameters. To guide that process we need a numeric measure that tells the model how good or bad a particular set of parameters is — this measure is provided by a loss function.

Continuing the house-price example: if a model predicts $700,000 but the home actually sold for $800,000, we need a precise way to quantify that mistake so the model can improve. A loss function converts prediction errors into a single scalar value that training algorithms can minimize.

## What is a loss function?

A loss function assigns a numerical penalty to a model’s prediction based on how far it deviates from the true value. Lower loss means better predictions. Different tasks and model behaviors call for different loss functions (regression vs classification, sensitivity to outliers, etc.).

## Error and why we transform it

The simplest measure of prediction error for a single example is:

`error = y_hat - y`

where `y_hat` is the prediction and `y` is the true value.

Example:

* `error = 700000 - 800000 = -100000` — the model underpredicted by 100,000.

Raw error has a drawback: positive and negative errors can cancel when averaged across many examples. For instance, an underprediction of 100,000 and an overprediction of 100,000 average to zero, which would misleadingly indicate perfect performance.

To avoid cancellation we transform the error. A common choice is squaring the error:

`L = (y_hat - y) ** 2`

Squaring has two desirable effects:

* Removes the sign so errors of opposite sign don't cancel.
* Penalizes larger errors more heavily (e.g., a 100,000 error is punished much more than a 5,000 error).

## Mean Squared Error (MSE)

When training on many examples we use the average loss across the dataset. For regression tasks, a widely used loss is Mean Squared Error (MSE):

`MSE = (1/N) * sum((y_hat_i - y_i) ** 2 for i in 1..N)`

MSE provides a single scalar feedback signal summarizing how wrong the model is on average. During training, a decreasing MSE generally indicates the model is improving.

Example properties:

* MSE has units of the target squared (e.g., dollars²).
* The square root of MSE is RMSE (root mean squared error), which returns to the original units.
* If you need less sensitivity to outliers, consider MAE (mean absolute error).

### Python example (NumPy)

This concise example computes per-example errors, squared errors, and MSE.

```python theme={null}
import numpy as np

# Example predictions and true values (in dollars)
y_hat = np.array([700_000, 900_000])    # model predictions
y = np.array([800_000, 800_000])        # actual sale prices

# Per-example error
errors = y_hat - y
# Squared errors
squared_errors = errors ** 2
# Mean Squared Error
mse = squared_errors.mean()

print("errors:", errors)               # [-100000   100000]
print("squared_errors:", squared_errors)  # [1e10 1e10]
print("MSE:", mse)                     # 10000000000.0
```

Note how the raw errors average to zero (`(-100000 + 100000) / 2 == 0`). If you used plain average error you would mistakenly conclude perfect predictions. MSE avoids that by squaring errors before averaging.

## Common loss functions (when to use what)

| Loss | Typical use case | Notes |
| - | - | - |
| MSE (Mean Squared Error) | Regression where large errors should be strongly penalized | Sensitive to outliers; units are squared |
| RMSE (Root Mean Squared Error) | Regression when you want error in original units | `sqrt(MSE)` |
| MAE (Mean Absolute Error) | Regression robust to outliers | Less sensitive to large mistakes than MSE |
| Cross-entropy (log loss) | Classification (binary or multiclass) | Works well with probabilistic outputs and softmax |

<Callout icon="lightbulb" color="#1CB2FE">
  Loss functions convert prediction errors into a numeric signal that an optimizer can minimize. Choose a loss based on the task (regression vs classification) and the behavior you want during training (e.g., sensitivity to outliers).
</Callout>

<Callout icon="warning" color="#FF6B6B">
  MSE squares the target units (e.g., dollars²). If interpretability in original units matters, use RMSE. If you want less sensitivity to extreme errors, consider MAE.
</Callout>

## From loss to learning

Computing loss is only half the story. Optimization algorithms such as gradient descent (and its variants) use the loss to compute parameter updates that reduce it. Tracking the chosen loss during training is the primary way to measure model progress.

## Links and references

* [Scikit-learn: mean\_squared\_error](https://scikit-learn.org/stable/modules/generated/sklearn.metrics.mean_squared_error.html)
* [Khan Academy: Mean squared error](https://www.khanacademy.org)
* Gradient descent overview: [https://en.wikipedia.org/wiki/Gradient\_descent](https://en.wikipedia.org/wiki/Gradient_descent)

<CardGroup>
  <Card title="Watch Video" icon="video" cta="Learn more" href="https://learn.kodekloud.com/user/courses/machine-learning-fundamentals/module/f4db2a23-638a-40de-b672-cbe3cae21cc2/lesson/f63de4e4-5cc2-4bc2-a971-dc0aafc23027" />
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.