> ## Documentation Index
> Fetch the complete documentation index at: https://notes.kodekloud.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Evaluation Metrics Regression Classification and LLMs

> Overview of evaluation metrics and best practices for regression, classification, and large language models, with examples, pitfalls, and a hands-on MNIST lab

Once a model is trained, we need objective ways to measure how well it performs. Evaluation metrics are numbers (or sets of numbers) that quantify model performance and help answer questions like: How close are predictions to the true values? How often is the model correct? Does the model generalize to new, unseen data?

Evaluation always ties back to how data is split into training, validation, and test sets. We care not only about performance on the training data but, crucially, about performance on held-out (validation/test) data.

***

## Regression (predicting continuous values)

Main question: how far are predictions from the true values? Regression metrics measure the size and distribution of errors.

Key metrics:

| Metric | Description | When to use |
| - | - | - |
| MAE (Mean Absolute Error) | Average absolute difference between predictions and targets | When you want a metric less sensitive to outliers |
| MSE (Mean Squared Error) | Mean of squared errors; penalizes large errors more | When larger errors should be penalized more heavily |
| RMSE (Root Mean Squared Error) | Square root of MSE; same units as the target | Easier interpretation in original units (e.g., dollars) |

Example implementations (numpy):

```python theme={null}
import numpy as np

def mae(y_true, y_pred):
    return np.mean(np.abs(y_true - y_pred))

def mse(y_true, y_pred):
    return np.mean((y_true - y_pred) ** 2)

def rmse(y_true, y_pred):
    return np.sqrt(mse(y_true, y_pred))
```

Best practices:

* Use MAE when you want a robust, interpretable error measure that is not dominated by a few large outliers.
* Use MSE/RMSE when large deviations should be penalized more aggressively.
* Always choose a metric aligned with business or scientific costs (e.g., underestimating vs. overestimating may have different consequences).

***

## Classification (predicting categories)

Main question: which class does the model predict? Common tasks include digit recognition (0–9), spam detection, or multi-class image labels.

Baseline metric:

* Accuracy = (number of correct predictions) / (total predictions)

Accuracy is easy to understand but can be misleading on imbalanced datasets (e.g., a model that always predicts the majority class may still show high accuracy).

Confusion matrix (binary case):

* True Positives (TP)
* False Positives (FP)
* True Negatives (TN)
* False Negatives (FN)

From these we compute:

* Precision = TP / (TP + FP) — of predicted positives, how many were correct?
* Recall = TP / (TP + FN) — of actual positives, how many did the model find?
* F1 score = 2 \* (precision \* recall) / (precision + recall) — harmonic mean of precision and recall

Example formulas (binary case):

```python theme={null}
def precision(tp, fp):
    return tp / (tp + fp) if (tp + fp) > 0 else 0.0

def recall(tp, fn):
    return tp / (tp + fn) if (tp + fn) > 0 else 0.0

def f1_score(prec, rec):
    return 2 * (prec * rec) / (prec + rec) if (prec + rec) > 0 else 0.0
```

Practical tips:

* For imbalanced datasets, prefer precision/recall/F1 and inspect the confusion matrix to see the types of errors the model makes.
* For multi-class problems, compute per-class precision/recall and consider macro/micro averages depending on your evaluation goals.

<Callout icon="lightbulb" color="#1CB2FE">
  Accuracy is a good quick check, but for imbalanced datasets or tasks where some errors are costlier than others, examine precision, recall, and the confusion matrix to understand the types of mistakes the model makes.
</Callout>

***

## Large Language Models (LLMs)

Evaluating LLMs is more complex because many tasks are open-ended and can have multiple valid outputs (e.g., explanations, creative writing, code). There is often no single ground truth.

Common approaches:

* Automated benchmarks and task suites (question answering, code generation, math problems)
* Human preference ratings and pairwise comparisons (which output do humans prefer?)
* Factuality and hallucination checks (does the model invent facts?)
* Task-specific tests (e.g., unit tests for generated code, step-by-step verification for math)
* Real-world feedback and A/B testing to measure user satisfaction

Important caveats:

* No single metric captures everything. A model can score highly on a benchmark but still hallucinate or fail to follow instructions reliably.
* Combine automated metrics with human evaluation and task-specific checks to obtain a fuller picture.

<Callout icon="warning" color="#FF6B6B">
  Be careful of over-relying on single benchmarks. Metrics optimized in isolation can lead to systems that game the metric without improving real-world usefulness. Also watch for data leakage, label artifacts, and distribution shift between training and deployment.
</Callout>

***

## Quick Reference: Which metric to choose?

| Problem Type | Primary Metric(s) | Notes |
| - | - | - |
| Regression | `MAE`, `RMSE` | Use RMSE to express error in original units; MAE for robustness to outliers |
| Binary classification | `Precision`, `Recall`, `F1`, `AUC`, `Confusion matrix` | Use F1 for balance; prioritize precision or recall depending on cost of FP vs FN |
| Multiclass classification | Per-class precision/recall, macro/micro averages | Inspect per-class confusion to find common confusions |
| LLM / generative | Benchmarks + human preference + task-specific tests | Use a mix of automated and human evaluations |

***

## Summary

* Evaluation metrics quantify how useful a trained model is.
* For regression, focus on error size: MAE (robust), MSE/RMSE (penalize large errors).
* For classification, accuracy is a starting point; use precision/recall/F1 and confusion matrices to understand error types, especially on imbalanced data.
* For LLMs, combine automated benchmarks, human judgments, and task-specific tests — expect trade-offs and gaps in any single metric.

***

## Lab: MNIST digit classification (hands-on)

In this lab you will:

1. Train a neural network to read handwritten digits from the MNIST dataset ([https://yann.lecun.com/exdb/mnist/](https://yann.lecun.com/exdb/mnist/)).
2. Observe training loss decrease epoch-by-epoch.
3. Evaluate the trained model on 2,000 unseen test examples.
4. Inspect the confusion matrix and identify the most confused digit pairs.
5. Examine one or more misclassified samples (e.g., print pixel values or render a small plot) to understand why the model erred — perhaps a sloppy "4" that looks like a "9".

Example terminal session (illustrative):

```bash theme={null}
Welcome to the KodeKloud Hands-On lab

All rights reserved

controlplane ~ on ☁ (us-east-1) → python3 train_mlp.py
Iteration 1, loss = 0.61842
Iteration 2, loss = 0.28417
Iteration 3, loss = 0.21063
...
Iteration 30, loss = 0.03791
test accuracy: 0.964

controlplane ~ on ☁ (us-east-1) → python3 evaluate.py
digit precision recall
4 0.95 0.93
9 0.94 0.95

most confused: 4 → 9
```

How to inspect misclassified examples:

* After evaluation, save indices of misclassified samples and print the corresponding labels and predictions.
* Visualize the image using your preferred plotting library (e.g., matplotlib) to see why the model confused the digits.
* Example snippet to display an MNIST sample:

```python theme={null}
import matplotlib.pyplot as plt

# x is a (28,28) numpy array for the image
plt.imshow(x, cmap='gray')
plt.title(f"true: {true_label}  pred: {pred_label}")
plt.axis('off')
plt.show()
```

This hands-on exploration helps connect metric numbers (e.g., recall or confusion pairs) to concrete model behavior and ambiguous inputs.

***

## Links and References

* MNIST dataset: [https://yann.lecun.com/exdb/mnist/](https://yann.lecun.com/exdb/mnist/)
* Practical advice on evaluation: [https://www.tensorflow.org/guide/keras/train\_and\_evaluate](https://www.tensorflow.org/guide/keras/train_and_evaluate)
* Precision / Recall / F1 explanation: [https://scikit-learn.org/stable/modules/model\_evaluation.html#precision-recall](https://scikit-learn.org/stable/modules/model_evaluation.html#precision-recall)

Good luck with the lab — inspect the confusion matrix closely and use visualization to understand and improve your model.

<CardGroup>
  <Card title="Watch Video" icon="video" cta="Learn more" href="https://learn.kodekloud.com/user/courses/machine-learning-fundamentals/module/926a0374-de9f-42b2-b43b-6a07b475c935/lesson/248c2f8b-4cce-47a5-ace1-ec448279e09c" />

  <Card title="Practice Lab" icon="flask-conical" cta="Learn more" href="https://learn.kodekloud.com/user/courses/machine-learning-fundamentals/module/926a0374-de9f-42b2-b43b-6a07b475c935/lesson/77fecfe9-8bb1-4f6a-8e84-12df922ab5c8" />
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.