Skip to main content
Once a model is trained, we need objective ways to measure how well it performs. Evaluation metrics are numbers (or sets of numbers) that quantify model performance and help answer questions like: How close are predictions to the true values? How often is the model correct? Does the model generalize to new, unseen data? Evaluation always ties back to how data is split into training, validation, and test sets. We care not only about performance on the training data but, crucially, about performance on held-out (validation/test) data.

Regression (predicting continuous values)

Main question: how far are predictions from the true values? Regression metrics measure the size and distribution of errors. Key metrics: Example implementations (numpy):
Best practices:
  • Use MAE when you want a robust, interpretable error measure that is not dominated by a few large outliers.
  • Use MSE/RMSE when large deviations should be penalized more aggressively.
  • Always choose a metric aligned with business or scientific costs (e.g., underestimating vs. overestimating may have different consequences).

Classification (predicting categories)

Main question: which class does the model predict? Common tasks include digit recognition (0–9), spam detection, or multi-class image labels. Baseline metric:
  • Accuracy = (number of correct predictions) / (total predictions)
Accuracy is easy to understand but can be misleading on imbalanced datasets (e.g., a model that always predicts the majority class may still show high accuracy). Confusion matrix (binary case):
  • True Positives (TP)
  • False Positives (FP)
  • True Negatives (TN)
  • False Negatives (FN)
From these we compute:
  • Precision = TP / (TP + FP) — of predicted positives, how many were correct?
  • Recall = TP / (TP + FN) — of actual positives, how many did the model find?
  • F1 score = 2 * (precision * recall) / (precision + recall) — harmonic mean of precision and recall
Example formulas (binary case):
Practical tips:
  • For imbalanced datasets, prefer precision/recall/F1 and inspect the confusion matrix to see the types of errors the model makes.
  • For multi-class problems, compute per-class precision/recall and consider macro/micro averages depending on your evaluation goals.
Accuracy is a good quick check, but for imbalanced datasets or tasks where some errors are costlier than others, examine precision, recall, and the confusion matrix to understand the types of mistakes the model makes.

Large Language Models (LLMs)

Evaluating LLMs is more complex because many tasks are open-ended and can have multiple valid outputs (e.g., explanations, creative writing, code). There is often no single ground truth. Common approaches:
  • Automated benchmarks and task suites (question answering, code generation, math problems)
  • Human preference ratings and pairwise comparisons (which output do humans prefer?)
  • Factuality and hallucination checks (does the model invent facts?)
  • Task-specific tests (e.g., unit tests for generated code, step-by-step verification for math)
  • Real-world feedback and A/B testing to measure user satisfaction
Important caveats:
  • No single metric captures everything. A model can score highly on a benchmark but still hallucinate or fail to follow instructions reliably.
  • Combine automated metrics with human evaluation and task-specific checks to obtain a fuller picture.
Be careful of over-relying on single benchmarks. Metrics optimized in isolation can lead to systems that game the metric without improving real-world usefulness. Also watch for data leakage, label artifacts, and distribution shift between training and deployment.

Quick Reference: Which metric to choose?


Summary

  • Evaluation metrics quantify how useful a trained model is.
  • For regression, focus on error size: MAE (robust), MSE/RMSE (penalize large errors).
  • For classification, accuracy is a starting point; use precision/recall/F1 and confusion matrices to understand error types, especially on imbalanced data.
  • For LLMs, combine automated benchmarks, human judgments, and task-specific tests — expect trade-offs and gaps in any single metric.

Lab: MNIST digit classification (hands-on)

In this lab you will:
  1. Train a neural network to read handwritten digits from the MNIST dataset (https://yann.lecun.com/exdb/mnist/).
  2. Observe training loss decrease epoch-by-epoch.
  3. Evaluate the trained model on 2,000 unseen test examples.
  4. Inspect the confusion matrix and identify the most confused digit pairs.
  5. Examine one or more misclassified samples (e.g., print pixel values or render a small plot) to understand why the model erred — perhaps a sloppy “4” that looks like a “9”.
Example terminal session (illustrative):
How to inspect misclassified examples:
  • After evaluation, save indices of misclassified samples and print the corresponding labels and predictions.
  • Visualize the image using your preferred plotting library (e.g., matplotlib) to see why the model confused the digits.
  • Example snippet to display an MNIST sample:
This hands-on exploration helps connect metric numbers (e.g., recall or confusion pairs) to concrete model behavior and ambiguous inputs.
Good luck with the lab — inspect the confusion matrix closely and use visualization to understand and improve your model.

Watch Video

Practice Lab