Skip to main content
Once you can measure how wrong a model is (via a loss function), you need a mechanism to change the model’s parameters so the loss decreases. Gradient descent is that mechanism. At a high level:
  • The loss function defines a landscape: higher points = worse predictions, lower points = better predictions.
  • Model parameters correspond to a position on that landscape.
  • The gradient gives the slope (direction of steepest increase) at the current position.
  • To reduce loss, we move parameters in the opposite direction of the gradient — this is gradient descent.
Analogy: imagine a ball on a hilly surface. The ball’s height is the loss. Gradient descent nudges the ball downhill toward a minimum. The gradient is the slope at the ball’s current position and tells us which way “down” is. Important concepts
  • Gradient: partial derivatives of the loss with respect to each parameter; it tells us whether changing a parameter increases or decreases the loss.
  • Learning rate (alpha): scales the update step. Too small → slow training. Too large → overshooting/divergence.
  • Batch (or mini-batch): a subset of the dataset used to compute a single gradient and update.
  • Epoch: one full pass over the entire training dataset.
A single-parameter update (conceptual) looks like:
  • dL_dW is the gradient of the loss w.r.t. parameter W.
  • alpha is the learning rate (step size).
An epoch is one full pass over the entire training dataset. A batch (or mini-batch) is a subset of the training data used for one parameter update.
Training loop (mini-batch SGD)
  • Split dataset into batches.
  • For each batch:
    1. Predict outputs for the batch.
    2. Compute loss across the batch.
    3. Compute gradient of the loss w.r.t. parameters.
    4. Update parameters using the gradient and learning rate.
  • One epoch completes after every batch has been processed.
  • Repeat for multiple epochs until convergence or stopping criteria.
Table: quick reference Why the learning rate matters
  • Small alpha: slow convergence; may get stuck on plateaus.
  • Large alpha: can overshoot minima, cause oscillation, or diverge.
Choosing a learning rate is critical. If training diverges or loss oscillates, reduce the learning rate. If progress is extremely slow, try increasing it or using learning rate schedules / adaptive optimizers (e.g., Adam, RMSprop).
Practical note on batches and epochs
  • Mini-batches smooth gradient estimates compared to single-example updates and make training more efficient on modern hardware.
  • More epochs give the model more opportunities to adjust, but too many epochs can overfit.
Example: a tiny two-parameter model you can run interactively
  • The following Python snippet shows a simple gradient descent on a linear model y = m*x + b with mean squared error. It illustrates the ideas: gradient computation, learning rate, batches, and epochs.
Run this script and tweak alpha, batch_size, and epochs to see how they change convergence and final error. Further reading and references
A handwritten diagram explaining gradient descent: a house-prices dataset table on the left with arrows for batches and epochs, and a neural network sketch on the right showing prediction, updates, and loss-driven parameter updates.
Summary
  • The loss function measures error; the gradient points where the loss increases fastest.
  • Gradient descent updates parameters in the opposite direction to reduce loss.
  • Learning rate, batch size, and number of epochs are key hyperparameters that control training dynamics.
  • Iterative updates across batches and epochs — not single-step fixes — is how a model learns from data.

Watch Video

Practice Lab