- The loss function defines a landscape: higher points = worse predictions, lower points = better predictions.
- Model parameters correspond to a position on that landscape.
- The gradient gives the slope (direction of steepest increase) at the current position.
- To reduce loss, we move parameters in the opposite direction of the gradient — this is gradient descent.
- Gradient: partial derivatives of the loss with respect to each parameter; it tells us whether changing a parameter increases or decreases the loss.
- Learning rate (alpha): scales the update step. Too small → slow training. Too large → overshooting/divergence.
- Batch (or mini-batch): a subset of the dataset used to compute a single gradient and update.
- Epoch: one full pass over the entire training dataset.
dL_dWis the gradient of the loss w.r.t. parameterW.alphais the learning rate (step size).
An epoch is one full pass over the entire training dataset. A batch (or mini-batch) is a subset of the training data used for one parameter update.
- Split dataset into batches.
- For each batch:
- Predict outputs for the batch.
- Compute loss across the batch.
- Compute gradient of the loss w.r.t. parameters.
- Update parameters using the gradient and learning rate.
- One epoch completes after every batch has been processed.
- Repeat for multiple epochs until convergence or stopping criteria.
Why the learning rate matters
- Small
alpha: slow convergence; may get stuck on plateaus. - Large
alpha: can overshoot minima, cause oscillation, or diverge.
Choosing a learning rate is critical. If training diverges or loss oscillates, reduce the learning rate. If progress is extremely slow, try increasing it or using learning rate schedules / adaptive optimizers (e.g., Adam, RMSprop).
- Mini-batches smooth gradient estimates compared to single-example updates and make training more efficient on modern hardware.
- More epochs give the model more opportunities to adjust, but too many epochs can overfit.
- The following Python snippet shows a simple gradient descent on a linear model y = m*x + b with mean squared error. It illustrates the ideas: gradient computation, learning rate, batches, and epochs.
alpha, batch_size, and epochs to see how they change convergence and final error.
Further reading and references
- Gradient Descent — Wikipedia
- Stochastic Gradient Descent (SGD) — Stanford CS231n notes
- [Adaptive optimizers (Adam, RMSprop) — Practical guides]

- The loss function measures error; the gradient points where the loss increases fastest.
- Gradient descent updates parameters in the opposite direction to reduce loss.
- Learning rate, batch size, and number of epochs are key hyperparameters that control training dynamics.
- Iterative updates across batches and epochs — not single-step fixes — is how a model learns from data.