> ## Documentation Index
> Fetch the complete documentation index at: https://notes.kodekloud.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Gradient Descent Learning Rate Batches and Epochs

> Explains gradient descent fundamentals, learning rate, batch sizes, epochs, and how they affect model training and convergence.

Once you can measure how wrong a model is (via a loss function), you need a mechanism to change the model’s parameters so the loss decreases. Gradient descent is that mechanism.

At a high level:

* The loss function defines a landscape: higher points = worse predictions, lower points = better predictions.
* Model parameters correspond to a position on that landscape.
* The gradient gives the slope (direction of steepest increase) at the current position.
* To reduce loss, we move parameters in the opposite direction of the gradient — this is gradient descent.

Analogy: imagine a ball on a hilly surface. The ball’s height is the loss. Gradient descent nudges the ball downhill toward a minimum. The gradient is the slope at the ball’s current position and tells us which way “down” is.

Important concepts

* Gradient: partial derivatives of the loss with respect to each parameter; it tells us whether changing a parameter increases or decreases the loss.
* Learning rate (alpha): scales the update step. Too small → slow training. Too large → overshooting/divergence.
* Batch (or mini-batch): a subset of the dataset used to compute a single gradient and update.
* Epoch: one full pass over the entire training dataset.

A single-parameter update (conceptual) looks like:

```python theme={null}
# gradient descent update (conceptual)
W_new = W_old - alpha * dL_dW
```

* `dL_dW` is the gradient of the loss w\.r.t. parameter `W`.
* `alpha` is the learning rate (step size).

<Callout icon="lightbulb" color="#1CB2FE">
  An epoch is one full pass over the entire training dataset. A batch (or mini-batch) is a subset of the training data used for one parameter update.
</Callout>

Training loop (mini-batch SGD)

* Split dataset into batches.
* For each batch:
  1. Predict outputs for the batch.
  2. Compute loss across the batch.
  3. Compute gradient of the loss w\.r.t. parameters.
  4. Update parameters using the gradient and learning rate.
* One epoch completes after every batch has been processed.
* Repeat for multiple epochs until convergence or stopping criteria.

Table: quick reference

| Term | Meaning | Typical choices / examples |
| - | - | - |
| Gradient | Direction and magnitude that increases the loss | Computed via backpropagation for neural nets |
| Learning rate (`alpha`) | Step size for updates | `1e-4`, `1e-3`, `1e-2` (tuned per problem) |
| Batch size | Number of examples per update | `1` (SGD), `32`, `64`, `128` (mini-batch) |
| Epoch | One full pass over training data | 10s–100s depending on data and model |

Why the learning rate matters

* Small `alpha`: slow convergence; may get stuck on plateaus.
* Large `alpha`: can overshoot minima, cause oscillation, or diverge.

<Callout icon="warning" color="#FF6B6B">
  Choosing a learning rate is critical. If training diverges or loss oscillates, reduce the learning rate. If progress is extremely slow, try increasing it or using learning rate schedules / adaptive optimizers (e.g., Adam, RMSprop).
</Callout>

Practical note on batches and epochs

* Mini-batches smooth gradient estimates compared to single-example updates and make training more efficient on modern hardware.
* More epochs give the model more opportunities to adjust, but too many epochs can overfit.

Example: a tiny two-parameter model you can run interactively

* The following Python snippet shows a simple gradient descent on a linear model y = m\*x + b with mean squared error. It illustrates the ideas: gradient computation, learning rate, batches, and epochs.

```python theme={null}
import random
import math

# synthetic data: y = 2*x + 1 with some noise
data = [(x, 2*x + 1 + random.uniform(-0.5, 0.5)) for x in range(-50, 51)]

# parameters (m = slope, b = intercept)
m, b = 0.0, 0.0
alpha = 0.01      # learning rate
batch_size = 16
epochs = 50

def mse_gradients(batch, m, b):
    # compute gradients dL/dm and dL/db for the batch
    dm, db = 0.0, 0.0
    n = len(batch)
    for x, y in batch:
        y_pred = m * x + b
        err = y_pred - y
        dm += (2 / n) * err * x
        db += (2 / n) * err
    return dm, db

for epoch in range(epochs):
    random.shuffle(data)
    # process mini-batches
    for i in range(0, len(data), batch_size):
        batch = data[i : i + batch_size]
        dm, db = mse_gradients(batch, m, b)
        # gradient descent update
        m -= alpha * dm
        b -= alpha * db
    # compute full-dataset MSE to monitor progress
    mse = sum((m*x + b - y)**2 for x, y in data) / len(data)
    if epoch % 10 == 0:
        print(f"epoch {epoch:3d}  m={m:.4f}  b={b:.4f}  mse={mse:.4f}")

print("final:", m, b)
```

Run this script and tweak `alpha`, `batch_size`, and `epochs` to see how they change convergence and final error.

Further reading and references

* [Gradient Descent — Wikipedia](https://en.wikipedia.org/wiki/Gradient_descent)
* [Stochastic Gradient Descent (SGD) — Stanford CS231n notes](http://cs231n.github.io/neural-networks-3/)
* \[Adaptive optimizers (Adam, RMSprop) — Practical guides]

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/qBe6x55gupKpUs7F/images/Machine-Learning-Fundamentals/How-Models-Learn/Gradient-Descent-Learning-Rate-Batches-and-Epochs/gradient-descent-neural-network-house-prices.jpg?fit=max&auto=format&n=qBe6x55gupKpUs7F&q=85&s=22c05ed3f2c0f9b5dd6f1d0119d79054" alt="A handwritten diagram explaining gradient descent: a house-prices dataset table on the left with arrows for batches and epochs, and a neural network sketch on the right showing prediction, updates, and loss-driven parameter updates." width="1920" height="1080" data-path="images/Machine-Learning-Fundamentals/How-Models-Learn/Gradient-Descent-Learning-Rate-Batches-and-Epochs/gradient-descent-neural-network-house-prices.jpg" />
</Frame>

Summary

* The loss function measures error; the gradient points where the loss increases fastest.
* Gradient descent updates parameters in the opposite direction to reduce loss.
* Learning rate, batch size, and number of epochs are key hyperparameters that control training dynamics.
* Iterative updates across batches and epochs — not single-step fixes — is how a model learns from data.

<CardGroup>
  <Card title="Watch Video" icon="video" cta="Learn more" href="https://learn.kodekloud.com/user/courses/machine-learning-fundamentals/module/f4db2a23-638a-40de-b672-cbe3cae21cc2/lesson/000da2d3-83ef-40f4-8105-90d1e51607e2" />

  <Card title="Practice Lab" icon="flask-conical" cta="Learn more" href="https://learn.kodekloud.com/user/courses/machine-learning-fundamentals/module/f4db2a23-638a-40de-b672-cbe3cae21cc2/lesson/fd933298-5078-4530-8d6d-079221442565" />
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.