> ## Documentation Index
> Fetch the complete documentation index at: https://notes.kodekloud.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Backpropagation and Deep Learning

> Overview of backpropagation, training loops, optimizers, hierarchical feature learning in deep networks, and how scale and architectures like transformers enable large language models.

But during training, a model must learn from its mistakes. How does that happen?

Suppose the correct label is 7 but the model predicts 1. The loss function quantifies how bad that prediction is: high loss for low probability on the correct class, low loss for high probability. In classification this is commonly implemented with a softmax output and cross-entropy loss ([softmax](https://en.wikipedia.org/wiki/Softmax_function), [cross-entropy](https://en.wikipedia.org/wiki/Cross_entropy)).

Backpropagation is the algorithm that propagates the error backwards through the network to compute the gradient of the loss with respect to every parameter (weights and biases). Concretely, it uses the chain rule to determine how each parameter contributed to the loss. Those gradients tell an optimizer which direction — and how much — to change each parameter to reduce the loss.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/qBe6x55gupKpUs7F/images/Machine-Learning-Fundamentals/Neural-Networks/Backpropagation-and-Deep-Learning/handwritten-digits-neural-network-relu-backprop.jpg?fit=max&auto=format&n=qBe6x55gupKpUs7F&q=85&s=ac0a6cfb10a435dce7509b6661947d3e" alt="A collage showing nine handwritten digit images on the left and a colorful hand-drawn diagram of a neural network on the right. The sketch labels layers, weights, activations (ReLU), backpropagation and the prediction/loss process." width="1920" height="1080" data-path="images/Machine-Learning-Fundamentals/Neural-Networks/Backpropagation-and-Deep-Learning/handwritten-digits-neural-network-relu-backprop.jpg" />
</Frame>

This scaling ability is critical because modern neural networks can contain thousands to trillions of parameters. Backpropagation provides the gradients that gradient-based optimizers (for example, SGD with momentum, RMSprop, or Adam) use to update parameters efficiently. Most deep learning frameworks implement this via automatic differentiation, so you rarely write the gradient math by hand ([automatic differentiation](https://en.wikipedia.org/wiki/Automatic_differentiation)).

Training loop — high-level steps

1. Forward pass: the model computes predictions for a batch of examples.
2. Compute loss: a loss function measures prediction error.
3. Backward pass: backpropagation computes gradients of the loss w\.r.t. parameters.
4. Update: an optimizer applies gradients (often scaled by a learning rate) to update parameters.
5. Repeat across many batches and epochs until the model converges.

Example pseudo-code for a basic training loop:

```python theme={null}
for epoch in range(num_epochs):
    for X_batch, y_batch in data_loader:
        preds = model.forward(X_batch)             # forward pass
        loss = loss_fn(preds, y_batch)            # compute loss
        grads = backprop(loss, model.parameters)  # backward pass
        optimizer.step(grads)                     # update parameters
```

Common optimizers and quick notes

| Optimizer | When to use | Characteristics |
| - | - | - |
| `SGD` | Simple problems, baseline | Stochastic updates; can use momentum; requires tuning learning rate |
| `SGD with momentum` | Improves convergence on noisy gradients | Accelerates along consistent gradient directions |
| `RMSprop` | Recurrent networks or nonstationary objectives | Adaptive per-parameter learning rates |
| `Adam` | Most common default | Combines momentum and adaptive rates; robust for many tasks |

Why depth matters: automatic feature learning

Traditional machine learning often depended on manual feature engineering — e.g., for house-price models you might craft features such as home age, number of rooms, or distance to downtown and feed them to a model.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/qBe6x55gupKpUs7F/images/Machine-Learning-Fundamentals/Neural-Networks/Backpropagation-and-Deep-Learning/house-prices-dataset-feature-engineering-model.jpg?fit=max&auto=format&n=qBe6x55gupKpUs7F&q=85&s=04b1ddae518cfc1f54c17b08b44cc108" alt="A blackboard-style diagram showing a table labeled &#x22;Dataset: House Prices&#x22; with columns like sqft, beds, area and $$$, with arrows pointing to a small neural-network sketch labeled &#x22;Model.&#x22; Below the table is a highlighted &#x22;feature engineering&#x22; note." width="1920" height="1080" data-path="images/Machine-Learning-Fundamentals/Neural-Networks/Backpropagation-and-Deep-Learning/house-prices-dataset-feature-engineering-model.jpg" />
</Frame>

Deep networks instead learn hierarchical internal features from raw inputs. For images:

* Early layers learn simple local patterns (edges, corners).
* Middle layers combine those into textures and motifs.
* Deep layers form high-level compositions (object parts or whole objects).

For MNIST, layers may learn pen strokes, loops, and junctions that distinguish digits. This layered composition is the essence of "deep learning" — “deep” refers to multiple stacked layers of computation, not human-like reasoning depth.

A short history: why deep learning became practical

For many years neural networks were limited by hardware, datasets, and optimization methods. Progress required several ingredients coming together: much larger labeled datasets, improved training algorithms, and faster compute (notably GPUs). A landmark was AlexNet (2012), a deep convolutional network that performed outstandingly on the ImageNet benchmark. ImageNet was far more challenging than MNIST, featuring many real-world object categories and much larger scale.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/qBe6x55gupKpUs7F/images/Machine-Learning-Fundamentals/Neural-Networks/Backpropagation-and-Deep-Learning/neural-network-alexnet-imagenet-not-scalable.jpg?fit=max&auto=format&n=qBe6x55gupKpUs7F&q=85&s=1680ab5f2620c54783c34c22cafa937a" alt="A hand-drawn diagram titled &#x22;Neural Network&#x22; showing a neural-net schematic on a timeline with &#x22;AlexNet&#x22; marked around 2010 and &#x22;ImageNet&#x22; noted. To the right are sketched people and checklist notes about weak hardware, small datasets, and &#x22;not scalable.&#x22;" width="1920" height="1080" data-path="images/Machine-Learning-Fundamentals/Neural-Networks/Backpropagation-and-Deep-Learning/neural-network-alexnet-imagenet-not-scalable.jpg" />
</Frame>

From image recognition, neural networks expanded into speech recognition, machine translation, recommendation systems, and — at much larger scale — language models.

Transformers, GPT, and scale

Models like GPT are still neural networks, but they use transformer architectures built from self-attention layers rather than simple fully connected or convolutional stacks. The core training loop remains the same: forward pass, compute loss, backpropagate, update parameters. The main differences are:

* Data type: tokens (words/subwords) instead of pixels.
* Architecture: transformer layers + self-attention mechanisms.
* Scale: far larger parameter counts, dataset sizes, and compute budgets.
* Objective: often next-token prediction or related language modeling losses.

<Callout icon="lightbulb" color="#1CB2FE">
  Parameter counts for large models (including GPT-series models) are not always publicly confirmed; reported values vary. Some sources have suggested very large counts (e.g., on the order of hundreds of billions to trillions), but exact numbers may differ between reports.
</Callout>

Summary

* Backpropagation + gradient-based optimizers enable neural networks to learn from errors by computing and applying parameter gradients.
* Deep networks automatically learn hierarchical features from raw data, reducing the need for manual feature engineering.
* Improvements in data, algorithms, and compute (especially GPUs) made training deep, large-scale networks practical.
* Modern large models (e.g., transformers and GPT) follow the same training principles but differ in architecture, data modality, and scale.

Links and references

* [Softmax function](https://en.wikipedia.org/wiki/Softmax_function)
* [Cross-entropy](https://en.wikipedia.org/wiki/Cross_entropy)
* [Stochastic gradient descent (SGD)](https://en.wikipedia.org/wiki/Stochastic_gradient_descent)
* [RMSprop](https://en.wikipedia.org/wiki/RMSprop)
* [Adam optimizer](https://en.wikipedia.org/wiki/Adam_\(optimization_algorithm\))
* [Automatic differentiation](https://en.wikipedia.org/wiki/Automatic_differentiation)
* [MNIST dataset](http://yann.lecun.com/exdb/mnist/)
* [AlexNet](https://en.wikipedia.org/wiki/AlexNet)
* [ImageNet](https://en.wikipedia.org/wiki/ImageNet)
* [Transformers](https://en.wikipedia.org/wiki/Transformer_\(machine_learning_model\))
* [Attention mechanisms](https://en.wikipedia.org/wiki/Attention_\(machine_learning\))
* [GPT (language model)](https://en.wikipedia.org/wiki/GPT_\(language_model\))

<CardGroup>
  <Card title="Watch Video" icon="video" cta="Learn more" href="https://learn.kodekloud.com/user/courses/machine-learning-fundamentals/module/1e6bfbc1-8908-45f8-a2bf-b37859eaea61/lesson/53fcd449-2f54-4c78-847b-4d6bfdf8395f" />
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.