> ## Documentation Index
> Fetch the complete documentation index at: https://notes.kodekloud.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Inside a Neural Network Neurons Layers and Activations

> Explains neural network basics including neurons, layers, activations, forward pass, training, and MNIST classification with ReLU and softmax.

Neural networks bring together many core machine-learning ideas: features, weights, biases, predictions, loss functions, gradient descent, classification, and regression — all inside a single model.

At a high level, a neural network is a machine-learning model composed of many small mathematical units connected together. These units are commonly called neurons, inspired loosely by biological brains. When organized into layers, these neurons can learn complex patterns that simple linear models cannot.

A brief history: simple artificial neurons were first modeled in the 1940s, and the practical training of multi-layer networks became possible with backpropagation in the 1980s. Today, neural networks are widely applied because of the availability of large datasets, faster compute, and improved training methods.

## The single artificial neuron

A neuron receives inputs, multiplies each input by a corresponding weight, sums the results, and adds a bias. Inputs are the features, weights scale each feature’s contribution, and the bias shifts the summed value. This is the same calculation used in linear models but expressed at the single-unit level.

For two inputs:
sum = w1 \* x1 + w2 \* x2 + b

Here x1 and x2 are input features, w1 and w2 are weights, and b is the bias. One neuron produces a single scalar output.

Example (Python-style) forward computation for a single neuron:

```python theme={null}
# Single neuron (forward computation)
def neuron(x1, x2, w1, w2, b):
    return w1 * x1 + w2 * x2 + b
```

One neuron is limited in expressiveness. The power of neural networks appears when many neurons are connected into layers.

## Layers and network architecture

A typical feedforward neural network consists of:

* Input layer: receives the raw data (features).
* Hidden layer(s): perform intermediate processing and feature extraction.
* Output layer: produces the final prediction or scores.

Stacking only linear transformations (weights + biases) results in another linear transformation. To learn complex, nonlinear relationships we must insert nonlinear activation functions between layers.

| Layer type | Role / Use case | Example |
| - | - | - |
| Input layer | Receives raw features (e.g., image pixels) | `flattened 28x28 image → 784 inputs` |
| Hidden layer | Learns intermediate representations; can be many layers | Dense layer with ReLU activation |
| Output layer | Produces final scores or predictions | 10 logits for MNIST digit classification |

## Activation functions

An activation function is a simple nonlinearity applied to a neuron’s output before passing it to the next layer. These nonlinearities let networks approximate functions beyond straight lines.

Common activation functions:

| Activation | Formula / Behavior | Typical use |
| - | - | - |
| ReLU | `ReLU(x) = max(0, x)` | Hidden layers for sparse activations and good empirical performance |
| Sigmoid | `σ(x) = 1 / (1 + e^{-x})` | Binary outputs or gating (less common in deep hidden layers today) |
| Tanh | `tanh(x)` | Centered outputs in \[-1, 1] |
| Softmax | `softmax(z)_i = e^{z_i} / sum_j e^{z_j}` | Convert logits to class probabilities (output layer) |

A very common activation is ReLU (Rectified Linear Unit): ReLU(x) = max(0, x). It zeroes out negatives and keeps positives unchanged:

* ReLU(-3) = 0
* ReLU(5) = 5

Simple ReLU implementation:

```python theme={null}
def relu(x):
    return max(0.0, x)
```

<Callout icon="lightbulb" color="#1CB2FE">
  Activation functions (like ReLU) are essential: without them, a multi-layer network collapses to an effective single linear transformation, no matter how many layers are stacked.
</Callout>

## A concrete example: MNIST digit classification

MNIST is a widely used dataset of handwritten digits. Each example is typically a 28×28 grayscale image of a digit 0–9. The neural network input is the pixel intensities; the output is a probability distribution over the 10 digit classes.

Typical forward pass (conceptual):

* Pixels are provided to the input layer (often flattened or encoded).
* Hidden layers compute weighted sums, add biases, and apply activation functions (e.g., ReLU).
* Output layer produces raw scores (logits) for each class.
* Softmax converts logits to probabilities.
* The predicted class is the one with the highest probability.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/ng62koO4R6C-1H0_/images/Machine-Learning-Fundamentals/Neural-Networks/Inside-a-Neural-Network-Neurons-Layers-and-Activations/mnist-handwritten-digits-neural-network-relu.jpg?fit=max&auto=format&n=ng62koO4R6C-1H0_&q=85&s=295193f43c05e5e02ce6eab128a73a8e" alt="A hand-drawn diagram showing MNIST handwritten digit images on the left being fed into a neural network on the right with a forward pass and ReLU activation. The network outputs class probabilities (example: 7 → 92%), illustrating digit classification." width="1920" height="1080" data-path="images/Machine-Learning-Fundamentals/Neural-Networks/Inside-a-Neural-Network-Neurons-Layers-and-Activations/mnist-handwritten-digits-neural-network-relu.jpg" />
</Frame>

Converting logits to probabilities (softmax) in code:

```python theme={null}
import math

def softmax(logits):
    exps = [math.exp(l) for l in logits]
    s = sum(exps)
    return [e / s for e in exps]

# Example logits and probabilities
logits = [1.2, 0.3, 2.1]  # scores for three classes
probs = softmax(logits)
# probs is a list of probabilities summing to 1
```

This is a classification problem: the model selects one of several discrete classes. During training, the network uses a loss function (e.g., cross-entropy for classification) and backpropagation plus an optimizer (e.g., SGD, Adam) to adjust weights and biases so predictions become more accurate.

## Summary

* A neural network is a stack of layers composed of many neurons; each neuron computes a weighted sum plus bias.
* Activation functions introduce the critical nonlinearities that enable networks to model complex relationships.
* The forward pass maps inputs to predictions; training with backpropagation and optimization updates weights and biases.
* For classification tasks like MNIST, the output logits are converted with softmax to class probabilities, and the class with the highest probability is selected.

## Links and references

* [MNIST dataset](https://yann.lecun.com/exdb/mnist/)
* [ReLU (Rectified Linear Unit)](https://en.wikipedia.org/wiki/Rectifier_\(neural_networks\))
* [Softmax function](https://en.wikipedia.org/wiki/Softmax_function)
* [Backpropagation](https://en.wikipedia.org/wiki/Backpropagation)
* [Kubernetes Documentation](https://kubernetes.io/docs/) (general reference)

<CardGroup>
  <Card title="Watch Video" icon="video" cta="Learn more" href="https://learn.kodekloud.com/user/courses/machine-learning-fundamentals/module/1e6bfbc1-8908-45f8-a2bf-b37859eaea61/lesson/17ef81db-f438-4902-9aa8-171231a1174a" />
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.