The single artificial neuron
A neuron receives inputs, multiplies each input by a corresponding weight, sums the results, and adds a bias. Inputs are the features, weights scale each feature’s contribution, and the bias shifts the summed value. This is the same calculation used in linear models but expressed at the single-unit level. For two inputs: sum = w1 * x1 + w2 * x2 + b Here x1 and x2 are input features, w1 and w2 are weights, and b is the bias. One neuron produces a single scalar output. Example (Python-style) forward computation for a single neuron:Layers and network architecture
A typical feedforward neural network consists of:- Input layer: receives the raw data (features).
- Hidden layer(s): perform intermediate processing and feature extraction.
- Output layer: produces the final prediction or scores.
Activation functions
An activation function is a simple nonlinearity applied to a neuron’s output before passing it to the next layer. These nonlinearities let networks approximate functions beyond straight lines. Common activation functions:
A very common activation is ReLU (Rectified Linear Unit): ReLU(x) = max(0, x). It zeroes out negatives and keeps positives unchanged:
- ReLU(-3) = 0
- ReLU(5) = 5
Activation functions (like ReLU) are essential: without them, a multi-layer network collapses to an effective single linear transformation, no matter how many layers are stacked.
A concrete example: MNIST digit classification
MNIST is a widely used dataset of handwritten digits. Each example is typically a 28×28 grayscale image of a digit 0–9. The neural network input is the pixel intensities; the output is a probability distribution over the 10 digit classes. Typical forward pass (conceptual):- Pixels are provided to the input layer (often flattened or encoded).
- Hidden layers compute weighted sums, add biases, and apply activation functions (e.g., ReLU).
- Output layer produces raw scores (logits) for each class.
- Softmax converts logits to probabilities.
- The predicted class is the one with the highest probability.

Summary
- A neural network is a stack of layers composed of many neurons; each neuron computes a weighted sum plus bias.
- Activation functions introduce the critical nonlinearities that enable networks to model complex relationships.
- The forward pass maps inputs to predictions; training with backpropagation and optimization updates weights and biases.
- For classification tasks like MNIST, the output logits are converted with softmax to class probabilities, and the class with the highest probability is selected.
Links and references
- MNIST dataset
- ReLU (Rectified Linear Unit)
- Softmax function
- Backpropagation
- Kubernetes Documentation (general reference)