Parameters: weights and bias
For the simplest case, suppose the model uses a single feature: square footage. A very basic model (a simple linear model) makes predictions with this equation: prediction (ŷ) = W × X + B where:- ŷ is the predicted value (predicted house price),
- X is the input feature (square footage),
- W is the weight (how strongly square footage affects price),
- B is the bias (a learned baseline offset).
Multiple features (vectorized view)
Real models usually use many features. If we include square footage, number of bedrooms, age, neighborhood, lot size, etc., the linear equation extends to: prediction = W1 × X1 + W2 × X2 + … + Wn × Xn + B Each feature Xi has its own weight Wi, and B remains a single bias term. All of these learned numbers (weights and bias) are the model’s parameters — when we say a model is learning, we usually mean its parameters are being updated. Vectorized form (conceptually):- X is the input vector [X1, X2, …, Xn]
- W is the weight vector [W1, W2, …, Wn]
- prediction = W · X + B
How model structure affects parameter count
Different model families represent the prediction function differently and therefore contain different kinds and numbers of parameters:- Linear models: one weight per feature plus a bias. A linear model with 79 variables typically has ~80 parameters (after encoding categorical variables, that count can increase).
- Decision trees: parameters are encoded by the tree structure (split thresholds, feature choices) rather than a simple weight vector.
- Neural networks: many layers of weighted sums and nonlinearities, leading to far larger parameter counts even for modest architectures.
- Large language models (LLMs): deep networks with billions of parameters.
Despite the wide range in scale, the core idea is the same: parameters are numeric values learned to shape predictions.
In short: parameters = learned internal numbers (weights and biases) inside the model. Learning means adjusting these parameters so predictions better match observed targets.
Training: loss and optimization
Having parameters does not make a model useful by itself. Training requires:- A loss function: a way to measure how wrong the model’s predictions are (for example, mean squared error for regression).
- An optimizer: an algorithm that updates parameters to reduce the loss (for example, gradient descent and its variants like Adam).
- Compute predictions using current parameters.
- Measure error with the loss function.
- Compute gradients (how to change parameters to reduce error).
- Update parameters via the optimizer. Repeat until convergence or until validation metrics indicate a good fit.
Further reading and references
- Kubernetes Documentation (general reference)
- GPT-3 research
- Llama 2 (Meta)