Pre-training builds a general-purpose language understanding by solving the statistically rich task of next-token prediction across billions of tokens. Later fine-tuning tailors that general knowledge to specific user-facing behaviors.
Tokens — how text becomes model input
Most models do not operate directly on raw characters or full words. Before text enters the model, it is broken into smaller units called tokens:- A token can be a whole word, a subword, punctuation, or even a single character depending on the tokenizer used.
- Common tokenization algorithms include Byte-Pair Encoding (BPE) and WordPiece; each chooses subword segments to balance vocabulary size and coverage.
- Tokens are still symbolic. The model requires numbers, so each token is mapped to a numerical representation called an embedding — a fixed-length vector of floats.
Examples of tokenizers:
- BPE: merges frequent symbol pairs into tokens to form a compact vocabulary.
- WordPiece: similar to BPE but uses a different likelihood objective during building.
- Unigram: uses probabilistic subword sampling.
Embeddings — numerical representations with semantic structure
Each token id is converted to an embedding vector before being processed by the model:- Embeddings place tokens into a continuous vector space where semantically similar tokens lie near each other.
- This geometric structure allows models to learn linguistic analogies and relationships (e.g.,
king≈queenshifted along a gender vector). - Embeddings are learned jointly with the rest of the model parameters during pre-training.
Next-token prediction — the core training objective
Pre-training commonly uses next-token prediction (autoregressive learning):- The model is trained to predict the next token given preceding tokens (the context).
- Example: given the context
"The capital of France is", the model should assign high probability to tokens forming"Paris". - Although the single-step objective is local, applying it at scale forces the model to internalize syntax, semantics, facts, formatting patterns (including code), and discourse behaviors.
- It encourages models to learn long-range dependencies and factual associations by minimizing prediction error across diverse corpora.
- The objective is simple to implement and scales linearly with data, making it effective for very large models trained on massive text sources.
Training loop — what actually happens
The high-level training loop resembles other supervised learning workflows:- Tokenize raw text to obtain token IDs for a batch.
- Map tokens to embeddings and run them forward through the model to produce logits (scores) over the vocabulary for each next-token position.
- Compute a loss (commonly cross-entropy) between predicted logits and the true next-token targets.
- Backpropagate gradients and update model parameters using an optimizer such as Adam.
- Repeat across many batches and many epochs over diverse text.
- Loss:
cross-entropy— measures how well the predicted distribution matches the true token. - Optimization:
Adam— adaptive gradient method commonly used for transformer pre-training. - Gradients: computed by
backpropagation.
Pre-training at scale requires large compute, massive datasets, and careful engineering. It can also amplify biases or memorize sensitive information from the training data, so dataset curation and privacy-aware practices are critical.
Why next-token prediction is powerful
- Predicting the next token forces the model to encode grammar, semantics, topical coherence, and frequent real-world facts.
- Trained across diverse sources (books, articles, code, web pages), the model internalizes many formats and styles, enabling strong zero-shot and few-shot generalization.
- Because the same objective is applied at huge scale, emergent capabilities often arise: the base model can be adapted to many downstream tasks with modest additional fine-tuning or prompt engineering.
Putting it together
- Pre-training pipeline summary:
- Raw text → tokenization → token ids.
- Token ids → learned embeddings.
- Model predicts next tokens, loss computed, gradients applied.
- Repeat across billions of tokens to tune billions of parameters.
- After pre-training, the base model can be fine-tuned for instruction following, safety alignment, or domain-specific tasks using supervised fine-tuning and techniques like RLHF.
- Byte-Pair Encoding (BPE)
- WordPiece
- Cross-entropy loss
- Backpropagation
- Adam optimizer
- Reinforcement Learning from Human Feedback (RLHF)