Skip to main content
Now that we’ve covered neural networks at a conceptual level, we can explain how large language models (LLMs) become useful assistants like ChatGPT. This article focuses on the first major phase of building an LLM: pre-training. Pre-training teaches a model broad language ability by exposing it to very large text corpora and training it to predict the next piece of text. After pre-training, models usually go through post-training steps such as instruction fine-tuning (to follow human directions) and techniques like Reinforcement Learning from Human Feedback (RLHF) to shape behavior toward helpfulness, safety, and alignment.
Pre-training builds a general-purpose language understanding by solving the statistically rich task of next-token prediction across billions of tokens. Later fine-tuning tailors that general knowledge to specific user-facing behaviors.

Tokens — how text becomes model input

Most models do not operate directly on raw characters or full words. Before text enters the model, it is broken into smaller units called tokens:
  • A token can be a whole word, a subword, punctuation, or even a single character depending on the tokenizer used.
  • Common tokenization algorithms include Byte-Pair Encoding (BPE) and WordPiece; each chooses subword segments to balance vocabulary size and coverage.
  • Tokens are still symbolic. The model requires numbers, so each token is mapped to a numerical representation called an embedding — a fixed-length vector of floats.
Table: Tokenization overview Examples of tokenizers:
  • BPE: merges frequent symbol pairs into tokens to form a compact vocabulary.
  • WordPiece: similar to BPE but uses a different likelihood objective during building.
  • Unigram: uses probabilistic subword sampling.

Embeddings — numerical representations with semantic structure

Each token id is converted to an embedding vector before being processed by the model:
Key points:
  • Embeddings place tokens into a continuous vector space where semantically similar tokens lie near each other.
  • This geometric structure allows models to learn linguistic analogies and relationships (e.g., king ≈ queen shifted along a gender vector).
  • Embeddings are learned jointly with the rest of the model parameters during pre-training.

Next-token prediction — the core training objective

Pre-training commonly uses next-token prediction (autoregressive learning):
  • The model is trained to predict the next token given preceding tokens (the context).
  • Example: given the context "The capital of France is", the model should assign high probability to tokens forming "Paris".
  • Although the single-step objective is local, applying it at scale forces the model to internalize syntax, semantics, facts, formatting patterns (including code), and discourse behaviors.
Why next-token prediction works:
  • It encourages models to learn long-range dependencies and factual associations by minimizing prediction error across diverse corpora.
  • The objective is simple to implement and scales linearly with data, making it effective for very large models trained on massive text sources.

Training loop — what actually happens

The high-level training loop resembles other supervised learning workflows:
  1. Tokenize raw text to obtain token IDs for a batch.
  2. Map tokens to embeddings and run them forward through the model to produce logits (scores) over the vocabulary for each next-token position.
  3. Compute a loss (commonly cross-entropy) between predicted logits and the true next-token targets.
  4. Backpropagate gradients and update model parameters using an optimizer such as Adam.
  5. Repeat across many batches and many epochs over diverse text.
A minimal pseudocode sketch for a next-token prediction training step:
Helpful references:
  • Loss: cross-entropy — measures how well the predicted distribution matches the true token.
  • Optimization: Adam — adaptive gradient method commonly used for transformer pre-training.
  • Gradients: computed by backpropagation.
Pre-training at scale requires large compute, massive datasets, and careful engineering. It can also amplify biases or memorize sensitive information from the training data, so dataset curation and privacy-aware practices are critical.

Why next-token prediction is powerful

  • Predicting the next token forces the model to encode grammar, semantics, topical coherence, and frequent real-world facts.
  • Trained across diverse sources (books, articles, code, web pages), the model internalizes many formats and styles, enabling strong zero-shot and few-shot generalization.
  • Because the same objective is applied at huge scale, emergent capabilities often arise: the base model can be adapted to many downstream tasks with modest additional fine-tuning or prompt engineering.

Putting it together

  • Pre-training pipeline summary:
    1. Raw text → tokenization → token ids.
    2. Token ids → learned embeddings.
    3. Model predicts next tokens, loss computed, gradients applied.
    4. Repeat across billions of tokens to tune billions of parameters.
  • After pre-training, the base model can be fine-tuned for instruction following, safety alignment, or domain-specific tasks using supervised fine-tuning and techniques like RLHF.
Further reading and references:

Watch Video