> ## Documentation Index
> Fetch the complete documentation index at: https://notes.kodekloud.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Pre training Tokens Embeddings and Next Token Prediction

> Explains LLM pre-training covering tokenization, embeddings, and next-token prediction including training loop, objectives, and implications for downstream fine-tuning and model capabilities.

Now that we've covered neural networks at a conceptual level, we can explain how large language models (LLMs) become useful assistants like ChatGPT. This article focuses on the first major phase of building an LLM: pre-training.

Pre-training teaches a model broad language ability by exposing it to very large text corpora and training it to predict the next piece of text. After pre-training, models usually go through post-training steps such as instruction fine-tuning (to follow human directions) and techniques like [Reinforcement Learning from Human Feedback (RLHF)](https://openai.com/research/fine-tuning-language-models-from-human-preferences) to shape behavior toward helpfulness, safety, and alignment.

<Callout icon="lightbulb" color="#1CB2FE">
  Pre-training builds a general-purpose language understanding by solving the statistically rich task of next-token prediction across billions of tokens. Later fine-tuning tailors that general knowledge to specific user-facing behaviors.
</Callout>

## Tokens — how text becomes model input

Most models do not operate directly on raw characters or full words. Before text enters the model, it is broken into smaller units called tokens:

* A token can be a whole word, a subword, punctuation, or even a single character depending on the tokenizer used.
* Common tokenization algorithms include Byte-Pair Encoding (BPE) and WordPiece; each chooses subword segments to balance vocabulary size and coverage.
* Tokens are still symbolic. The model requires numbers, so each token is mapped to a numerical representation called an embedding — a fixed-length vector of floats.

Table: Tokenization overview

| Concept | Purpose | Examples / Notes |
| - | - | - |
| Token | Atomic input unit for the model | Can be word, subword, punctuation |
| Tokenizer | Converts raw text → tokens | `BPE` ([Byte-Pair Encoding](https://en.wikipedia.org/wiki/Byte_pair_encoding)), `WordPiece` ([WordPiece](https://en.wikipedia.org/wiki/WordPiece)) |
| Embedding | Numeric vector representing a token | Typical sizes: 768, 1024, 2048, ... |

Examples of tokenizers:

* BPE: merges frequent symbol pairs into tokens to form a compact vocabulary.
* WordPiece: similar to BPE but uses a different likelihood objective during building.
* Unigram: uses probabilistic subword sampling.

## Embeddings — numerical representations with semantic structure

Each token id is converted to an embedding vector before being processed by the model:

```python theme={null}
# Illustrative embedding (not real model values)
embedding = [0.12, -0.03, 0.78, ...]  # length typically 768, 1024, 2048, etc.
```

Key points:

* Embeddings place tokens into a continuous vector space where semantically similar tokens lie near each other.
* This geometric structure allows models to learn linguistic analogies and relationships (e.g., `king` ≈ `queen` shifted along a gender vector).
* Embeddings are learned jointly with the rest of the model parameters during pre-training.

## Next-token prediction — the core training objective

Pre-training commonly uses next-token prediction (autoregressive learning):

* The model is trained to predict the next token given preceding tokens (the context).
* Example: given the context `"The capital of France is"`, the model should assign high probability to tokens forming `"Paris"`.
* Although the single-step objective is local, applying it at scale forces the model to internalize syntax, semantics, facts, formatting patterns (including code), and discourse behaviors.

Why next-token prediction works:

* It encourages models to learn long-range dependencies and factual associations by minimizing prediction error across diverse corpora.
* The objective is simple to implement and scales linearly with data, making it effective for very large models trained on massive text sources.

## Training loop — what actually happens

The high-level training loop resembles other supervised learning workflows:

1. Tokenize raw text to obtain token IDs for a batch.
2. Map tokens to embeddings and run them forward through the model to produce logits (scores) over the vocabulary for each next-token position.
3. Compute a loss (commonly cross-entropy) between predicted logits and the true next-token targets.
4. Backpropagate gradients and update model parameters using an optimizer such as Adam.
5. Repeat across many batches and many epochs over diverse text.

A minimal pseudocode sketch for a next-token prediction training step:

```python theme={null}
# Pseudocode for a next-token prediction training step
inputs = tokenize(batch_of_text)              # convert text -> token ids
embeddings = model.embed(inputs)              # token ids -> vectors
logits = model.forward(embeddings)            # model produces scores over vocabulary
loss = cross_entropy_loss(logits, targets)    # measure prediction error
optimizer.zero_grad()                         # clear previous gradients
loss.backward()                               # compute gradients
optimizer.step()                              # apply parameter updates
```

Helpful references:

* Loss: `cross-entropy` — measures how well the predicted distribution matches the true token.
* Optimization: `Adam` — adaptive gradient method commonly used for transformer pre-training.
* Gradients: computed by `backpropagation`.

<Callout icon="warning" color="#FF6B6B">
  Pre-training at scale requires large compute, massive datasets, and careful engineering. It can also amplify biases or memorize sensitive information from the training data, so dataset curation and privacy-aware practices are critical.
</Callout>

## Why next-token prediction is powerful

* Predicting the next token forces the model to encode grammar, semantics, topical coherence, and frequent real-world facts.
* Trained across diverse sources (books, articles, code, web pages), the model internalizes many formats and styles, enabling strong zero-shot and few-shot generalization.
* Because the same objective is applied at huge scale, emergent capabilities often arise: the base model can be adapted to many downstream tasks with modest additional fine-tuning or prompt engineering.

## Putting it together

* Pre-training pipeline summary:
  1. Raw text → tokenization → token ids.
  2. Token ids → learned embeddings.
  3. Model predicts next tokens, loss computed, gradients applied.
  4. Repeat across billions of tokens to tune billions of parameters.
* After pre-training, the base model can be fine-tuned for instruction following, safety alignment, or domain-specific tasks using supervised fine-tuning and techniques like RLHF.

Further reading and references:

* [Byte-Pair Encoding (BPE)](https://en.wikipedia.org/wiki/Byte_pair_encoding)
* [WordPiece](https://en.wikipedia.org/wiki/WordPiece)
* [Cross-entropy loss](https://en.wikipedia.org/wiki/Cross_entropy)
* [Backpropagation](https://en.wikipedia.org/wiki/Backpropagation)
* [Adam optimizer](https://en.wikipedia.org/wiki/Adam_\(optimization_algorithm\))
* [Reinforcement Learning from Human Feedback (RLHF)](https://openai.com/research/fine-tuning-language-models-from-human-preferences)

<CardGroup>
  <Card title="Watch Video" icon="video" cta="Learn more" href="https://learn.kodekloud.com/user/courses/machine-learning-fundamentals/module/926a0374-de9f-42b2-b43b-6a07b475c935/lesson/f45bce90-7ceb-4e80-a685-688350b1d516" />
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.