
Attention (self-attention): the key idea
The technical advance that enabled transformers is attention—specifically self-attention. Self-attention lets the model consider every token (word or subword) in the input and compute how strongly each token should influence the representation of every other token. Modeling these pairwise relationships across the whole sequence is what lets transformers resolve many language phenomena earlier models found difficult. Example — pronoun resolution: “The dog chased its tail because it was bored.” Which does “it” refer to — the dog or the tail? Humans resolve this instantly (we know “it” refers to the dog). Transformers are much better at this because they can attend to the relationships among all words simultaneously.
Pre-trained: learning from massive text corpora
“Pre-trained” means the model learns statistical patterns from a very large and diverse collection of text before it ever interacts with users. During pre-training, the model reads books, web pages, articles, code, and other text to learn how language typically flows. The primary objective used for GPT-style models is next-token prediction: given a context, predict the next token. Through billions of examples the model adjusts its parameters to improve next-token prediction across many contexts. This is pattern learning, not memorization of facts. Examples of next-token prediction patterns:- Given “The capital of France is”, a likely next token is “Paris”.
- Given “The cat sat on the”, a likely next token is “mat”.
- Given “Water boils at 100 degrees”, a likely next token is “Celsius”.
Pre-training teaches the model statistical patterns across language. The model does not store a structured database of facts to “look up.” Instead it generates text by sampling likely next tokens given the context it has learned.
Generative: producing new text token-by-token
“Generative” simply means the model composes responses one token at a time rather than returning a pre-written answer. At inference, each next token is chosen according to the model’s learned probabilities and the chosen decoding strategy (greedy, beam search, sampling, etc.). Because the model generates text from learned patterns rather than consulting an authoritative source, it can produce incorrect or fabricated statements. These incorrect outputs are commonly called hallucinations — when the model’s most probable continuation does not reflect reality or overgeneralizes from noisy training data.
GPT family and other LLMs
GPT is a family of models that evolved over time. Each new generation typically used larger model sizes, more diverse data, and more compute during training to improve the model’s ability to predict plausible next tokens in many contexts.
GPT is one family among many modern large language models (LLMs). Other notable models include Anthropic’s Claude, Google’s Gemini, and Meta’s Llama (open source). They share the same high-level ingredients: transformer architecture + large-scale pre-training + generative decoding.
Transformers are used beyond text: variants support models for audio, images, and video. While some models are explicitly multimodal, the L in LLM denotes a specialization in language tasks.

Key takeaways
- Transformers and self-attention allow models to model relationships among all tokens in a sequence.
- Pre-training on massive corpora with a next-token objective teaches statistical language patterns.
- Because generation is pattern-based (not a lookup), models can hallucinate — verify critical facts.
- The GPT family evolved through successive generations; other LLMs use similar core ideas.
Because GPT-style models generate text from patterns learned during training, always verify critical facts against authoritative sources—especially for code, health, legal, or other high-stakes decisions.
Links and references
- Attention Is All You Need (Vaswani et al., 2017)
- GPT-4 research overview
- Anthropic Claude
- Google Gemini
- Meta Llama