
- Latency depends on token count: generating hundreds of tokens requires hundreds of prediction steps, so long outputs take longer than short ones.
- Compute and cost scale with token work: a short Q&A might be trivial, while summarizing a 100-page contract can be thousands of times heavier because it involves far more tokens.

- Prefill (prompt processing)
- Streaming (autoregressive decoding)
- Tokenization: the prompt (user text, system messages, instructions) is converted into the model’s token IDs.
- Embedding and forward pass: those tokens are embedded and passed through the Transformer’s layers. During this pass the model computes internal representations and, for each layer, produces key (K) and value (V) vectors used for attention.
- KV cache population: the keys and values for each prompt token are stored in a KV cache. This cache is essential for efficient subsequent decoding.
- Autoregressive loop: decoding generates one token at a time. For each generated token:
- The model uses the KV cache (built from the prompt and previously generated tokens) as context to compute attention for the new token.
- It computes logits for the next-token distribution, selects (or samples) a token, and emits it.
- The generated token is appended and its K/V vectors are added to the KV cache.
- Efficiency through caching: because the KV cache stores prior keys and values, the model doesn’t re-run the full forward pass over the entire prompt for each new token. Each decoding step only needs to compute attention for the new token against cached keys/values, plus the new token’s own projections. That substantially reduces per-token work compared to naively recomputing the entire context each step.
A rough mental model: prefill pays the cost to fully ingest the prompt and prepare KV caches; decoding then pays a small per-token cost leveraging those caches to keep generation tractable. This is also why streaming token-by-token output is both practical and expected in chat interfaces.
Tokens affect cost, latency, and maximum context. APIs typically count both input (prompt) tokens and output (response) tokens toward billing and context limits. Keep prompts concise for fast, low-cost responses; expand the prompt when more context is truly needed.
- Measure both input and output tokens when estimating cost and latency.
- For long context needs, prefer architectures and providers that support larger context windows and efficient KV caching.
- Use streaming when you want incremental results (reduced perceived latency) or to begin processing output before generation finishes.
- Keep prompt design concise and focused to save compute and money.
- “Attention Is All You Need” (Transformer paper): https://arxiv.org/abs/1706.03762
- Transformer overview: https://en.wikipedia.org/wiki/Transformer_(machine_learning_model)
- Tokenization tools and examples: https://huggingface.co/docs/tokenizers