Skip to main content
Before diving in, you need to understand one central concept: tokens. A language model does not read or produce whole words in the same way humans do. Instead it breaks text into pieces called tokens. A token can be a whole word, a subword (prefix/suffix), or sometimes a single character. On average a token is roughly three-quarters of a word. Everything the model processes—what it can read, what you are billed for, how long it takes, and how output is generated—is measured in tokens. For example, the sentence “Serving LLMs is not like serving web apps.” contains eight words, but tokenization may yield about nine tokens. Tokens matter when you measure prompt size, response length, cost, or latency.
A presentation slide titled "It Works in Tokens" with colored token-shaped boxes spelling "Serving LLMs is not like serving web apps." Below the boxes it says "8 words → 9 tokens," with a KodeKloud logo and a small circular webcam image of a person in the lower right.
At its core, a large language model is a next-token predictor. Given all the text so far (the “context”), it computes which token is most likely to come next, appends that token to the context, and repeats this process. A 100-token response is not a single monolithic computation — it is the model performing that prediction step 100 times, once per token. This is why chat UIs appear to “type” token by token: each displayed token arrives when one prediction step completes. Streaming partial output to the client naturally follows from this token-by-token generation. Two practical consequences follow from this behavior:
  • Latency depends on token count: generating hundreds of tokens requires hundreds of prediction steps, so long outputs take longer than short ones.
  • Compute and cost scale with token work: a short Q&A might be trivial, while summarizing a 100-page contract can be thousands of times heavier because it involves far more tokens.
An illustrated infographic titled "No Two Requests Are the Same" showing a simple query ("capital of France?") and a large request ("summarize this contract") sent via the same API to a server. The diagram highlights different token sizes and workloads (e.g., 100 pages, tens of thousands of tokens, "1,000x more work") and includes a small circular photo of a person with a microphone.
A closer look at the token-by-token loop reveals two conceptual halves that make up generation:
  • Prefill (prompt processing)
  • Streaming (autoregressive decoding)
Both are part of the same pipeline but have different behaviors and computational characteristics. Prefill (prompt processing)
  • Tokenization: the prompt (user text, system messages, instructions) is converted into the model’s token IDs.
  • Embedding and forward pass: those tokens are embedded and passed through the Transformer’s layers. During this pass the model computes internal representations and, for each layer, produces key (K) and value (V) vectors used for attention.
  • KV cache population: the keys and values for each prompt token are stored in a KV cache. This cache is essential for efficient subsequent decoding.
Prefill usually processes the entire prompt at once. Because self-attention conceptually compares every token to every other token in the context, its computation scales roughly with the square of the prompt length (O(L^2)) in a naive analysis — hence very long prompts cost disproportionately more. Streaming (autoregressive decoding)
  • Autoregressive loop: decoding generates one token at a time. For each generated token:
    1. The model uses the KV cache (built from the prompt and previously generated tokens) as context to compute attention for the new token.
    2. It computes logits for the next-token distribution, selects (or samples) a token, and emits it.
    3. The generated token is appended and its K/V vectors are added to the KV cache.
  • Efficiency through caching: because the KV cache stores prior keys and values, the model doesn’t re-run the full forward pass over the entire prompt for each new token. Each decoding step only needs to compute attention for the new token against cached keys/values, plus the new token’s own projections. That substantially reduces per-token work compared to naively recomputing the entire context each step.
Comparison: prefill vs streaming A rough mental model: prefill pays the cost to fully ingest the prompt and prepare KV caches; decoding then pays a small per-token cost leveraging those caches to keep generation tractable. This is also why streaming token-by-token output is both practical and expected in chat interfaces.
Tokens affect cost, latency, and maximum context. APIs typically count both input (prompt) tokens and output (response) tokens toward billing and context limits. Keep prompts concise for fast, low-cost responses; expand the prompt when more context is truly needed.
Practical takeaways
  • Measure both input and output tokens when estimating cost and latency.
  • For long context needs, prefer architectures and providers that support larger context windows and efficient KV caching.
  • Use streaming when you want incremental results (reduced perceived latency) or to begin processing output before generation finishes.
  • Keep prompt design concise and focused to save compute and money.
Further reading and references Summary An LLM writes text by predicting one token at a time. The prompt is processed in a prefill phase that builds a KV cache, and the response is generated token-by-token using that cache during decoding. This split explains why long prompts and long outputs have different performance and cost characteristics and why streaming incremental output is efficient and common.

Watch Video