Skip to main content
When an autoregressive large language model (LLM) generates text it performs two distinct steps for each new token:
  • A heavy prefill (or forward) pass that processes the entire input context up to the current token.
  • A lighter decode step that computes the next token using the outputs from that prefill.
If the model reran the expensive prefill for every token, responses would pause before every word. Instead, chat apps stream tokens quickly: one initial pause while the model does the heavy work, then continuous streaming output for subsequent tokens.
A hand-drawn slide comparing two chat approaches: "The naive way" shows a pause before every word when serving an LLM that needs a GPU, while "But chat is fast" shows one initial pause then streaming text. A small presenter inset and a KodeKloud logo appear at the bottom.
How is streaming achieved? The prefill performs many matrix multiplies across layers, which is computationally expensive. Rather than discarding those intermediates after producing the first token, the model saves them in GPU memory. These saved tensors are called the KV cache (key/value cache). For each attention layer the cache stores the keys and values computed from the context so they can be reused on subsequent decode steps. So: prefill runs once for the provided context and its outputs are stored in the KV cache. As the model generates tokens, it reuses the KV cache and only computes the new token on top of that cached state — avoiding the full prefill on each step.
A hand-drawn flowchart titled "Saving its work" that shows a prompt being summarized, sent through a "Prefill" into a "KV cache," and then reused during token decoding with notes like "runs once" and "each token reuses the cache." There’s also a small circular inset showing a speaker and a KodeKloud logo in the corner.
In short: you pay the expensive pause once per request (the initial prefill), not once per token, because the KV cache preserves the intermediate computations. Multi-turn chats and why latency grows That explanation covers a single message: one question in, one answer out. By default, when a reply finishes the request is complete and the server discards the KV cache. The server generally does not retain memory across separate API calls. When you send a second message in the same conversation, the model no longer has the previous turn’s KV cache. Chat applications solve this by resending the full conversation history (user messages and assistant replies) in the next request. Because the server lacks cached keys/values from earlier turns, it must re-run the prefill across the entire thread again before responding. As a conversation grows, so does the repeated work — increasing latency and compute cost per turn. What if the server preserved that cached work? Prefix caching If the server keeps cached KV blocks and indexes them by the exact text that produced them, it can reuse cached work for future turns that share the same text prefix. When a later turn arrives, the model can detect that everything up to the new message already has cached keys/values and only needs to prefill the new material. With prefix caching enabled, later turns cost roughly the same as earlier turns because the cached prefix avoids redoing the heavy computation.
A hand-drawn infographic titled "Prefix caching" showing chat-style message bubbles on the left and a diagram of per-turn prefill, cache, and kept items on the right. There's also a small circular photo of a man speaking in the bottom-right corner.
Sharing cached prefixes across users Cached work depends only on the text it was computed from — not on which user or which session produced it. If many conversations begin with the same long system instruction or identical prompt, the server can compute and store that prefix once and reuse it across many chats. This is effectively prefix caching shared across users: one prompt, many chats.
A diagram titled "One prompt, many chats" showing a single "Same system prompt" saved once and routed to three boxes labeled Chat A, Chat B, and Chat C. A small circular photo of a man appears in the bottom-right corner.
Quick summary table How to think about using prefix caching
  1. Avoid sending the same long system prompts repeatedly when possible — share or cache them.
  2. If your provider supports prefix caching, enable it for predictable prefixes (e.g., fixed system instructions).
  3. For dynamic user content you can still benefit by caching common beginnings or boilerplate text that many users share.
KV cache (keys and values) stores layer-wise intermediate tensors so the model can avoid repeating expensive attention computations. Prefix caching leverages those cached tensors for identical text prefixes, reducing latency and compute cost per token for long or repeated prompts.
Links and references

Watch Video