> ## Documentation Index
> Fetch the complete documentation index at: https://notes.kodekloud.com/llms.txt
> Use this file to discover all available pages before exploring further.

# The Pause and the Stream

> Describes LLM latency phases, prefill initial pause from prompt processing, and decode streaming limited by VRAM bandwidth, plus optimization strategies.

What are the two halves of an LLM? When you paste text into ChatGPT and press Enter you typically notice two distinct latency behaviors: an initial pause, then streaming tokens. Those correspond to two different phases of model execution — *prefill* and *decode* — and each phase has different resource bottlenecks and performance characteristics. Understanding both helps explain why long prompts slow the first response and why per-token throughput is limited during streaming.

## Prefill (the initial pause)

Before the model emits a single token, it must read and process the entire prompt. The model consumes the prompt tokens in parallel: the tokens flow through the full network in one large, compute-heavy operation. If the model weights are not already resident on the device, they must be loaded into GPU VRAM; once available, many compute cores run a burst of matrix math to transform the prompt into the initial model state and produce the first output token.

That waiting time is called TTFT — time to first token. TTFT grows with the size (token count) of the prompt: a short question produces an imperceptible pause, while a very long document can create a noticeable delay before any output appears.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/1VOAquXcXcLfyOTx/images/AI-Infrastructure-LLM-D-vLLM-and-GPUs/What-Happens-When-You-Interact-with-LLM/The-Pause-and-the-Stream/handdrawn-model-prefill-inference-vram-burst.jpg?fit=max&auto=format&n=1VOAquXcXcLfyOTx&q=85&s=868eaa55117ae870810dfe948ebf6245" alt="A hand-drawn diagram explaining model prefill and inference, showing tokens being read at once, VRAM/the whole model loaded, and a grid labeled &#x22;one big burst of math&#x22; representing many cores firing. There's also a small circular video inset of a person speaking in the corner." width="1920" height="1080" data-path="images/AI-Infrastructure-LLM-D-vLLM-and-GPUs/What-Happens-When-You-Interact-with-LLM/The-Pause-and-the-Stream/handdrawn-model-prefill-inference-vram-burst.jpg" />
</Frame>

## Decode (the streaming output)

After the model produces the first token it enters the token-by-token loop that generates the remainder of the response. Each new token requires a full forward pass conditioned on the current context (the previous tokens). For large models that do not fit entirely in the GPU’s fastest caches, this means model weights are streamed repeatedly from VRAM into on‑chip memory and compute units — roughly one streaming pass per output token.

Because decode repeatedly moves weights, the bottleneck shifts from raw compute to memory bandwidth. Producing a 200-token reply can require streaming most of the model from VRAM roughly 200 times. During decode, compute cores perform useful work on each pass, but the rate at which weights can be supplied from VRAM limits token throughput.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/1VOAquXcXcLfyOTx/images/AI-Infrastructure-LLM-D-vLLM-and-GPUs/What-Happens-When-You-Interact-with-LLM/The-Pause-and-the-Stream/vram-streaming-repeated-loads-200x.jpg?fit=max&auto=format&n=1VOAquXcXcLfyOTx&q=85&s=de42be053e6f454ee2b1d511ccdf338b" alt="A hand-drawn diagram titled &#x22;The stream&#x22; showing a VRAM box labeled &#x22;the whole model&#x22; with arrows loading parts of it into a larger compute area labeled &#x22;barely any work,&#x22; illustrating repeated loads per token. A badge reading &#x22;×200&#x22; and the text &#x22;Serving an LLM needs a GPU&#x22; appear, with a small circular webcam thumbnail in the corner." width="1920" height="1080" data-path="images/AI-Infrastructure-LLM-D-vLLM-and-GPUs/What-Happens-When-You-Interact-with-LLM/The-Pause-and-the-Stream/vram-streaming-repeated-loads-200x.jpg" />
</Frame>

## The speed limit and TPOT

Why does memory bandwidth dominate during decode? Model weights are far too large to live in tiny register files — they remain in VRAM and must be streamed into the compute units. High-end GPUs can read their VRAM at multi‑terabyte-per-second rates. For example, with \~3 TB/s effective VRAM read throughput, a 16 GB model could be read entirely about:

`3 TB/s ÷ 16 GB ≈ 187 reads/s` (≈ 200 reads/s)

That rough calculation corresponds to approximately 200 tokens per second — one token every \~5 ms — in an optimistic steady-state. Real-world per-token latency (TPOT, time per output token) is higher because of kernel launch overheads, attention recomputation, token sampling, and other runtime costs.

Benchmarks typically report TPOT because it reflects the steady-state gap between successive tokens during decoding. Together, TTFT and TPOT describe the two main latency components you’ll see when interacting with an LLM.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/1VOAquXcXcLfyOTx/images/AI-Infrastructure-LLM-D-vLLM-and-GPUs/What-Happens-When-You-Interact-with-LLM/The-Pause-and-the-Stream/speed-limit-vram-flow-3tbps-200tps.jpg?fit=max&auto=format&n=1VOAquXcXcLfyOTx&q=85&s=dd709e9c959c7fed7b8920b742725b4d" alt="A schematic titled &#x22;The speed limit&#x22; showing VRAM, model data flow, and a calculation that lists 3 TB/s memory read speed ≈ 200 tokens/sec with TPOT ≈ 5 ms. A small circular inset photo of a person appears in the lower-right." width="1920" height="1080" data-path="images/AI-Infrastructure-LLM-D-vLLM-and-GPUs/What-Happens-When-You-Interact-with-LLM/The-Pause-and-the-Stream/speed-limit-vram-flow-3tbps-200tps.jpg" />
</Frame>

## Quick comparison: Prefill vs Decode

| Phase | What it does | Main bottleneck | Latency metric |
| - | - | - | - |
| Prefill | Read and process the entire prompt in one parallel computation | Compute (large matrix multiplications) and model load time | TTFT (time to first token) |
| Decode | Generate tokens one-by-one in a sequential loop | Memory bandwidth (VRAM → on-chip) and per-token overhead | TPOT (time per output token) |

## Practical implications

* Long prompts increase TTFT. For latency-sensitive apps, minimize prompt token count or precompute embeddings/contexts where possible.
* TPOT limits streaming throughput. High VRAM bandwidth, model quantization, attention compression, or caching frequently accessed weight blocks can improve TPOT.
* Serving architectures often try to optimize both: place model weights in GPU memory to reduce prefill overhead and use optimizations (e.g., kernel fusion, quantized kernels, model sharding) to increase effective read throughput.
* Tools and libraries that target high-throughput inference (for example, vLLM or DeepSpeed inference) focus on reducing both prefill and decode costs through smarter memory management and kernel optimizations.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/1VOAquXcXcLfyOTx/images/AI-Infrastructure-LLM-D-vLLM-and-GPUs/What-Happens-When-You-Interact-with-LLM/The-Pause-and-the-Stream/two-jobs-prefill-decode-compute-memory.jpg?fit=max&auto=format&n=1VOAquXcXcLfyOTx&q=85&s=7241228b74334e87ff67b12caf2acb3f" alt="A slide titled &#x22;Two different jobs&#x22; shows two rounded boxes labeled &#x22;Prefill&#x22; (read the prompt; one big burst of math; compute heavy) and &#x22;Decode&#x22; (write the answer; re‑read the model over and over; memory bandwidth heavy). A small circular photo of a person appears in the bottom-right and a KodeKloud logo is in the bottom-left." width="1920" height="1080" data-path="images/AI-Infrastructure-LLM-D-vLLM-and-GPUs/What-Happens-When-You-Interact-with-LLM/The-Pause-and-the-Stream/two-jobs-prefill-decode-compute-memory.jpg" />
</Frame>

<Callout icon="lightbulb" color="#1CB2FE">
  TTFT (time to first token) measures the initial pause while the prompt is processed (prefill). TPOT (time per output token) measures the steady per-token latency during decoding, typically limited by VRAM read bandwidth.
</Callout>

## Links and references

* vLLM (high-performance LLM serving): [https://github.com/vllm-project/vllm](https://github.com/vllm-project/vllm)
* DeepSpeed Inference: [https://github.com/microsoft/DeepSpeed](https://github.com/microsoft/DeepSpeed)
* NVIDIA: GPU memory and bandwidth concepts — [https://developer.nvidia.com/what-is-memory-bandwidth](https://developer.nvidia.com/what-is-memory-bandwidth)
* Tokenization and prompts: [https://huggingface.co/docs/tokenizers/index](https://huggingface.co/docs/tokenizers/index)

<CardGroup>
  <Card title="Watch Video" icon="video" cta="Learn more" href="https://learn.kodekloud.com/user/courses/ai-infrastructure-llm-d-vllm-and-gpus/module/cbb32ed8-e080-4f06-a745-33e8099f5157/lesson/259bc9f6-d1e3-4879-8efb-66b9816c14e5" />
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.