Skip to main content
In this lesson we’ll explain how a single GPU can serve many concurrent users, why batching improves throughput, and what ultimately limits how many simultaneous sessions a GPU can support. Serving a large language model (LLM) is often dominated by memory access rather than raw math. For every token generated, the GPU repeatedly reads model weights (and associated state) from memory. Those reads — and the memory bandwidth they consume — are the expensive part of inference. Because the model weights are accessed repeatedly for each token, a single model read can be reused to produce tokens for many independent requests at once. Packing multiple users’ token-generation requests into a single GPU invocation is called batching. Batching raises throughput — tokens produced per second — with only a small per-user added queueing delay while the batch is formed.
  • Benefit: Throughput increases roughly in proportion to batch size.
  • Cost: Small queuing latency while waiting for the batch to fill.
  • Per-user compute: Remains roughly the same per token; the savings are amortized across the batch.
Here are a few example independent requests you might batch together:
Tip: Measure both throughput (tokens/sec) and tail latency (95th/99th percentile latency). Increasing batch size improves throughput but can worsen tail latency for interactive users.
However, batching is limited by GPU memory. Every user in a generation session requires their own per-session state on the GPU, commonly called the KV cache (key/value cache) or attention cache. The KV cache stores intermediate activations needed to continue generation across tokens. Each cache must remain resident in GPU memory for the duration of that user’s session. Two primary consumers of GPU memory during inference:
  • The model weights (usually kept resident for low-latency inference).
  • All simultaneously active users’ KV caches.
Once the available GPU memory is exhausted, you cannot increase the number of concurrent users on that GPU without eviction or additional hardware. That’s the memory ceiling: even when the GPU still has spare FLOPS, the memory capacity determines how many sessions can be hosted concurrently.
An illustrated infographic titled "The Ceiling: Memory" showing a GPU memory container split into a large gray block labeled "the model's weights" and many colored squares labeled "everyone's KV caches" to indicate memory filling up. A small round webcam-style photo of a person appears in the lower right.
This memory constraint has practical business impact:
  • Overprovisioning: Teams often buy additional GPUs not because they need more compute, but because they need extra memory capacity to hold more KV caches. Those added GPUs may sit partially idle for compute while still incurring cost.
  • Capacity signals: When a platform reports “at capacity,” it is often the GPU memory that’s exhausted, not the math units.
  • User experience: If memory is exhausted, incoming sessions may be queued, rejected, or served with degraded performance after evicting other sessions.
Warning: Evicting active KV caches to free memory can cause users to lose session context, increase response latency, or require expensive recomputation. Prefer architectural solutions (sharding, offloading, quantization) before eviction when possible.
What to do when a model or its KV caches don’t fit on a single GPU
  • Model too big for one GPU: The largest models may not fit entirely on any single GPU. In that case you must use techniques that trade memory for other resources.
  • Common strategies:
    • Model sharding (model parallelism): split model weights across multiple GPUs.
    • Offloading: move parts of the model or KV cache to CPU RAM or NVMe.
    • Quantization: reduce weight precision (e.g., 8-bit or 4-bit) to shrink model size and KV cache size.
    • Memory-optimized runtimes: specialized serving frameworks that compress or stream KV caches.
Use the table below to compare approaches quickly: Practical recommendations
  • Profile memory usage per-session and KV cache size at your target sequence lengths to estimate max concurrent sessions per GPU.
  • Benchmark throughput and tail latency at different batch sizes for real traffic patterns.
  • Consider hybrid strategies: quantize weights, offload cold parts to CPU, and shard hot layers across GPUs.
  • Use autoscaling policies that react to memory pressure, not just compute utilization.
Further reading and references By treating GPU memory as the primary limiting resource for concurrent LLM serving, you can design architectures and capacity planning strategies that maximize throughput while controlling cost.

Watch Video