1) Stateful requests and the KV cache
When a conversation is processed, the model computes per-token key/value pairs and stores them in the KV (key-value) cache. This cached state is the saved work that makes generation of subsequent tokens fast. If message A is routed to server-2 and fills server-2’s KV cache, then message B in the same conversation must land on server-2 to reuse that cached work. If the next message is instead routed to server-5, server-5 has no KV cache for that conversation and must re-run a full prefill over the entire context to rebuild the keys and values. That recomputation wastes GPU cycles, memory, and latency — effectively throwing away perfectly good cached work.
2) Requests vary enormously in compute cost
LLM requests are not uniformly expensive: a short prompt like “hello” is cheap, while summarizing 50 pages or generating long streams is extremely heavy. Classic load balancers (L4/L7) typically make decisions based on network metadata and headers; they don’t see per-server KV cache contents, queue lengths, or real-time GPU utilization. They also can’t reliably predict a request’s internal compute cost. As a result, a balancer may send a very heavy request to a server that is already busy streaming many responses, creating a hotspot. This causes long-tail latencies and head-of-line blocking that a naive balancer cannot avoid.
Quick comparison
Because the balancer cannot see the KV cache, queues, or per-request compute cost, it cannot make the routing decisions needed to preserve cached state or avoid overloaded GPUs. Serving LLMs effectively therefore requires a smarter front layer — one that is both cache-aware and load-aware — rather than treating every server as interchangeable.
Use sticky or cache-aware routing (for example, hashing requests by conversation ID) and a scheduler that factors in GPU utilization and estimated request cost. This reduces wasted prefill work, prevents hotspots, and improves end-to-end latency for LLM serving.