Skip to main content
When you serve a real large language model (LLM) across a fleet of GPU-backed machines, the routing layer becomes critical. The instinctive fix—putting a traditional load balancer in front of the fleet and round-robin-ing requests—works for stateless web traffic but breaks LLM serving for two fundamental reasons: stateful caching (the model’s KV cache) and highly variable request cost. A typical deployment runs many identical servers, each holding a copy of the model and its own GPU. The router’s job is to map each incoming request to one of those servers. For conventional web apps this is fine; for LLMs, it discards important contextual state and fails to account for per-request compute variability.

1) Stateful requests and the KV cache

When a conversation is processed, the model computes per-token key/value pairs and stores them in the KV (key-value) cache. This cached state is the saved work that makes generation of subsequent tokens fast. If message A is routed to server-2 and fills server-2’s KV cache, then message B in the same conversation must land on server-2 to reuse that cached work. If the next message is instead routed to server-5, server-5 has no KV cache for that conversation and must re-run a full prefill over the entire context to rebuild the keys and values. That recomputation wastes GPU cycles, memory, and latency — effectively throwing away perfectly good cached work.
A hand-drawn diagram showing a load balancer routing a message to multiple servers, with one server marked "saved" and another labeled "no cache." The image illustrates how saved work gets scattered across servers (KodeKloud branding and a speaker inset appear).

2) Requests vary enormously in compute cost

LLM requests are not uniformly expensive: a short prompt like “hello” is cheap, while summarizing 50 pages or generating long streams is extremely heavy. Classic load balancers (L4/L7) typically make decisions based on network metadata and headers; they don’t see per-server KV cache contents, queue lengths, or real-time GPU utilization. They also can’t reliably predict a request’s internal compute cost. As a result, a balancer may send a very heavy request to a server that is already busy streaming many responses, creating a hotspot. This causes long-tail latencies and head-of-line blocking that a naive balancer cannot avoid.
A sketched diagram titled "Not every request is equal" illustrating a load balancer routing a tiny request and a giant request to multiple servers, with one server shown as busy. A small circular video-feed of a person appears in the bottom-right and a KodeKloud logo is at the bottom-left.

Quick comparison

Because the balancer cannot see the KV cache, queues, or per-request compute cost, it cannot make the routing decisions needed to preserve cached state or avoid overloaded GPUs. Serving LLMs effectively therefore requires a smarter front layer — one that is both cache-aware and load-aware — rather than treating every server as interchangeable.
Use sticky or cache-aware routing (for example, hashing requests by conversation ID) and a scheduler that factors in GPU utilization and estimated request cost. This reduces wasted prefill work, prevents hotspots, and improves end-to-end latency for LLM serving.

Watch Video