> ## Documentation Index
> Fetch the complete documentation index at: https://notes.kodekloud.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Why Load Balancing Breaks

> Explains why traditional load balancers fail for LLM serving due to KV-cache state and variable request costs, urging cache- and load-aware routing and scheduling

When you serve a real large language model (LLM) across a fleet of GPU-backed machines, the routing layer becomes critical. The instinctive fix—putting a traditional load balancer in front of the fleet and round-robin-ing requests—works for stateless web traffic but breaks LLM serving for two fundamental reasons: stateful caching (the model’s KV cache) and highly variable request cost.

A typical deployment runs many identical servers, each holding a copy of the model and its own GPU. The router’s job is to map each incoming request to one of those servers. For conventional web apps this is fine; for LLMs, it discards important contextual state and fails to account for per-request compute variability.

## 1) Stateful requests and the KV cache

When a conversation is processed, the model computes per-token key/value pairs and stores them in the KV (key-value) cache. This cached state is the saved work that makes generation of subsequent tokens fast. If message A is routed to server-2 and fills server-2’s KV cache, then message B in the same conversation must land on server-2 to reuse that cached work.

If the next message is instead routed to server-5, server-5 has no KV cache for that conversation and must re-run a full prefill over the entire context to rebuild the keys and values. That recomputation wastes GPU cycles, memory, and latency — effectively throwing away perfectly good cached work.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/1VOAquXcXcLfyOTx/images/AI-Infrastructure-LLM-D-vLLM-and-GPUs/Why-Serving-Gets-Hard/Why-Load-Balancing-Breaks/load-balancer-saved-no-cache-scatter.jpg?fit=max&auto=format&n=1VOAquXcXcLfyOTx&q=85&s=4c1dd1e5636a2f24bae72ebf1cb2b862" alt="A hand-drawn diagram showing a load balancer routing a message to multiple servers, with one server marked &#x22;saved&#x22; and another labeled &#x22;no cache.&#x22; The image illustrates how saved work gets scattered across servers (KodeKloud branding and a speaker inset appear)." width="1920" height="1080" data-path="images/AI-Infrastructure-LLM-D-vLLM-and-GPUs/Why-Serving-Gets-Hard/Why-Load-Balancing-Breaks/load-balancer-saved-no-cache-scatter.jpg" />
</Frame>

## 2) Requests vary enormously in compute cost

LLM requests are not uniformly expensive: a short prompt like “hello” is cheap, while summarizing 50 pages or generating long streams is extremely heavy. Classic load balancers (L4/L7) typically make decisions based on network metadata and headers; they don’t see per-server KV cache contents, queue lengths, or real-time GPU utilization. They also can’t reliably predict a request’s internal compute cost.

As a result, a balancer may send a very heavy request to a server that is already busy streaming many responses, creating a hotspot. This causes long-tail latencies and head-of-line blocking that a naive balancer cannot avoid.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/1VOAquXcXcLfyOTx/images/AI-Infrastructure-LLM-D-vLLM-and-GPUs/Why-Serving-Gets-Hard/Why-Load-Balancing-Breaks/request-size-imbalance-load-balancer-diagram.jpg?fit=max&auto=format&n=1VOAquXcXcLfyOTx&q=85&s=86d64fe11f660558ddd6c95e5d6a82e4" alt="A sketched diagram titled &#x22;Not every request is equal&#x22; illustrating a load balancer routing a tiny request and a giant request to multiple servers, with one server shown as busy. A small circular video-feed of a person appears in the bottom-right and a KodeKloud logo is at the bottom-left." width="1920" height="1080" data-path="images/AI-Infrastructure-LLM-D-vLLM-and-GPUs/Why-Serving-Gets-Hard/Why-Load-Balancing-Breaks/request-size-imbalance-load-balancer-diagram.jpg" />
</Frame>

## Quick comparison

| Concern | Traditional load balancer | Cache- and load-aware router |
| - | -: | - |
| Knows per-conversation KV cache location | No | Yes |
| Observes server GPU utilization and queues | No | Yes |
| Predicts per-request computational cost | No | Yes (via heuristics or profiling) |
| Avoids wasted prefill recomputation | No | Yes (sticky/conversation-aware routing) |
| Mitigates head-of-line blocking from heavy requests | No | Yes (scheduling, admission control) |

Because the balancer cannot see the KV cache, queues, or per-request compute cost, it cannot make the routing decisions needed to preserve cached state or avoid overloaded GPUs. Serving LLMs effectively therefore requires a smarter front layer — one that is both cache-aware and load-aware — rather than treating every server as interchangeable.

<Callout icon="lightbulb" color="#1CB2FE">
  Use sticky or cache-aware routing (for example, hashing requests by conversation ID) and a scheduler that factors in GPU utilization and estimated request cost. This reduces wasted prefill work, prevents hotspots, and improves end-to-end latency for LLM serving.
</Callout>

<CardGroup>
  <Card title="Watch Video" icon="video" cta="Learn more" href="https://learn.kodekloud.com/user/courses/ai-infrastructure-llm-d-vllm-and-gpus/module/d17e9822-30ea-4580-96b1-303eafcaae97/lesson/4b92a8e7-1e9a-44e7-af80-4fc6121192f5" />
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.