Skip to main content
This document explains how llm-d runs inside a Kubernetes cluster: how the Helm releases map to Kubernetes objects, how model servers (vLLM) are deployed and scaled, and how routing and sharded models are coordinated. Two Helm releases set up a model deployment:
  • A Helm release that installs the llm-d router plus the small control-plane components (inference gateway + scheduler).
  • A Helm release that installs the model service (the vLLM-based servers).
Example Helm commands:
Start from a Kubernetes cluster with available worker nodes. Initially there are no model pods running. When you deploy a model, each model server is a vLLM instance (see https://github.com/vllm-project/vllm). In llm-d, each vLLM process runs inside a pod scheduled onto nodes with available GPU/CPU resources. A group of identical pods serving the same role is represented by a Deployment. llm-d uses two conceptual pools for each model deployment:
  • Prefill pool: pods that read prompts and build/restore model execution state (prepare cache).
  • Decode pool: pods that produce tokens and stream outputs.
Prefill pods and decode pods are scaled independently so you can tune read-heavy vs. decode-heavy workloads. So what about very large models that are too big for one machine? You might consider StatefulSets because they provide stable pod identities and an ordinal startup order. That pattern is useful for databases and other services with independent replicas.
A simple slide showing four empty rounded boxes labeled shard-0 through shard-3 under the heading "StatefulSet?" with the caption "but an LLM is one server, in pieces." A small circular video thumbnail of a person speaking appears in the lower-right corner and a KodeKloud logo is in the bottom-left.
StatefulSets model each pod as an independent replica. For a sharded LLM, shards are not independent — they combine to form one logical server. Losing a single shard makes the whole model unusable, so scale and healing must treat the shard group as a unit.
llm-d addresses this by grouping a leader pod with worker pods into a single logical unit (commonly referred to as a LeaderWorkerSet). That grouping is scaled and healed together, so a sharded model behaves like a single server composed of coordinated parts.
A simple hand-drawn diagram titled "LeaderWorkerSet" showing a rounded rectangle with a 2x2 grid of boxes, the top-left labeled "lead." There's a small circular video inset of a person speaking in the bottom-right and a KodeKloud logo in the lower-left.
All components — prefill pods, decode pods, gateway, scheduler, and any sharded groups — are expressed to Kubernetes via manifests generated by Helm. The values.yaml you pass to Helm declares the model artifacts and the replica counts for each pool. Updating those values and reapplying the Helm release changes the deployed fleet topology. Example values.yaml fragment (model artifacts and pool sizes):
After you install the Helm releases, llm-d will create the prefill and decode pods, deploy the inference gateway and scheduler, and wire the pieces together. Use kubectl to inspect the running pods:
Because this runs on Kubernetes, you get standard platform guarantees: crashed pods are restarted, pods on failed nodes are rescheduled, and you can scale prefill and decode pools independently. llm-d handles model-aware routing, caching, and request orchestration while Kubernetes handles liveness, scheduling, and rescheduling.
An illustrated diagram titled "Kubernetes Keeps Them Alive" showing a cluster with a gateway, scheduler, and multiple prefill and decode pods handling incoming requests. There's also a small circular presenter video overlay in the bottom-right corner.
llm-d exposes a single public endpoint via the inference gateway. Clients send all requests to that gateway — you never send requests directly to individual pods. The scheduler routes requests to the appropriate prefill and decode pods based on cache locality and current load.
Get the gateway address with kubectl:
When a client request hits the gateway, the scheduler selects a pod (or shard group) that already holds the needed cached state when possible. The prefill pod reads the prompt and prepares model state; the decode pod generates tokens and streams them, token-by-token, back through the gateway to the client.
A hand-drawn diagram titled "Behind One Line" showing a request flow for serving an LLM from a browser input through points labeled the gateway, scheduler picked the pod, prefill read it, decode wrote it token-by-token, ending at a GPU. A small circular video thumbnail of a man speaking and a KodeKloud logo appear in the corners.
For massive sharded models, token generation is distributed across GPUs on multiple machines. The LeaderWorkerSet coordinates those shards so they behave as a single logical model server while preserving efficient cross-machine computation. Component summary Links and References

Watch Video