Describes how llm-d runs vLLM model servers on Kubernetes via Helm, using prefill and decode pools, an inference gateway and scheduler, and coordinated sharded LeaderWorkerSets
This document explains how llm-d runs inside a Kubernetes cluster: how the Helm releases map to Kubernetes objects, how model servers (vLLM) are deployed and scaled, and how routing and sharded models are coordinated.Two Helm releases set up a model deployment:
A Helm release that installs the llm-d router plus the small control-plane components (inference gateway + scheduler).
A Helm release that installs the model service (the vLLM-based servers).
Start from a Kubernetes cluster with available worker nodes. Initially there are no model pods running. When you deploy a model, each model server is a vLLM instance (see https://github.com/vllm-project/vllm). In llm-d, each vLLM process runs inside a pod scheduled onto nodes with available GPU/CPU resources.A group of identical pods serving the same role is represented by a Deployment. llm-d uses two conceptual pools for each model deployment:
Prefill pool: pods that read prompts and build/restore model execution state (prepare cache).
Decode pool: pods that produce tokens and stream outputs.
Prefill pods and decode pods are scaled independently so you can tune read-heavy vs. decode-heavy workloads.So what about very large models that are too big for one machine? You might consider StatefulSets because they provide stable pod identities and an ordinal startup order. That pattern is useful for databases and other services with independent replicas.
StatefulSets model each pod as an independent replica. For a sharded LLM, shards are not independent — they combine to form one logical server. Losing a single shard makes the whole model unusable, so scale and healing must treat the shard group as a unit.
llm-d addresses this by grouping a leader pod with worker pods into a single logical unit (commonly referred to as a LeaderWorkerSet). That grouping is scaled and healed together, so a sharded model behaves like a single server composed of coordinated parts.
All components — prefill pods, decode pods, gateway, scheduler, and any sharded groups — are expressed to Kubernetes via manifests generated by Helm. The values.yaml you pass to Helm declares the model artifacts and the replica counts for each pool. Updating those values and reapplying the Helm release changes the deployed fleet topology.Example values.yaml fragment (model artifacts and pool sizes):
# the fleetmodelArtifacts: uri: hf://meta-llama/llama-3-70Bprefill: replicas: 4 # pods that read promptsdecode: replicas: 1 # pods that write tokens
After you install the Helm releases, llm-d will create the prefill and decode pods, deploy the inference gateway and scheduler, and wire the pieces together. Use kubectl to inspect the running pods:
Because this runs on Kubernetes, you get standard platform guarantees: crashed pods are restarted, pods on failed nodes are rescheduled, and you can scale prefill and decode pools independently. llm-d handles model-aware routing, caching, and request orchestration while Kubernetes handles liveness, scheduling, and rescheduling.
llm-d exposes a single public endpoint via the inference gateway. Clients send all requests to that gateway — you never send requests directly to individual pods. The scheduler routes requests to the appropriate prefill and decode pods based on cache locality and current load.
Get the gateway address with kubectl:
$ kubectl get gatewayNAME ADDRESSinference-gateway 203.0.113.10
When a client request hits the gateway, the scheduler selects a pod (or shard group) that already holds the needed cached state when possible. The prefill pod reads the prompt and prepares model state; the decode pod generates tokens and streams them, token-by-token, back through the gateway to the client.
For massive sharded models, token generation is distributed across GPUs on multiple machines. The LeaderWorkerSet coordinates those shards so they behave as a single logical model server while preserving efficient cross-machine computation.Component summary