> ## Documentation Index
> Fetch the complete documentation index at: https://notes.kodekloud.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Inside the Cluster

> Describes how llm-d runs vLLM model servers on Kubernetes via Helm, using prefill and decode pools, an inference gateway and scheduler, and coordinated sharded LeaderWorkerSets

This document explains how llm-d runs inside a Kubernetes cluster: how the Helm releases map to Kubernetes objects, how model servers (vLLM) are deployed and scaled, and how routing and sharded models are coordinated.

Two Helm releases set up a model deployment:

* A Helm release that installs the llm-d router plus the small control-plane components (inference gateway + scheduler).
* A Helm release that installs the model service (the vLLM-based servers).

Example Helm commands:

```bash theme={null}
$ helm install router llm-d/router \
  -f router.values.yaml

$ helm install llama llm-d/modelservice \
  -f llama-70b.values.yaml
```

Start from a Kubernetes cluster with available worker nodes. Initially there are no model pods running. When you deploy a model, each model server is a vLLM instance (see [https://github.com/vllm-project/vllm](https://github.com/vllm-project/vllm)). In llm-d, each vLLM process runs inside a pod scheduled onto nodes with available GPU/CPU resources.

A group of identical pods serving the same role is represented by a Deployment. llm-d uses two conceptual pools for each model deployment:

* Prefill pool: pods that read prompts and build/restore model execution state (prepare cache).
* Decode pool: pods that produce tokens and stream outputs.

Prefill pods and decode pods are scaled independently so you can tune read-heavy vs. decode-heavy workloads.

So what about very large models that are too big for one machine? You might consider StatefulSets because they provide stable pod identities and an ordinal startup order. That pattern is useful for databases and other services with independent replicas.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/1VOAquXcXcLfyOTx/images/AI-Infrastructure-LLM-D-vLLM-and-GPUs/Meet-llm-d/Inside-the-Cluster/statefulset-llm-shards-slide-kodekloud.jpg?fit=max&auto=format&n=1VOAquXcXcLfyOTx&q=85&s=d373a75073841029e0726dce702c7b6a" alt="A simple slide showing four empty rounded boxes labeled shard-0 through shard-3 under the heading &#x22;StatefulSet?&#x22; with the caption &#x22;but an LLM is one server, in pieces.&#x22; A small circular video thumbnail of a person speaking appears in the lower-right corner and a KodeKloud logo is in the bottom-left." width="1920" height="1080" data-path="images/AI-Infrastructure-LLM-D-vLLM-and-GPUs/Meet-llm-d/Inside-the-Cluster/statefulset-llm-shards-slide-kodekloud.jpg" />
</Frame>

<Callout icon="warning" color="#FF6B6B">
  StatefulSets model each pod as an independent replica. For a sharded LLM, shards are not independent — they combine to form one logical server. Losing a single shard makes the whole model unusable, so scale and healing must treat the shard group as a unit.
</Callout>

llm-d addresses this by grouping a leader pod with worker pods into a single logical unit (commonly referred to as a LeaderWorkerSet). That grouping is scaled and healed together, so a sharded model behaves like a single server composed of coordinated parts.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/1VOAquXcXcLfyOTx/images/AI-Infrastructure-LLM-D-vLLM-and-GPUs/Meet-llm-d/Inside-the-Cluster/leaderworkerset-2x2-lead-video-kodekloud.jpg?fit=max&auto=format&n=1VOAquXcXcLfyOTx&q=85&s=7735994aee95f0b6a95e89040a66dbd6" alt="A simple hand-drawn diagram titled &#x22;LeaderWorkerSet&#x22; showing a rounded rectangle with a 2x2 grid of boxes, the top-left labeled &#x22;lead.&#x22; There's a small circular video inset of a person speaking in the bottom-right and a KodeKloud logo in the lower-left." width="1920" height="1080" data-path="images/AI-Infrastructure-LLM-D-vLLM-and-GPUs/Meet-llm-d/Inside-the-Cluster/leaderworkerset-2x2-lead-video-kodekloud.jpg" />
</Frame>

All components — prefill pods, decode pods, gateway, scheduler, and any sharded groups — are expressed to Kubernetes via manifests generated by Helm. The `values.yaml` you pass to Helm declares the model artifacts and the replica counts for each pool. Updating those values and reapplying the Helm release changes the deployed fleet topology.

Example `values.yaml` fragment (model artifacts and pool sizes):

```yaml theme={null}
# the fleet
modelArtifacts:
  uri: hf://meta-llama/llama-3-70B

prefill:
  replicas: 4  # pods that read prompts

decode:
  replicas: 1  # pods that write tokens
```

After you install the Helm releases, llm-d will create the prefill and decode pods, deploy the inference gateway and scheduler, and wire the pieces together. Use `kubectl` to inspect the running pods:

```bash theme={null}
$ kubectl get pods
NAME                      STATUS
llama-70b-prefill-0       Running
llama-70b-prefill-1       Running
llama-70b-prefill-2       Running
llama-70b-prefill-3       Running
llama-70b-decode-0        Running
inference-gateway-...     Running
```

Because this runs on Kubernetes, you get standard platform guarantees: crashed pods are restarted, pods on failed nodes are rescheduled, and you can scale prefill and decode pools independently. llm-d handles model-aware routing, caching, and request orchestration while Kubernetes handles liveness, scheduling, and rescheduling.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/1VOAquXcXcLfyOTx/images/AI-Infrastructure-LLM-D-vLLM-and-GPUs/Meet-llm-d/Inside-the-Cluster/kubernetes-keeps-them-alive-cluster-diagram.jpg?fit=max&auto=format&n=1VOAquXcXcLfyOTx&q=85&s=8aa13f914eff3125bb1275bd6db61d92" alt="An illustrated diagram titled &#x22;Kubernetes Keeps Them Alive&#x22; showing a cluster with a gateway, scheduler, and multiple prefill and decode pods handling incoming requests. There's also a small circular presenter video overlay in the bottom-right corner." width="1920" height="1080" data-path="images/AI-Infrastructure-LLM-D-vLLM-and-GPUs/Meet-llm-d/Inside-the-Cluster/kubernetes-keeps-them-alive-cluster-diagram.jpg" />
</Frame>

<Callout icon="lightbulb" color="#1CB2FE">
  llm-d exposes a single public endpoint via the inference gateway. Clients send all requests to that gateway — you never send requests directly to individual pods. The scheduler routes requests to the appropriate prefill and decode pods based on cache locality and current load.
</Callout>

Get the gateway address with `kubectl`:

```bash theme={null}
$ kubectl get gateway
NAME                 ADDRESS
inference-gateway    203.0.113.10
```

When a client request hits the gateway, the scheduler selects a pod (or shard group) that already holds the needed cached state when possible. The prefill pod reads the prompt and prepares model state; the decode pod generates tokens and streams them, token-by-token, back through the gateway to the client.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/1VOAquXcXcLfyOTx/images/AI-Infrastructure-LLM-D-vLLM-and-GPUs/Meet-llm-d/Inside-the-Cluster/behind-oneline-llm-request-flow-gpu.jpg?fit=max&auto=format&n=1VOAquXcXcLfyOTx&q=85&s=f894ed7d02f5c3d7eb9083455a9a2a12" alt="A hand-drawn diagram titled &#x22;Behind One Line&#x22; showing a request flow for serving an LLM from a browser input through points labeled the gateway, scheduler picked the pod, prefill read it, decode wrote it token-by-token, ending at a GPU. A small circular video thumbnail of a man speaking and a KodeKloud logo appear in the corners." width="1920" height="1080" data-path="images/AI-Infrastructure-LLM-D-vLLM-and-GPUs/Meet-llm-d/Inside-the-Cluster/behind-oneline-llm-request-flow-gpu.jpg" />
</Frame>

For massive sharded models, token generation is distributed across GPUs on multiple machines. The LeaderWorkerSet coordinates those shards so they behave as a single logical model server while preserving efficient cross-machine computation.

Component summary

| Component | Role | Example / Notes |
| - | - | - |
| Inference gateway | Single public endpoint for clients | `inference-gateway` with external IP |
| Scheduler | Routes requests to pods based on cache/load | Coordinates prefill→decode flow |
| Prefill pods | Read prompts and restore model state | Scaled independently (`prefill.replicas`) |
| Decode pods | Generate tokens and stream outputs | Scaled independently (`decode.replicas`) |
| LeaderWorkerSet | Group of pods representing a sharded model | Scaled and healed as a unit (not a StatefulSet) |
| vLLM | Inference server runtime inside pods | [https://github.com/vllm-project/vllm](https://github.com/vllm-project/vllm) |

Links and References

* llm-d project: [https://llm-d.ai](https://llm-d.ai)
* vLLM (inference server runtime): [https://github.com/vllm-project/vllm](https://github.com/vllm-project/vllm)
* Kubernetes documentation: [https://kubernetes.io/docs/](https://kubernetes.io/docs/)
* Helm documentation: [https://helm.sh/docs/](https://helm.sh/docs/)

<CardGroup>
  <Card title="Watch Video" icon="video" cta="Learn more" href="https://learn.kodekloud.com/user/courses/ai-infrastructure-llm-d-vllm-and-gpus/module/5fe7e764-aa57-4834-8150-905e8fdac59f/lesson/ec5aaf3c-4108-4c76-90e9-16c1994e5564" />
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.