
-
Optimized baseline
- Topology: one pool of identical servers.
- Routing: cache-aware routing steers requests toward nodes that already hold useful cached state.
- Sharding/splitting: none — the service is homogeneous.
- When to choose: small teams or early-stage deployments that need a simple, low-friction setup. This path gets you running quickly and provides a stable baseline for later tuning.
-
Prefill / Decode split
- Topology: two specialized pools — a prefill pool to read and tokenize prompts, and a decode pool to generate answers.
- Routing: a scheduler forwards work between prefill and decode pools, with sizing tuned independently.
- When to choose: long prompts, heavy context handling, or workloads where separating memory-heavy prefill from latency-sensitive decoding improves utilization and throughput.
-
Multi-GPU model sharding
- Topology: a single logical model is spread across multiple GPUs (often across machines), and the group acts like one server.
- Routing: remains the same as other paths — the router’s behavior does not change, only the fleet shape does.
- When to choose: models too large for a single GPU or when you need more aggregate model capacity than a single device can provide.
Start with the Optimized Baseline to get a running system quickly. Monitor latency, throughput, and GPU utilization, then evolve to the Prefill/Decode split or Multi-GPU sharding as model size and traffic patterns demand it.
Common tunable knobs (and what they affect)
- Router cache preference — pushes requests toward nodes with cached state (affects latency and cache hit rate).
- Prefill vs. decode pool sizing — balances memory-intensive prompt processing vs. compute-intensive decoding (affects throughput and utilization).
- GPUs per model / sharding strategy — allows models larger than a single GPU to run (affects feasibility and latency).
- Batch size per server — increases throughput at the cost of tail latency when batches wait to fill.
Tuning multiple knobs simultaneously without measurement can hide regressions. Always iterate with metrics: p50/p95 latency, throughput (tokens/sec), GPU utilization, and cost-per-request.
- Deploy the Optimized Baseline well-lit path.
- Collect metrics: latency (p50/p95), throughput, GPU/memory utilization, and cache hit rates.
- Identify bottlenecks: high decode CPU/GPU usage, low cache hit rate, or large-model memory pressure.
- If prompts are long or utilization is uneven, test the Prefill / Decode split.
- If the model does not fit on one GPU, plan Multi-GPU sharding and re-evaluate routing and batching.
- Iterate: change one knob at a time, measure impact, and keep cost/performance targets in view.
- Kubernetes Basics
- llm-d repository / docs (check your project-specific docs for exact configs and well-lit path names)