- Layer-wise splitting (pipeline parallelism): assign contiguous whole layers to each GPU. GPU 1 runs the first group of layers, GPU 2 runs the next group, and so on. Each GPU processes full layers and sends relatively small activation tensors to the next device.
- Within-layer splitting (tensor or model parallelism): split the internal computation of a single layer across several GPUs — for example, partition matrices or attention heads. Each GPU computes a piece of the layer’s output; the pieces are combined to produce the full result.

- Inside one machine, GPUs are often connected by NVLink (or similar high-bandwidth interconnects), which supports frequent, high-throughput exchanges. In this setting, chatty within-layer (tensor) parallelism performs well: multiple GPUs act like one tightly-coupled device, exchanging partial results rapidly. It’s common to pack 4–8 NVLink-connected GPUs into a single server and shard a giant model across them.
- Across machines, GPUs communicate over the network (Ethernet/InfiniBand) which is typically slower and higher-latency than NVLink. For multi-node setups, layer-wise (pipeline) parallelism is usually better: each server owns whole layers and only needs to send relatively small activation tensors to the next server, minimizing cross-machine traffic.
Rule of thumb: keep the high-bandwidth, chatty communication inside a machine (NVLink) and use lighter-weight communication between machines (network).
References and further reading:
- vLLM model server: https://github.com/vllm-project/vllm
- NVIDIA NVLink overview: https://developer.nvidia.com/nvlink
- Pipeline and tensor parallelism concepts: see transformer model parallelism literature
