GPU inference throughput depends on more than accelerator generation or count.
Memory bandwidth, model parallelism, cache configuration, and the load generator itself all influence measured throughput.
Federico Iezzi, Customer Engineer at Google Cloud, explains how his team achieved 1 million output tokens per second using Qwen 3.5 27B, vLLM, GKE Autopilot, and NVIDIA B200 GPUs.
The discussion covers:
Why memory bandwidth limits decode performance
How Federico chose between tensor and data parallelism
What changed after enabling multi-token prediction and reducing the KV cache footprint with FP8 quantization.
Podden och tillhörande omslagsbild på den här sidan tillhör
KubeFM. Innehållet i podden är skapat av KubeFM och inte av,
eller tillsammans med, Poddtoppen.