KubeFM
Avsnitt

1 Million Tokens Per Second on Kubernetes, with Federico Iezzi

Dela

GPU inference throughput depends on more than accelerator generation or count.

Memory bandwidth, model parallelism, cache configuration, and the load generator itself all influence measured throughput.

Federico Iezzi, Customer Engineer at Google Cloud, explains how his team achieved 1 million output tokens per second using Qwen 3.5 27B, vLLM, GKE Autopilot, and NVIDIA B200 GPUs.

The discussion covers:

  1. Why memory bandwidth limits decode performance

  2. How Federico chose between tensor and data parallelism

  3. What changed after enabling multi-token prediction and reducing the KV cache footprint with FP8 quantization.

Sponsor

This episode is sponsored by LearnKube. Download the free book, The Technical Guide to Kubernetes Rightsizing, to understand what Prometheus and Grafana cannot tell you about safely reducing requests and limits.

More info

Podden och tillhörande omslagsbild på den här sidan tillhör KubeFM. Innehållet i podden är skapat av KubeFM och inte av, eller tillsammans med, Poddtoppen.