AI Post Transformers
Avsnitt

Small Collectives, Big TLB Cost: Reverse Address Translation in Scale-Up GPU Pods

Dela

This episode examines the "last mile" problem in multi-GPU scale-up fabrics: when a remote memory request arrives at a destination GPU carrying a Network Physical Address, that address means nothing locally until it's converted back to a System Physical Address through Reverse Address Translation. The discussion traces why this destination-side translation problem is genuinely new — decades of TLB optimization research has assumed the initiating processor controls the access pattern, while NVLink and UALink fabrics flip that model, forcing the receiving GPU to translate incoming requests with no warning and no control. The hosts connect this seemingly niche hardware detail to a concrete workload: Mixture-of-Experts models rely on All-to-All dispatch and gather collectives, implemented in libraries like NCCL and RCCL, that cross this translation step twice per layer across dozens of layers. To quantify the impact, the paper extends the ASTRA-sim2.0 simulator with an Omnet++ network backend to model packet-level UALink Clos topologies, generating realistic All-to-All traffic via Microsoft's MSCCLang. Listeners interested in the unglamorous plumbing beneath large-scale AI infrastructure will find a compelling case that address translation, not just compute or bandwidth, may be a hidden bottleneck for the collective communication patterns underpinning today's largest inference deployments.

Sources: 1. Amel Fatima's Reverse Address Translation in GPU Fabrics https://arxiv.org/pdf/2604.02473 2. Trans-FW: Short Circuiting Page Table Walk in Multi-GPU Systems via Remote Forwarding — Bingyao Li, Jieming Yin, Anup Holey, Youtao Zhang, Jun Yang, Xulong Tang, 2023 https://scholar.google.com/scholar?q=Trans-FW%3A+Short+Circuiting+Page+Table+Walk+in+Multi-GPU+Systems+via+Remote+Forwarding 3. ASTRA-sim2.0: Modeling Hierarchical Networks and Disaggregated Systems for Large-model Training at Scale — William Won, Taekyung Heo, Saeed Rashidi, Srinivas Sridharan, Sudarshan Srinivasan, Tushar Krishna, 2023 https://scholar.google.com/scholar?q=ASTRA-sim2.0%3A+Modeling+Hierarchical+Networks+and+Disaggregated+Systems+for+Large-model+Training+at+Scale 4. MSCCLang: Microsoft Collective Communication Language — Meghan Cowan, Saeed Maleki, Madanlal Musuvathi, Olli Saarikivi, Yifan Xiong, 2023 https://scholar.google.com/scholar?q=MSCCLang%3A+Microsoft+Collective+Communication+Language 5. Tutel: Adaptive Mixture-of-Experts at Scale — Changho Hwang, Wei Cui, Yifan Xiong, Ziyue Yang, Ze Liu, Han Hu, Zilong Wang, Rafael Salas, Jithin Jose, Prabhat Ram, et al., 2023 https://scholar.google.com/scholar?q=Tutel%3A+Adaptive+Mixture-of-Experts+at+Scale 6. Optimizing distributed ML communication with fused computation-collective operations — Kishore Punniyamurthy, Khaled Hamidouche, Bradford M. Beckmann, 2024 https://scholar.google.com/scholar?q=Optimizing+distributed+ML+communication+with+fused+computation-collective+operations 7. TokenWeave: Efficient Compute-Communication Overlap for Distributed LLM Inference — Raja Gond, Nipun Kwatra, Ramachandran Ramjee, 2025 https://scholar.google.com/scholar?q=TokenWeave%3A+Efficient+Compute-Communication+Overlap+for+Distributed+LLM+Inference

Interactive Visualization: Small Collectives, Big TLB Cost: Reverse Address Translation in Scale-Up GPU Pods

Podden och tillhörande omslagsbild på den här sidan tillhör mcgrof. Innehållet i podden är skapat av mcgrof och inte av, eller tillsammans med, Poddtoppen.