This episode examines the "last mile" problem in multi-GPU scale-up fabrics: when a remote memory request arrives at a destination GPU carrying a Network Physical Address, that address means nothing locally until it's converted back to a System Physical Address through Reverse Address Translation. The discussion traces why this destination-side translation problem is genuinely new — decades of TLB optimization research has assumed the initiating processor controls the access pattern, while NVLink and UALink fabrics flip that model, forcing the receiving GPU to translate incoming requests with no warning and no control. The hosts connect this seemingly niche hardware detail to a concrete workload: Mixture-of-Experts models rely on All-to-All dispatch and gather collectives, implemented in libraries like NCCL and RCCL, that cross this translation step twice per layer across dozens of layers. To quantify the impact, the paper extends the ASTRA-sim2.0 simulator with an Omnet++ network backend to model packet-level UALink Clos topologies, generating realistic All-to-All traffic via Microsoft's MSCCLang. Listeners interested in the unglamorous plumbing beneath large-scale AI infrastructure will find a compelling case that address translation, not just compute or bandwidth, may be a hidden bottleneck for the collective communication patterns underpinning today's largest inference deployments.
Sources:
1. Amel Fatima's Reverse Address Translation in GPU Fabrics
https://arxiv.org/pdf/2604.02473
2. Trans-FW: Short Circuiting Page Table Walk in Multi-GPU Systems via Remote Forwarding — Bingyao Li, Jieming Yin, Anup Holey, Youtao Zhang, Jun Yang, Xulong Tang, 2023
https://scholar.google.com/scholar?q=Trans-FW%3A+Short+Circuiting+Page+Table+Walk+in+Multi-GPU+Systems+via+Remote+Forwarding
3. ASTRA-sim2.0: Modeling Hierarchical Networks and Disaggregated Systems for Large-model Training at Scale — William Won, Taekyung Heo, Saeed Rashidi, Srinivas Sridharan, Sudarshan Srinivasan, Tushar Krishna, 2023
https://scholar.google.com/scholar?q=ASTRA-sim2.0%3A+Modeling+Hierarchical+Networks+and+Disaggregated+Systems+for+Large-model+Training+at+Scale
4. MSCCLang: Microsoft Collective Communication Language — Meghan Cowan, Saeed Maleki, Madanlal Musuvathi, Olli Saarikivi, Yifan Xiong, 2023
https://scholar.google.com/scholar?q=MSCCLang%3A+Microsoft+Collective+Communication+Language
5. Tutel: Adaptive Mixture-of-Experts at Scale — Changho Hwang, Wei Cui, Yifan Xiong, Ziyue Yang, Ze Liu, Han Hu, Zilong Wang, Rafael Salas, Jithin Jose, Prabhat Ram, et al., 2023
https://scholar.google.com/scholar?q=Tutel%3A+Adaptive+Mixture-of-Experts+at+Scale
6. Optimizing distributed ML communication with fused computation-collective operations — Kishore Punniyamurthy, Khaled Hamidouche, Bradford M. Beckmann, 2024
https://scholar.google.com/scholar?q=Optimizing+distributed+ML+communication+with+fused+computation-collective+operations
7. TokenWeave: Efficient Compute-Communication Overlap for Distributed LLM Inference — Raja Gond, Nipun Kwatra, Ramachandran Ramjee, 2025
https://scholar.google.com/scholar?q=TokenWeave%3A+Efficient+Compute-Communication+Overlap+for+Distributed+LLM+Inference
Interactive Visualization: Small Collectives, Big TLB Cost: Reverse Address Translation in Scale-Up GPU Pods