This episode explores EventTensor, a compiler abstraction from Carnegie Mellon and collaborators (presented at MLSys 2026) that treats synchronization events as first-class tensors for compiling GPU megakernels. The discussion covers how encoding true data dependencies—rather than waiting for entire kernels to finish—enables fine-grained scheduling, illustrated through a split-K summation example and a symbolic batch-size template that avoids recompilation when shapes change. A key focus is how the system handles Mixture-of-Experts routing, where dependencies aren't known until runtime, via data-dependent event counters and task triggering computed from router outputs. The hosts also unpack the tradeoffs between static and dynamic scheduling, showing that static wins on predictable dense workloads while dynamic pays off only under genuine irregularity like MoE. Benchmark results show up to 1.40x speedups over cuBLAS+NCCL, 1.23x over Triton/FlashInfer on MoE layers, and end-to-end gains of 1.48x over vLLM, making this a concrete look at how compile-time and runtime scheduling can be unified without sacrificing performance.

Sources: 1. Event Tensor: A Unified Abstraction for Compiling Dynamic Megakernel — Hongyi Jin, Bohan Hou, Guanjie Wang, Ruihang Lai, Jinqi Chen, Zihao Ye, Yaxing Cai, Yixin Dong, Xinhao Cheng, Zhihao Zhang, Yilong Zhao, Yingyi Huang, Lijie Yang, Jinchen Jiang, Gabriele Oliaro, Jianan Ji, Xupeng Miao, Vinod Grover, Todd C. Mowry, Zhihao Jia, Tianqi Chen, 2026 http://arxiv.org/abs/2604.13327 2. Legion: Expressing Locality and Independence with Logical Regions — Michael Bauer, Sean Treichler, Elliott Slaughter, Alex Aiken, 2012 https://scholar.google.com/scholar?q=Legion%3A+Expressing+Locality+and+Independence+with+Logical+Regions 3. StarPU: A Unified Platform for Task Scheduling on Heterogeneous Multicore Architectures — Cédric Augonnet, Samuel Thibault, Raymond Namyst, Pierre-André Wacrenier, 2011 https://scholar.google.com/scholar?q=StarPU%3A+A+Unified+Platform+for+Task+Scheduling+on+Heterogeneous+Multicore+Architectures 4. Dynamic Control Flow in Large-Scale Machine Learning — Yuan Yu, Martín Abadi, Paul Barham, Eugene Brevdo, Mike Burrows, Andy Davis, Jeff Dean, Sanjay Ghemawat, Tim Harley, Peter Hawkins, Mark Hong, Rajat Monga, Derek Murray, Xiaoqiang Zheng, and others (Google Brain), 2018 https://scholar.google.com/scholar?q=Dynamic+Control+Flow+in+Large-Scale+Machine+Learning 5. Ray: A Distributed Framework for Emerging AI Applications — Philipp Moritz, Robert Nishihara, Stephanie Wang, Alexey Tumanov, Richard Liaw, Eric Liang, Melih Elibol, Zongheng Yang, William Paul, Michael I. Jordan, Ion Stoica, 2018 https://scholar.google.com/scholar?q=Ray%3A+A+Distributed+Framework+for+Emerging+AI+Applications 6. Mirage Persistent Kernel: A Compiler and Runtime for Mega-Kernelizing Tensor Programs — Cheng, X., Zhang, Z., Zhou, Y., Ji, J., Jiang, J., Zhao, Z., et al. (overlapping author list with this paper), 2025 https://scholar.google.com/scholar?q=Mirage+Persistent+Kernel%3A+A+Compiler+and+Runtime+for+Mega-Kernelizing+Tensor+Programs 7. Look ma, no bubbles! Designing a low-latency megakernel for Llama-1B — Spector, B., Juravsky, J., Sul, S., Dugan, O., Lim, D., Fu, D., Arora, S., R, C., 2025 https://scholar.google.com/scholar?q=Look+ma%2C+no+bubbles%21+Designing+a+low-latency+megakernel+for+Llama-1B 8. A Framework for Fine-Grained Synchronization of Dependent GPU Kernels (CuSync) — Jangda, A., Maleki, S., Dehnavi, M. M., Musuvathi, M., Saarikivi, O., 2024 https://scholar.google.com/scholar?q=A+Framework+for+Fine-Grained+Synchronization+of+Dependent+GPU+Kernels+%28CuSync%29 9. Graphene: An IR for Optimized Tensor Computations on GPUs — Hagedorn, B., Fan, B., Chen, H., Cecka, C., Garland, M., Grover, V., 2023 https://scholar.google.com/scholar?q=Graphene%3A+An+IR+for+Optimized+Tensor+Computations+on+GPUs 10. FlashMoE: Fast Distributed MoE in a Single Kernel — Aimuyo, O. J., Oh, B., Singh, R., 2025 https://scholar.google.com/scholar?q=FlashMoE%3A+Fast+Distributed+MoE+in+a+Single+Kernel

Interactive Visualization: Prompt Boundary-Aware Scheduling with Event Tensors for Dynamic Kernels

Podden och tillhörande omslagsbild på den här sidan tillhör mcgrof. Innehållet i podden är skapat av mcgrof och inte av, eller tillsammans med, Poddtoppen.