This episode explores STEEL, a sparsity-aware fused attention design for running long-sequence prefill inference efficiently on AMD's XDNA neural processing unit. The discussion contrasts spatial-dataflow NPU architectures, where compute tiles are explicitly scheduled with no dynamic cache management, against GPU SIMT execution, and explains how the causal attention mask creates load imbalance that a fixed pipeline can't easily absorb the way a GPU scheduler can. Building on FlashAttention-2's tiling and online-softmax approach, the paper restructures the computation into a three-stage pipeline across dedicated compute cores to address that imbalance directly. The hosts walk through why this matters for on-device AI agents that need low latency, privacy, and battery efficiency without offloading to cloud GPUs. Reported results include over 9.5x latency reduction versus prior state-of-the-art NPU implementations, over 9x energy savings against a CPU baseline, and more than 22x speedup over a naive layer-by-layer approach.

Sources: 1. STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU — Victor J. B. Jung, Gagandeep Singh, Joseph Melber, Kristof Denolf, Francesco Conti, Luca Benini, 2026 http://arxiv.org/abs/2607.09385v1 2. FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness — Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, Christopher Ré, 2022 https://scholar.google.com/scholar?q=FlashAttention%3A+Fast+and+Memory-Efficient+Exact+Attention+with+IO-Awareness 3. Eyeriss: An Energy-Efficient Reconfigurable Accelerator for Deep Convolutional Neural Networks — Yu-Hsin Chen, Joel Emer, Vivienne Sze, 2016 https://scholar.google.com/scholar?q=Eyeriss%3A+An+Energy-Efficient+Reconfigurable+Accelerator+for+Deep+Convolutional+Neural+Networks 4. In-Datacenter Performance Analysis of a Tensor Processing Unit — Norman P. Jouppi et al. (Google), 2017 https://scholar.google.com/scholar?q=In-Datacenter+Performance+Analysis+of+a+Tensor+Processing+Unit 5. Plasticine: A Reconfigurable Architecture for Parallel Patterns — Raghu Prabhakar, Yaqi Zhang, David Koeplinger, Matt Feldman, Tian Zhao, Stefan Hadjis, Ardavan Pedram, Christos Kozyrakis, Kunle Olukotun, 2017 https://scholar.google.com/scholar?q=Plasticine%3A+A+Reconfigurable+Architecture+for+Parallel+Patterns 6. Efficient Memory Management for Large Language Model Serving with PagedAttention — Kwon et al., 2023 https://scholar.google.com/scholar?q=Efficient+Memory+Management+for+Large+Language+Model+Serving+with+PagedAttention 7. GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints — Ainslie et al., 2023 https://scholar.google.com/scholar?q=GQA%3A+Training+Generalized+Multi-Query+Transformer+Models+from+Multi-Head+Checkpoints 8. Efficiently Scaling Transformer Inference — Pope et al., 2022 https://scholar.google.com/scholar?q=Efficiently+Scaling+Transformer+Inference 9. FlashDecoding++: Faster Large Language Model Inference on GPUs — Hong et al., 2024 https://scholar.google.com/scholar?q=FlashDecoding%2B%2B%3A+Faster+Large+Language+Model+Inference+on+GPUs 10. NITRO: LLM Inference on Intel Laptop NPUs — Fei and Abdelfattah, 2024 https://scholar.google.com/scholar?q=NITRO%3A+LLM+Inference+on+Intel+Laptop+NPUs

Interactive Visualization: AMD XDNA NPU Sparsity-Aware Attention for Long-Context Prefill

Podden och tillhörande omslagsbild på den här sidan tillhör mcgrof. Innehållet i podden är skapat av mcgrof och inte av, eller tillsammans med, Poddtoppen.