Efforts to make vision-language models efficient on edge devices have focused on reducing visual tokens, assuming visual processing dominates energy costs. This paper's systematic energy profiling across multiple models and hardware platforms overturns that assumption: inference power is nearly constant regardless of input, while output token count - driven by slower per-token decode time - is the true energy driver, with image complexity affecting energy mainly through longer generated responses. Applications include redesigning edge AI efficiency strategies to prioritize controlling output length over visual token pruning, informing hardware and software optimization for battery-powered robots, drones, and mobile AI assistants.
Podden och tillhörande omslagsbild på den här sidan tillhör
Craig Spencer Smith. Innehållet i podden är skapat av Craig Spencer Smith och inte av,
eller tillsammans med, Poddtoppen.