A 30B parameter model runs on a MacBook because only 3B parameters fire per token. Mixture of Experts splits memory cost from compute cost, and that changes everything about where AI can run.
Podden och tillhörande omslagsbild på den här sidan tillhör
Mo Bhuiyan. Innehållet i podden är skapat av Mo Bhuiyan och inte av,
eller tillsammans med, Poddtoppen.