This episode explores DeltaProduct, a linear-RNN architecture that improves state-tracking by generalizing DeltaNet's single Householder reflection into a product of multiple reflections per token. The discussion traces the theoretical foundation: transformers and diagonal linear RNNs like Mamba face a proven complexity-class ceiling (TC0 vs NC1) that prevents them from tracking permutation-group state such as parity, while DeltaNet's recurrence—reinterpreted as one step of online gradient descent on an associative recall objective—naturally produces a Householder reflection as its state-transition matrix. DeltaProduct extends this by taking multiple gradient steps per token, and the hosts unpack why this matters via the Cartan-Dieudonné theorem: composing two reflections yields a full rotation, not just a bigger reflection, making the jump from one to two steps a qualitative leap in expressive power rather than an incremental one. Listeners get a rare example of a paper that derives a geometric theorem before running experiments, then tests whether empirical results match the prediction, offering a genuine dial between computational efficiency and expressivity instead of a fixed architectural tradeoff.
Sources:
1. DeltaProduct: Improving State-Tracking in Linear RNNs via Householder Products — Julien Siems, Timur Carstensen, Arber Zela, Frank Hutter, Massimiliano Pontil, Riccardo Grazzi, 2025
http://arxiv.org/abs/2502.10297
2. Unlocking State-Tracking in Linear RNNs Through Negative Eigenvalues — R. Grazzi, J. Siems, A. Zela, J. Franke, F. Hutter, M. Pontil, 2025 (ICLR)
https://scholar.google.com/scholar?q=Unlocking+State-Tracking+in+Linear+RNNs+Through+Negative+Eigenvalues
3. Parallelizing Linear Transformers with the Delta Rule over Sequence Length — S. Yang, B. Wang, Y. Zhang, Y. Shen, Y. Kim, 2024 (NeurIPS)
https://scholar.google.com/scholar?q=Parallelizing+Linear+Transformers+with+the+Delta+Rule+over+Sequence+Length
4. Gated Delta Networks: Improving Mamba2 with the Delta Rule — S. Yang, J. Kautz, A. Hatamizadeh, 2025 (ICLR)
https://scholar.google.com/scholar?q=Gated+Delta+Networks%3A+Improving+Mamba2+with+the+Delta+Rule
5. RWKV-7 'Goose' with Expressive Dynamic State Evolution — B. Peng, R. Zhang, D. Goldstein, et al., 2025
https://scholar.google.com/scholar?q=RWKV-7+%27Goose%27+with+Expressive+Dynamic+State+Evolution
6. The Illusion of State in State-Space Models — W. Merrill, J. Petty, A. Sabharwal, 2024 (ICML)
https://scholar.google.com/scholar?q=The+Illusion+of+State+in+State-Space+Models
7. Fixed-Point RNNs: From Diagonal to Dense in a Few Iterations — S. Movahedi, F. Sarnthein, N. Muca Cirone, A. Orvieto, 2025
https://scholar.google.com/scholar?q=Fixed-Point+RNNs%3A+From+Diagonal+to+Dense+in+a+Few+Iterations
Interactive Visualization: DeltaProduct: Extending DeltaNet's State-Tracking via Householder Products