This episode explores Parcae, a paper on scaling laws for stable looped language models, where instead of stacking distinct transformer layers, a single block is applied repeatedly through the residual stream — echoing Universal Transformers and ALBERT's weight-tying but tackling the training instability that has historically plagued the approach. The discussion centers on reframing looped inference as a linear time-invariant dynamical system, showing that the spectral norm of the transition matrix A determines whether the residual stream stays bounded or explodes exponentially — turning a mysterious loss-spike failure mode into a measurable, checkable quantity. It also covers how the authors diagnose this concretely by examining the eigenvalues of A (contrasting how different prelude-embedding injection methods affect stability), and how they extend Chinchilla-style isoFLOP curve-fitting with a third axis — recurrence depth — to find the FLOP-optimal number of loops at a given compute budget. Listeners interested in efficient inference, edge deployment, or the mechanics of why prior recurrent-depth models like RDM needed fragile tuning will find the control-theory framing a clarifying, math-grounded alternative to typical trial-and-error architecture papers.
Sources:
1. Parcae: Scaling Laws For Stable Looped Language Models — Hayden Prairie, Zachary Novack, Taylor Berg-Kirkpatrick, Daniel Y. Fu, 2026
http://arxiv.org/abs/2604.12946
2. Universal Transformers — Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, Lukasz Kaiser, 2018
https://scholar.google.com/scholar?q=Universal+Transformers
3. ALBERT: A Lite BERT for Self-supervised Learning of Language Representations — Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, Radu Soricut, 2019
https://scholar.google.com/scholar?q=ALBERT%3A+A+Lite+BERT+for+Self-supervised+Learning+of+Language+Representations
4. Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach — Jonas Geiping et al., 2025
https://scholar.google.com/scholar?q=Scaling+up+Test-Time+Compute+with+Latent+Reasoning%3A+A+Recurrent+Depth+Approach
5. Relaxed Recursive Transformers: Effective Parameter Sharing with Layer-wise LoRA — Sangmin Bae et al., 2024
https://scholar.google.com/scholar?q=Relaxed+Recursive+Transformers%3A+Effective+Parameter+Sharing+with+Layer-wise+LoRA
6. A Proposal on Machine Learning via Dynamical Systems — Weinan E, 2017
https://scholar.google.com/scholar?q=A+Proposal+on+Machine+Learning+via+Dynamical+Systems
7. Neural Ordinary Differential Equations — Ricky T. Q. Chen, Yulia Rubanova, Jesse Bettencourt, David Duvenaud, 2018
https://scholar.google.com/scholar?q=Neural+Ordinary+Differential+Equations
8. Stable Architectures for Deep Neural Networks — Eldad Haber, Lars Ruthotto, 2017
https://scholar.google.com/scholar?q=Stable+Architectures+for+Deep+Neural+Networks
9. Mamba: Linear-Time Sequence Modeling with Selective State Spaces — Albert Gu, Tri Dao, 2023
https://scholar.google.com/scholar?q=Mamba%3A+Linear-Time+Sequence+Modeling+with+Selective+State+Spaces
10. Transformers are SSMs: Generalized Models and Efficient Algorithms through Structured State Space Duality — Tri Dao, Albert Gu, 2024
https://scholar.google.com/scholar?q=Transformers+are+SSMs%3A+Generalized+Models+and+Efficient+Algorithms+through+Structured+State+Space+Duality
11. Mixture-of-Recursions: Learning Dynamic Recursive Depths for Adaptive Token-Level Computation — Sangmin Bae, Yujin Kim, Reza Bayat, et al., 2025
https://scholar.google.com/scholar?q=Mixture-of-Recursions%3A+Learning+Dynamic+Recursive+Depths+for+Adaptive+Token-Level+Computation
12. Scaling Latent Reasoning via Looped Language Models — Rui-Jie Zhu, Zixuan Wang, Kai Hua, et al., 2025
https://scholar.google.com/scholar?q=Scaling+Latent+Reasoning+via+Looped+Language+Models
13. Teaching Pretrained Language Models to Think Deeper with Retrofitted Recurrence — Sean McLeish, Ang Li, John Kirchenbauer, et al., 2025
https://scholar.google.com/scholar?q=Teaching+Pretrained+Language+Models+to+Think+Deeper+with+Retrofitted+Recurrence
14. Reasoning with Latent Thoughts: On the Power of Looped Transformers — Nikunj Saunshi, Nishanth Dikkala, Zhiyuan Li, Sanjiv Kumar, Sashank J. Reddi, 2025
https://scholar.google.com/scholar?q=Reasoning+with+Latent+Thoughts%3A+On+the+Power+of+Looped+Transformers
Interactive Visualization: Parcae: Stabilizing Looped Language Models with Control Theory