Separating Prediction and Memory Improves Transformers

Posted: July 14, 2026

An AI-MI–supported study proposes that standard Transformers overload a single computation stream to both predict the next token and store state for future tokens—and that disentangling these roles yields more efficient, higher-performing models. Giovanni Monea, Nathan Godey, Kianté Brantley (Harvard), and Yoav Artzi (Cornell) introduce the “state-prediction separation hypothesis” and a two-stream Transformer variant that validates it, achieving 2–3 percentage points better performance on downstream tasks across pretraining scales compared to standard Transformers.

Figure 1 from Monea et al. arXiv:2607.01218 — state-prediction separation architecture

From Monea, Godey, Brantley & Artzi, arXiv:2607.01218 (2026). Used under arXiv non-exclusive license.

Source: arXiv:2607.01218

More News

AI-MI at ICML 2026, Part II

AI-MI researchers continued to make their mark at the International Conference on Machine Learning (ICML 2026) in Seoul. Kilian Weinberger, serving as ICML 2016 program co-chair, led the selection process for this year’s Test of Time Award; his…