An AI-MI–supported study proposes that standard Transformers overload a single computation stream to both predict the next token and store state for future tokens—and that disentangling these roles yields more efficient, higher-performing models. Giovanni Monea, Nathan Godey, Kianté Brantley (Harvard), and Yoav Artzi (Cornell) introduce the “state-prediction separation hypothesis” and a two-stream Transformer variant that validates it, achieving 2–3 percentage points better performance on downstream tasks across pretraining scales compared to standard Transformers.

From Monea, Godey, Brantley & Artzi, arXiv:2607.01218 (2026). Used under arXiv non-exclusive license.
Source: arXiv:2607.01218

