ChatPaper.aiChatPaper

De Toestand-Voorspellingsscheidingshypothese

The State-Prediction Separation Hypothesis

July 1, 2026
Auteurs: Giovanni Monea, Nathan Godey, Kianté Brantley, Yoav Artzi
cs.AI

Samenvatting

Transformers gebruiken dezelfde voorwaartse rekenstroom om zowel het volgende token te voorspellen als nuttige toestand op te slaan voor toekomstige tokenvoorspellingen. We formuleren de toestand-voorspellingsscheidingshypothese: het ontwarren van de twee rollen leidt tot betere taalmodelleringsprestaties. We ontwerpen een Transformer-variant die twee rekenstromen gebruikt om de twee functies te scheiden, en voeren voortrainingsexperimenten uit op verschillende schalen. Onze experimenten tonen aan dat toestand-voorspellingsscheiding consistent betere data- en rekenefficiëntie biedt, het validatieverlies verbetert en gemiddeld 2 tot 3 procentpunt beter presteert dan standaard Transformers op downstream taken. We voeren ook uitgebreide empirische analyse uit die mogelijke verstorende factoren uitsluit en het fundamentele verschil aantoont in de gradienten die ons ontwerp met zich meebrengt.
English
Transformers use the same forward computation stream to both predict the next token and store useful state for future token predictions. We formulate the state-prediction separation hypothesis: disentangling the two roles yields better language modeling performance. We design a Transformer variant that uses two computation streams to separate the two functions, and conduct pretraining experiments across various scales. Our experiments show that state-prediction separation consistently offers better data and compute efficiencies, improving validation loss and outperforming standard Transformers by 2--3 percentage points on average on downstream tasks. We also conduct extensive empirical analysis that rules out potential confounders and demonstrates the fundamental difference in the gradients our design entails.