ChatPaper.aiChatPaper

邁向多模態預訓練的物理機制:知識流動、模態協同、早期統一與實作法則

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes

August 5, 2026
作者: Junlin Han, Shengbang Tong, David Fan, Minghao Chen, Philip Torr, Filippos Kokkinos, Mike Lewis
cs.AI

摘要

視覺為推進基礎模型提供了一條關鍵軸線,驅動著向原生統一多模態預訓練的轉變。儘管有這樣的發展勢頭,關於模態在統一訓練期間如何互動的設計空間與基本機制仍未被充分探索。我們透過對多模態預訓練的系統性探索提供了經驗性的清晰理解。我們在合成資料集與大規模真實世界資料集上進行的受控實驗,產出了關於多模態預訓練物理機制之四大關鍵洞見:(i) 知識流動:我們釐清了語言、視覺理解與視覺生成如何在各模態之間傳遞知識,揭示了截然不同的影響模式與不對稱性;(ii) 協同與競爭:我們表明資料的「複雜度」在很大程度上決定了模態之間是否具有協同作用,並識別出能促進協同作用的架構選擇,例如共享注意力與正規化搭配模態特定的前饋層,同時發現這些行為能推廣至不同的視覺分詞器設計;(iii) 早期統一:從非常早期的階段就統一各模態並進行聯合訓練,被證明比晚期對齊或循序訓練更為有效。此過程揭示了「視覺惰性」現象,即延遲整合會導致模型依賴語言先驗;(iv) 配方:我們推導出高效的預訓練配方,僅使用 5% 的計算預算即可實現強大的生成性能。這些核心發現隨後透過在多個 13.5B 混合專家模型上以 2T 詞元進行訓練而在大規模上得到驗證。我們希望這項研究能為理解與擴展多模態預訓練提供原則性的基礎。
English
Vision offers a critical axis for advancing foundation models, driving a shift towards natively unified multimodal pretraining. Despite this momentum, the design space and the fundamental mechanisms of how modalities interact during unified training remain underexplored. We provide empirical clarity through a systematic exploration of multimodal pretraining. Our controlled experiments on both synthetic and large-scale real-world datasets yield four key insights into the physics of multimodal pretraining: (i) Knowledge Flow: We disentangle how language, visual understanding, and visual generation transfer knowledge across modalities, revealing distinct patterns of influence and asymmetry; (ii) Synergy vs. Competition: We show that data "complexity" largely determines whether modalities are synergistic, identify architectural choices that promote synergy: such as shared attention and normalization with modality-specific feed-forward layers, and find that these behaviors generalize across different visual tokenizer designs; (iii) Early Unification: Unifying modalities from the very early stages and training them jointly is shown to be more effective than late alignment or sequential training. This process uncovers a vision laziness phenomenon, where delayed integration leads models to rely on language priors; (iv) Recipes: We derive efficient pretraining recipes that achieve strong generative performance using only 5% of the compute budget. These core findings are subsequently validated at scale by training multiple 13.5B MoE models on 2T tokens. We hope this study provides a principled foundation for understanding and scaling multimodal pretraining.