迈向多模态预训练的物理学:知识流、模态协同、早期统一与配方
Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes
August 5, 2026
作者: Junlin Han, Shengbang Tong, David Fan, Minghao Chen, Philip Torr, Filippos Kokkinos, Mike Lewis
cs.AI
摘要
视觉为推进基础模型提供了一个关键轴心,推动其向原生统一的多模态预训练转变。尽管势头强劲,但统一训练过程中模态交互的设计空间与基本机制仍未得到充分探索。我们通过对多模态预训练的系统性探索,提供了实证层面的清晰认识。我们在合成数据集和大规模真实世界数据集上进行的受控实验,得出了关于多模态预训练机理的四点关键洞见:(i) 知识流动:我们厘清了语言、视觉理解与视觉生成如何跨模态传递知识,揭示了不同的影响模式与不对称性;(ii) 协同与竞争:我们表明,数据的"复杂度"在很大程度上决定了模态之间是协同还是竞争,识别出能够促进协同的架构选择——例如共享注意力与归一化配合模态特定的前馈层——并发现这些行为可泛化至不同的视觉分词器设计;(iii) 早期统一:从最初阶段就统一模态并进行联合训练,被证明比晚期对齐或顺序训练更为有效。这一过程揭示了"视觉惰性"现象,即延迟整合会导致模型依赖语言先验;(iv) 配方:我们推导出了高效的预训练配方,仅用5%的计算预算即可实现强生成性能。这些核心发现随后通过训练多个参数量为13.5B的MoE模型、在2T token上进行验证。我们希望这项研究能够为理解和扩展多模态预训练提供一个有原则的基础。
English
Vision offers a critical axis for advancing foundation models, driving a shift towards natively unified multimodal pretraining. Despite this momentum, the design space and the fundamental mechanisms of how modalities interact during unified training remain underexplored. We provide empirical clarity through a systematic exploration of multimodal pretraining. Our controlled experiments on both synthetic and large-scale real-world datasets yield four key insights into the physics of multimodal pretraining: (i) Knowledge Flow: We disentangle how language, visual understanding, and visual generation transfer knowledge across modalities, revealing distinct patterns of influence and asymmetry; (ii) Synergy vs. Competition: We show that data "complexity" largely determines whether modalities are synergistic, identify architectural choices that promote synergy: such as shared attention and normalization with modality-specific feed-forward layers, and find that these behaviors generalize across different visual tokenizer designs; (iii) Early Unification: Unifying modalities from the very early stages and training them jointly is shown to be more effective than late alignment or sequential training. This process uncovers a vision laziness phenomenon, where delayed integration leads models to rely on language priors; (iv) Recipes: We derive efficient pretraining recipes that achieve strong generative performance using only 5% of the compute budget. These core findings are subsequently validated at scale by training multiple 13.5B MoE models on 2T tokens. We hope this study provides a principled foundation for understanding and scaling multimodal pretraining.