ChatPaper.aiChatPaper

マルチモーダル事前学習の物理学に向けて:知識フロー、モダリティの相乗効果、早期統合、そしてレシピ

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes

August 5, 2026
著者: Junlin Han, Shengbang Tong, David Fan, Minghao Chen, Philip Torr, Filippos Kokkinos, Mike Lewis
cs.AI

要旨

視覚は、基盤モデルを発展させるための重要な軸を提供し、ネイティブに統合されたマルチモーダル事前学習への移行を推進する。この勢いにもかかわらず、統合訓練中におけるモダリティ間の相互作用についての設計空間と基本的なメカニズムは、十分に探求されていない。我々は、マルチモーダル事前学習の体系的な探求を通じて、実証的な明確さを提供する。合成データセットと大規模な実世界データセットの両方を用いた我々の制御実験は、マルチモーダル事前学習の物理に関する4つの重要な洞察をもたらす。(i) 知識の流れ:言語、視覚的理解、視覚生成がモダリティ間でどのように知識を伝達するかを解きほぐし、影響と非対称性の明確なパターンを明らかにする。(ii) 相乗効果と競合:データの「複雑さ」が、モダリティが相乗的であるかどうかを主に決定することを示す。また、相乗効果を促進するアーキテクチャ上の選択、例えば共有のアテンションと正規化、およびモダリティ固有のフィードフォワード層を特定し、これらの挙動が異なる視覚トークナイザ設計にわたって一般化されることを見出す。(iii) 早期統合:モダリティを非常に初期の段階から統合し、それらを共同で訓練することは、後期のアライメントや逐次訓練よりも効果的であることが示される。このプロセスは、統合の遅延がモデルに言語の事前知識への依存をもたらす「視覚の怠惰」現象を明らかにする。(iv) レシピ:計算予算のわずか5%で強力な生成性能を達成する効率的な事前学習レシピを導出する。これらの中心的な発見は、その後、2Tトークンを用いて複数の13.5B MoEモデルを訓練することにより、大規模に検証される。我々は、この研究がマルチモーダル事前学習の理解とスケーリングのための原理に基づく基盤を提供することを願っている。
English
Vision offers a critical axis for advancing foundation models, driving a shift towards natively unified multimodal pretraining. Despite this momentum, the design space and the fundamental mechanisms of how modalities interact during unified training remain underexplored. We provide empirical clarity through a systematic exploration of multimodal pretraining. Our controlled experiments on both synthetic and large-scale real-world datasets yield four key insights into the physics of multimodal pretraining: (i) Knowledge Flow: We disentangle how language, visual understanding, and visual generation transfer knowledge across modalities, revealing distinct patterns of influence and asymmetry; (ii) Synergy vs. Competition: We show that data "complexity" largely determines whether modalities are synergistic, identify architectural choices that promote synergy: such as shared attention and normalization with modality-specific feed-forward layers, and find that these behaviors generalize across different visual tokenizer designs; (iii) Early Unification: Unifying modalities from the very early stages and training them jointly is shown to be more effective than late alignment or sequential training. This process uncovers a vision laziness phenomenon, where delayed integration leads models to rely on language priors; (iv) Recipes: We derive efficient pretraining recipes that achieve strong generative performance using only 5% of the compute budget. These core findings are subsequently validated at scale by training multiple 13.5B MoE models on 2T tokens. We hope this study provides a principled foundation for understanding and scaling multimodal pretraining.