멀티모달 사전학습의 물리학을 향하여: 지식 흐름, 모달리티 시너지, 조기 통합, 그리고 레시피
Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes
August 5, 2026
저자: Junlin Han, Shengbang Tong, David Fan, Minghao Chen, Philip Torr, Filippos Kokkinos, Mike Lewis
cs.AI
초록
비전은 파운데이션 모델을 발전시키는 중요한 축을 제공하며, 본질적으로 통합된 멀티모달 사전학습으로의 전환을 주도한다. 이러한 추진력에도 불구하고, 통합 학습 중 모달리티 간 상호작용의 설계 공간과 기본 메커니즘은 여전히 충분히 탐구되지 않았다. 우리는 멀티모달 사전학습에 대한 체계적 탐구를 통해 경험적 명확성을 제공한다. 합성 데이터와 대규모 실세계 데이터셋 모두에 대한 통제된 실험을 통해 멀티모달 사전학습의 물리학에 관한 네 가지 핵심 통찰을 도출한다: (i) 지식 흐름: 언어, 시각 이해, 시각 생성이 모달리티 간에 지식을 어떻게 전이하는지를 분리하여, 영향력과 비대칭성의 뚜렷한 패턴을 밝힌다; (ii) 시너지 대 경쟁: 데이터의 "복잡성"이 모달리티가 시너지 효과를 갖는지를 크게 결정함을 보여주고, 시너지를 촉진하는 아키텍처 선택(예: 모달리티별 피드포워드 계층을 갖춘 공유 어텐션 및 정규화)을 식별하며, 이러한 행동이 다양한 시각 토크나이저 설계에 걸쳐 일반화됨을 발견한다; (iii) 조기 통합: 모달리티를 매우 초기 단계부터 통합하고 공동으로 학습하는 것이 후기 정렬이나 순차적 학습보다 더 효과적임을 보여준다. 이 과정에서 지연된 통합이 모델이 언어 사전 지식에 의존하게 만드는 비전 나태 현상을 발견한다; (iv) 레시피: 컴퓨팅 예산의 5%만 사용하여 강력한 생성 성능을 달성하는 효율적인 사전학습 레시피를 도출한다. 이러한 핵심 발견은 이후 2T 토큰으로 13.5B MoE 모델 여러 개를 학습시켜 대규모로 검증된다. 본 연구가 멀티모달 사전학습을 이해하고 확장하기 위한 원리적 기반을 제공하기를 기대한다.
English
Vision offers a critical axis for advancing foundation models, driving a shift towards natively unified multimodal pretraining. Despite this momentum, the design space and the fundamental mechanisms of how modalities interact during unified training remain underexplored. We provide empirical clarity through a systematic exploration of multimodal pretraining. Our controlled experiments on both synthetic and large-scale real-world datasets yield four key insights into the physics of multimodal pretraining: (i) Knowledge Flow: We disentangle how language, visual understanding, and visual generation transfer knowledge across modalities, revealing distinct patterns of influence and asymmetry; (ii) Synergy vs. Competition: We show that data "complexity" largely determines whether modalities are synergistic, identify architectural choices that promote synergy: such as shared attention and normalization with modality-specific feed-forward layers, and find that these behaviors generalize across different visual tokenizer designs; (iii) Early Unification: Unifying modalities from the very early stages and training them jointly is shown to be more effective than late alignment or sequential training. This process uncovers a vision laziness phenomenon, where delayed integration leads models to rely on language priors; (iv) Recipes: We derive efficient pretraining recipes that achieve strong generative performance using only 5% of the compute budget. These core findings are subsequently validated at scale by training multiple 13.5B MoE models on 2T tokens. We hope this study provides a principled foundation for understanding and scaling multimodal pretraining.