네이티브 통합 멀티모달 모델의 이해-생성 시너지 규명: 표현에서 태스크, 시스템까지
Uncovering Understanding-Generation Synergy in Native Unified Multimodal Models: From Representation, Task to System
September 1, 2026
저자: Penghao Wu, Haiwen Diao, Weichen Fan, Lewei Lu, Dahua Lin, Ziwei Liu
cs.AI
초록
통합 멀티모달 모델(UMMs)은 단일 모델 내에서 시각 이해와 생성을 공동으로 수행하지만, 기능적 통합이 학습 시너지를 보장하지는 않는다. 두 목표는 서로를 강화하거나, 용량을 두고 경쟁하거나, 단순히 공존할 수도 있다. 우리는 사전 훈련된 시각 사전 지식 없이, 구조적으로 고유한 통제된 환경에서 이 관계를 표현, 과제, 시스템 수준으로 나누어 조사한다. 표현 수준에서는 각 목표가 다른 목표에 유용한 신호를 제공한다는 것을 발견한다. 생성은 이해를 위해 학습된 시각적 특징을 풍부하게 하고, 이해는 생성을 위한 시각-언어 정렬을 강화한다. 그러나 두 목표가 동일한 계산 경로를 통해 강제될 때 한쪽이 지배하는 경향이 있다. 상충하는 시각 계산을 특화하면서도 의미론적 상호작용을 유지하는 과제 분리 아키텍처는 이러한 비대칭적 성능 저하를 피한다. 과제 수준에서는 세 가지 사례 연구를 통해 이해 및 생성 과제가 공유 지식에 의존할 때 긍정적인 양방향 전이가 발생함을 확인한다. 시스템 수준에서는 종단간 UMM이 이미지 이해와 생성을 모두 명시적으로 요구하는 복잡한 과제에서 이에 대응되는 플래너-실행자 파이프라인보다 우수한 성능을 보임을 입증한다. 종합하면, 이러한 결과는 UMM의 가치가 통합된 인터페이스 이상에 있음을 보여준다. 적절한 특화, 공유된 과제 지식, 종단간 최적화는 공존을 시너지로 전환할 수 있다.
English
While unified multimodal models (UMMs) jointly perform visual understanding and generation within a single model, functional unification does not guarantee learning synergy: the two objectives may reinforce each other, compete for capacity, or merely coexist. We investigate their relationship at the representation, task, and system levels in a controlled, structurally native setting without pretrained vision priors. At the representation level, we find that each objective provides useful signal to the other: generation enriches the visual features learned for understanding, while understanding strengthens vision--language alignment for generation. However, when both objectives are forced through the same computation path, one tends to dominate. A task-decoupled architecture that specializes conflicting visual computation while preserving semantic interaction avoids this asymmetric degradation. At the task level, through three case studies, we find positive bidirectional transfer when understanding and generation tasks rely on shared knowledge. At the system level, we show that an end-to-end UMM outperforms a matched planner--executor pipeline on complex tasks that explicitly require both image understanding and generation. Together, these results show that the value of UMMs extends beyond a unified interface: appropriate specialization, shared task knowledge, and end-to-end optimization can turn coexistence into synergy.