ネイティブ統合マルチモーダルモデルにおける理解と生成のシナジーの解明:表現、タスク、システムの観点から
Uncovering Understanding-Generation Synergy in Native Unified Multimodal Models: From Representation, Task to System
September 1, 2026
著者: Penghao Wu, Haiwen Diao, Weichen Fan, Lewei Lu, Dahua Lin, Ziwei Liu
cs.AI
要旨
統合マルチモーダルモデル(UMMs)は、単一のモデル内で視覚的理解と生成を同時に実行する一方、機能的な統合は学習における相乗効果を保証するものではない。すなわち、これら二つの目的は、互いに強化し合うこともあれば、モデル容量を競い合うこともあれば、単に共存することもある。我々は、事前学習された視覚的先行知識を持たない、制御された構造的にネイティブな設定において、表現レベル・タスクレベル・システムレベルでのそれらの関係を調査する。表現レベルでは、各目的が他方にとって有用な信号を提供することがわかる。生成は、理解のために学習される視覚的特徴を豊かにし、理解は、生成のための視覚・言語アライメントを強化する。しかし、両目的が同じ計算経路を通るように強制される場合、一方が支配的になる傾向がある。意味的相互作用を維持しつつ、競合する視覚計算を専用化するタスク分離型アーキテクチャは、この非対称な性能劣化を回避する。タスクレベルでは、3つのケーススタディを通じて、理解タスクと生成タスクが共有知識に依存する場合に、正の双方向転移が見られることが確認された。システムレベルでは、エンドツーエンドのUMMが、画像理解と生成の両方を明示的に必要とする複雑なタスクにおいて、同等のプランナー・実行器パイプラインを上回ることを示す。これらの結果は、UMMの価値が統一的なインターフェースを超えて及ぶことを示している。適切な特化、共有されたタスク知識、およびエンドツーエンドの最適化は、共存を相乗効果へと転換できるのである。
English
While unified multimodal models (UMMs) jointly perform visual understanding and generation within a single model, functional unification does not guarantee learning synergy: the two objectives may reinforce each other, compete for capacity, or merely coexist. We investigate their relationship at the representation, task, and system levels in a controlled, structurally native setting without pretrained vision priors. At the representation level, we find that each objective provides useful signal to the other: generation enriches the visual features learned for understanding, while understanding strengthens vision--language alignment for generation. However, when both objectives are forced through the same computation path, one tends to dominate. A task-decoupled architecture that specializes conflicting visual computation while preserving semantic interaction avoids this asymmetric degradation. At the task level, through three case studies, we find positive bidirectional transfer when understanding and generation tasks rely on shared knowledge. At the system level, we show that an end-to-end UMM outperforms a matched planner--executor pipeline on complex tasks that explicitly require both image understanding and generation. Together, these results show that the value of UMMs extends beyond a unified interface: appropriate specialization, shared task knowledge, and end-to-end optimization can turn coexistence into synergy.