生成を補助監視として:分離埋め込み予測によるゼロ推論オーバーヘッドでの視覚的理解の向上
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction
August 12, 2026
著者: Zhongbin Guo, Jiahao Xie, Dongling Xiao, Qianle Wang, Ruiqi Lu, Xiaomin He, Wanxuan Sun, Cheng Yang
cs.AI
要旨
マルチモーダル大規模言語モデル(MLLMs)は目覚ましい進歩を達成している一方で、視覚的理解と生成は通常、 divergent な目的として扱われている。既存の統一フレームワークは、離散的な視覚トークン化や拡散目的に依存することが多く、その生成ターゲットは視覚理解モデルが消費する連続表現とは異なるため、既存の事前学習済みMLLMを強化するための直接的な転移は容易ではない。本研究では、視覚生成を表現学習のための補助的監督として再解釈する、生成誘導型訓練フレームワークGASを提案する。具体的には、GASは分離型Mixture-of-Transformers(MoT)アーキテクチャ内で、次エンベディング予測(NEP)をクロスモーダル生成パラダイムとして適用する。共有下部トランクと並列上部層を維持することにより、GASは生成損失が共有視覚経路をより細かい空間精度とより強い視覚保持で豊かにする一方で、上部の理解層を生成勾配から直接的には遮蔽する。この相乗効果を最大化するため、さらに、単なる一般的な合成ではなく深い認知基盤を必要とする高度に相関した生成タスクを構築する。モデル規模と訓練段階を通じて、GASは総合的なマルチモーダル理解を改善し、特に知覚と空間理解において最も確実な利得をもたらす。重要なことに、補助的な生成ブランチは訓練後に破棄されるため、これらの利得は推論オーバーヘッドを一切生じない。広範な対照比較と表現レベルの分析により、生成誘導型訓練が理解にいつ、なぜ有益であるかがさらに解明され、生成誘導型訓練がより強力なマルチモーダル理解への実用的な経路としての実現可能性が示される。
English
While Multimodal Large Language Models (MLLMs) have achieved remarkable progress, visual understanding and generation are typically treated as divergent objectives. Existing unified frameworks often rely on discrete visual tokenization or diffusion objectives whose generative targets differ from the continuous representations consumed by visual understanding models, making direct transfer to enhance existing pretrained MLLMs non-trivial. In this work, we present GAS, a generation-guided training framework that reinterprets visual generation as auxiliary supervision for representation learning. Concretely, GAS adapts Next Embedding Prediction (NEP) as a cross-modal generation paradigm within a decoupled Mixture-of-Transformers (MoT) architecture. By maintaining a shared lower trunk and parallel upper layers, GAS lets generation losses enrich the shared visual pathway with finer spatial precision and stronger visual retention while shielding the upper understanding layers from direct generation gradients. To maximize this synergy, we further construct highly correlated generation tasks that demand deep cognitive grounding rather than generic synthesis alone. Across model scales and training stages, GAS improves aggregate multimodal understanding, with its most reliable gains on perception and spatial comprehension. Crucially, because the auxiliary generation branch is discarded after training, these gains incur zero inference overhead. Extensive controlled comparisons and representation-level analyses further clarify when and why generation-guided training benefits understanding, and demonstrate the feasibility of generation-guided training as a practical route to stronger multimodal understanding.