ChatPaper.aiChatPaper

以生成作為輔助監督:透過解耦嵌入預測,在零推理開銷下增強視覺理解

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction

August 12, 2026
作者: Zhongbin Guo, Jiahao Xie, Dongling Xiao, Qianle Wang, Ruiqi Lu, Xiaomin He, Wanxuan Sun, Cheng Yang
cs.AI

摘要

儘管多模態大型語言模型(MLLMs)已取得顯著進展,視覺理解與視覺生成通常仍被視為分歧的目標。現有的統一框架往往依賴離散視覺標記化或擴散目標,而這些生成目標與視覺理解模型所使用的連續表徵不同,使得直接遷移以增強現有預訓練MLLMs並非易事。在本工作中,我們提出GAS,一個生成引導訓練框架,將視覺生成重新詮釋為表徵學習的輔助監督。具體而言,GAS在解耦的Mixture-of-Transformers(MoT)架構中採用下一嵌入預測(NEP)作為跨模態生成範式。透過維持共享的下層主幹與並行的上層網路,GAS讓生成損失以更精細的空間精度與更強的視覺保留來豐富共享的視覺路徑,同時隔離上層理解網路使其不受直接生成梯度的影響。為了最大化這種協同作用,我們進一步構建高度相關的生成任務,這些任務需要深層認知基礎而非僅是通用合成。在不同模型規模與訓練階段中,GAS均能提升整體多模態理解能力,其中在感知與空間理解上的增益最為可靠。至關重要的是,由於輔助生成分支在訓練後即被丟棄,這些增益不會產生任何推論開銷。廣泛的受控比較與表徵層級分析進一步闡明了生成引導訓練何時及為何有助於理解,並證實生成引導訓練作為增強多模態理解的實用途徑之可行性。
English
While Multimodal Large Language Models (MLLMs) have achieved remarkable progress, visual understanding and generation are typically treated as divergent objectives. Existing unified frameworks often rely on discrete visual tokenization or diffusion objectives whose generative targets differ from the continuous representations consumed by visual understanding models, making direct transfer to enhance existing pretrained MLLMs non-trivial. In this work, we present GAS, a generation-guided training framework that reinterprets visual generation as auxiliary supervision for representation learning. Concretely, GAS adapts Next Embedding Prediction (NEP) as a cross-modal generation paradigm within a decoupled Mixture-of-Transformers (MoT) architecture. By maintaining a shared lower trunk and parallel upper layers, GAS lets generation losses enrich the shared visual pathway with finer spatial precision and stronger visual retention while shielding the upper understanding layers from direct generation gradients. To maximize this synergy, we further construct highly correlated generation tasks that demand deep cognitive grounding rather than generic synthesis alone. Across model scales and training stages, GAS improves aggregate multimodal understanding, with its most reliable gains on perception and spatial comprehension. Crucially, because the auxiliary generation branch is discarded after training, these gains incur zero inference overhead. Extensive controlled comparisons and representation-level analyses further clarify when and why generation-guided training benefits understanding, and demonstrate the feasibility of generation-guided training as a practical route to stronger multimodal understanding.