ChatPaper.aiChatPaper

生成作为辅助监督:通过解耦嵌入预测在零推理开销下增强视觉理解

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction

August 12, 2026
作者: Zhongbin Guo, Jiahao Xie, Dongling Xiao, Qianle Wang, Ruiqi Lu, Xiaomin He, Wanxuan Sun, Cheng Yang
cs.AI

摘要

尽管多模态大语言模型(MLLMs)已取得显著进展,视觉理解与视觉生成通常仍被视为相互分歧的目标。现有统一框架往往依赖离散视觉标记化或扩散目标,其生成目标与视觉理解模型所消耗的连续表示不同,使得直接迁移以增强现有预训练MLLMs并非易事。在本工作中,我们提出GAS,一种将视觉生成重新诠释为表征学习辅助监督的生成引导训练框架。具体而言,GAS在解耦的混合Transformer(MoT)架构中采用下一嵌入预测(NEP)作为跨模态生成范式。通过保持共享的下层主干与并行的上层分支,GAS使生成损失能够以更精细的空间精度和更强的视觉保持能力丰富共享视觉通路,同时保护上层理解分支免受直接生成梯度的干扰。为最大化这种协同效应,我们进一步构建了高度相关的生成任务,这些任务要求深层认知基础而不仅仅是通用合成。跨模型规模和训练阶段,GAS均能提升多模态综合理解能力,其在感知和空间理解方面的收益最为稳定。至关重要的是,由于辅助生成分支在训练后被丢弃,这些收益不产生任何推理开销。大规模受控比较和表征层面分析进一步阐明了生成引导训练何时以及为何有益于理解,并验证了生成引导训练作为实现更强多模态理解的实际路径的可行性。
English
While Multimodal Large Language Models (MLLMs) have achieved remarkable progress, visual understanding and generation are typically treated as divergent objectives. Existing unified frameworks often rely on discrete visual tokenization or diffusion objectives whose generative targets differ from the continuous representations consumed by visual understanding models, making direct transfer to enhance existing pretrained MLLMs non-trivial. In this work, we present GAS, a generation-guided training framework that reinterprets visual generation as auxiliary supervision for representation learning. Concretely, GAS adapts Next Embedding Prediction (NEP) as a cross-modal generation paradigm within a decoupled Mixture-of-Transformers (MoT) architecture. By maintaining a shared lower trunk and parallel upper layers, GAS lets generation losses enrich the shared visual pathway with finer spatial precision and stronger visual retention while shielding the upper understanding layers from direct generation gradients. To maximize this synergy, we further construct highly correlated generation tasks that demand deep cognitive grounding rather than generic synthesis alone. Across model scales and training stages, GAS improves aggregate multimodal understanding, with its most reliable gains on perception and spatial comprehension. Crucially, because the auxiliary generation branch is discarded after training, these gains incur zero inference overhead. Extensive controlled comparisons and representation-level analyses further clarify when and why generation-guided training benefits understanding, and demonstrate the feasibility of generation-guided training as a practical route to stronger multimodal understanding.