ChatPaper.aiChatPaper

생성을 보조 감독으로 활용: 분리된 임베딩 예측을 통한 제로 추론 오버헤드 시각적 이해 향상

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction

August 12, 2026
저자: Zhongbin Guo, Jiahao Xie, Dongling Xiao, Qianle Wang, Ruiqi Lu, Xiaomin He, Wanxuan Sun, Cheng Yang
cs.AI

초록

다중모달 대규모 언어 모델(MLLM)이 놀라운 발전을 이루었지만, 시각 이해와 생성은 일반적으로 서로 다른 목표로 취급된다. 기존의 통합 프레임워크는 종종 이산 시각 토큰화나 확산 목표에 의존하는데, 이러한 생성 목표는 시각 이해 모델이 소비하는 연속 표현과 다르므로 사전학습된 MLLM을 향상시키기 위한 직접 전이를 어렵게 만든다. 본 연구에서는 시각 생성을 표현 학습을 위한 보조 감독으로 재해석하는 생성 유도 훈련 프레임워크인 GAS를 제시한다. 구체적으로, GAS는 분리된 Mixture-of-Transformers(MoT) 아키텍처 내에서 교차 모달 생성 패러다임으로서 다음 임베딩 예측(Next Embedding Prediction, NEP)을 적용한다. GAS는 공유 하부 트렁크와 병렬 상위 레이어를 유지하여, 생성 손실이 공유 시각 경로를 더 세밀한 공간 정밀도와 더 강한 시각 유지력으로 풍부하게 하면서도 상위 이해 레이어를 직접적인 생성 그라디언트로부터 보호한다. 이러한 시너지를 극대화하기 위해, 우리는 단순한 일반적 합성보다 깊은 인지적 근거를 요구하는 고도로 연관된 생성 작업을 추가로 구성한다. 모델 규모와 훈련 단계에 걸쳐 GAS는 종합적 다중모달 이해를 개선하며, 지각과 공간 이해에서 가장 안정적인 향상을 보인다. 중요한 점은 보조 생성 분기가 훈련 후 폐기되므로 이러한 성능 향상이 추론 오버헤드를 전혀 발생시키지 않는다는 것이다. 광범위한 통제 비교와 표현 수준 분석은 생성 유도 훈련이 이해에 이점을 주는 시기와 이유를 추가로 명확히 하며, 보다 강력한 다중모달 이해를 위한 실용적 경로로서의 실현 가능성을 입증한다.
English
While Multimodal Large Language Models (MLLMs) have achieved remarkable progress, visual understanding and generation are typically treated as divergent objectives. Existing unified frameworks often rely on discrete visual tokenization or diffusion objectives whose generative targets differ from the continuous representations consumed by visual understanding models, making direct transfer to enhance existing pretrained MLLMs non-trivial. In this work, we present GAS, a generation-guided training framework that reinterprets visual generation as auxiliary supervision for representation learning. Concretely, GAS adapts Next Embedding Prediction (NEP) as a cross-modal generation paradigm within a decoupled Mixture-of-Transformers (MoT) architecture. By maintaining a shared lower trunk and parallel upper layers, GAS lets generation losses enrich the shared visual pathway with finer spatial precision and stronger visual retention while shielding the upper understanding layers from direct generation gradients. To maximize this synergy, we further construct highly correlated generation tasks that demand deep cognitive grounding rather than generic synthesis alone. Across model scales and training stages, GAS improves aggregate multimodal understanding, with its most reliable gains on perception and spatial comprehension. Crucially, because the auxiliary generation branch is discarded after training, these gains incur zero inference overhead. Extensive controlled comparisons and representation-level analyses further clarify when and why generation-guided training benefits understanding, and demonstrate the feasibility of generation-guided training as a practical route to stronger multimodal understanding.