ChatPaper.aiChatPaper

Nemotron-Labs-Diffusion: 자동회귀, 확산 및 자기추측 디코딩을 통합하는 삼중 모드 언어 모델

Nemotron-Labs-Diffusion: A Tri-Mode Language Model Unifying Autoregressive, Diffusion, and Self-Speculation Decoding

July 7, 2026
저자: Yonggan Fu, Lexington Whalen, Abhinav Garg, Chengyue Wu, Maksim Khadkevich, Nicolai Oswald, Enze Xie, Daniel Egert, Sharath Turuvekere Sreenivas, Shizhe Diao, Chenhan Yu, Ye Yu, Weijia Chen, Sajad Norouzi, Jingyu Liu, Shiyi Lan, Ligeng Zhu, Jin Wang, Jindong Jiang, Morteza Mardani, Mehran Maghoumi, Song Han, Ante Jukić, Nima Tajbakhsh, Jan Kautz, Pavlo Molchanov
cs.AI

초록

여기 Nemotron-Labs-Diffusion, 즉 단일 아키텍처 내에서 자기회귀(AR), 확산(diffusion), 자기 추측 디코딩(self-speculation decoding)을 통합한 삼중 모드 언어 모델을 소개한다. 공동 자기회귀-확산 목적 함수로 학습된 Nemotron-Labs-Diffusion은 배포 환경과 동시성 수준에 따라 모드를 전환하여 높은 처리량을 유지할 수 있다. 본 연구는 다음을 보여준다. (1) 자기회귀와 확산 목적 함수는 상호 보완적이다. 확산은 선견 계획을 개선하고, 자기회귀는 좌측에서 우측으로의 언어적 사전 정보를 제공한다. (2) 자기 추측 모드에서 확산이 초안을 작성하고 자기회귀가 검증할 때, 다중 토큰 예측(MTP) 방법보다 수용률과 실제 장치 효율성 모두에서 우수한 성능을 보인다. (3) 광속 분석은 확산의 장기적 잠재력을 추가로 입증하며, 최적 샘플러 사용 시 순방향 전달당 자기 추측보다 최대 76.5% 더 많은 토큰을 생성한다. 3B, 8B, 14B 파라미터로 확장한 당사의 Nemotron-Labs-Diffusion 제품군(기본 모델, 명령어 모델, 시각-언어 모델 포함)은 정확도와 속도 모두에서 최첨단 오픈소스 자기회귀 및 확산 언어 모델을 지속적으로 능가한다. 예를 들어, Nemotron-Labs-Diffusion-8B는 Qwen3-8B와 비슷한 정확도로 순방향 전달당 6배 더 많은 토큰을 디코딩하며, GB200 GPU에서 SGLang을 사용한 SPEED-Bench에서 4배 더 높은 처리량으로 이어진다.
English
We introduce Nemotron-Labs-Diffusion, a tri-mode language model (LM) that unifies AR, diffusion, and self-speculation decoding within a single architecture. Trained with a joint AR-diffusion objective, Nemotron-Labs-Diffusion can switch modes to sustain high throughput across deployment settings and concurrency levels. Our study shows that (1) AR and diffusion objectives are complementary: diffusion improves lookahead planning, while AR provides left-to-right linguistic priors. (2) In self-speculation mode, diffusion drafts while AR verifies, outperforming multi-token prediction (MTP) methods in both acceptance rate and real-device efficiency. (3) A speed-of-light analysis further demonstrates diffusion's long-term potential, with up to 76.5% more tokens per forward pass than self-speculation under an optimal sampler. Scaling to 3B, 8B, and 14B parameters, our Nemotron-Labs-Diffusion family, including base, instruct, and vision-language models, consistently outperforms state-of-the-art open-source AR and diffusion LMs in both accuracy and speed. For example, Nemotron-Labs-Diffusion-8B decodes 6x more tokens per forward than Qwen3-8B with comparable accuracy, translating to 4x higher throughput on SPEED-Bench with SGLang on a GB200 GPU.