ChatPaper.aiChatPaper

Nemotron-Labs-Diffusion:一個統一自迴歸、擴散與自推測解碼的三模態語言模型

Nemotron-Labs-Diffusion: A Tri-Mode Language Model Unifying Autoregressive, Diffusion, and Self-Speculation Decoding

July 7, 2026
作者: Yonggan Fu, Lexington Whalen, Abhinav Garg, Chengyue Wu, Maksim Khadkevich, Nicolai Oswald, Enze Xie, Daniel Egert, Sharath Turuvekere Sreenivas, Shizhe Diao, Chenhan Yu, Ye Yu, Weijia Chen, Sajad Norouzi, Jingyu Liu, Shiyi Lan, Ligeng Zhu, Jin Wang, Jindong Jiang, Morteza Mardani, Mehran Maghoumi, Song Han, Ante Jukić, Nima Tajbakhsh, Jan Kautz, Pavlo Molchanov
cs.AI

摘要

我們推出了 Nemotron-Labs-Diffusion,這是一個三模式語言模型(LM),可將自迴歸(AR)、擴散模型與自推測解碼整合在單一架構中。透過聯合 AR-擴散的訓練目標,Nemotron-Labs-Diffusion 能動態切換模式,在不同部署環境與並行程度下維持高吞吐量。我們的研究顯示:(1) AR 與擴散目標具有互補性:擴散改進了前瞻規劃能力,而 AR 則提供了從左至右的語言先驗。(2) 在自推測模式下,擴散模型負責生成草稿,AR 負責驗證,其接受率與實際設備效率均優於多令牌預測(MTP)方法。(3) 從光速分析進一步證明了擴散模型的長期潛力,在最佳取樣器下,每次前向傳遞可較自推測模式多產生 76.5% 的令牌。將模型規模擴展至 3B、8B 與 14B 參數後,我們的 Nemotron-Labs-Diffusion 系列(包含基礎版、指令版及視覺語言模型)在準確度與速度上均持續優於最先進的開源 AR 與擴散語言模型。例如,Nemotron-Labs-Diffusion-8B 每次前向傳遞可解碼的令牌數量為 Qwen3-8B 的 6 倍,同時在 GB200 GPU 上搭配 SGLang 於 SPEED-Bench 基準測試中,吞吐量更提升達 4 倍。
English
We introduce Nemotron-Labs-Diffusion, a tri-mode language model (LM) that unifies AR, diffusion, and self-speculation decoding within a single architecture. Trained with a joint AR-diffusion objective, Nemotron-Labs-Diffusion can switch modes to sustain high throughput across deployment settings and concurrency levels. Our study shows that (1) AR and diffusion objectives are complementary: diffusion improves lookahead planning, while AR provides left-to-right linguistic priors. (2) In self-speculation mode, diffusion drafts while AR verifies, outperforming multi-token prediction (MTP) methods in both acceptance rate and real-device efficiency. (3) A speed-of-light analysis further demonstrates diffusion's long-term potential, with up to 76.5% more tokens per forward pass than self-speculation under an optimal sampler. Scaling to 3B, 8B, and 14B parameters, our Nemotron-Labs-Diffusion family, including base, instruct, and vision-language models, consistently outperforms state-of-the-art open-source AR and diffusion LMs in both accuracy and speed. For example, Nemotron-Labs-Diffusion-8B decodes 6x more tokens per forward than Qwen3-8B with comparable accuracy, translating to 4x higher throughput on SPEED-Bench with SGLang on a GB200 GPU.