ChatPaper.aiChatPaper

Nemotron-Labs-Diffusion: 自己回帰、拡散、自己推測デコードを統合した三モード言語モデル

Nemotron-Labs-Diffusion: A Tri-Mode Language Model Unifying Autoregressive, Diffusion, and Self-Speculation Decoding

July 7, 2026
著者: Yonggan Fu, Lexington Whalen, Abhinav Garg, Chengyue Wu, Maksim Khadkevich, Nicolai Oswald, Enze Xie, Daniel Egert, Sharath Turuvekere Sreenivas, Shizhe Diao, Chenhan Yu, Ye Yu, Weijia Chen, Sajad Norouzi, Jingyu Liu, Shiyi Lan, Ligeng Zhu, Jin Wang, Jindong Jiang, Morteza Mardani, Mehran Maghoumi, Song Han, Ante Jukić, Nima Tajbakhsh, Jan Kautz, Pavlo Molchanov
cs.AI

要旨

当社は、AR(自己回帰)、拡散モデル、自己投機的デコードを単一のアーキテクチャで統合した3モード言語モデル「Nemotron-Labs-Diffusion」を発表します。ARと拡散の統合目的関数で学習されたNemotron-Labs-Diffusionは、デプロイ環境や同時実行レベルに応じてモードを切り替え、高いスループットを維持できます。本研究では以下の知見を得ました。(1)ARと拡散の目的関数は相補的であり、拡散は先読み計画を改善し、ARは左から右への言語的先行知識を提供します。(2)自己投機モードでは、拡散がドラフト(下書き)を生成しARが検証を行う方式が、マルチトークン予測(MTP)手法よりも受容率と実デバイス効率の両方で優れています。(3)限界速度解析により、拡散の長期的な可能性がさらに実証され、最適サンプラー使用時には自己投機よりもフォワードパスあたり最大76.5%多いトークンを生成可能です。パラメータ規模を3B、8B、14Bにスケーリングした当社のNemotron-Labs-Diffusionファミリー(ベースモデル、指示モデル、視覚言語モデルを含む)は、精度と速度の両方で最先端のオープンソースARおよび拡散言語モデルを一貫して上回ります。例えば、Nemotron-Labs-Diffusion-8Bは、Qwen3-8Bと同等の精度でフォワードパスあたり6倍多くのトークンをデコードし、GB200 GPU上のSGLangを用いたSPEED-Benchでは4倍のスループット向上を実現します。
English
We introduce Nemotron-Labs-Diffusion, a tri-mode language model (LM) that unifies AR, diffusion, and self-speculation decoding within a single architecture. Trained with a joint AR-diffusion objective, Nemotron-Labs-Diffusion can switch modes to sustain high throughput across deployment settings and concurrency levels. Our study shows that (1) AR and diffusion objectives are complementary: diffusion improves lookahead planning, while AR provides left-to-right linguistic priors. (2) In self-speculation mode, diffusion drafts while AR verifies, outperforming multi-token prediction (MTP) methods in both acceptance rate and real-device efficiency. (3) A speed-of-light analysis further demonstrates diffusion's long-term potential, with up to 76.5% more tokens per forward pass than self-speculation under an optimal sampler. Scaling to 3B, 8B, and 14B parameters, our Nemotron-Labs-Diffusion family, including base, instruct, and vision-language models, consistently outperforms state-of-the-art open-source AR and diffusion LMs in both accuracy and speed. For example, Nemotron-Labs-Diffusion-8B decodes 6x more tokens per forward than Qwen3-8B with comparable accuracy, translating to 4x higher throughput on SPEED-Bench with SGLang on a GB200 GPU.