Nemotron-Labs-Diffusion: Een drie-modus taalmodel dat autoregressieve, diffusie- en zelfspeculatie-decodering verenigt
Nemotron-Labs-Diffusion: A Tri-Mode Language Model Unifying Autoregressive, Diffusion, and Self-Speculation Decoding
July 7, 2026
Auteurs: Yonggan Fu, Lexington Whalen, Abhinav Garg, Chengyue Wu, Maksim Khadkevich, Nicolai Oswald, Enze Xie, Daniel Egert, Sharath Turuvekere Sreenivas, Shizhe Diao, Chenhan Yu, Ye Yu, Weijia Chen, Sajad Norouzi, Jingyu Liu, Shiyi Lan, Ligeng Zhu, Jin Wang, Jindong Jiang, Morteza Mardani, Mehran Maghoumi, Song Han, Ante Jukić, Nima Tajbakhsh, Jan Kautz, Pavlo Molchanov
cs.AI
Samenvatting
We introduceren Nemotron-Labs-Diffusion, een driemodus-taalmodel (TM) dat AR, diffusie en zelfspeculatie-decodering verenigt in één architectuur. Getraind met een gezamenlijke AR-diffusiedoelstelling, kan Nemotron-Labs-Diffusion van modus wisselen om een hoge doorvoer te handhaven in verschillende implementatieomgevingen en concurrrentieniveaus. Ons onderzoek toont aan dat (1) AR- en diffusiedoelstellingen complementair zijn: diffusie verbetert de vooruitkijkplanning, terwijl AR links-naar-rechts taalkundige priori's biedt. (2) In de zelfspeculatiemodus stelt diffusie concepten op terwijl AR verifieert, wat zowel in acceptatiegraad als in praktische efficiëntie op apparaten beter presteert dan methoden voor meerdere-tokenvoorspelling (MTP). (3) Een lichtsnelheidsanalyse toont verder het langetermijnpotentieel van diffusie aan, met tot 76,5% meer tokens per forwardpass dan zelfspeculatie onder een optimale sampler. Bij opschaling naar 3B, 8B en 14B parameters presteert onze Nemotron-Labs-Diffusion-familie, inclusief basis-, instructie- en visie-taalmodellen, consistent beter dan de modernste open-source AR- en diffusie-TM's in zowel nauwkeurigheid als snelheid. Zo decodeert Nemotron-Labs-Diffusion-8B 6x meer tokens per forwardpass dan Qwen3-8B met vergelijkbare nauwkeurigheid, wat resulteert in een 4x hogere doorvoer op SPEED-Bench met SGLang op een GB200-GPU.
English
We introduce Nemotron-Labs-Diffusion, a tri-mode language model (LM) that unifies AR, diffusion, and self-speculation decoding within a single architecture. Trained with a joint AR-diffusion objective, Nemotron-Labs-Diffusion can switch modes to sustain high throughput across deployment settings and concurrency levels. Our study shows that (1) AR and diffusion objectives are complementary: diffusion improves lookahead planning, while AR provides left-to-right linguistic priors. (2) In self-speculation mode, diffusion drafts while AR verifies, outperforming multi-token prediction (MTP) methods in both acceptance rate and real-device efficiency. (3) A speed-of-light analysis further demonstrates diffusion's long-term potential, with up to 76.5% more tokens per forward pass than self-speculation under an optimal sampler. Scaling to 3B, 8B, and 14B parameters, our Nemotron-Labs-Diffusion family, including base, instruct, and vision-language models, consistently outperforms state-of-the-art open-source AR and diffusion LMs in both accuracy and speed. For example, Nemotron-Labs-Diffusion-8B decodes 6x more tokens per forward than Qwen3-8B with comparable accuracy, translating to 4x higher throughput on SPEED-Bench with SGLang on a GB200 GPU.