ChatPaper.aiChatPaper

Nemotron-Labs-Diffusion : un modèle de langage tri-mode unifiant le décodage autorégressif, par diffusion et auto-spéculatif

Nemotron-Labs-Diffusion: A Tri-Mode Language Model Unifying Autoregressive, Diffusion, and Self-Speculation Decoding

July 7, 2026
Auteurs: Yonggan Fu, Lexington Whalen, Abhinav Garg, Chengyue Wu, Maksim Khadkevich, Nicolai Oswald, Enze Xie, Daniel Egert, Sharath Turuvekere Sreenivas, Shizhe Diao, Chenhan Yu, Ye Yu, Weijia Chen, Sajad Norouzi, Jingyu Liu, Shiyi Lan, Ligeng Zhu, Jin Wang, Jindong Jiang, Morteza Mardani, Mehran Maghoumi, Song Han, Ante Jukić, Nima Tajbakhsh, Jan Kautz, Pavlo Molchanov
cs.AI

Résumé

Nous présentons Nemotron-Labs-Diffusion, un modèle de langage (LM) à trois modes qui unifie le décodage AR, la diffusion et l’auto-spéculation au sein d’une seule architecture. Entraîné avec un objectif conjoint AR-diffusion, Nemotron-Labs-Diffusion peut changer de mode pour maintenir un débit élevé dans différents contextes de déploiement et niveaux de concurrence. Notre étude montre que (1) les objectifs AR et diffusion sont complémentaires : la diffusion améliore la planification anticipée, tandis que l’AR fournit des priors linguistiques de gauche à droite. (2) En mode auto-spéculation, la diffusion génère des ébauches tandis que l’AR les vérifie, surpassant les méthodes de prédiction multi-tokens (MTP) tant en taux d’acceptation qu’en efficacité réelle sur appareil. (3) Une analyse de vitesse limite révèle en outre le potentiel à long terme de la diffusion, avec jusqu’à 76,5 % de tokens supplémentaires par passage avant par rapport à l’auto-spéculation sous un échantillonneur optimal. En montant en échelle jusqu’à 3B, 8B et 14B paramètres, notre famille Nemotron-Labs-Diffusion, incluant les modèles de base, instruct et vision-langage, surpasse systématiquement les LM AR et de diffusion open source de pointe en précision comme en vitesse. Par exemple, Nemotron-Labs-Diffusion-8B décode 6 fois plus de tokens par passage avant que Qwen3-8B avec une précision comparable, ce qui se traduit par un débit 4 fois supérieur sur SPEED-Bench avec SGLang sur un GPU GB200.
English
We introduce Nemotron-Labs-Diffusion, a tri-mode language model (LM) that unifies AR, diffusion, and self-speculation decoding within a single architecture. Trained with a joint AR-diffusion objective, Nemotron-Labs-Diffusion can switch modes to sustain high throughput across deployment settings and concurrency levels. Our study shows that (1) AR and diffusion objectives are complementary: diffusion improves lookahead planning, while AR provides left-to-right linguistic priors. (2) In self-speculation mode, diffusion drafts while AR verifies, outperforming multi-token prediction (MTP) methods in both acceptance rate and real-device efficiency. (3) A speed-of-light analysis further demonstrates diffusion's long-term potential, with up to 76.5% more tokens per forward pass than self-speculation under an optimal sampler. Scaling to 3B, 8B, and 14B parameters, our Nemotron-Labs-Diffusion family, including base, instruct, and vision-language models, consistently outperforms state-of-the-art open-source AR and diffusion LMs in both accuracy and speed. For example, Nemotron-Labs-Diffusion-8B decodes 6x more tokens per forward than Qwen3-8B with comparable accuracy, translating to 4x higher throughput on SPEED-Bench with SGLang on a GB200 GPU.