ChatPaper.aiChatPaper

DiffusionGemma 技术报告

DiffusionGemma Technical Report

July 31, 2026
作者: DiffusionGemma Team, Adrien Ali Taïga, James Assiene, Daniele Calandriello, Rahma Chaabouni, João Gante, Tamara von Glehn, Nate Keating, Chris Knutsen, Martin Kukla, Tianlin Liu, Ivan Lobov, Ofir Nabati, João Gabriel Oliveira, Nicolas Perez-Nieves, Nastasia Prutianova, Bobak Shahriari, Jean Tarbouriech, Pavel Tyletski, Çağlar Ünlü, Cindy Wu, Glenn Cameron, Jerome Connor, Sertan Girgin, Maarten Grootendorst, Alon Levkovitch, Eliya Nachmani, Omar Sanseviero, Piotr Stanczyk, Quentin Berthet, Andrew Campbell, Clément Crepy, Valentin De Bortoli, Arnaud Doucet, Romuald Elie, Alexandre Galashov, Klaus Greff, Alexis Jacq, David Ruhe, Yu-Han Wu, Sebastian Flennerhag, Brendan O'Donoghue, George Scrivener, Shantanu Thakoor
cs.AI

摘要

我们推出 DiffusionGemma,一个实验性的开放权重语言模型,使用离散扩散以极高速度生成文本。与逐 token 解码不同,DiffusionGemma 并行迭代精炼 256 个 token 的块,从而绕过了传统自回归(AR)大语言模型顺序解码的瓶颈。我们并非从零训练,而是通过对 mixture-of-experts 架构的 Gemma 4 模型进行微调得到 DiffusionGemma,其激活参数为 38 亿,总参数为 252 亿。我们计算高效的两阶段训练流程仅使用了初始自回归模型总训练 token 预算的不到 10%。第一阶段使用监督微调来教授双向去噪,第二阶段将强化学习与采样器蒸馏相结合,共同提升生成质量和推理效率。DiffusionGemma 在生成速度与模型能力之间的权衡中建立了新的 Pareto 前沿。在我们的完整评测套件上平均来看,它每次前向传播约生成 20 个 token,并在单个 NVIDIA H100 GPU 上达到每秒约 1,500 个输出 token,这显著快于即使采用了最先进推测解码的自回归模型。DiffusionGemma 还保留了初始模型对思考模式、多模态输入和长上下文的支持。尽管经过了扩散微调,它仍能以仅有轻微性能下降的方式进行自回归生成,这为混合扩散-自回归解码指明了一条路径。
English
We introduce DiffusionGemma, an experimental open-weight language model that uses discrete diffusion to generate text at exceptionally high speed. Rather than decoding one token at a time, DiffusionGemma iteratively refines blocks of 256 tokens in parallel, avoiding the sequential decoding bottleneck of conventional autoregressive (AR) large language models. Instead of training from scratch, we obtain DiffusionGemma by fine-tuning the mixture-of-experts Gemma 4 model with 3.8B activated and 25.2B total parameters. Our compute-efficient two-stage training pipeline uses fewer than 10% of the starting AR model's total training token budget. The first stage uses supervised fine-tuning to teach bidirectional denoising, while the second stage combines reinforcement learning with sampler distillation to jointly improve generation quality and inference efficiency. DiffusionGemma establishes a new Pareto frontier for the trade-off between generation speed and model capability. Averaged across our full evaluation suite, it generates around 20 tokens per forward pass and achieves roughly 1,500 output tokens per second on a single NVIDIA H100 GPU, which is substantially faster than AR models even with state-of-the-art speculative decoding. DiffusionGemma also retains the starting model's support for thinking mode, multimodal inputs, and long contexts. Despite diffusion fine-tuning, it remains capable of AR generation with only minor performance degradation, suggesting a path toward hybrid diffusion-AR decoding.