ChatPaper.aiChatPaper

DiffusionGemma 技術報告

DiffusionGemma Technical Report

July 31, 2026
作者: DiffusionGemma Team, Adrien Ali Taïga, James Assiene, Daniele Calandriello, Rahma Chaabouni, João Gante, Tamara von Glehn, Nate Keating, Chris Knutsen, Martin Kukla, Tianlin Liu, Ivan Lobov, Ofir Nabati, João Gabriel Oliveira, Nicolas Perez-Nieves, Nastasia Prutianova, Bobak Shahriari, Jean Tarbouriech, Pavel Tyletski, Çağlar Ünlü, Cindy Wu, Glenn Cameron, Jerome Connor, Sertan Girgin, Maarten Grootendorst, Alon Levkovitch, Eliya Nachmani, Omar Sanseviero, Piotr Stanczyk, Quentin Berthet, Andrew Campbell, Clément Crepy, Valentin De Bortoli, Arnaud Doucet, Romuald Elie, Alexandre Galashov, Klaus Greff, Alexis Jacq, David Ruhe, Yu-Han Wu, Sebastian Flennerhag, Brendan O'Donoghue, George Scrivener, Shantanu Thakoor
cs.AI

摘要

我們推出 DiffusionGemma,這是一個實驗性的開放權重語言模型,利用離散擴散以極高速度生成文字。DiffusionGemma 並非逐個 token 解碼,而是並行地迭代精煉包含 256 個 token 的區塊,避開了傳統自迴歸(AR)大型語言模型循序解碼的瓶頸。我們並非從零開始訓練,而是透過微調 Gemma 4 混合專家模型取得 DiffusionGemma,該模型啟用 38 億個參數,總參數達 252 億。我們計算效率高的兩階段訓練流程,僅使用了原始 AR 模型總訓練 token 預算的不到 10%。第一階段使用監督式微調來教導雙向去噪,第二階段則結合強化學習與取樣器蒸餾,共同提升生成品質與推論效率。DiffusionGemma 在生成速度與模型能力之間的取捨上,建立了新的帕累托前沿。在我們完整的評估套件中平均而言,它每次前向傳播約生成 20 個 token,並在單張 NVIDIA H100 GPU 上達到每秒約 1,500 個輸出 token,即使與採用最新推論加速技術的自迴歸模型相比,速度也大幅領先。DiffusionGemma 也保留了原始模型對思考模式、多模態輸入與長上下文的支援。儘管經過擴散微調,它仍能以僅有輕微效能衰減的方式進行自迴歸生成,這暗示了通往混合擴散-自迴歸解碼的途徑。
English
We introduce DiffusionGemma, an experimental open-weight language model that uses discrete diffusion to generate text at exceptionally high speed. Rather than decoding one token at a time, DiffusionGemma iteratively refines blocks of 256 tokens in parallel, avoiding the sequential decoding bottleneck of conventional autoregressive (AR) large language models. Instead of training from scratch, we obtain DiffusionGemma by fine-tuning the mixture-of-experts Gemma 4 model with 3.8B activated and 25.2B total parameters. Our compute-efficient two-stage training pipeline uses fewer than 10% of the starting AR model's total training token budget. The first stage uses supervised fine-tuning to teach bidirectional denoising, while the second stage combines reinforcement learning with sampler distillation to jointly improve generation quality and inference efficiency. DiffusionGemma establishes a new Pareto frontier for the trade-off between generation speed and model capability. Averaged across our full evaluation suite, it generates around 20 tokens per forward pass and achieves roughly 1,500 output tokens per second on a single NVIDIA H100 GPU, which is substantially faster than AR models even with state-of-the-art speculative decoding. DiffusionGemma also retains the starting model's support for thinking mode, multimodal inputs, and long contexts. Despite diffusion fine-tuning, it remains capable of AR generation with only minor performance degradation, suggesting a path toward hybrid diffusion-AR decoding.