ChatPaper.aiChatPaper

DiffusionGemma 技術レポート

DiffusionGemma Technical Report

July 31, 2026
著者: DiffusionGemma Team, Adrien Ali Taïga, James Assiene, Daniele Calandriello, Rahma Chaabouni, João Gante, Tamara von Glehn, Nate Keating, Chris Knutsen, Martin Kukla, Tianlin Liu, Ivan Lobov, Ofir Nabati, João Gabriel Oliveira, Nicolas Perez-Nieves, Nastasia Prutianova, Bobak Shahriari, Jean Tarbouriech, Pavel Tyletski, Çağlar Ünlü, Cindy Wu, Glenn Cameron, Jerome Connor, Sertan Girgin, Maarten Grootendorst, Alon Levkovitch, Eliya Nachmani, Omar Sanseviero, Piotr Stanczyk, Quentin Berthet, Andrew Campbell, Clément Crepy, Valentin De Bortoli, Arnaud Doucet, Romuald Elie, Alexandre Galashov, Klaus Greff, Alexis Jacq, David Ruhe, Yu-Han Wu, Sebastian Flennerhag, Brendan O'Donoghue, George Scrivener, Shantanu Thakoor
cs.AI

要旨

DiffusionGemmaを紹介する。これは、離散拡散を用いてテキストを極めて高速に生成する実験的なオープンウェイト言語モデルである。DiffusionGemmaは一度に1トークンずつデコードするのではなく、256トークンのブロックを並列に反復的に精緻化することで、従来の自己回帰(AR)大規模言語モデルにおける逐次デコードのボトルネックを回避する。DiffusionGemmaはゼロから学習するのではなく、活性化3.8B・総25.2Bパラメータを備えた混合専門家(MoE)モデルであるGemma 4をファインチューニングすることで得られる。計算効率の高い2段階トレーニングパイプラインは、元のARモデルの総トレーニングトークン予算の10%未満しか使用しない。第1段階では教師ありファインチューニングにより双方向ノイズ除去を学習し、第2段階では強化学習とサンプラー蒸留を組み合わせることで、生成品質と推論効率を同時に向上させる。DiffusionGemmaは、生成速度とモデル性能のトレードオフにおける新たなパレート最前線を確立する。全評価スイートにわたる平均では、フォワードパスあたり約20トークンを生成し、単一のNVIDIA H100 GPU上で毎秒約1,500出力トークンを達成する。これは、最先端の投機的デコードを用いたARモデルよりも大幅に高速である。また、DiffusionGemmaは元のモデルの思考モード、マルチモーダル入力、長いコンテキストへの対応も維持している。拡散ファインチューニングにもかかわらず、わずかな性能低下のみでAR生成も可能であり、ハイブリッド拡散-ARデコードへの道筋が示唆される。
English
We introduce DiffusionGemma, an experimental open-weight language model that uses discrete diffusion to generate text at exceptionally high speed. Rather than decoding one token at a time, DiffusionGemma iteratively refines blocks of 256 tokens in parallel, avoiding the sequential decoding bottleneck of conventional autoregressive (AR) large language models. Instead of training from scratch, we obtain DiffusionGemma by fine-tuning the mixture-of-experts Gemma 4 model with 3.8B activated and 25.2B total parameters. Our compute-efficient two-stage training pipeline uses fewer than 10% of the starting AR model's total training token budget. The first stage uses supervised fine-tuning to teach bidirectional denoising, while the second stage combines reinforcement learning with sampler distillation to jointly improve generation quality and inference efficiency. DiffusionGemma establishes a new Pareto frontier for the trade-off between generation speed and model capability. Averaged across our full evaluation suite, it generates around 20 tokens per forward pass and achieves roughly 1,500 output tokens per second on a single NVIDIA H100 GPU, which is substantially faster than AR models even with state-of-the-art speculative decoding. DiffusionGemma also retains the starting model's support for thinking mode, multimodal inputs, and long contexts. Despite diffusion fine-tuning, it remains capable of AR generation with only minor performance degradation, suggesting a path toward hybrid diffusion-AR decoding.