DiffusionGemma 기술 보고서
DiffusionGemma Technical Report
July 31, 2026
저자: DiffusionGemma Team, Adrien Ali Taïga, James Assiene, Daniele Calandriello, Rahma Chaabouni, João Gante, Tamara von Glehn, Nate Keating, Chris Knutsen, Martin Kukla, Tianlin Liu, Ivan Lobov, Ofir Nabati, João Gabriel Oliveira, Nicolas Perez-Nieves, Nastasia Prutianova, Bobak Shahriari, Jean Tarbouriech, Pavel Tyletski, Çağlar Ünlü, Cindy Wu, Glenn Cameron, Jerome Connor, Sertan Girgin, Maarten Grootendorst, Alon Levkovitch, Eliya Nachmani, Omar Sanseviero, Piotr Stanczyk, Quentin Berthet, Andrew Campbell, Clément Crepy, Valentin De Bortoli, Arnaud Doucet, Romuald Elie, Alexandre Galashov, Klaus Greff, Alexis Jacq, David Ruhe, Yu-Han Wu, Sebastian Flennerhag, Brendan O'Donoghue, George Scrivener, Shantanu Thakoor
cs.AI
초록
우리는 DiffusionGemma를 소개한다. 이는 이산 확산(discrete diffusion)을 사용하여 비할 데 없이 빠른 속도로 텍스트를 생성하는 실험적 오픈 가중치 언어 모델이다. 한 번에 하나의 토큰을 디코딩하는 대신, DiffusionGemma는 256개 토큰 블록을 병렬로 반복적으로 정제하여 기존 자기회귀(AR) 대규모 언어 모델의 순차적 디코딩 병목 현상을 피한다. DiffusionGemma는 처음부터 훈련하는 대신, 활성 매개변수 38억 개, 총 매개변수 252억 개를 가진 혼합 전문가(MoE) Gemma 4 모델을 미세 조정하여 얻었다. 우리의 계산 효율적인 2단계 훈련 파이프라인은 초기 AR 모델의 총 훈련 토큰 예산의 10% 미만을 사용한다. 첫 번째 단계는 지도 미세 조정을 통해 양방향 노이즈 제거를 학습시키고, 두 번째 단계는 강화 학습과 샘플러 증류(sampler distillation)를 결합하여 생성 품질과 추론 효율성을 동시에 개선한다. DiffusionGemma는 생성 속도와 모델 성능 간의 절충에 대한 새로운 파레토 경계를 구축한다. 전체 평가 스위트를 평균한 결과, 순전파당 약 20개의 토큰을 생성하고 단일 NVIDIA H100 GPU에서 초당 약 1,500개의 출력 토큰을 달성하는데, 이는 최첨단 투기적 디코딩을 사용하는 AR 모델보다 훨씬 빠른 속도이다. DiffusionGemma는 또한 초기 모델의 사고 모드, 다중 모달 입력, 긴 컨텍스트 지원을 유지한다. 확산 미세 조정에도 불구하고, 약간의 성능 저하만으로 AR 생성이 가능하여 하이브리드 확산-AR 디코딩으로 가는 길을 시사한다.
English
We introduce DiffusionGemma, an experimental open-weight language model that uses discrete diffusion to generate text at exceptionally high speed. Rather than decoding one token at a time, DiffusionGemma iteratively refines blocks of 256 tokens in parallel, avoiding the sequential decoding bottleneck of conventional autoregressive (AR) large language models. Instead of training from scratch, we obtain DiffusionGemma by fine-tuning the mixture-of-experts Gemma 4 model with 3.8B activated and 25.2B total parameters. Our compute-efficient two-stage training pipeline uses fewer than 10% of the starting AR model's total training token budget. The first stage uses supervised fine-tuning to teach bidirectional denoising, while the second stage combines reinforcement learning with sampler distillation to jointly improve generation quality and inference efficiency. DiffusionGemma establishes a new Pareto frontier for the trade-off between generation speed and model capability. Averaged across our full evaluation suite, it generates around 20 tokens per forward pass and achieves roughly 1,500 output tokens per second on a single NVIDIA H100 GPU, which is substantially faster than AR models even with state-of-the-art speculative decoding. DiffusionGemma also retains the starting model's support for thinking mode, multimodal inputs, and long contexts. Despite diffusion fine-tuning, it remains capable of AR generation with only minor performance degradation, suggesting a path toward hybrid diffusion-AR decoding.