DiffusionGemma Technical Report
DiffusionGemma Team, Adrien Ali Taïga, James Assiene, Daniele Calandriello, Rahma Chaabouni, João Gante, Tamara von Glehn, Nate Keating, Chris Knutsen, Martin Kukla, Tianlin Liu, Ivan Lobov, Ofir Nabati, João Gabriel Oliveira, Nicolas Perez-Nieves, Nastasia Prutianova, Bobak Shahriari, Jean Tarbouriech, Pavel Tyletski, Çağlar Ünlü, Cindy Wu, Glenn Cameron, Jerome Connor, Sertan Girgin, Maarten Grootendorst, Alon Levkovitch, Eliya Nachmani, Omar Sanseviero, Piotr Stanczyk, Quentin Berthet, Andrew Campbell, Clément Crepy, Valentin De Bortoli, Arnaud Doucet, Romuald Elie, Alexandre Galashov, Klaus Greff, Alexis Jacq, David Ruhe, Yu-Han Wu, Sebastian Flennerhag, Brendan O'Donoghue, George Scrivener, Shantanu Thakoor
arXiv:2608.00146 · 2026-08-04 공개 · arXiv · PDF
reinforcement-learning diffusion-models long-context language-models mixture-of-experts speculative-decoding denoising gemma
Abstract
We introduce DiffusionGemma, an experimental open-weight language model that uses discrete diffusion to generate text at exceptionally high speed. Rather than decoding one token at a time, DiffusionGemma iteratively refines blocks of 256 tokens in parallel, avoiding the sequential decoding bottleneck of conventional autoregressive (AR) large language models. Instead of training from scratch, we obtain DiffusionGemma by fine-tuning the mixture-of-experts Gemma 4 model with 3.8B activated and 25.2B total parameters. Our compute-efficient two-stage training pipeline uses fewer than 10% of the starting AR model's total training token budget. The first stage uses supervised fine-tuning to teach bidirectional denoising, while the second stage combines reinforcement learning with sampler distillation to jointly improve generation quality and inference efficiency. DiffusionGemma establishes a new Pareto frontier for the trade-off between generation speed and model capability. Averaged across our full evaluation suite, it generates around 20 tokens per forward pass and achieves roughly 1,500 output tokens per second on a single NVIDIA H100 GPU, which is substantially faster than AR models even with state-of-the-art speculative decoding. DiffusionGemma also retains the starting model's support for thinking mode, multimodal inputs, and long contexts. Despite diffusion fine-tuning, it remains capable of AR generation with only minor performance degradation, suggesting a path toward hybrid diffusion-AR decoding.
한국어 요약
한 줄 요약
DiffusionGemma는 고속 텍스트 생성을 위한 이산 확산 기반 언어 모델로, 기존 오토회귀 모델 대비 1,500 토큰/초의 속도를 달성한다.
핵심 기여도
- 256 토큰 블록을 병렬로 반복적으로 정제하는 방식으로 순차 디코딩 병목을 회피.
- Gemma 4 모델(3.8B 활성화, 25.2B 전체 파라미터)을 미세 조정하여 구축.
- 2단계 학습 파이프라인으로, 전체 토큰 예산의 10% 미만으로 훈련.
- 1,500 토큰/초의 생성 속도를 NVIDIA H100 GPU에서 달성.
핵심 아이디어
DiffusionGemma는 기존 오토회귀 언어 모델이 토큰을 하나씩 생성하는 방식 대신, 256개 토큰 블록을 동시에 정제하는 이산 확산 기법을 도입한다. 이는 병렬 처리를 통해 생성 속도를 극대화하는 동시에, 품질 저하를 최소화하는 것이 핵심이다. 두 번째 핵심 아이디어는 기존 Gemma 4 모델을 기반으로 하여, 훈련 데이터 사용량을 10% 미만으로 줄인 효율적인 2단계 학습 파이프라인을 설계한 점이다. 이는 지도 학습과 강화 학습, 샘플러 디스틸레이션을 결합한 방식으로, 생성 품질과 추론 효율성을 동시에 개선한다.
기술적 접근법
주요 결과
- 평균적으로 20 토큰/포워드 패스 생성.
- NVIDIA H100 GPU에서 1,500 토큰/초 생성 속도 달성.
- 기존 최신 추론 최적화 오토회귀 모델 대비 훨씬 빠른 속도.
- AR 생성 능력 유지 (성능 저하 최소).
의의 및 한계
DiffusionGemma는 생성 속도와 모델 능력 간의 트레이드오프에서 새로운 최전선(Pareto frontier)을 제시한다. 특히, 기존 오토회귀 모델의 디코딩 병목을 회피하면서도, 생각 모드, 멀티모달 입력, 긴 컨텍스트 처리를 유지하는 점에서 실용적 가치가 높다. 다만, 확산 기반 디코딩이 모든 사용 사례에 적합하지 않을 수 있으며, AR 모드로의 전환 시 일부 성능 저하가 발생할 수 있다는 한계가 있다.
실용적 활용
DiffusionGemma는 대규모 텍스트 생성이 필요한 실시간 응용(예: 챗봇, 자동 번역, 콘텐츠 생성)에 유용하게 활용될 수 있다. 또한, 멀티모달 및 긴 컨텍스트 처리가 필요한 연구 분야에서도 적용 가능하다.