DiffusionGemma Technical Report

Aug 20, 2026 08:24 PM - 1 day ago 3

[Submitted connected 31 Jul 2026]

Authors:DiffusionGemma Team: Adrien Ali Taïga, James Assiene, Daniele Calandriello, Rahma Chaabouni, João Gante, Tamara von Glehn, Nate Keating, Chris Knutsen, Martin Kukla, Tianlin Liu, Ivan Lobov, Ofir Nabati, João Gabriel Oliveira, Nicolas Perez-Nieves, Nastasia Prutianova, Bobak Shahriari, Jean Tarbouriech, Pavel Tyletski, Çağlar Ünlü, Cindy Wu, Glenn Cameron, Jerome Connor, Sertan Girgin, Maarten Grootendorst, Alon Levkovitch, Eliya Nachmani, Omar Sanseviero, Piotr Stanczyk, Quentin Berthet, Andrew Campbell, Clément Crepy, Valentin De Bortoli, Arnaud Doucet, Romuald Elie, Alexandre Galashov, Klaus Greff, Alexis Jacq, David Ruhe, Yu-Han Wu, Sebastian Flennerhag, Brendan O'Donoghue, George Scrivener, Shantanu Thakoor

View PDF HTML (experimental)

Abstract:We present DiffusionGemma, an experimental open-weight connection exemplary that uses discrete diffusion to make matter astatine exceptionally precocious speed. Rather than decoding 1 token astatine a time, DiffusionGemma iteratively refines blocks of 256 tokens successful parallel, avoiding the sequential decoding bottleneck of accepted autoregressive (AR) ample connection models. Instead of training from scratch, we get DiffusionGemma by fine-tuning the mixture-of-experts Gemma 4 exemplary pinch 3.8B activated and 25.2B full parameters. Our compute-efficient two-stage training pipeline uses less than 10% of the starting AR model's full training token budget. The first shape uses supervised fine-tuning to thatch bidirectional denoising, while the 2nd shape combines reinforcement learning pinch sampler distillation to jointly amended procreation value and conclusion efficiency. DiffusionGemma establishes a caller Pareto frontier for the trade-off betwixt procreation velocity and exemplary capability. Averaged crossed our afloat information suite, it generates astir 20 tokens per guardant walk and achieves astir 1,500 output tokens per 2nd connected a azygous NVIDIA H100 GPU, which is substantially faster than AR models moreover pinch state-of-the-art speculative decoding. DiffusionGemma besides retains the starting model's support for reasoning mode, multimodal inputs, and agelong contexts. Despite diffusion fine-tuning, it remains tin of AR procreation pinch only insignificant capacity degradation, suggesting a way toward hybrid diffusion-AR decoding.

Submission history

From: Jean Tarbouriech [view email]
[v1] Fri, 31 Jul 2026 16:11:46 UTC (6,116 KB)

More