Models · The Decoder ·

Google's DiffusionGemma proves you don't need to train from scratch to build a text diffusion model

Google's DiffusionGemma proves you don't need to train from scratch to build a text diffusion model

Google DeepMind converted Gemma 4 into a text diffusion model using less than 10% of its original training budget. DiffusionGemma generates 256 tokens in parallel at about 1,500 tokens per second, but trails the original model on benchmarks, particularly reasoning tasks.

Read the full story at The Decoder →