Google Releases DiffusionGemma, a 26B Model That Generates Text Four Times Faster
Google's experimental DiffusionGemma drafts 256-token blocks in parallel, hitting 1,000+ tokens per second on an H100 — at the cost of output quality versus standard Gemma 4.

Updated
Why it matters
- DiffusionGemma is a 26B-parameter MoE model (3.8B active) released under Apache 2.0 that generates 256 tokens in parallel, reaching 1,000+ tokens/second on a single NVIDIA H100 and 700+ on an RTX 5090
- The model fits within 18GB VRAM when quantized and targets local, low-concurrency workflows; Google says output quality is lower than standard Gemma 4 and recommends the autoregressive models for production
- Launch support includes vLLM (Red Hat-supported integration), MLX, Hugging Face Transformers, fine-tuning via Hackable Diffusion, Unsloth and NVIDIA NeMo, with llama.cpp support arriving soon
Google has released DiffusionGemma, an experimental open model that generates text up to four times faster on GPUs by drafting entire 256-token blocks simultaneously instead of predicting one token at a time. The 26B-parameter Mixture of Experts model, available now on Hugging Face under an Apache 2.0 license, hits 1,000+ tokens per second on a single NVIDIA H100 and 700+ tokens per second on a consumer GeForce RTX 5090.
The release matters because it brings text diffusion — a technique the AI research community has explored for years but struggled to scale — into a large, downloadable model built on the intelligence-per-parameter foundation of the Gemma 4 family and Google's Gemini Diffusion research. For developers, it offers a concrete answer to a persistent problem: local inference latency.
A different hardware trade-off
Autoregressive language models behave like typewriters, generating tokens left to right. That works well in the cloud, where servers batch thousands of requests to share hardware load. Run locally for a single user, the sequential approach leaves a GPU or TPU underutilized — the chip mostly waits for the next token.
DiffusionGemma attacks that inefficiency directly. It drafts a full 256-token paragraph in each forward pass, shifting the decode bottleneck from memory bandwidth to compute. Google's framing: the model upgrades inference "from a single, sequential typewriter to a massive printing press that stamps the entire block of text simultaneously."
The speedup is specific to its context. In high-QPS cloud serving, autoregressive models can saturate compute efficiently, so DiffusionGemma's parallel decoding offers diminishing returns and can raise serving costs. The throughput advantage is strongest at low-to-medium batch sizes on a single accelerator.
The specifications
The model activates only 3.8B of its 26B total parameters during inference, and when quantized it fits within the 18GB VRAM envelope of high-end consumer GPUs. Generating 256 tokens in parallel gives every token visibility into all others — bi-directional attention that Google says provides significant advantages for non-linear domains such as in-line editing, code infilling, amino acid sequences, and mathematical graphs. The model also iteratively refines its own output, evaluating the entire text block at once to fix mistakes in real time.
Google demonstrated the fine-tuning angle with Unsloth, which trained DiffusionGemma to play Sudoku — a task autoregressive models struggle with because each token depends on future tokens. DiffusionGemma's bi-directional attention makes the problem much easier.
How the diffusion process works
The mechanics mirror AI image generators that start with visual static and refine it into a picture:
- The canvas: The model starts with a canvas of random placeholder tokens.
- Iterative refinement: The model makes multiple passes, locking in correct tokens and using them as context clues to refine the rest.
- Final polish: The text converges into high-quality output.
Because the model processes the whole paragraph while generating, it unlocks new behavior patterns, such as perfectly closing complex markdown formatting or generating and rendering code in near real time. Hugging Face built a text-to-3D SVG demo showing the step-by-step generation.
The quality caveat
Google is explicit about the trade-off. Because DiffusionGemma prioritizes speed and parallel generation, its overall output quality is lower than standard Gemma 4, which the company recommends for applications that demand maximum quality. Autoregressive Gemma 4 models remain the standard for high-quality production outputs; DiffusionGemma targets researchers and developers building speed-critical, interactive local workflows such as in-line editing and rapid iteration.
Ecosystem support at launch
The release arrives with a broad tooling stack. Developers can serve the model with MLX, vLLM (with integration supported by Red Hat), and Hugging Face Transformers, with official llama.cpp support coming soon. For fine-tuning, Google is releasing a tutorial using Hackable Diffusion, a modular JAX toolbox designed for composability, and users can also fine-tune with Unsloth and NVIDIA NeMo.
Google worked with NVIDIA to optimize across its hardware stack: quantized builds target consumer setups on GeForce RTX 5090 and 4090 GPUs, while enterprise systems run on Hopper and Blackwell with advanced NVFP4 kernels — including NVIDIA DGX Spark and DGX Station for deskside deployment and RTX PRO for AI professionals. Native NVFP4 (4-bit floating-point) support accelerates compute throughput with near-lossless accuracy, according to Google.
The model runs on desktop GPUs or in the cloud through the Gemini Enterprise Agent Platform Model Garden or NVIDIA NIM. Google is also publishing a developer guide and "A Visual Guide to DiffusionGemma" explaining the mechanics.
Whether text diffusion moves beyond an experimental niche will depend on fine-tuning closing the quality gap with autoregressive models — the Sudoku result suggests the architecture's advantages are real but task-specific, and Google itself is telling production users to stay with standard Gemma 4 for now.
Original: unsloth.ai
More from Marcus Bennett
Show full bio
Senior reporter covering consumer brands and retail at AI In Context.
108 articles