Google Converts Gemma 2 Into Encoder-Decoder Models With T5Gemma
Google's T5Gemma converts pretrained decoder-only Gemma 2 models into encoder-decoder LLMs, beating Gemma 2 on GSM8K, DROP and MMLU at similar latency, with checkpoints now released.

Updated
Why it matters
- T5Gemma adapts pretrained decoder-only Gemma 2 models (2B and 9B) into encoder-decoder LLMs via UL2 or PrefixLM-based pre-training
- T5Gemma 2B-2B IT improves GSM8K from 58.0% to 70.7% and gains nearly 12 points on MMLU over Gemma 2 2B
- Unbalanced configurations are possible, e.g. a 9B encoder paired with a 2B decoder, matching Gemma 2 2B latency while boosting accuracy
Google has released T5Gemma, a collection of encoder-decoder large language models built by converting pretrained decoder-only Gemma 2 models — and the adapted models outperform their decoder-only counterparts on benchmarks including GSM8K, DROP and MMLU.
The release includes adapted Gemma 2 2B and 9B models plus newly trained T5-sized variants (Small, Base, Large and XL), available in both pretrained and instruction-tuned versions. Google frames the work as a direct answer to the question: "can we build top-tier encoder-decoder models based on pretrained decoder-only models?"
The technique behind T5Gemma is model adaptation. It initializes the parameters of an encoder-decoder model using the weights of an already pretrained decoder-only model, then further adapts them through UL2 or PrefixLM-based pre-training. The result is a family of models rooted in the Gemma 2 framework but structured around the classic encoder-decoder architecture popularized by T5, the Text-to-Text Transfer Transformer.
The stakes are architectural as well as commercial. Decoder-only designs dominate today's LLM market, but encoder-decoder models retain real advantages for production workloads: high inference efficiency, design flexibility, and richer encoder representations for understanding input. Tasks like summarization, translation and question answering — where comprehension of the input matters as much as generation — remain natural fits for the older architecture. As Google puts it, encoder-decoder models "often excel at summarization, translation, QA, and more due to their high inference efficiency, design flexibility, and richer encoder representation for understanding input," yet the architecture "has received little relative attention."
Adaptation also unlocks a design option unavailable to standard decoder-only models: unbalanced configurations. Developers can pair a large encoder with a small decoder — for example, a 9B encoder with a 2B decoder. This lets engineers tune the quality-efficiency trade-off per task, which matters for workloads like summarization, where, as Google notes, "a deep understanding of the input is more critical than the complexity of the generated output."
Benchmark results
The headline numbers favor the adapted models. In Google's experiments, T5Gemma models achieve "comparable or better performance than their decoder-only Gemma counterparts, nearly dominating the quality-inference efficiency pareto frontier across several benchmarks," including SuperGLUE, which measures the quality of learned representations.
The gains appear before any instruction tuning. T5Gemma 9B-9B scores more than 9 points higher than the original Gemma 2 9B on GSM8K, a math reasoning benchmark, and 4 points higher on DROP, a reading comprehension benchmark. Google reads this pattern as evidence that "the encoder-decoder architecture, when initialized via adaptation, has the potential to create a more capable, performant foundational model."
Instruction tuning widens the gap further. T5Gemma 2B-2B IT gains nearly 12 points on MMLU over Gemma 2 2B, and its GSM8K score rises from 58.0% to 70.7%. In Google's assessment, "the adapted architecture not only potentially provides a better starting point but also responds more effectively to instruction-tuning, ultimately leading to a substantially more capable and helpful final model."
Latency, not just accuracy
The performance claims extend to measured wall-clock speed on GSM8K. T5Gemma 9B-9B achieves higher accuracy than Gemma 2 9B at similar latency, according to Google's measurements. The unbalanced configuration delivers the more striking result: T5Gemma 9B-2B provides "a significant accuracy boost over the 2B-2B model" while its latency stays "nearly identical to the much smaller Gemma 2 2B model."
That combination — near-small-model speed with accuracy above the larger decoder-only baseline — is the practical case for the architecture. As Google summarizes the experiments, encoder-decoder adaptation "offers a flexible, powerful way to balance across quality and inference speed," a trade-off that directly shapes serving costs for anyone deploying models in production.
Open release
Google has published a suite of T5Gemma checkpoints to the community, covering the pretrained and instruction-tuned variants. The company positions the release as a research resource: "We hope these checkpoints will provide a valuable resource for investigating model architecture, efficiency, and performance."
The release invites a concrete line of follow-up work: whether the adaptation recipe — initialized from an already pretrained decoder-only model rather than trained from scratch — generalizes beyond Gemma 2 to other model families, and whether unbalanced encoder-decoder configurations become a standard lever for latency-sensitive deployments.
Original: arxiv.org
More from Sophie Lindqvist
Show full bio
Staff writer covering marketplaces and e-commerce at AI In Context.
115 articles