Models

Google Ships Gemma 4 12B, an Encoder-Free Multimodal Model for Laptops

Google's Gemma 4 12B drops multimodal encoders entirely, runs on 16GB of memory, handles audio natively and ships under Apache 2.0 as Gemma 4 passes 150 million downloads.

Introducing Gemma 4 12B: a unified, encoder-free multimodal model
Introducing Gemma 4 12B: a unified, encoder-free multimodal modelAI-generated
By Sophie Lindqvist3 min read

Updated

Why it matters

  • Gemma 4 12B is an encoder-free multimodal model that runs locally on laptops with 16GB of VRAM or unified memory.
  • Gemma 4 models have crossed 150 million downloads, per Google.
  • The model ships under Apache 2.0 with Multi-Token Prediction drafters, native audio input, and benchmark performance nearing the 26B MoE model at less than half the memory footprint.

Google has released Gemma 4 12B, an open-weight multimodal model that runs on consumer laptops with 16GB of VRAM or unified memory. The release fills the gap between the edge-focused Gemma 4 E4B and the larger 26B Mixture of Experts (MoE) model in the Gemma 4 lineup, and it is the company's first mid-sized model with native audio inputs.

The launch lands as Gemma 4 models cross 150 million downloads, according to Google. The company credits the developer community with building applications ranging from "wearable robotic arms for physical assistance to enterprise-grade AI security."

Gemma 4 12B's defining feature is its unified, encoder-free architecture. Traditional multimodal models use separate encoders to translate images and audio before passing representations to the language model. Google says those split encoders add latency and increase memory usage, so the company trained Gemma 4 12B to integrate audio and vision input directly into the LLM backbone.

The design choices are concrete. For vision, Google replaced Gemma 4's vision encoder with "a lightweight embedding module consisting of a single matrix multiplication, positional embedding and normalizations," handing visual processing to the LLM backbone. For audio, the company simplified further: it removed the audio encoder entirely and projects the raw audio signal into the same dimensional space as text tokens.

The trade-off matters for the local-inference market. Encoder-free designs promise lower latency and smaller memory footprints, which determines whether a multimodal model fits on a laptop rather than a workstation or cloud instance. Google claims Gemma 4 12B delivers benchmark performance nearing its 26B MoE model at less than half the total memory footprint, enabling multi-step reasoning and agentic workflows on everyday hardware.

The model also ships with Multi-Token Prediction (MTP) drafters, which Google says reduce latency. MTP drafters support speculative decoding, a technique that generates multiple candidate tokens per step to accelerate inference without changing the final output.

Gemma 4 12B is released under an Apache 2.0 license, continuing Google's practice of open-sourcing the Gemma family for commercial and research use. Weights for pre-trained and instruction-tuned checkpoints are available on Hugging Face and Kaggle. The release positions Gemma against open-weight competitors in the sub-13B class, where Meta, Mistral and Qwen are also competing for local and edge deployments.

The tooling ecosystem is broad at launch. Users can experiment with the model in LM Studio, Ollama, the Google AI Edge Gallery App, the Google AI Edge Eloquent app and the LiteRT-LM CLI. Local inference pipelines work with Hugging Face Transformers, llama.cpp, MLX, SGLang and vLLM, and fine-tuning is supported through Unsloth. Google published a companion Gemma 4 12B Developer Guide with a technical breakdown of the architecture, plus developer documentation and a quick start notebook.

For agentic development, Google is releasing an official Skills Repository, a library of skills "designed specifically to enable agents to build with Gemma models." For production deployments, the model is available through Google Cloud via the Gemini Enterprise Agent Platform Model Garden, Cloud Run and GKE.

With Gemma 4 12B, Google is betting that encoder-free architectures can bring near-flagship multimodal reasoning to the 16GB-memory tier of hardware. The 150-million-download milestone for the Gemma 4 family suggests demand for locally runnable models is real, and this release extends that reach to audio-input workloads on mid-range laptops.

Original: blog.google

Share this article:

More from Sophie Lindqvist

Sophie Lindqvist

Show full bio

Staff writer covering marketplaces and e-commerce at AI In Context.

136 articles

Related articles

  1. Google Ships Gemma 4, Its Smallest-to-Strongest Open Model Family
  2. Google Converts Gemma 2 Into Encoder-Decoder Models With T5Gemma
  3. Google Launches Gemini 3.1 Flash TTS With Audio Tags and 70+ Languages
  4. OpenAI Launches gpt-realtime Speech-to-Speech Model With MCP and Phone Support
  5. Google Ships Upgraded Gemini 2.5 Flash Native Audio and Live Translation

« Previous articleNext article »