Models

Google Ships Gemma 4, Its Smallest-to-Strongest Open Model Family

Google's Gemma 4 lands under Apache 2.0 with four sizes from 2B to 31B; the 31B model ranks #3 among open models on Arena AI while outcompeting rivals 20x its size.

Gemma 4: Byte for byte, the most capable open models
Gemma 4: Byte for byte, the most capable open modelsAI-generated
By James Calloway4 min read

Updated

Why it matters

  • Gemma 4 ships in four sizes — E2B, E4B, 26B MoE, and 31B Dense — under a commercially permissive Apache 2.0 license.
  • The 31B model ranks #3 among open models on the Arena AI text leaderboard; the 26B MoE ranks #6, activating only 3.8B parameters during inference.
  • Gemma has been downloaded over 400 million times, with more than 100,000 community variants; edge models run offline on phones, Raspberry Pi, and Jetson Orin Nano with up to 256K context on larger models.

Google has released Gemma 4, a four-model family that puts its 31B Dense model at #3 on the Arena AI text leaderboard — a spot where it outcompetes open models 20 times its size, according to Google. The 26B Mixture of Experts variant holds #6 on the same board.

The release matters for two reasons. First, it pushes frontier-adjacent reasoning onto consumer and edge hardware at a moment when developers and enterprises want local control over models and data. Second, Google has switched Gemma 4 to a commercially permissive Apache 2.0 license, responding to community feedback and removing barriers that constrained earlier open-model deployments in regulated and sovereign environments.

Four sizes, one family

Gemma 4 ships in four configurations: Effective 2B (E2B), Effective 4B (E4B), 26B Mixture of Experts, and 31B Dense. Google built the family from the same research and technology as Gemini 3, positioning Gemma as the open complement to its proprietary Gemini line — what Google calls "the industry's most powerful combination of both open and proprietary tools."

The two larger models target workstations and personal computers. Their unquantized bfloat16 weights fit on a single 80GB NVIDIA H100 GPU, and quantized versions run natively on consumer GPUs. The split between them is deliberate: the 26B MoE activates only 3.8 billion of its total parameters during inference to maximize tokens-per-second, while the 31B Dense trades latency for raw quality and serves as a foundation for fine-tuning.

The edge models take a different path. E2B and E4B activate effective 2-billion and 4-billion parameter footprints to preserve RAM and battery life, and Google engineered them with its Pixel team and hardware partners Qualcomm Technologies and MediaTek. They run completely offline with near-zero latency on phones, Raspberry Pi, and NVIDIA Jetson Orin Nano.

Capability profile

Google claims several concrete advances across the family:

  • Reasoning and agents. The models handle multi-step planning and deep logic, with improvements on math and instruction-following benchmarks. Native function-calling, structured JSON output, and system instructions support autonomous agent workflows.
  • Multimodality. All models process video and images at variable resolutions, with strong performance on OCR and chart understanding. E2B and E4B add native audio input for speech recognition.
  • Long context. Edge models carry a 128K context window; larger models reach 256K, enough to pass entire repositories or long documents in a single prompt.
  • Languages. Native training on over 140 languages.
  • Code. Offline code generation turns a workstation into a local-first coding assistant, Google says.

Momentum behind the license change

The Apache 2.0 decision follows sustained community growth. Since the first generation launched, developers have downloaded Gemma more than 400 million times and produced a "Gemmaverse" of more than 100,000 variants. Google frames the license as a foundation for "complete developer flexibility and digital sovereignty," granting control over data, infrastructure, and models, with deployment on-premises or in the cloud.

The company also says Gemma 4 undergoes the same infrastructure security protocols as its proprietary models — a signal aimed squarely at enterprises and sovereign organizations that need auditable, transparent foundations for regulated workloads.

Fine-tuning on accessible hardware has already produced results in prior generations. INSAIT built BgGPT, a Bulgarian-first language model, and Google worked with Yale University on Cell2Sentence-Scale to discover new pathways for cancer therapy.

Day-one ecosystem

Gemma 4 arrives with broad tooling support from the start: Hugging Face (Transformers, TRL, Transformers.js, Candle), LiteRT-LM, vLLM, llama.cpp, MLX, Ollama, NVIDIA NIM and NeMo, LM Studio, Unsloth, SGLang, Cactus, Baseten, Docker, MaxText, Tunix, and Keras. Weights are available on Hugging Face, Kaggle, and Ollama.

Experimentation paths include Google AI Studio for the 31B and 26B MoE models and Google AI Edge Gallery for E4B and E2B. Android developers can prototype agentic flows in the AICore Developer Preview today, with forward-compatibility promised for Gemini Nano 4 — an early look at how Google plans to unify its edge AI stack. Agent Mode in Android Studio and the ML Kit GenAI Prompt API round out the Android story.

For production scale, Google Cloud offers deployment through Vertex AI, Cloud Run, GKE, Sovereign Cloud, and TPU-accelerated serving, with what Google describes as the highest compliance guarantees for regulated workloads. Hardware optimization spans NVIDIA's range from Jetson Orin Nano to Blackwell GPUs, AMD GPUs via the open-source ROCm stack, and Trillium and Ironwood TPUs.

Google is also launching the Gemma 4 Good Challenge on Kaggle to encourage products with positive social impact.

The open-model race now hinges on intelligence per parameter rather than raw scale, and Google has just made its strongest argument yet that the frontier can run on a laptop GPU.

Original: arena.ai

Share this article:

More from James Calloway

James Calloway

Show full bio

News editor covering industry trends and analytics at AI In Context.

121 articles

Related articles

  1. Google Releases DiffusionGemma, a 26B Model That Generates Text Four Times Faster
  2. Google Ships Gemini 3.1 Pro, More Than Doubling Reasoning Score
  3. Google Launches Gemini 3 Pro at $2/Million Input Tokens
  4. Google Ships Gemini 3.5 Flash, Promises Pro Model Next Month
  5. Google Upgrades Gemini 3 Deep Think With Record Benchmark Runs

« Previous articleNext article »