Safety & Security

Google releases Gemma Scope 2, the largest open interpretability suite

Google's Gemma Scope 2 spans all Gemma 3 models from 270M to 27B parameters, with ~110 petabytes of data — the largest open interpretability release by an AI lab.

By James Calloway4 min read

Updated

Why it matters

  • Google released Gemma Scope 2, calling it the largest ever open-source interpretability release by an AI lab.
  • The suite covers all Gemma 3 model sizes, from 270M to 27B parameters.
  • Producing it required storing ~110 petabytes of data and training over 1 trillion total parameters.
  • It uses Matryoshka-trained SAEs plus skip- and cross-layer transcoders on every Gemma 3 layer.
  • An interactive demo is available via Neuronpedia.

Google has released Gemma Scope 2, an open suite of interpretability tools covering every model in the Gemma 3 family — from 270 million to 27 billion parameters — which the company describes as "the largest ever open-source release of interpretability tools by an AI lab to date."

The release required storing approximately 110 petabytes of data and training sparse autoencoders and transcoders totaling more than 1 trillion parameters, according to Google's announcement. The scale matters: interpretability researchers have long lacked tooling for large production-grade models, and Google is betting that open access will let the safety community audit behaviors that only emerge at scale.

An interactive demo of Gemma Scope 2 is already available, built by Neuronpedia.

Why does interpretability matter right now?

Large language models can perform what Google calls "incredible feats of reasoning, yet their internal decision-making processes remain largely opaque." When a system misbehaves, that opacity makes it hard to pinpoint the cause.

Interpretability research aims to open the black box. As Google puts it, "As AI becomes increasingly more capable and complex, interpretability is crucial for building AI that is safe and reliable."

The stakes are concrete. Google says it expects the research community to use Gemma Scope 2 to:

  • Debug emergent model behaviors
  • Audit and debug AI agents
  • Accelerate "practical and robust safety interventions against issues like jailbreaks, hallucinations and sycophancy"

Gemma Scope 2 follows the original Gemma Scope, released last year for the Gemma 2 family. That toolkit already supported research into model hallucination, identifying secrets known by a model, and training safer models.

What does the new suite actually include?

Gemma Scope 2 works as a microscope for the Gemma model family. It combines sparse autoencoders (SAEs) and transcoders so researchers can inspect what a model is "thinking about" and how those internal states connect to observable behavior. That includes studying discrepancies between what a model says it is reasoning about and what its internal state actually shows — a problem known as chain-of-thought faithfulness.

The upgrade over last year's release rests on four pillars, per Google:

  • Full coverage at scale. The suite spans the entire Gemma 3 family, up to 27B parameters. Google argues this is essential for studying emergent behaviors that appear only in large models, citing the 27B-parameter C2S Scale model that previously helped discover a new potential cancer therapy pathway. Gemma Scope 2 is not trained on that model, Google notes, but the finding illustrates the kind of emergent behavior the tools could help explain.
  • More refined tools for complex behaviors. SAEs and transcoders are trained on every layer of the Gemma 3 models. New skip-transcoders and cross-layer transcoders make it easier to decipher multi-step computations and algorithms distributed across the model.
  • Advanced training techniques. The suite uses the Matryoshka training technique, which Google says helps SAEs detect more useful concepts and fixes certain flaws discovered in the original Gemma Scope.
  • Chatbot behavior analysis. Interpretability tools target the chat-tuned versions of Gemma 3, enabling analysis of complex, multi-step behaviors such as jailbreaks, refusal mechanisms, and chain-of-thought faithfulness.

What's the context for the release?

Google frames the release as infrastructure for the broader AI safety community rather than an internal research artifact. "By releasing Gemma Scope 2, we aim to enable the AI safety research community to push the field forward using a suite of cutting-edge interpretability tools," the company said.

The argument is that real-world safety problems arise only in larger, modern LLMs — and that without open tooling at that scale, external researchers cannot study them. "This new level of access is crucial for tackling real-world safety problems that only arise in larger, modern LLMs," Google said.

The original Gemma Scope, launched last year for Gemma 2, already produced results in several safety-relevant areas. Google credits it with enabling research on hallucination, on identifying secrets a model has memorized, and on methods for training safer models. Gemma Scope 2 extends that coverage to a newer, larger, and more capable model generation.

Who can use it, and how?

The tools are open and available now. The easiest entry point is the interactive demo hosted by Neuronpedia, which lets users explore the SAEs and transcoders without running the underlying models themselves.

For researchers, the release covers both base and chat-tuned Gemma 3 models. That distinction is significant for safety work: jailbreaks, refusals, and reasoning faithfulness are behaviors of tuned chat systems, and analyzing them requires tooling trained specifically on those variants rather than only on base models.

What happens next?

Google says the success of the release depends on what the community builds with it. The company's stated roadmap for the tooling runs from debugging emergent behavior, through auditing AI agents, to practical defenses against jailbreaks, hallucinations, and sycophancy. With 110 petabytes of storage behind it and coverage up to 27B parameters, Gemma Scope 2 gives outside researchers, for the first time at this scale, a way to test whether those safety interventions actually work inside a model's "brain."

Original: storage.googleapis.com

Share this article:

More from James Calloway

James Calloway

Show full bio

News editor covering industry trends and analytics at AI In Context.

200 articles

Related articles

  1. Google DeepMind Puts $10M Toward Multi-Agent AI Safety
  2. Google Ships Gemma 4, Its Smallest-to-Strongest Open Model Family
  3. Google Runs First Double-Blind Evaluation of a Frontier AI Model
  4. Google Upgrades Gemini 3 Deep Think With Record Benchmark Runs
  5. Goodfire Opens Silico Platform to Peek Inside AI Models

« Previous articleNext article »