Chips & Compute

Google's Decoupled DiLoCo Trains LLMs Across Data Centers 20x Faster

Google trained a 12B-parameter model across four U.S. regions over 2-5 Gbps links, 20x faster than sync methods, using self-healing asynchronous training.

Decoupled DiLoCo: A new frontier for resilient, distributed AI training
Decoupled DiLoCo: A new frontier for resilient, distributed AI trainingjurvetson / Openverse
By James Calloway3 min read

Updated

Why it matters

  • Decoupled DiLoCo trained a 12 billion parameter model across four U.S. regions using 2-5 Gbps wide-area networking, more than 20x faster than conventional synchronization methods.
  • Chaos-engineering tests showed the system survived the loss of entire learner units and reintegrated them, matching benchmarked ML performance of traditional training on Gemma 4 models.
  • The architecture mixes TPU v6e and TPU v5p chips in a single run with no loss in ML performance, extending the life of older hardware.

Google says it has trained a 12 billion parameter model across four separate U.S. regions using just 2-5 Gbps of wide-area networking — and finished more than 20 times faster than conventional synchronization methods. The result comes from Decoupled DiLoCo, a new distributed training architecture described in a paper published today by a team spanning Google DeepMind and Google Research.

The work targets a structural problem in frontier AI training. Today's state-of-the-art models depend on large, tightly coupled systems in which identical chips must stay in near-perfect synchronization. Google says that as models scale toward future generations, maintaining that synchronization across thousands of chips becomes a significant logistical challenge.

Decoupled DiLoCo (Distributed Low-Communication) splits large training runs across decoupled "islands" of compute, called learner units, with asynchronous data flowing between them. The architecture isolates local disruptions so other parts of the system keep learning, and it avoids the communication delays that made earlier distributed methods like Data-Parallel impractical at global scale.

The system builds on two prior Google advances: Pathways, a distributed AI system based on asynchronous data flow, and DiLoCo, which dramatically reduced the bandwidth required between data centers and made training large language models across distant locations practical. Decoupled DiLoCo runs on top of Pathways and adds fault tolerance through what the team calls a self-healing design.

To test that claim, the researchers used "chaos engineering" — deliberately injecting artificial hardware failures during live training runs. Decoupled DiLoCo continued training after losing entire learner units, then reintegrated them when they came back online. Testing with Gemma 4 models showed the system maintained greater availability of learning clusters than traditional training methods while ultimately delivering the same benchmarked ML performance.

The speed gain comes from scheduling: the system folds required communication into longer periods of computation, avoiding the "blocking" bottlenecks where one part of the system must wait for another. The 2-5 Gbps bandwidth figure matters because it is achievable with existing internet connectivity between data center facilities, rather than requiring new custom network infrastructure — a meaningful cost consideration as AI labs race to expand compute capacity.

The architecture also unlocks heterogeneous hardware. In Google's experiments, training runs that mixed TPU v6e and TPU v5p chips — different generations running at different speeds — matched the ML performance of single-chip-type runs. That extends the useful life of existing hardware, increases total available compute, and eases a recurring logistics problem: new hardware generations do not arrive everywhere at once.

Google frames the result as part of a full-stack approach to AI training spanning hardware, software infrastructure, and research, with gains increasingly coming from rethinking how those layers fit together. Because Decoupled DiLoCo operates at internet-scale bandwidth, it can tap unused compute wherever it sits, turning stranded resources into useful capacity.

The paper's lead authors and core contributors are Arthur Douillard, Keith Rush, Yani Donchev, Zachary Charles, Ayush Dubey, Blake Woodworth, Ionel Gog, Josef Dean, Nova Fallen, and Zachary Garrett, with operational support from Nate Keating and Jenny Bishop. Advisers included Jeff Dean, Paul Barham, Michael Isard, and Marc'Aurelio Ranzato, among others.

Google says it is continuing to explore resilient systems of this kind to unlock the next generation of AI — a signal that asynchronous, heterogeneous training may become a standard lever as frontier models outgrow any single tightly synchronized cluster.

Original: arxiv.org

Share this article:

More from James Calloway

James Calloway

Show full bio

News editor covering industry trends and analytics at AI In Context.

121 articles

Related articles

  1. Google Releases DiffusionGemma, a 26B Model That Generates Text Four Times Faster
  2. Google DeepMind Joins DOE's Genesis Mission to Bring AI to 17 National Labs
  3. OpenAI open-sources MRC, a networking protocol for AI training clusters
  4. Google DeepMind Launches European Robotics Accelerator

« Previous articleNext article »