Google's Decoupled DiLoCo Trains LLMs Across Data Centers 20x Faster
Google trained a 12B-parameter model across four U.S. regions over 2-5 Gbps links, 20x faster than sync methods, using self-healing asynchronous training.
Topic
Topic
Google trained a 12B-parameter model across four U.S. regions over 2-5 Gbps links, 20x faster than sync methods, using self-healing asynchronous training.