Models

Google Launches Gemini Robotics ER 2 With Multi-Robot Collaboration

Google's Gemini Robotics ER 2 ships publicly today, adding video-based progress tracking, 91.3% moment-finding accuracy, multi-robot collaboration, and top safety benchmark scores.

Gemini Robotics ER 2: powering robotics with video understanding, task orchestration, and multi-robot collaboration
Gemini Robotics ER 2: powering robotics with video understanding, task orchestration, and multi-robot collaborationAI-generated
By Sophie Lindqvist6 min read

Updated

Why it matters

  • Gemini Robotics ER 2 achieves 91.3% accuracy and 0.96s mean absolute distance on moment-finding tasks, at 4x the execution speed and a fraction of the compute cost of larger model categories.
  • The model reaches 57.4% accuracy on five-level progress classification and outperforms ER 1.6 and other frontier models on Safety Instruction Following and Human Proximity benchmarks.
  • Available now via the Gemini API and Google AI Studio, with demos spanning Boston Dynamics' Spot, Apptronik's Apollo 2, and the Franka F3 Duo in multi-robot collaboration.

Google has launched Gemini Robotics ER 2, its most capable "embodied reasoning" model for robotics, adding continuous video understanding, task orchestration, multi-robot collaboration, and what the company calls its strongest safety performance to date. The model is publicly available to developers now via the Gemini API and Google AI Studio, and in private preview on the Gemini Enterprise Agent Platform.

The release matters because it targets the two hardest problems standing between lab demos and useful robots: knowing when a physical task is actually finished, and reasoning fast enough to keep pace with the real world. "For robots to assist humans in everyday environments, accurate spatial reasoning is not enough," Google writes in its announcement. "Robots must also think fast, timing their decisions and reasoning with the real-time speed of the physical world."

A high-level brain that delegates the body

Google positions Gemini Robotics ER 2 as a "high-level brain for robots." The model chats with humans, understands the physical world, and plans multi-step tasks, then hands off motor execution to any lower-level vision-language-action (VLA) model. It can also natively call tools like Google Search or any user-defined function. Crucially, its architecture lets a robot "think" about what comes next while simultaneously performing actions — eliminating the stop-and-think pauses that have plagued chain-of-thought robotics systems.

Developers build agentic setups by declaring low-level control interfaces — VLA models or navigation APIs — as tools, then streaming multimodal video, audio, or text directly into the model. Google says Gemini Robotics ER 2 consistently outperforms its predecessor, ER 1.6, for tool orchestration across three control modes: real VLA, simulated VLA, and human teleoperation.

To demonstrate the system, Google built a demo with Spot, the quadruped robot from partner Boston Dynamics. Gemini Robotics ER 2 orchestrates Spot's APIs, including navigation and manipulator movement, to create an interactive robot that fetches objects — such as a popcorn snack — on a natural language command. The code is available on GitHub alongside other examples.

For latency-critical work, the model integrates into the Gemini Live API through a bidirectional streaming endpoint optimized for latency-sensitive tasks. The result, according to Google, is "fluid orchestration": multi-step task completion without jarring pauses between reasoning and action.

Temporal intelligence: knowing when a task is done

The headline technical claim in this release concerns what Google calls "temporal intelligence." One of robotics' hardest challenges is knowing when a task is done. Gemini Robotics ER 2 brings what Google describes as a step-change in video understanding and progress tracking, verifying that complex tasks — such as tightening a light bulb or tying a trash bag — are complete to specification before the robot moves on.

Google reports progress on two foundational capabilities.

Continuous progress classification. The model tracks how far along a task is in real time. In Google's evaluations, each frame in a video feed is assigned to one of five progress levels (0–20%, 20–40%, 40–60%, 60–80%, 80–100%). Gemini Robotics ER 2 achieves 57.4% accuracy on progress classification tasks, outperforming previous-generation models and competing frontier models. By quantifying task progress, the model gives robots real-time situational awareness, letting them adjust actions on the fly or retry failed steps without restarting an entire workflow.

Precision moment-finding. This measures a model's ability to identify the exact video frame where a critical event occurs — for example, when to stop pouring coffee into a cup. Here Gemini Robotics ER 2 achieves 91.3% accuracy and a 0.96-second mean absolute distance. Google says the model competes closely with much larger model categories while delivering that precision at a fraction of the compute cost and 4x the execution speed. That sub-second latency, the company argues, is what safely operating physical robots in the real world actually requires.

Robots working together

The release also introduces multi-robot collaboration, a first for the Gemini Robotics line. The logic is straightforward: no single robot fits every task. A wheeled rover excels indoors; a humanoid robot may handle uneven terrain better. Gemini Robotics ER 2 lets diverse machines communicate via a shared semantic understanding to hand off and complete complex workflows that a single robot could not do alone.

Google demonstrates the capability with Apptronik's Apollo 2 humanoid working alongside a Franka F3 Duo robot arm.

Stronger spatial reasoning

Gemini Robotics ER 2 also advances the model's core spatial reasoning, measured across three benchmarks:

  • Success/failure detection now operates on raw video feeds rather than static snapshots, catching mid-execution failures like spills, slips, or misalignments.
  • General instrument reading extends beyond circular dials and sight glasses to include digital displays, linear scales, rulers, and liquid thermometers, tested across 10 different instrument types.
  • Enhanced spatial VQA improves Visual Question Answering through Gemini's multi-modal understanding advances.

Google says the model consistently achieves the highest accuracy across all core capabilities, with highlights including success detection on image and video, question answering on the ERQA benchmark, and generalized instrument reading.

Safety gains and a new benchmark

Google calls Gemini Robotics ER 2 its safest robotics model. It achieves significant gains on the Safety Instruction Following and Human Proximity benchmarks, which evaluate how well a model adheres to physical constraints during reasoning tasks and its spatial awareness for detecting humans. In testing, Google found the model successfully halts a humanoid robot when a person is nearby and autonomously resumes work only once the area is clear.

The company is also introducing a new benchmark that evaluates a foundation model's ability to act as a safe VLA orchestrator by testing its capacity to enforce safety constraints, monitor the environment, assess physical feasibility, and seek human clarification. Google says Gemini Robotics ER 2 outperforms ER 1.6 and other frontier models on both safety benchmarks. Full details appear in the company's safety technical report.

Why it matters

The announcement pushes Google deeper into a contested market. Robotics foundation models have attracted heavy investment across the industry, and Google — with its Boston Dynamics and Apptronik partnerships, its Gemini model family, and now an open API path for developers — is betting that a reasoning layer decoupled from motor control will scale faster than monolithic VLA systems. The benchmark numbers matter here: 91.3% moment-finding accuracy with 0.96-second latency at a fraction of competitors' compute cost is a concrete argument that general-purpose models can run the control loop physical robots need.

The strategy echoes how large language models absorbed tool use and agentic workflows, now transposed to hardware: declare control interfaces as tools, stream video in, let the model orchestrate. Whether that architecture wins depends on real-world reliability beyond demos — light bulbs that are actually tight and trash bags that are actually tied.

Google says its plans are to push these models toward even more complex tasks to accelerate the development of helpful robots and support the robotics community. With public availability through the Gemini API and Google AI Studio starting today, developers can now test that claim themselves.

Original: ai.google.dev

Share this article:

More from Sophie Lindqvist

Sophie Lindqvist

Show full bio

Staff writer covering marketplaces and e-commerce at AI In Context.

136 articles

Related articles

  1. Google DeepMind's Gemini Robotics 2 Gives Humanoids Whole-Body Control
  2. Google DeepMind Launches European Robotics Accelerator
  3. Google Ships Gemini 3.5 Flash, Promises Pro Model Next Month
  4. Google Launches Gemini 3 Pro at $2/Million Input Tokens
  5. Google DeepMind's SIMA 2 Turns AI Into a Gaming Companion

« Previous articleNext article »