Google DeepMind's Gemini Robotics 2 Gives Humanoids Whole-Body Control
Google DeepMind ships Gemini Robotics 2: whole-body humanoid control, 22-DOF dexterity, multi-robot teamwork, and on-device adaptation to new robot bodies in under 200 examples.

Updated
Why it matters
- Gemini Robotics 2 controls full humanoids from feet to fingertips, demonstrated on Apptronik's Apollo 2 walking to fetch and place objects.
- Gemini Robotics On-Device 2 adapts to completely new robot embodiments with a few hours of data and typically fewer than 200 examples.
- Gemini Robotics ER 2 is available now on Google AI Studio and in private preview on the Gemini Enterprise Agent Platform; VLA models are early-access only.
Google DeepMind has released Gemini Robotics 2, an AI model suite that for the first time controls entire humanoid robots — from feet to fingertips — and adapts to entirely new robot bodies with just a few hours of data and fewer than 200 examples.
The company frames the release as the intelligence layer for what it calls "truly adaptable robots." Where most industrial and research robots today are pre-programmed or teleoperated for narrow, repetitive sequences, Gemini Robotics 2 is built to reason through each movement, generalize to unfamiliar environments, and — new in this generation — coordinate with other robots as a team.
The stakes are straightforward. Robot learning has struggled for years with two problems: transferring skills between different robot bodies, and executing long, multi-step tasks in messy physical spaces. DeepMind claims progress on both fronts, and it is shipping the results through concrete channels rather than demos alone.
Three models, three jobs
The release actually contains three distinct models:
Gemini Robotics 2 is DeepMind's most advanced vision-language-action (VLA) model. It converts vision and language input directly into motor control, and it can control full humanoids and other bi-arm robots. It also brings what the company calls "a new level of dexterous manipulation" to both hands and grippers.
Gemini Robotics ER 2 is the embodied reasoning model — a vision-language model that acts as the robot's high-level agent. It communicates with humans, understands the physical world, plans multi-step tasks lasting several minutes, and now supports multi-robot collaboration.
Gemini Robotics On-Device 2 is the efficiency play: a VLA model optimized to run locally on robotic hardware, without network latency or internet connectivity. DeepMind says it can adapt to completely new robot embodiments with a few hours of data.
Availability is tiered. Gemini Robotics ER 2 is live now on Google AI Studio and in private preview on the Gemini Enterprise Agent Platform. The VLA and On-Device models are restricted to early-access partners, with hardware integration instructions on Google's Developer blog.
Humanoids that walk to work
The headline capability is whole-body control. Previous Gemini Robotics models controlled only a humanoid's upper body for tabletop tasks. Gemini Robotics 2 extends physical AI to full-body motion — walking, crouching, stretching, and manipulating objects in the same continuous task.
DeepMind demonstrated the system on Apptronik's Apollo 2 humanoid. Given the instruction "put the watering can into the green bin in the bottom shelf," the robot processes the request, walks to the table, picks up the watering can, steps over to the shelves, and places it at the destination. The company concedes the limitation plainly: "While our robots have more to advance in movement speed, this is an important step towards the skills needed to complete more complex, real-world tasks that require whole-body coordination."
That matters because human environments are built for human bodies. Tasks in homes and warehouses require reaching, bending, and balancing in tight, cluttered spaces — a profile that tabletop-only arms cannot match.
Dexterity: from 22-DOF hands to commodity grippers
The second advance is fine manipulation. Gemini Robotics 2 can drive the five-fingered, 22 degree-of-freedom SharpaWave hand mounted on Apollo 2 to perform delicate actions such as tying knots or sealing a ziplock bag. The same model also operates standard two-fingered parallel grippers on a Franka Duo platform for tasks like tight packing.
Supporting both anthropomorphic hands and off-the-shelf grippers is a deliberate design choice. It lets hardware partners adopt the model without redesigning their end effectors, and DeepMind says it is "continuing to advance the level of precision and speed to achieve human-level dexterity."
An agent that plans for minutes, not seconds
Real tasks take time. Gemini Robotics ER 2 addresses this by acting as the robot's planner: it observes the room, breaks a user's instruction into steps, delegates each action to the VLA model, and tracks progress until completion. If a step fails, the system can self-correct, and it generalizes to novel situations and goals.
In this release, DeepMind extended the reliable execution horizon. Robots can now run task sequences lasting several minutes and involving hundreds of decisions. The reasoning model also understands when tasks begin and end, and can pinpoint the moment key events occur — which DeepMind characterizes as "a step change in progress understanding."
Multi-robot collaboration is the other new capability here. Different types of robots can now communicate and work together on workflows that no single robot could complete alone — a prerequisite for warehouse, logistics, and eventually household scenarios where specialized machines must share a workspace.
On-device adaptation in hours
The most technically consequential claim concerns embodiment transfer. Retraining a learned policy for a new robot body has historically taken weeks of data collection. Gemini Robotics On-Device 2 is natively multi-embodiment and inherits the "motion transfer" techniques introduced with Gemini Robotics 1.5.
The result: adaptation to a new bi-arm robot in a few hours, typically with fewer than 200 examples. DeepMind says this works even for embodiments with drastically different shapes, sensors, and degrees of freedom, and demonstrated the approach on three platforms — Dexmate, SO101, and Trossen — performing diverse tasks.
Running locally also sidesteps latency and connectivity constraints that rule out cloud inference in many industrial and field settings.
Safety: a new benchmark and a safer reasoning model
As robots gain physical capability, safety becomes the gating factor for deployment. DeepMind says it pairs traditional physical safety measures with AI safety frameworks in every release, and Gemini Robotics 2 advances both.
The company introduced ASIMOV-Agentic, a new benchmark for agentic safety orchestration and uncertainty resolution. It measures whether the embodied reasoning agent will refuse unsafe tool calls from a VLA, predict whether a task is possible, and proactively request human intervention when uncertain.
DeepMind also calls Gemini Robotics ER 2 "our safest robotics model to date" on safety constraint following and human proximity benchmarks. The model better detects nearby humans, triggers safety tool calls, and brings the robot to a safe stop if someone approaches too closely — behavior the company notes is "a key requirement in collaborative safety standards." Full details appear in the Gemini Robotics 2 Safety Technical Report.
The bigger picture
DeepMind positions the release as "an important milestone on the path toward solving AGI in the physical world." The company's stated thesis: unlocking robotics requires moving past single-task automation toward general-purpose intelligence, built as a core layer that hardware makers can adopt.
The partner list signals who Google expects to build on that layer. Apptronik, Boston Dynamics, and Agile Robots are credited as partners on the work, and the Apollo 2 demonstration makes Apptronik the most visible hardware reference so far.
The competitive context is hard to ignore. Figure AI, Tesla, Physical Intelligence, and several well-funded startups are racing to ship general-purpose robot foundation models, and NVIDIA is pushing its own robotics stack. Google's differentiators in this release are specific: full-body humanoid control rather than arms-only, sub-200-example embodiment transfer, verified multi-robot coordination, and a published safety benchmark.
For now, the VLA models remain early-access, which means the claims will be tested by a small circle of hardware partners first. The pace of that testing — and whether the few-hour adaptation claim holds outside DeepMind's demos — will determine whether Gemini Robotics 2 becomes the default intelligence layer for the humanoid industry or another impressive research milestone.
Original: ai.dev
More from James Calloway
Show full bio
News editor covering industry trends and analytics at AI In Context.
144 articles