Models

Gemini Robotics-ER 1.6 adds instrument reading for Boston Dynamics' Spot

Google ships Gemini Robotics-ER 1.6, adding analog gauge reading for Boston Dynamics' Spot and reporting +6% on text and +10% on video injury-risk perception over Gemini 3.0 Flash.

Gemini Robotics-ER 1.6: Powering real-world robotics tasks through enhanced embodied reasoning
Gemini Robotics-ER 1.6: Powering real-world robotics tasks through enhanced embodied reasoningAI-generated
By Sophie Lindqvist4 min read

Updated

Why it matters

  • Gemini Robotics-ER 1.6 released today via the Gemini API and Google AI Studio.
  • New instrument-reading capability co-developed with Boston Dynamics for the Spot quadruped.
  • Outperforms both Gemini Robotics-ER 1.5 and Gemini 3.0 Flash on Google's internal spatial and physical reasoning benchmarks.
  • +6% over Gemini 3.0 Flash on text-based injury-risk perception; +10% on video-based injury-risk perception.
  • Developer Colab and a feedback pipeline accepting 10–50 labeled images per submission are live.

Gemini Robotics-ER 1.6 ships today as Google's most capable embodied-reasoning model yet, the company announced, adding an analog gauge-reading capability built with Boston Dynamics for the Spot quadruped robot.

The model is live now in the Gemini API and Google AI Studio. Google said it improves on both Gemini Robotics-ER 1.5 and Gemini 3.0 Flash in spatial reasoning, multi-view perception, pointing and success detection. The instrument-reading skill is new.

What does the model actually do?

Google built Gemini Robotics-ER 1.6 as a high-level reasoning layer. It calls vision-language-action models, Google Search and developer-defined functions before any robot actuator moves. Google describes the design as "reasoning-first": the model plans, points, counts and verifies outcomes.

Pointing underpins most of those calls. The model marks pixels in an image to express spatial relations, motion trajectories, comparisons and constraints. In one Google example, the model points at "every object small enough to fit inside the blue cup" after reasoning about size.

That pointing drives success detection. The model decides when a task is finished, and whether to retry. Google argues this loop is the difference between a robot that executes blindly and one that adapts.

Why does instrument reading matter?

Industrial plants still rely on analog pressure gauges, thermometers and chemical sight glasses that human operators check by walking the floor. Boston Dynamics' Spot can reach those instruments and photograph them. Reading a tilted needle or a liquid column through curved glass takes layered reasoning.

The model zooms into the image, runs pointing and code execution to estimate proportions, then applies world knowledge to convert a needle angle into a unit-correct value. Google calls this pipeline "agentic vision."

"Capabilities like instrument reading and more reliable task reasoning will enable Spot to see, understand, and react to real-world challenges completely autonomously," Google wrote in its announcement.

The use case points to a commercial direction. Plants could retrofit older facilities with autonomous patrols without ripping out analog infrastructure. Boston Dynamics gains a software reasoner to sell alongside its hardware.

How does it improve on prior models?

Google reports gains in four areas: pointing, multi-view reasoning, instrument reading and safety. The new model beats both Gemini Robotics-ER 1.5 and Gemini 3.0 Flash on the company's internal spatial and physical reasoning benchmarks, though Google has not yet published the full tables.

Multi-view reasoning is the technical jump Google leans on. Robots stream from overhead, wrist and ego cameras at once. The model fuses those feeds into a coherent world state even when targets are occluded or lighting changes.

Agentic vision powers instrument reading. The model writes and runs code to crop, measure and compute, then re-grounds its answer against visual evidence. It treats the image as a working object, not a static input.

What safety gains did Google report?

Google called Gemini Robotics-ER 1.6 its safest robotics model to date. The model improves compliance with Gemini safety policies on adversarial spatial-reasoning tasks compared with all earlier generations.

Specific numbers Google shared:

  • +6% over Gemini 3.0 Flash on text-based injury-risk perception
  • +10% over Gemini 3.0 Flash on video-based injury-risk perception
  • Better adherence to physical safety constraints, including rules like "don't pick up objects heavier than 20kg" and "don't handle liquids"

The gains come from training the model to refuse unsafe grasps and reroute when prompts request hazardous manipulation. Google framed the work as aligning the reasoning layer with gripper and material limits before commands reach motors.

How can developers start?

The model is available today through the Gemini API and Google AI Studio. Google published a Colab notebook with configuration and prompting examples for embodied-reasoning tasks.

Developers who hit capability ceilings can submit 10–50 labeled images showing specific failure modes through a Google form. Google said it would use those submissions to prioritize features in future releases.

The submission pipeline signals an active-learning loop. Google wants to extend instrument reading to other domains where Spot and third-party robots operate, from warehouse picking to surgical assistance.

What does this mean for the embodied-AI race?

Gemini Robotics-ER 1.6 lands amid Google's broader effort to put foundation models in front of physical hardware. Boston Dynamics contributes the deployable platform; Google contributes the high-level planner.

The instrument-reading release matters because a foundation model does useful industrial work today, not in a controlled demo. If Spot reliably logs gauge readings during routine patrols, plant operators gain a template for AI supervision without ripping out existing hardware.

Open questions remain. Google has not disclosed how many Boston Dynamics robots will run the new model in production, nor how the company will price the reasoning layer relative to Gemini API tokens. The Colab and open API access now put pressure on the rest of the robotics field to publish comparable spatial-reasoning benchmarks, where shared evaluation suites are still missing.

For developers, the practical entry point is the Gemini API endpoint combined with the Colab examples. For enterprises, the next milestone is whether instrument reading becomes a one-off showcase or the opening move in a sustained Google push into facility automation.

Original: developers.googleblog.com

Share this article:

More from Sophie Lindqvist

Sophie Lindqvist

Show full bio

Staff writer covering marketplaces and e-commerce at AI In Context.

209 articles

Related articles

  1. Google Launches Gemini Robotics ER 2 With Multi-Robot Collaboration
  2. Google DeepMind's Gemini Robotics 2 Gives Humanoids Whole-Body Control
  3. Google Ships Gemini 3.5 Flash, Promises Pro Model Next Month
  4. Google Upgrades Gemini 3 Deep Think With Record Benchmark Runs
  5. Google Recaps 2025: Gemini 3, AlphaFold Milestones and a Physics Nobel

« Previous article