Google has announced Gemini Robotics 2, a set of multimodal models intended to serve as an intelligence layer for adaptable physical robots. The company says these models convert vision and language inputs into motor control, enabling whole‑body coordination, improved dexterity and multi‑robot collaboration.
Capabilities overview
Gemini Robotics 2 is designed to let robots reason about each movement, enabling a broader range of tasks. Demonstrations include humanoid robots performing household chores — walking, crouching, reaching and manipulating objects to tidy a room — and teams of robots cooperating to complete jobs more efficiently.
The three models
Google describes three core models:
-
Gemini Robotics 2: the company’s most advanced vision–language–action (VLA) model, which maps visual and language signals into motor commands. It is intended to control full humanoids (from feet to fingertips) and bi‑arm robots, and introduces more dexterous manipulation for both hands and grippers.
-
Gemini Robotics ER 2: the most capable embodied reasoning (ER) model, a vision–language model that acts as an agent. It communicates with humans, understands the physical environment, plans multi‑step tasks that can last several minutes, and now supports multi‑robot teamwork.
-
Gemini Robotics On‑Device 2: the most efficient VLA, optimized to run locally on robotic devices. It is natively multi‑embodiment and inherits motion‑transfer techniques from Gemini Robotics 1.5, allowing adaptation to new robot embodiments in a few hours, typically with fewer than 200 examples.
Whole‑body control on humanoids
The update expands prior capabilities — which focused on upper‑body, tabletop tasks — into whole‑body motion control. Google demonstrated an example with Apptronik’s Apollo 2 humanoid: given an instruction such as “put the watering can into the green bin in the bottom shelf,” the robot interprets the command, walks to the table, picks up the watering can, walks to the shelf and places it accurately. Google notes movement speed still needs improvement but highlights whole‑body coordination as a step toward more complex real‑world skills.
Improved dexterity for hands and grippers
Gemini Robotics 2 introduces finer manipulation across different end effectors. The model can control the SharpaWave five‑finger hand (22 degrees of freedom) on Apollo 2 for delicate tasks like tying knots or sealing a ziplock bag. It can also operate two‑finger parallel grippers on a Franka Duo platform for complex dexterous operations such as tight packing. Google states it continues to work on increasing precision and speed toward human‑level dexterity.
Agentic reasoning and multi‑robot collaboration
Gemini Robotics ER 2 functions as a high‑level planner: it processes user instructions, observes the environment, reasons about the steps required, coordinates with the VLA to perform actions, and tracks progress. The update improves the reliability of executing longer sequences that can last minutes and include hundreds of decisions, with better understanding of task boundaries and key event timing.
A new capability is multi‑robot collaboration, where different types of robots can communicate and work together to solve workflows that a single robot could not handle alone.
Fast on‑device adaptation for different robots
Many robotics use cases require operation without network latency. Gemini Robotics On‑Device 2 is optimized for such constraints. The model supports multiple embodiments and, using motion transfer, can adapt to new bi‑arm robot bodies in a few hours with typically under 200 examples. Google shows this working on diverse platforms including Dexmate, SO101 and Trossen.
Safety and responsible development
Google emphasizes safety as foundational. Gemini Robotics 2 adds capabilities to handle real‑world uncertainty and close human collaboration. The company introduces ASIMOV‑Agentic, a benchmark for agentic safety orchestration and uncertainty resolution — for example, measuring whether an embodied reasoning agent can refuse unsafe tool calls from a VLA or request human intervention when uncertain.
Google also reports that Gemini Robotics ER 2 achieves its best results to date on safety constraint following and human‑proximity benchmarks, with improved detection of nearby humans, triggering of safety tool calls and the ability to bring the robot to a safe stop if someone approaches too closely. The company refers readers to the Gemini Robotics 2: Safety Technical Report for more details.
Availability and further information
Gemini Robotics ER 2 is available on Google AI Studio and in private preview on the Gemini Enterprise Agent Platform. The VLA and On‑Device models are available to early‑access partners. Google’s Developer blog provides guidance on deploying the models to hardware.
What this means going forward
Google positions Gemini Robotics 2 as a milestone toward bringing more general‑purpose AI into the physical world, moving beyond single‑task automation to systems that can collaborate with humans and solve more complex, real‑world challenges.



