Google: Gemini Robotics ER 2 Adds Multi-Robot Collaboration And 91.3% Moment-Finding Accuracy

Google Gemini Robotics ER 2 has been introduced as a new embodied reasoning model designed to help robots understand their surroundings, communicate with people, and coordinate complex physical tasks.

The model functions as a high-level reasoning system that can plan a series of actions and then direct lower-level vision-language-action models or robotics interfaces to perform the physical movements.

Gemini Robotics ER 2 can process continuous streams of video, audio, and text. It can also call external tools, including Google Search and user-defined functions, when additional information is required.

The model is designed to reason about upcoming steps while a robot is already performing its current action, reducing the pauses that can occur when a system alternates between planning and execution.

Gemini Robotics ER 2 represents an upgrade from Gemini Robotics ER 1.6, with stronger capabilities in video understanding, task orchestration, progress tracking, and self-correction.

By analyzing continuous video feeds, robots can monitor their own progress, identify when a task is going wrong, and determine when to advance to the next step.

The model can orchestrate low-level control tools such as navigation systems and vision-language-action models. Developers can configure these tools and stream multimodal information directly into Gemini Robotics ER 2.

Testing showed that the model outperformed Gemini Robotics ER 1.6 in tool orchestration across real-world vision-language-action systems, simulated environments and human-controlled robotic systems.

Gemini Robotics ER 2 also integrates with the Gemini Live API through a bidirectional streaming connection designed for latency-sensitive robotics applications.

A demonstration with Boston Dynamics’ Spot robot showed Gemini Robotics ER 2 directing navigation and manipulator systems to retrieve an object in response to a natural-language command.

One of the model’s primary improvements involves determining how far a robot has progressed through a physical task.

During progress-classification testing, each video frame was assigned to one of five completion ranges, extending from zero to 20% through 80% to 100%.

Gemini Robotics ER 2 achieved 57.4% accuracy on progress-classification tasks, outperforming previous-generation and competing frontier models.

The model also achieved 91.3% accuracy on moment-finding tasks, which measure whether a system can identify the precise video frame when an important event occurs.

Examples could include determining exactly when a container is full, when a light bulb has been tightened or when another physical action has been completed.

Gemini Robotics ER 2 recorded a mean absolute distance of 0.96 seconds in moment-finding evaluations. The developers said it delivered this performance with four times the execution speed of larger model categories.

The model also introduces multi-robot collaboration. Different robots can use a shared semantic understanding to coordinate actions, transfer tasks and complete workflows that a single machine could not perform alone.

One example involves collaboration between Apptronik’s Apollo 2 humanoid robot and a Franka F3 Duo system.

Other improvements include detecting task failures from live video, reading digital displays and other instruments, and answering spatial questions about physical environments.

Safety capabilities were also expanded. Gemini Robotics ER 2 can detect nearby people, stop a humanoid robot when someone enters its working area and resume only after the area is clear.

The model outperformed Gemini Robotics ER 1.6 and other evaluated frontier models on safety instruction-following and human-proximity benchmarks.

Gemini Robotics ER 2 is publicly available to developers through the Gemini API and Google AI Studio. It is also available in private preview through the Gemini Enterprise Agent Platform.