Skip to content
KoishiAI
ไทย
← Back to all articles
gemini-robotics deepmind robotics-ai embodied-ai multi-robot

Google DeepMind Gemini Robotics ER 2: Multi-Robot AI

Discover Google DeepMind's Gemini Robotics ER 2. This AI enables multi-robot collaboration and real-time video understanding for complex task orchestration.

AI-drafted from cited sources, fact-checked and reviewed by a human editor. How we work · Standards · Report an error
High-tech automated warehouse system featuring a green robotic arm handling blue storage crates.
Photo by Peter Xie on Pexels

TL;DR: Google DeepMind released Gemini Robotics ER 2 on July 30, 2026, an AI model that enables robots to plan tasks and collaborate in real-time using video understanding. The system achieves 91.3% accuracy in moment-finding and runs four times faster than competing models, allowing multiple robot types to divide complex workflows efficiently.

Key facts

  • Google DeepMind launched Gemini Robotics ER 2 on July 30, 2026, as a high-level ‘embodied reasoning’ model for multi-step task planning.
  • ER 2 achieves 91.3% accuracy on moment-finding tasks with a mean absolute distance error of 0.96 seconds.
  • The model reaches 57.4% accuracy in continuous progress classification by binning video frames into five completion bands.
  • Gemini Robotics ER 2 runs at four times the execution speed of larger competing models on identical benchmarks.
  • ER 2 enables multi-robot collaboration, allowing diverse robot types like humanoids and rovers to share semantic understanding and divide workflows.
  • The system uses a bidirectional streaming endpoint via the Gemini Live API to plan upcoming steps while performing current actions.
  • ER 2 is available via the Gemini API, Google AI Studio, and in private preview on the Gemini Enterprise Agent Platform.

Google DeepMind has released Gemini Robotics ER 2, an advanced AI model designed to serve as the central planning brain for physical robots. Launched on July 30, 2026, this update represents a significant shift from simple object manipulation to complex task orchestration and multi-robot collaboration [1][2]. By integrating high-level reasoning with real-time video understanding, ER 2 allows robots to plan multiple steps ahead while simultaneously executing physical actions, effectively eliminating the ‘stop-and-think’ delays that plagued earlier systems [4].

The core innovation of ER 2 lies in its ability to process continuous streams of visual and audio data to maintain situational awareness. Unlike previous models that required discrete inputs for each action, ER 2 uses a bidirectional streaming endpoint via the Gemini Live API [1][5]. This architecture enables the robot to ‘think’ about upcoming steps while performing current tasks, creating a fluid workflow where perception and planning happen in parallel rather than sequentially [2]. The model can also natively call external tools, such as Google Search or user-defined functions, to enhance its decision-making process without interrupting physical operations [1][5].

A critical advancement is the model’s improved temporal intelligence. ER 2 achieves 91.3% accuracy on moment-finding tasks, identifying specific video frames where critical events occur with a mean absolute distance error of just 0.96 seconds [2][4]. This precision allows robots to track progress continuously by binning live video frames into five completion bands (0-20%, 20-40%, 40-60%, 60-80%, and 80-100%), achieving a 57.4% accuracy rate in classifying task status [2][4]. This capability enables self-correction; if a step fails, the robot can retry without restarting the entire workflow, adapting to novel situations on the fly [4].

The system also introduces robust multi-robot collaboration capabilities. ER 2 allows diverse robot types, such as wheeled rovers and humanoids, to communicate through shared semantic understanding [5][6]. This enables them to divide complex workflows dynamically, with each unit handling specific sub-tasks while contributing to a larger goal [7]. The model operates at four times the execution speed of larger competing models on similar benchmarks, ensuring that high-level planning does not bottleneck physical performance [4].

Safety is integrated directly into the reasoning process. ER 2 includes protocols to halt humanoid robots when humans are nearby and resume operations only after the area is clear [2]. This feature significantly outperforms its predecessor, ER 1.6, on safety instruction benchmarks, addressing one of the primary concerns in deploying AI in shared human environments [2].

ER 2 is part of the broader ‘Gemini Robotics 2’ suite, which also includes a Vision-Language-Action (VLA) model for whole-body control and an On-Device 2 model for local adaptation [6][7]. While ER 2 focuses on high-level planning and coordination, it works in tandem with these other models to move physical AI beyond tabletop tasks into complex real-world environments like logistics and manufacturing [8].

Developers can access Gemini Robotics ER 2 via the Gemini API, Google AI Studio, and through a private preview on the Gemini Enterprise Agent Platform [1][4][8]. The release marks a pivotal moment in embodied AI, transitioning robots from isolated agents to coordinated components of larger, intelligent systems.

Sources

  1. Introducing Gemini Robotics ER 2 (deepmind.google) — 2026-07-30
  2. Google DeepMind Launches Gemini Robotics ER 2 With Multi-Robot Collaboration (www.iclarified.com) — 2026-07-30
  3. Google DeepMind ships Gemini Robotics ER 2 as a video-native brain for robots (www.aichatdaily.com) — 2026-07-30
  4. Gemini Robotics ER 2 (deepmind.google) — 2000-01-01
  5. Google DeepMind (@GoogleDeepMind) on X (x.com) — 2026-07-30
  6. Gemini Robotics 2 brings whole body intelligence to robots (deepmind.google) — 2026-07-30
  7. Introducing Gemini Robotics ER 2 (blog.google) — 2026-07-30

Frequently asked questions

When was Google DeepMind's Gemini Robotics ER 2 released?
Google DeepMind released Gemini Robotics ER 2 on July 30, 2026. This AI model serves as a central planning brain for physical robots, enabling multi-step task orchestration and real-time collaboration.
How accurate is Gemini Robotics ER 2 at understanding video and timing?
The system achieves 91.3% accuracy on moment-finding tasks with a mean absolute distance error of just 0.96 seconds. It also reaches 57.4% accuracy in continuous progress classification by binning video frames into five completion bands.
How does ER 2 enable real-time planning without stopping?
ER 2 uses a bidirectional streaming endpoint via the Gemini Live API to process visual and audio data continuously. This allows robots to plan upcoming steps while simultaneously executing current physical actions, eliminating traditional 'stop-and-think' delays.
Can different types of robots collaborate using Gemini Robotics ER 2?
ER 2 allows diverse robot types, such as humanoids and rovers, to share semantic understanding and divide complex workflows dynamically. It runs at four times the execution speed of larger competing models on identical benchmarks.
How does Gemini Robotics ER 2 handle safety around humans?
ER 2 includes safety protocols that automatically halt humanoid robots when humans are nearby and resume operations only after the area is clear. This feature significantly outperforms its predecessor, ER 1.6, on safety instruction benchmarks.
Where can I access or use Gemini Robotics ER 2?
The model is available via the Gemini API and Google AI Studio. It is also accessible in private preview on the Gemini Enterprise Agent Platform for enterprise users.