Humanoid Robotics Technology
hrt awards banner
Home » Google DeepMind Unveils Gemini Robotics: Bringing AI to the Physical World

Google DeepMind Unveils Gemini Robotics: Bringing AI to the Physical World

google-deepmind-unveils-gemini-robotics-for-humanoids

Google DeepMind has introduced Gemini Robotics, a Gemini 2.0-based model designed for robotics.

DeepMind has been making progress in how its Gemini models solve complex problems through multimodal reasoning across text, images, audio, and video. However, these abilities have largely been confined to the digital realm. For AI to be useful and helpful in the physical world, it must demonstrate “embodied” reasoning—the human-like ability to comprehend and react to the surrounding environment—as well as safely take action to complete tasks.

Two new AI models, based on Gemini 2.0, lay the foundation for a new generation of helpful robots.

The first, Gemini Robotics, is an advanced vision-language-action (VLA) model built on Gemini 2.0 with the addition of physical actions as a new output modality, enabling direct robot control. The second, Gemini Robotics-ER, is a Gemini model with advanced spatial understanding, allowing roboticists to run their own programs using Gemini’s embodied reasoning (ER) capabilities.

Both models enable a variety of robots to perform a wider range of real-world tasks than ever before. As part of these efforts, DeepMind is partnering with Apptronik to build the next generation of humanoid robots with Gemini 2.0. Additionally, a select number of trusted testers are providing insights to guide the future of Gemini Robotics-ER.

DeepMind continues to explore the models’ capabilities and develop them further toward real-world applications.

Gemini Robotics: An Advanced Vision-Language-Action Model

For AI models in robotics to be useful and helpful, they must possess three key qualities: generality, interactivity, and dexterity.

Generality

Gemini Robotics leverages Gemini’s world understanding to adapt to novel situations and solve a wide variety of tasks, including those it has never encountered in training. It is adept at handling new objects, diverse instructions, and unfamiliar environments. A technical report shows that Gemini Robotics more than doubles performance on a comprehensive generalization benchmark compared to other state-of-the-art vision-language-action models.

Interactivity

To function in dynamic, physical environments, robots must seamlessly interact with people and their surroundings while adapting to real-time changes.

Built on Gemini 2.0, Gemini Robotics is highly interactive, utilizing advanced language understanding to process and respond to natural language commands in multiple languages. It continuously monitors its surroundings, detects environmental changes, and adjusts its actions accordingly. This level of “steerability” enhances human-robot collaboration across home, workplace, and industrial settings.

Dexterity

A key pillar of effective robotics is dexterous movement. Many everyday tasks requiring fine motor skills remain challenging for robots. Gemini Robotics demonstrates significant progress in this area, performing intricate, multi-step tasks such as folding origami or packing a snack into a Ziploc bag.

Adaptability to Multiple Embodiments

Robots come in diverse forms, and Gemini Robotics was designed to be adaptable across various platforms. Initially trained on the bi-arm robotic platform ALOHA 2, the model has also demonstrated control over a bi-arm platform based on Franka arms used in academic research. Additionally, Gemini Robotics has been specialized for humanoid robots, such as the Apollo robot developed by Apptronik, with the goal of executing real-world tasks.

Gemini Robotics-ER: Enhancing Gemini’s World Understanding

Alongside Gemini Robotics, DeepMind has introduced Gemini Robotics-ER (“embodied reasoning”), an advanced vision-language model designed to enhance Gemini’s spatial reasoning and world understanding for robotics applications. This model enables roboticists to integrate Gemini’s capabilities with their existing low-level controllers.

Gemini Robotics-ER significantly improves Gemini 2.0’s abilities in areas such as pointing and 3D detection. By combining spatial reasoning and code-generation capabilities, it can develop new functionalities dynamically. For instance, when presented with a coffee mug, the model can determine an appropriate two-finger grasp for the handle and generate a safe trajectory for picking it up.

Gemini Robotics-ER operates in an end-to-end manner, covering perception, state estimation, spatial understanding, planning, and code generation. In this setting, it achieves a 2x-3x success rate compared to Gemini 2.0. Additionally, when direct code generation is insufficient, Gemini Robotics-ER can leverage in-context learning, using a small set of human demonstrations to infer the correct solution.

The model excels in embodied reasoning tasks, such as object detection, pointing, multi-view correspondence, and 3D object recognition, further enhancing robotic capabilities.

Advancing AI and Robotics Responsibly

As AI-driven robotics continues to evolve, DeepMind is implementing a layered, holistic approach to research safety, addressing both low-level motor control and high-level semantic understanding.

Physical safety in robotics has long been a foundational concern, with traditional safeguards such as collision avoidance, force limits, and dynamic stability controls. Gemini Robotics-ER can interface with these safety-critical controllers while also leveraging Gemini’s core safety features to assess whether an action is safe in a given context.

To further robotics safety research, DeepMind is releasing a new dataset aimed at evaluating and improving semantic safety in embodied AI. Previous research introduced a “Robot Constitution” inspired by Isaac Asimov’s Three Laws of Robotics, guiding AI models toward safer decision-making. This approach has since evolved into an automated, data-driven framework for generating constitutions—rules expressed in natural language that steer a robot’s behavior. The new ASIMOV dataset will help researchers rigorously measure safety considerations in real-world robotic applications.

DeepMind collaborates with experts in its Responsible Development and Innovation team, as well as the Responsibility and Safety Council, an internal review group dedicated to the ethical advancement of AI. External specialists also provide insights on the broader implications of embodied AI in robotics.

Beyond its partnership with Apptronik, Gemini Robotics-ER is being evaluated by trusted testers, including Agile Robots, Agility Robotics, Boston Dynamics, and Enchanted Tools. These collaborations will further refine the model’s capabilities and pave the way for the next generation of AI-driven robotics.

SOURCE: Google DeepMind

HRT Badge

Humanoid Robotics Technology

  • 11701 Stonehollow 4, STE 150Austin, TX 78758
  • (512) 937-2112
Visit Website View Profile
Join thousands of Humanoid Robotics Experts and get the latest updates straight to your inbox!