Gemini Robotics 2 icon

Gemini Robotics 2

Gemini Robotics 2 is Google DeepMind's robotics model family for whole-body control, dexterous manipulation, and task reasoning across different robot embodiments. It includes an on-device option for local execution and adaptation when connectivity or latency is a constraint.

Gemini Robotics 2

Overview

Gemini Robotics 2 is Google DeepMind's robotics model family for physical AI. The announcement centers on three models: Gemini Robotics 2 for vision-language-action control, Gemini Robotics ER 2 for embodied reasoning, and Gemini Robotics On-Device 2 for local execution on robotic hardware.

The system is intended to help robots perceive instructions, reason about tasks, and interact with the physical world across a broad range of embodiments. The source shows use on humanoids, bi-arm robots, hands, and grippers, with examples that include whole-body movement, dexterous object handling, and collaboration between robots.

Capabilities

General robot control

Controls robots of different shapes and sizes, from tabletop systems to full humanoids, so the same model can scale across embodiments.

Vision-language-action control

Translates vision and language input into motor actions, enabling a robot to act on instructions in the physical world.

Whole-body coordination

Extends control to whole-body motion on humanoid robots, including stepping, squatting, bending, and balancing in cluttered spaces.

Dexterous manipulation

Supports dexterous manipulation with hands and grippers, including five-finger, 22-degree-of-freedom hands and standard two-finger grippers.

Multi-step task planning

Uses an embodied reasoning model to plan multi-step tasks, communicate with humans, and track progress over several minutes.

Local deployment and adaptation

Includes an on-device model optimized for local execution and fast adaptation to new robot embodiments with a few hours of data.

Use Cases

  • Whole-body household or workspace tasks

    A humanoid robot can follow a task that requires walking to an object, crouching or bending to reach it, and placing it in a target location in a cluttered room.

  • Dexterous object handling

    A robot with a five-finger hand or parallel gripper can perform delicate manipulation such as tying knots, sealing a ziplock bag, or packing tightly into a container.

  • Long-horizon task execution

    A reasoning layer can break a multi-step instruction into smaller actions, monitor progress, and recover when a step fails during longer tasks.

  • Multi-robot collaboration

    Multiple robots can coordinate on a shared workflow when a single robot cannot complete the job alone, using the reasoning model to organize collaboration.

  • Local robotics deployment

    A robot platform that must run without internet connectivity or low-latency cloud access can use the on-device model for local control and adaptation.

Pros and Cons

Pros

  • Covers a broad set of robot embodiments, including tabletop systems, bi-arm robots, humanoids, hands, and grippers.
  • Combines control and reasoning models so robots can plan, act, and track longer task sequences.
  • Supports whole-body motion that is more suitable for cluttered, human-built environments than static arm-only control.
  • Includes an on-device option for local execution when latency or connectivity are constraints.

Cons

  • The source says whole-body motion is still moving toward higher speed and more advanced manipulation remains challenging, especially for multi-finger dexterous tasks.
  • The public information is incomplete on integrations, deployment details, and general availability beyond early-access and preview access.

FAQ

What is the difference between Gemini Robotics 2, ER 2, and On-Device 2?

Gemini Robotics 2 is a vision-language-action model for robot control. Gemini Robotics ER 2 is the embodied reasoning model that plans multi-step tasks and communicates with humans, while Gemini Robotics On-Device 2 is the local, efficient VLA model for running on robotic hardware without relying on network latency or internet connectivity.

How is Gemini Robotics 2 made available?

Google DeepMind says Gemini Robotics ER 2 is available on Google AI Studio and in private preview on Gemini Enterprise Agent Platform. The VLA and On-Device models are available to early-access partners.

What kinds of robots can it control?

The page describes Gemini Robotics 2 as a model for whole-body control and dexterous manipulation on robots ranging from tabletop robots to full humanoids and bi-arm systems.

What kinds of tasks does it support?

The announcement highlights whole-body movements such as walking, crouching, stretching, balancing, and manipulating objects, as well as dexterous actions with hands and grippers like tying knots, sealing a ziplock bag, and tight packing.

When would the on-device model be useful?

The source says the on-device model is designed for cases that need to run locally, including situations without network latency or internet connectivity. It can adapt to new robot embodiments with a few hours of data.

Quick Facts

Category
World models & physical AI
Product family
Gemini Robotics 2
Source domain
deepmind.google
Primary model types
VLA, embodied reasoning (VLM), on-device VLA
Availability
ER 2 on Google AI Studio and private preview in Gemini Enterprise Agent Platform; VLA and On-Device to early-access partners
Typical robot forms
Tabletop robots, bi-arm robots, humanoids, hands, grippers

Gemini Robotics 2 Alternativen

ByteAsk icon

ByteAsk

ByteAsk is a terminal-first AI coding agent for C and C++ that edits repositories and verifies changes with the real compiler, debugger, sanitizers, and tests before showing a diff. It offers a free tier plus paid plans, with editor connectors and zero-retention handling described in the source.

Lasso icon

Lasso

Lasso is an ecommerce product data platform for enriching catalog records, processing supplier files, generating product content, and monitoring competitors. It combines a web app with a REST API, SDK, and MCP server for teams and developers.

CreateOS Sandbox icon

CreateOS Sandbox

CreateOS Sandbox is an isolated compute environment for running code and agent workloads inside Firecracker micro-VMs. It is designed for workflows that need machine-level isolation, private networking between sandboxes, and programmatic control through SDK, CLI, or MCP.

hob icon

hob

hob is an independent workspace for coding agents that keeps agent sessions, terminals, history, and follow-up work organized around the tools and providers you already use. It is aimed at developers who want local control over routing, history, and workspace structure rather than a bundled model stack.

Ably Chat icon

Ably Chat

Ably Chat is a chat API platform for building custom realtime chat applications. It supports room-based messaging, typing indicators, presence, reactions, and message updates, with usage-based pricing options for different deployment stages.

Botacts icon

Botacts

Botacts is a web directory for finding AI bots and agents you can contact by phone, email, SMS, WhatsApp, Telegram, or Signal. It also lets creators suggest bots for manual review before publication.