Key takeaways
- From feet to fingertips — we are teaching robots intelligent whole-body control, fine dexterity, and teamwork to complete a broad range of…
- They lack the ability to truly learn for themselves or adapt to unpredictable environments.
- We demonstrated how Gemini's multimodal understanding could drive real-world action with Gemini Robotics.
What happened
From feet to fingertips — we are teaching robots intelligent whole-body control, fine dexterity, and teamwork to complete a broad range of complex tasks For decades, we’ve dreamed of robots that can seamlessly step into our world and lend a hand. Now, that vision takes a significant stride forward. Most robots are pre-programmed or teleoperated for narrow, repetitive task sequences.
Gemini Robotics 2 unlocks a new level of physical dexterity across different end effectors, whether a robot is using hands or grippers, enabling robots to be more useful than ever before. The model can now control the five-fingered, 22 degree-of-freedom SharpaWave hand on the Apollo 2 robot to complete delicate actions like tying knots or sealing a ziplock bag. g. tight packing).
We are continuing to advance the level of precision and speed to achieve human-level dexterity. Most real-world tasks require multiple steps over an extended period of time. To manage this complexity, our embodied reasoning (ER) model, Gemini Robotics ER 2, serves as the robot’s high-level brain, processing user instructions and communicating with humans.
It observes the room, reasons about the steps needed to complete the task, coordinates with the VLA to carry out the actions, and tracks progress until the task is done. This setup allows robots to execute complex multi-step tasks, self-correct if a step fails, and generalize to novel situations and goals.
In this update, we are enabling robots to more reliably execute longer task sequences, lasting several minutes and involving hundreds of decisions. Gemini Robotics ER 2 now understands when tasks begin and end, and can pinpoint the moment key events occur, marking a step change in progress understanding. Furthermore, we are introducing multi-robot collaboration.
Why it matters
They lack the ability to truly learn for themselves or adapt to unpredictable environments. Moreover, transferring learned skills from one robot body to another remains incredibly difficult. To take on the hardest problems at scale, robots of every shape and size need AI models giving them the ability to think, act, and interact intelligently to safely complete tasks.
We demonstrated how Gemini's multimodal understanding could drive real-world action with Gemini Robotics. Today, we are introducing Gemini Robotics 2 - the intelligence layer powering the next generation of truly adaptable robots. As it takes its first literal steps, this major advance unlocks intelligent whole-body control, advanced dexterity, and multi-robot collaboration. Gemini Robotics 2 enables robots to reason through every movement, unlocking a broad range of tasks.
For example, it can enable a humanoid to walk, crouch, stretch, and manipulate objects to clean up a cluttered room. It can even team up with other robots to finish the job faster. And this profound intelligence can also run locally on-device while seamlessly adapting to entirely new robotic bodies in just a few hours.
Gemini Robotics ER 2, our reasoning model, is now available on Google AI Studio and in private preview on Gemini Enterprise Agent Platform. Our VLA and On-Device models are available to early-access partners. Read how to bring these models to your hardware on our Developer blog. The world is built for human movements; it requires us to reach, bend, and balance in tight, cluttered spaces.
While our previous models controlled the humanoid’s upper-body to achieve table-top tasks, Gemini Robotics 2 expands physical AI into whole-body motions. For the first time, our model can now control entire humanoid robots, translating intent into intelligent whole-body control. ” Apollo processes the instruction, walks to the table, and picks up the watering can, takes a few steps to the shelves, and places it precisely in its destination.
While our robots have more to advance in movement speed, this is an important step towards the skills needed to complete more complex, real-world tasks that require whole-body coordination. To be genuinely useful in our homes and workplaces, robots need finesse.
What to watch
This enables different types of robots to communicate and work together to solve complex workflows a single robot could not do alone. Many robotic applications need to operate without network latency or internet connectivity. Gemini Robotics On-Device 2 is built specifically to handle these constraints — it is our most-efficient vision-language-action model (VLA) optimized to run locally on robotic devices. 5.
We can now adapt to new bi-arm robot embodiments with just a few hours of adaptation time, typically with less than 200 examples. This works even with new embodiments with drastically different shapes, sensors and degrees of freedom, as shown below with a diverse set of tasks being performed by the Dexmate, SO101, and Trossen platforms. Safety is foundational to our robotics research. As robots gain more physical capabilities, we are committed to ensuring end-to-end safety and alignment.


