Skip to content

Gemini Robotics 2 puts whole-body control into humanoids

Original: Gemini Robotics 2 brings whole-body control to humanoids View original →

Read in other languages: 한국어日本語
Humanoid Robots Aug 1, 2026 By Insights AI (Twitter) 1 min read 1 views Source

A robotics model aimed at many bodies

The hard part of robot AI is often not language understanding, but reliable control of an entire body in the physical world. On July 30, 2026, Google DeepMind framed Gemini Robotics 2 around that problem. The key tweet said: One brain. For any robot. Gemini Robotics 2 brings full body intelligence to humanoids, advanced dexterity, multi-robot teamwork and more. The source post is available on X.

Follow-up material made the architecture more concrete. DeepMind described three models: Gemini Robotics 2, a vision-language-action model for controlling humanoids from feet to fingertips; Gemini Robotics ER 2, a model for real-world video understanding and complex multi-step planning; and On-Device 2, which runs locally and can adapt to new robot bodies in a few hours. The accompanying DeepMind blog, published the same day, argues that useful robots need models that can think, act, and interact safely in unpredictable environments rather than execute narrow pre-programmed sequences.

Google DeepMind’s account is one of the primary channels for research and product results around Gemini, robotics, and science models. The important shift here is not just that a humanoid can perform a staged task. The company is presenting dexterous hand control, whole-body movement, and coordination between different robot types as one system. Related posts cite examples such as tying a knot, screwing in a lightbulb, packing objects with grippers, and having Apollo and Duo divide work in a messy garage.

The next thing to watch is generalization outside polished demos. DeepMind will need to show how far the “few hours” adaptation claim extends across hardware, how multi-robot teamwork recovers from mistakes, and whether the ER 2 planning layer can keep long household or workplace tasks stable when objects, lighting, and human interruptions change mid-task.

Share: Long

Related Articles