Labs

DeepMind's Gemini Robotics
2 gives humanoids
a full body,
from feet to
fingertips

A three-model release on July 30 extends the Gemini stack from tabletop arms to walking, crouching Apollo humanoids — with a new ASIMOV-Agentic safety benchmark bolted on and a lightbulb-unscrewing success rate of 92%.

Google DeepMind released Gemini Robotics 2 on July 30, a three-model suite that pushes the company’s physical-AI stack past tabletop manipulation and into whole-body humanoid control. The move formalizes what the last twelve months of Apollo demos had been quietly forecasting: DeepMind wants to be the model layer, not the robot maker, and it wants every serious humanoid program building on Gemini before OpenAI or Nvidia’s competing efforts crystallize.

The suite has three parts. Gemini Robotics 2 is the flagship vision-language-action model. Gemini Robotics ER 2 handles embodied reasoning, planning, and (per SiliconANGLE) the ASIMOV-Agentic safety benchmark that checks collision avoidance and whether the reasoner flags impossible tasks or asks a human for help. Gemini Robotics On-Device 2 is a smaller VLA meant to run locally; TNW reports it can be adapted to a new robot body with fewer than 200 examples, in a few hours.

In a pre-taped demo shown to reporters on July 28, Apptronik’s Apollo 2 walked to a table, picked up a watering can, crossed to a set of shelves, and placed it, following the instruction “put the watering can into the green bin in the bottom shelf.” The same model checkpoint, DeepMind says, drives three different embodiments: Apollo 2 with SharpaWave hands, Apollo 2 with Inspire hands, and a Franka Duo with a Robotiq gripper. Hardware-agnostic by design.

The dexterity numbers are where the pitch gets sharper. DeepMind researchers, via Bloomberg, claim a 92% success rate unscrewing a lightbulb. TNW reports the model can drive Apollo’s five-fingered, 22-joint hand to tie knots and seal a ziplock bag. Kanishka Rao, DeepMind’s director of robotics, cautioned that true dexterity remains distant and that robot movements stay slow and deliberate. The caveat is doing real work in the framing.

“Our goal is to bring AI into the physical world and then build the intelligence layer that can be used by every robot,” said Carolina Parada, DeepMind’s vice president of robotics. She told Wired that with humanoids “the safety question is even more pressing because you’re putting them in a lot of other situations.” ER 2, per TNW, halts when a person enters its workspace and resumes only when clear.

Distribution is tiered. ER 2 goes out broadly through the Gemini API, Google AI Studio, and the Gemini Enterprise Agent Platform in private preview. The full VLA and On-Device 2 go to named partners: Apptronik, Boston Dynamics, Agile Robots, and more than 100 trusted testers.

The competitive frame is unmistakable. OpenAI and Nvidia are building their own robot models, and per an Axios report cited by TNW, the US has moved to ban future sales of Chinese-made robots on security grounds. That’s a supply-chain wrinkle for a hardware-agnostic strategy that only pays off if enough Western humanoids reach the floor to run it on.