ROBOTNESS
Product5 min readROBOTNESS DeskUnited States

Google DeepMind ships Gemini Robotics 2, taking its robot models to whole-body humanoid control

Google DeepMind on July 30 released Gemini Robotics 2, a family of three models covering whole-body humanoid control, high-level embodied reasoning and on-device inference. The reasoning model ER 2 is available to developers through the Gemini API and Google AI Studio, while the action models remain limited to early-access partners and trusted testers.

Summary

Google DeepMind has moved its robotics models from tabletop arms to full humanoid bodies. On July 30, 2026 the Alphabet research unit released Gemini Robotics 2, a three-model family that it says can drive a humanoid through walking, crouching, reaching and manipulating within a single policy, and that can be adapted to a new two-armed robot with a few hours of data.

The family has three members. Gemini Robotics 2 is a vision-language-action (VLA) model that turns camera input and natural-language instructions into motor commands for humanoid and bi-arm robots. Gemini Robotics ER 2 is an embodied-reasoning vision-language model meant to sit above the action layer as a planner and coordinator. Gemini Robotics On-Device 2 is a smaller VLA designed to run locally on the robot. According to the model card, On-Device 2 is built on Gemini Robotics 1.5 technology and on-device Gemma models and was trained on Google's Tensor Processing Units.

DeepMind published task-level success rates rather than a single headline score. On Apptronik's Apollo 2 humanoid, whole-body picking succeeded 68.4% of the time from a table, 45.7% from the floor and 76.3% from a shelf. With the 22-degree-of-freedom SharpaWave hand fitted to Apollo, the model unscrewed a light bulb in 92% of trials but screwed one in only 36%, tied a trash bag in 44%, used a dustpan in 32% and handled a ziplock bag in 40%. On a Franka Duo with Robotiq grippers it reached 89.6% on precise insertion, 78.9% on diverse tool kitting and 74.2% on general pick and place.

Gemini Robotics 2: published task success rates
  • Pick from table
    Robot platform
    Apptronik Apollo 2
    Success rate (%)
    68.4
  • Pick from floor
    Robot platform
    Apptronik Apollo 2
    Success rate (%)
    45.7
  • Pick from shelf
    Robot platform
    Apptronik Apollo 2
    Success rate (%)
    76.3
  • Unscrew bulb
    Robot platform
    Apollo + SharpaWave hand
    Success rate (%)
    92
  • Screw bulb
    Robot platform
    Apollo + SharpaWave hand
    Success rate (%)
    36
  • Tie trash bag
    Robot platform
    Apollo + SharpaWave hand
    Success rate (%)
    44
  • Precise insertion
    Robot platform
    Franka Duo
    Success rate (%)
    89.6
  • Diverse tool kitting
    Robot platform
    Franka Duo
    Success rate (%)
    78.9
  • General pick and place
    Robot platform
    Franka Duo
    Success rate (%)
    74.2

Success rates as disclosed by Google DeepMind on July 30, 2026; selected tasks, no estimates.

As of Oct 1, 2026

The company said the VLA can carry out multi-step jobs lasting several minutes and involving hundreds of decisions. It also said the model adapts to new bi-arm embodiments with a few hours of adaptation time, typically with fewer than 200 examples, and named Dexmate, SO101 and Trossen platforms as the test cases. For the on-device variant, the model card reports success on SO101 tasks rising to 53.3% from 6.7% for the first version, and on Dexmate tasks to 75.6% from 33.3%, after extended training.

ER 2 is the part developers can touch today. Google said it is available through the Gemini API and Google AI Studio, with a private preview on the Gemini Enterprise Agent Platform and code samples on GitHub. The model watches continuous video rather than still frames to judge whether a step has succeeded, scores 57.4% on sorting video frames into five task-progress bands and 91.3% on locating a specific moment in a video, with a mean error of 0.96 seconds. Google said it reads ten types of instruments, from digital displays to liquid thermometers, calls tools such as Google Search natively and streams through the Gemini Live API. The VLA and On-Device models are offered only through an early-access partner program and a trusted tester sign-up.

DeepMind thanked Apptronik, Boston Dynamics and Agile Robots as partners. Google's demonstrations showed ER 2 directing a Boston Dynamics Spot to fetch objects by voice command and coordinating Apollo 2 with a Franka F3 Duo arm in a shared task, which the company described as robots handing off work through a common semantic understanding. Apptronik, founded in 2016, closed a Series A extension of $520 million on February 11, 2026, according to its own announcement.

The release lands in a crowded quarter for robot foundation models. NVIDIA added commercial licensing to its open Isaac GR00T N1.7 humanoid model in April and put a 4-billion-parameter Cosmos 3 Edge world model on Hugging Face on July 20. Skild AI introduced its S1 model on August 18. Physical Intelligence, which raised a Series B in November 2025, has published a series of VLA releases since its first π0 policy in October 2024. DeepMind's distinguishing choice is to split reasoning and action into separate models and to expose only the reasoning layer as a paid cloud product.

On-Device model, version 1 vs version 2 on new robot platforms
  • SO101
    On-Device v1 success (%)
    6.7
    On-Device 2 success (%)
    53.3
  • Dexmate
    On-Device v1 success (%)
    33.3
    On-Device 2 success (%)
    75.6

Figures from the Gemini Robotics On-Device 2 model card, after extended training; disclosed, not estimated.

As of Oct 1, 2026

Technically, the split mirrors how a human supervisor and a skilled operator divide labour. ER 2 decides what should happen next, checks progress against video and refuses unsafe tool calls. The VLA translates each instruction into joint-level motion at control rates a cloud model cannot reach. DeepMind introduced a new benchmark, ASIMOV-Agentic, to measure whether the reasoning agent declines unsafe tool calls and correctly judges whether a task is possible, and said ER 2 stops a humanoid when a person comes close.

The open questions are reliability and access. Success rates between 32% and 92% on household-style dexterous tasks are research numbers, not factory uptime figures. The On-Device model card itself warns that the model is limited in generalising to out-of-distribution tasks and in controlling high-degree-of-freedom robots, and that mobile and whole-body risks are outside its current scope. Google has not disclosed pricing for ER 2 or a date for general availability of the action models.

ROBOTNESS analysis

Google is positioning Gemini as the reasoning layer every robot maker can rent, while keeping the motor layer as a partner-only asset that builds lock-in.

The evidence is in the distribution choices. ER 2 is on the same API and AI Studio surface as Google's consumer and enterprise Gemini models, while the VLA and On-Device models sit behind trusted tester gates. The partner list of Apptronik, Boston Dynamics and Agile Robots gives Google demonstrations on humanoid, legged and industrial arm hardware without building its own robot.

The strongest counter-argument is that open models are moving faster on the motor layer. NVIDIA ships GR00T N1.7 weights with a commercial licence, and a robot maker that wants control of its stack may prefer weights it can fine-tune on its own servers over a planner that lives in Google's cloud.

Bull case: ER 2 becomes the default task planner for mixed fleets because it already handles video, tools and safety refusals, and robot makers pay per call much as software developers pay for language model APIs. The action models then graduate from trusted tester to general release with Apptronik as the reference humanoid.

Bear case: latency, connectivity and data-sovereignty concerns keep factories from routing robot decisions through a cloud model. The 36% bulb-screwing and 32% dustpan figures show dexterity remains far from production grade, and partners hedge by running open models.

Signals to watch:

  • Whether Google publishes ER 2 pricing or a general-availability date for the Gemini Robotics 2 VLA before the end of 2026.
  • Apptronik disclosures on Apollo 2 commercial deployments that name Gemini Robotics as the control stack.
  • New ASIMOV-Agentic results from other labs, which would show whether the benchmark becomes an industry yardstick.
Key facts
Models
Gemini Robotics 2 (VLA), Gemini Robotics ER 2, Gemini Robotics On-Device 2
Announced
July 30, 2026
Named partners
Apptronik, Boston Dynamics, Agile Robots
VLA availability
Early-access partners and trusted testers only
ER 2 availability
Gemini API, Google AI Studio; private preview on Gemini Enterprise Agent Platform
New safety benchmark
ASIMOV-Agentic
Best dexterous result
92% unscrew bulb with 22-DoF SharpaWave hand
Adaptation to new bi-arm robot
A few hours, typically fewer than 200 examples
Whole-body pick success (Apollo 2)
68.4% table, 45.7% floor, 76.3% shelf
Sources
ROBOTNESS Intelligence
  1. 01

    Why it matters

    With Gemini Robotics 2, Google DeepMind publishes whole-body humanoid results from a learned policy and offers its robot reasoning model through the same API as its mainstream Gemini models. That combination turns robotics from a research showcase into a product line with a developer funnel.

    For robot makers, the practical change is that a high-level planner with video understanding, tool calling and safety refusal can now be called from code without training one in house. Startups that lack the budget to build a reasoning model can wire ER 2 into an existing low-level controller and focus spending on hardware and data.

  2. 02

    Rival analysis

    NVIDIA's approach is the opposite on openness. GR00T N1.7 ships as open weights with a licence that NVIDIA says supports production use, and Cosmos 3 Edge runs on Jetson Thor. NVIDIA earns on the compute rather than the model. Google earns on API calls and cloud, and keeps the motor-control weights closed.

    Skild AI and Physical Intelligence sit between the two. Skild sells deployed solutions and reports $100 million in annual recurring revenue; Physical Intelligence has open-sourced π0 and publishes research openly. Google's advantage over both is that ER 2 sits on the same Gemini API and AI Studio surface developers already use for its general models; its disadvantage is that it does not build or sell robots and depends on partners such as Apptronik for embodiment.

  3. 03

    Valuation context

    Alphabet does not break out DeepMind robotics revenue, so the release cannot be valued directly. The more useful lens is partner capital. Apptronik, the main humanoid partner, closed a $520 million Series A extension on February 11, 2026, after a $350 million Series A in February 2025, according to its own releases.

    Skild AI was valued at $14 billion post-money in its January 2026 Series C and Figure AI at $39 billion in September 2025. Those figures set the private-market price for owning a robot brain. Google's choice to rent rather than own the hardware layer means it competes for that value through software margin instead of equity in a robot maker.

  4. 04

    Supply-chain implications

    The demonstrations name specific hardware: Apptronik Apollo 2, the 22-degree-of-freedom SharpaWave hand, Inspire hands, Franka Duo arms with Robotiq grippers, Boston Dynamics Spot, Dexmate, SO101 and Trossen platforms. Hand suppliers benefit most, because the published dexterity tasks are where the gap between 92% and 32% success is widest.

    On compute, the model card says training ran on Google TPUs. The on-device variant is aimed at local robot hardware, but Google has not named a target chip. That leaves room for NVIDIA Jetson, Qualcomm or custom silicon on the robot, and is an open procurement question for any integrator adopting the stack.

  5. 05

    Signals to watch

    First, any pricing page for Gemini Robotics ER 2 on the Gemini API will show whether Google treats robot reasoning as a premium tier or a loss leader. Second, a general-availability notice for the Gemini Robotics 2 VLA would mark the shift from trusted tester to commercial product.

    Third, watch for Apptronik or Agile Robots naming Gemini Robotics in customer deployments, and for Boston Dynamics product notes that reference ER 2 on Spot. Fourth, third-party results on the ASIMOV-Agentic benchmark would show whether Google can set the safety yardstick for embodied agents.

  6. 06

    Analyst view

    Thesis: Google is building a cloud reasoning franchise for robots, and ER 2 is the wedge. Confidence: medium. The API availability and the partner roster are confirmed on Google's own pages, which supports the thesis. Confidence is not high because Google has disclosed no pricing, no customer count and no timeline for the action models.

    The dexterity numbers argue for patience. A 36% success rate on screwing a bulb is a credible research result but would not pass a factory acceptance test. We expect the reasoning layer to monetise well before the action layer does.

  7. 07

    Questions you should be asking

    Will Google release the Gemini Robotics 2 VLA weights to partners for on-premises use, or only as a hosted service? How will ER 2 be priced per video minute or per call, and will robot fleets get volume terms?

    Which chip is On-Device 2 optimised for, and what latency does it reach on that hardware? Does the partnership with Apptronik include exclusivity on humanoid form factors, and what role do Boston Dynamics and Agile Robots play beyond demonstrations?