ROBOTNESS
Industry8 min readROBOTNESS DeskUnited States

What is a vision-language-action (VLA) model, and who is building them?

A vision-language-action (VLA) model is one neural network that reads camera images and a plain-language instruction and outputs a robot's next motor commands, built by training a vision-language model on robot demonstrations. Google DeepMind, Physical Intelligence, NVIDIA, Figure AI, Skild AI and AgiBot lead, and the commercial race now turns on the cost of action data rather than model design.

What is a vision-language-action (VLA) model, and who is building them? (Illustration by ROBOTNESS)
Summary

A vision-language-action (VLA) model is a single neural network that takes camera images and a plain-language instruction and returns the motor commands a robot should execute next. It is made by taking a vision-language model, the kind of AI that can describe a photo, and training it further on robot demonstrations so that its output is movement rather than text.

The term entered robotics in July 2023, when Google DeepMind described RT-2 in an arXiv preprint and reported that a model trained on web images and robot data roughly doubled success on unseen objects and scenes against its earlier RT-1 controller, at about 62% versus 32%. Three years later the label covers models from Google DeepMind, Physical Intelligence, NVIDIA, Figure AI, Skild AI and AgiBot, and the companies behind them have raised some of the largest private rounds in robotics. This explainer sets out how the models work, how they differ from older robot software and from world models, who leads, and why the cost of data, more than model design, now decides the commercial race.

How a VLA turns pixels and words into motor commands

Most VLAs have three parts. A vision encoder turns each camera frame into a sequence of visual tokens. A language backbone, the pretrained vision-language model, reads those tokens together with the instruction and the robot's own joint readings. An action head then converts the backbone's internal state into numbers for joints, grippers or wheels. NVIDIA's model card for GR00T N1 lays this out plainly: a SigLIP 2 vision transformer reading 224 by 224 pixel frames, a T5 text encoder, a small network for the robot's state that changes with each robot body, and a diffusion transformer that produces the actions.

Starting from a vision-language model buys transfer. A backbone trained on vast numbers of captioned web images already knows what most household and factory objects look like, so a robot can act on things it never saw during robot training. The RT-2 authors reported that this inherited knowledge let a robot pick the smallest object on a table or move a can onto a number it had only met on the web, with success on such emergent tasks of 60% against 17% for RT-1. The harder question is how the network should express a movement at all.

Tokens or a flow: two ways to output an action

The first VLAs treated movement as words. RT-2 split each of eight action dimensions, six for the gripper's position and rotation, one for opening the gripper and one for ending the task, into 256 bins, so a motion became a short string of tokens that the language model emitted like text, according to the paper. OpenVLA, a 7-billion-parameter model from a Stanford-led academic team released with open weights in June 2024, used the same approach and reported a success rate 16.5 percentage points higher than Google's 55-billion-parameter RT-2-X across 29 tasks.

Tokens are slow and coarse for dexterous work. The largest RT-2 ran at 1 to 3 Hz, its paper says,. Physical Intelligence's π0, described in an October 31, 2024 preprint, attached a separate 300-million-parameter action expert to a 3-billion-parameter PaliGemma backbone and trained it with flow matching, a relative of the diffusion method behind image generators. The expert starts from random noise and refines it over 10 steps into a chunk of 50 future actions, which lets π0 control robots at up to 50 Hz. GR00T N1 does the same job with a diffusion transformer.

A third design separates thinking from reflexes. Figure AI's Helix, announced on February 20, 2025, pairs a 7-billion-parameter vision-language model running at 7 to 9 Hz with an 80-million-parameter policy running at 200 Hz, steering 35 degrees of freedom on two onboard GPUs, the company said. NVIDIA calls the same split System 2 and System 1, and AgiBot's GO-2, released on September 8, 2026, uses an asynchronous dual system in which a slower semantic planner hands action sequences to a faster module that corrects for noise. That layering is also what separates a VLA from the robot software factories already run.

Why a VLA is neither robot programming nor a world model

Classical industrial robots are programmed point by point. An integrator teaches waypoints with a pendant or writes the path offline, and the program repeats unchanged until a part shifts or the line is retooled. A VLA replaces that script with a learned policy that reads the scene every cycle, so it can cope with clutter and new objects, but it gives up the strict repeatability that makes scripted cells easy to certify. Japanese makers are trying to keep both. Yaskawa Electric said in July 2026 that its MOTOMAN NEXT system uses Google DeepMind's Gemini Robotics ER 1.6 to plan tasks from high-level instructions without detailed teaching, while the robot's own vision, path planning and force sensing stay underneath.

A world model answers a different question. It predicts how a scene will change, usually as video, given what the robot does, and it is used to generate synthetic training data or to test policies before they touch hardware. Runway said on September 30, 2026 that policy results inside its world model correlate with real-world results at 0.95. The two ideas are converging. Runway calls its new Praxis-1 a world action model that outputs robot control, and Generalist AI pretrains GEN-1 on more than 500,000 hours of human interaction data recorded with wearables, with no robot data at all, according to the company.

Who builds VLAs, and who gives them away

Google DeepMind launched Gemini Robotics, built on Gemini 2.0, on March 12, 2025 with Apptronik as its humanoid partner, and on July 30, 2026 released Gemini Robotics 2 for whole-body humanoid control; its action models remain limited to partners and trusted testers. Physical Intelligence went the other way. Its openpi repository carries π0, π0-FAST and, since September 2025, π0.5 under an Apache 2.0 licence. NVIDIA published GR00T N1 on March 18, 2025 under a non-commercial licence, and its 3-billion-parameter GR00T N1.7 ships under the NVIDIA Open Model License inside Hugging Face's LeRobot library. Figure keeps Helix in-house, Skild AI sells its Skild Brain and S1 models through deployments, and AgiBot delivers GO-2 through its Genie Studio platform.

Vision-language-action models compared
  • RT-2
    Developer
    Google DeepMind
    First public release
    2023-07-28
    Parameters (bn)
    55
    Open weights
    No (partners only)
  • OpenVLA
    Developer
    Stanford-led team
    First public release
    2024-06-13
    Parameters (bn)
    7
    Open weights
    Yes
  • π0
    Developer
    Physical Intelligence
    First public release
    2024-10-31
    Parameters (bn)
    3.3
    Open weights
    Yes
  • Helix
    Developer
    Figure AI
    First public release
    2025-02-20
    Parameters (bn)
    7.08
    Open weights
    No data
  • Gemini Robotics
    Developer
    Google DeepMind
    First public release
    2025-03-12
    Parameters (bn)
    No data
    Open weights
    No (partners only)
  • GR00T N1
    Developer
    NVIDIA
    First public release
    2025-03-18
    Parameters (bn)
    2
    Open weights
    Yes (non-commercial)
  • π0.5
    Developer
    Physical Intelligence
    First public release
    2025-04-22
    Parameters (bn)
    No data
    Open weights
    Yes
  • Skild Brain
    Developer
    Skild AI
    First public release
    2025-07-29
    Parameters (bn)
    No data
    Open weights
    No data
  • Gemini Robotics 2
    Developer
    Google DeepMind
    First public release
    2026-07-30
    Parameters (bn)
    No data
    Open weights
    No (partners only)
  • S1
    Developer
    Skild AI
    First public release
    2026-08-18
    Parameters (bn)
    No data
    Open weights
    No data
  • GO-2
    Developer
    AgiBot
    First public release
    2026-09-08
    Parameters (bn)
    No data
    Open weights
    No data

First public release = first paper, model card or company post. Helix = 7 bn System 2 plus 80 m System 1. π0.5 weights via openpi since September 2025. null = not disclosed. Figures disclosed, not estimated.

As of Oct 1, 2026

Openness is a business decision. Open weights let universities and small hardware makers adapt a model on one workstation; openpi says low-rank fine-tuning needs more than 22.5 GB of GPU memory, within reach of an RTX 4090. Closed models keep the data flywheel and the customer relationship with the owner. Where the two camps really part is in how they pay for data.

Data, not architecture, is the cost line

Unlike chatbots, VLAs cannot scrape their training data from the web. Robot demonstrations are mostly gathered by teleoperation, with a person steering a robot through a task. Figure said Helix learned from about 500 hours of teleoperated behaviour. Physical Intelligence used more than 10,000 hours across seven robot configurations and 68 tasks for π0. Skild AI said in August 2026 that collecting 380 demonstrations of a task longer than four minutes takes 50 to 100 hours of teleoperation, and that its in-context S1 model gets comparable value from a single video.

Companies are attacking that cost from three sides. NVIDIA generated 780,000 synthetic trajectories, equal by its count to about 6,500 hours of human demonstrations, in 11 hours, and said mixing them with real data lifted GR00T N1's performance by 40%. Generalist and General Intuition are betting on human and gameplay video. Figure built Index, a crowdsourced video programme that had gathered 16 million videos by August 2026, and committed $3.5 billion to Nscale for GPUs to train Helix on it.

How much data the models say they used
  • Helix
    Developer
    Figure AI
    Data type
    Teleoperation
    Hours (disclosed)
    500
  • π0
    Developer
    Physical Intelligence
    Data type
    Robot demonstrations
    Hours (disclosed)
    10,000
  • π0.5
    Developer
    Physical Intelligence
    Data type
    Mobile manipulation (part of mix)
    Hours (disclosed)
    400
  • GR00T N1
    Developer
    NVIDIA
    Data type
    Synthetic trajectories (hour equivalent)
    Hours (disclosed)
    6,500
  • LBM
    Developer
    Toyota Research Institute
    Data type
    Robot demonstrations
    Hours (disclosed)
    1,700
  • GEN-1
    Developer
    Generalist AI
    Data type
    Human wearable recordings, no robot data
    Hours (disclosed)
    500,000

Company posts and papers (TRI figure from arXiv preprint 2507.05331). Helix and π0.5 about; π0, GEN-1 more than. GR00T N1: 780,000 trajectories generated in 11 hours, which NVIDIA equates to about 6,500 demonstration hours, so about 590 demonstration hours per hour of generation (6,500 / 11).

As of Oct 1, 2026

The cost of moving a model to a new robot is falling as well. Google DeepMind said Gemini Robotics 2 adapts to a new two-armed robot in a few hours, typically with fewer than 200 examples. That is the figure a factory buyer should watch, because it decides whether a model can move from one cell to the next without a fresh data campaign, and it is what investors have been paying for.

Where the money has gone

Skild AI raised $1.4 billion in January 2026 at a valuation of more than $14 billion, Figure AI raised $1 billion at $39 billion post-money in September 2025, and Generalist AI raised $400 million in June 2026, according to the companies. Skild is the only one of the group to put revenue on record, saying on September 10, 2026 that it had passed $100 million in annual recurring revenue with more than 60 paying customers.

Robot-model developers: latest disclosed rounds
  • Skild AI
    Round
    Series C
    Amount ($m)
    1400
    Post-money ($bn)
    14
    Date
    2026-01-14
  • Figure AI
    Round
    Series C
    Amount ($m)
    1000
    Post-money ($bn)
    39
    Date
    2025-09-16
  • Generalist AI
    Round
    Growth
    Amount ($m)
    400
    Post-money ($bn)
    No data
    Date
    2026-06-04
  • General Intuition
    Round
    Venture
    Amount ($m)
    220
    Post-money ($bn)
    No data
    Date
    2026-09-29
  • Dyna Robotics
    Round
    Series A
    Amount ($m)
    120
    Post-money ($bn)
    No data
    Date
    2025-09-15
  • Genesis AI
    Round
    Launch round
    Amount ($m)
    105
    Post-money ($bn)
    No data
    Date
    2025-06-30
  • Physical Intelligence
    Round
    Series B
    Amount ($m)
    No data
    Post-money ($bn)
    No data
    Date
    2025-11-20

Latest round in the ROBOTNESS database, from company or investor releases. Skild AI valuation is stated as more than $14 bn. Skild step-up = $14 bn / $1.5 bn Series A post-money (July 2024) = 9.3x. null = not disclosed.

As of Oct 1, 2026

ROBOTNESS analysis

The VLA itself is becoming a shared layer; lasting value will sit with whoever owns the cheapest stream of relevant action data and the deployments that keep producing it.

The evidence is in the tables. Backbones are converging on a handful of public vision-language models, and capable open VLAs between 2 and 7 billion parameters can be downloaded free. The data behind them, by contrast, ranges from about 500 hours of teleoperation for Helix to more than 500,000 hours of human recordings for GEN-1. Skild's valuation rests on deployments more than on architecture: at more than $14 billion it is priced at about 140 times its stated recurring revenue ($14 bn / $100 m), after a 9.3-fold rise in 18 months from its $1.5 billion Series A valuation in July 2024.

The strongest counter-argument is that reliability, which is still a modelling problem, has not been solved by anyone. Gemini Robotics 2 succeeded in 68.4% of whole-body picks from a table and 45.7% from the floor on Apptronik's Apollo 2, by Google DeepMind's own count, numbers no plant manager would accept. A lab that closes that gap through a better architecture could win regardless of who holds more data, and world-action models from Runway or Generalist could make today's VLA design a transition step.

Bull case. Data costs keep falling through synthetic trajectories, wearables and in-context learning, open models spread to thousands of integrators, and the companies with fleets in the field compound their lead because every shift produces training data. Revenue like Skild's becomes common by 2028.

Bear case. Success rates stall in the 60% to 80% range on unstructured tasks, buyers stay with scripted cells plus narrow AI, and the valuations above $10 billion prove to have priced a research lead rather than a business. Consolidation follows.

What would change our view is a closed model showing above 90% task success at unseen customer sites without customer-specific data. We will track three dated signals.

  • Runway's public release of Praxis-1 weights, promised in the months after September 30, 2026.
  • LG's bipedal humanoid built on NVIDIA Jetson Thor and Isaac GR00T, due to be unveiled in the first quarter of 2027.
  • Figure's first Vera Rubin GPU deployment with Nscale for Helix training, scheduled for the second half of 2027.
Key facts
Definition
One network mapping camera images plus a text instruction to robot motor commands
Open models
OpenVLA (7 bn), π0 and π0.5 (openpi, Apache 2.0), GR00T N1 and N1.7
Flow approach
π0: 3 bn PaliGemma plus 300 m action expert, 50-step chunks, up to 50 Hz
Synthetic data
NVIDIA: 780,000 trajectories in 11 hours, about 6,500 demonstration hours
Token approach
RT-2 splits 8 action dimensions into 256 bins each
Term introduced
RT-2, Google DeepMind, arXiv preprint of July 28, 2023
Largest VLA round
Skild AI, $1.4 bn Series C at more than $14 bn, January 2026
Revenue disclosed
Skild AI, $100 m ARR and 60+ customers, September 2026
Teleoperation data
About 500 hours for Helix; more than 10,000 hours for π0
Dual-system approach
Helix: 7 bn model at 7 to 9 Hz plus 80 m policy at 200 Hz
Sources
ROBOTNESS Intelligence
  1. 01

    Why it matters

    VLA models decide whether general-purpose robots become a software business or stay a custom-integration business. If one model can be adapted to a new robot in a few hours with fewer than 200 examples, as Google DeepMind now claims for Gemini Robotics 2, the cost of each new deployment falls from an engineering project to a data-collection task measured in days. That shift would move value from integrators and robot programmers towards model owners and whoever controls deployment data.

    For investors, the explainer matters because nearly every large private round in robotics since 2024 is, at bottom, a bet on this model class. Skild AI, Figure AI, Physical Intelligence, Generalist AI and Dyna Robotics are all priced on the assumption that their models will generalise. Understanding the mechanics, from tokenised actions to flow-matching experts, is how one separates engineering progress from marketing.

  2. 02

    Rival analysis

    Google DeepMind has the strongest backbone and the broadest reasoning layer, with Gemini Robotics ER models offered through the Gemini API and action models kept with partners such as Apptronik, Boston Dynamics and Agile Robots. Its weakness is distribution: it does not build robots, and it depends on partners to generate field data. Physical Intelligence has the strongest open research footprint through openpi, which gives it mindshare among labs and small hardware makers, but it has not disclosed revenue.

    NVIDIA plays a different game. GR00T is a means to sell Jetson Thor computers, Isaac simulation and Cosmos world models, so it can afford to give models away, and LG, FANUC and Hugging Face have all built on it. Figure and Skild are vertically integrated in different ways: Figure trains Helix only for its own humanoid and has committed $3.5 billion to compute, while Skild sells a cross-embodiment brain and is the only one to report $100 million in recurring revenue. In China AgiBot pairs its GO models with its own hardware volume, which it put above 10,000 robots produced by April 2026.

  3. 03

    Valuation context

    On disclosed numbers, Skild AI's valuation of more than $14 billion equals about 140 times its stated $100 million annual recurring revenue, and its valuation rose 9.3 times in the 18 months from its July 2024 Series A. Figure AI's $39 billion post-money from September 2025 carries no disclosed revenue, so no comparable multiple can be computed. Physical Intelligence's November 2025 Series B was announced without an amount or valuation in the investor post we hold.

    The comparison that matters is not between these companies but between them and robot makers with revenue. Agility Robotics, which plans to list through a SPAC at a $2.5 billion pre-money value, reported $1.8 million of 2025 sales in its S-4. Investors are paying far more for the model layer than for hardware businesses with deployed fleets, which is the clearest sign the market has priced the thesis that intelligence, not hardware, captures margin.

  4. 04

    Supply-chain implications

    A VLA depends on three supply chains. Compute for training sits with NVIDIA GPUs and Google TPUs; Figure's Nscale commitment covers up to 100,000 Vera Rubin GPUs in Texas. Compute on the robot is increasingly NVIDIA's Jetson Thor, which LG, FANUC and NVIDIA's own reference humanoid use, leaving robot brains heavily exposed to one US supplier and to US export rules.

    The third chain is data. Teleoperation rigs, motion-capture labs, wearable recorders and synthetic data pipelines are becoming a procurement category in their own right. LG has said its Yangjae data factory aims to produce 100,000 hours of data by the end of 2026, and China has ordered provinces and central state firms to open real-world training scenes. Countries with large manufacturing bases can turn factory floors into data sources, which favours China, Korea, Japan and Germany if their makers organise to share it.

  5. 05

    Signals to watch

    The first signal is reliability at unseen sites. Watch for any developer reporting task success above 90% at customer locations it did not train on; Figure's Helix 2.5 claim of 56% zero-shot success across 30 unseen homes is the current public reference.

    The second is open-weight momentum: Runway's promised Praxis-1 weights, further GR00T releases and any open release of π-series successors. The third is revenue: a second company after Skild disclosing recurring revenue above $50 million would confirm that buyers pay for general models, while LG's GR00T-based bipedal humanoid in the first quarter of 2027 will show whether large manufacturers prefer to rent a model rather than build one.

  6. 06

    Analyst view

    Thesis: the model layer is commoditising faster than the data layer, so the winners will be the companies that own deployments and data flywheels, not those with the cleverest action head. Confidence: medium.

    Why medium and not high: architectures still matter. The move from tokenised actions at 1 to 3 Hz to flow-matching experts at 50 Hz and dual systems at 200 Hz was an architectural gain, not a data gain, and similar steps could reset the race. Why not low: every disclosed data point shows data volume and adaptation cost, not parameter count, explaining the gap between leaders and followers, and open models have already closed much of the architecture gap.

  7. 07

    Questions you should be asking

    For Skild AI: what share of the $100 million recurring revenue comes from software licences rather than robots or services, and what is gross margin per deployment? For Figure: how much of Helix's improvement since 2025 comes from Index video versus teleoperation, and what does each hour cost?

    For Physical Intelligence: what is the commercial model behind openpi, and which customers pay for closed successors? For Google DeepMind: when will Gemini Robotics action models be offered outside the trusted tester programme, and on what terms? For NVIDIA: will GR00T remain free once Jetson Thor volumes scale, and how is liability handled when an open model fails on a customer line?