What is a vision-language-action (VLA) model, and who is building them?
A vision-language-action (VLA) model is one neural network that reads camera images and a plain-language instruction and outputs a robot's next motor commands, built by training a vision-language model on robot demonstrations. Google DeepMind, Physical Intelligence, NVIDIA, Figure AI, Skild AI and AgiBot lead, and the commercial race now turns on the cost of action data rather than model design.

A vision-language-action (VLA) model is a single neural network that takes camera images and a plain-language instruction and returns the motor commands a robot should execute next. It is made by taking a vision-language model, the kind of AI that can describe a photo, and training it further on robot demonstrations so that its output is movement rather than text.
The term entered robotics in July 2023, when Google DeepMind described RT-2 in an arXiv preprint and reported that a model trained on web images and robot data roughly doubled success on unseen objects and scenes against its earlier RT-1 controller, at about 62% versus 32%. Three years later the label covers models from Google DeepMind, Physical Intelligence, NVIDIA, Figure AI, Skild AI and AgiBot, and the companies behind them have raised some of the largest private rounds in robotics. This explainer sets out how the models work, how they differ from older robot software and from world models, who leads, and why the cost of data, more than model design, now decides the commercial race.
How a VLA turns pixels and words into motor commands
Most VLAs have three parts. A vision encoder turns each camera frame into a sequence of visual tokens. A language backbone, the pretrained vision-language model, reads those tokens together with the instruction and the robot's own joint readings. An action head then converts the backbone's internal state into numbers for joints, grippers or wheels. NVIDIA's model card for GR00T N1 lays this out plainly: a SigLIP 2 vision transformer reading 224 by 224 pixel frames, a T5 text encoder, a small network for the robot's state that changes with each robot body, and a diffusion transformer that produces the actions.
Starting from a vision-language model buys transfer. A backbone trained on vast numbers of captioned web images already knows what most household and factory objects look like, so a robot can act on things it never saw during robot training. The RT-2 authors reported that this inherited knowledge let a robot pick the smallest object on a table or move a can onto a number it had only met on the web, with success on such emergent tasks of 60% against 17% for RT-1. The harder question is how the network should express a movement at all.
Tokens or a flow: two ways to output an action
The first VLAs treated movement as words. RT-2 split each of eight action dimensions, six for the gripper's position and rotation, one for opening the gripper and one for ending the task, into 256 bins, so a motion became a short string of tokens that the language model emitted like text, according to the paper. OpenVLA, a 7-billion-parameter model from a Stanford-led academic team released with open weights in June 2024, used the same approach and reported a success rate 16.5 percentage points higher than Google's 55-billion-parameter RT-2-X across 29 tasks.
Tokens are slow and coarse for dexterous work. The largest RT-2 ran at 1 to 3 Hz, its paper says,. Physical Intelligence's π0, described in an October 31, 2024 preprint, attached a separate 300-million-parameter action expert to a 3-billion-parameter PaliGemma backbone and trained it with flow matching, a relative of the diffusion method behind image generators. The expert starts from random noise and refines it over 10 steps into a chunk of 50 future actions, which lets π0 control robots at up to 50 Hz. GR00T N1 does the same job with a diffusion transformer.
A third design separates thinking from reflexes. Figure AI's Helix, announced on February 20, 2025, pairs a 7-billion-parameter vision-language model running at 7 to 9 Hz with an 80-million-parameter policy running at 200 Hz, steering 35 degrees of freedom on two onboard GPUs, the company said. NVIDIA calls the same split System 2 and System 1, and AgiBot's GO-2, released on September 8, 2026, uses an asynchronous dual system in which a slower semantic planner hands action sequences to a faster module that corrects for noise. That layering is also what separates a VLA from the robot software factories already run.
Why a VLA is neither robot programming nor a world model
Classical industrial robots are programmed point by point. An integrator teaches waypoints with a pendant or writes the path offline, and the program repeats unchanged until a part shifts or the line is retooled. A VLA replaces that script with a learned policy that reads the scene every cycle, so it can cope with clutter and new objects, but it gives up the strict repeatability that makes scripted cells easy to certify. Japanese makers are trying to keep both. Yaskawa Electric said in July 2026 that its MOTOMAN NEXT system uses Google DeepMind's Gemini Robotics ER 1.6 to plan tasks from high-level instructions without detailed teaching, while the robot's own vision, path planning and force sensing stay underneath.
A world model answers a different question. It predicts how a scene will change, usually as video, given what the robot does, and it is used to generate synthetic training data or to test policies before they touch hardware. Runway said on September 30, 2026 that policy results inside its world model correlate with real-world results at 0.95. The two ideas are converging. Runway calls its new Praxis-1 a world action model that outputs robot control, and Generalist AI pretrains GEN-1 on more than 500,000 hours of human interaction data recorded with wearables, with no robot data at all, according to the company.
Who builds VLAs, and who gives them away
Google DeepMind launched Gemini Robotics, built on Gemini 2.0, on March 12, 2025 with Apptronik as its humanoid partner, and on July 30, 2026 released Gemini Robotics 2 for whole-body humanoid control; its action models remain limited to partners and trusted testers. Physical Intelligence went the other way. Its openpi repository carries π0, π0-FAST and, since September 2025, π0.5 under an Apache 2.0 licence. NVIDIA published GR00T N1 on March 18, 2025 under a non-commercial licence, and its 3-billion-parameter GR00T N1.7 ships under the NVIDIA Open Model License inside Hugging Face's LeRobot library. Figure keeps Helix in-house, Skild AI sells its Skild Brain and S1 models through deployments, and AgiBot delivers GO-2 through its Genie Studio platform.
- RT-2
- Google DeepMind
- 2023-07-28
- 55
- No (partners only)
- OpenVLA
- Stanford-led team
- 2024-06-13
- 7
- Yes
- π0
- Physical Intelligence
- 2024-10-31
- 3.3
- Yes
- Helix
- Figure AI
- 2025-02-20
- 7.08
- No data
- Gemini Robotics
- Google DeepMind
- 2025-03-12
- No data
- No (partners only)
- GR00T N1
- NVIDIA
- 2025-03-18
- 2
- Yes (non-commercial)
- π0.5
- Physical Intelligence
- 2025-04-22
- No data
- Yes
- Skild Brain
- Skild AI
- 2025-07-29
- No data
- No data
- Gemini Robotics 2
- Google DeepMind
- 2026-07-30
- No data
- No (partners only)
- S1
- Skild AI
- 2026-08-18
- No data
- No data
- GO-2
- AgiBot
- 2026-09-08
- No data
- No data
First public release = first paper, model card or company post. Helix = 7 bn System 2 plus 80 m System 1. π0.5 weights via openpi since September 2025. null = not disclosed. Figures disclosed, not estimated.
As of Oct 1, 2026
Openness is a business decision. Open weights let universities and small hardware makers adapt a model on one workstation; openpi says low-rank fine-tuning needs more than 22.5 GB of GPU memory, within reach of an RTX 4090. Closed models keep the data flywheel and the customer relationship with the owner. Where the two camps really part is in how they pay for data.
Data, not architecture, is the cost line
Unlike chatbots, VLAs cannot scrape their training data from the web. Robot demonstrations are mostly gathered by teleoperation, with a person steering a robot through a task. Figure said Helix learned from about 500 hours of teleoperated behaviour. Physical Intelligence used more than 10,000 hours across seven robot configurations and 68 tasks for π0. Skild AI said in August 2026 that collecting 380 demonstrations of a task longer than four minutes takes 50 to 100 hours of teleoperation, and that its in-context S1 model gets comparable value from a single video.
Companies are attacking that cost from three sides. NVIDIA generated 780,000 synthetic trajectories, equal by its count to about 6,500 hours of human demonstrations, in 11 hours, and said mixing them with real data lifted GR00T N1's performance by 40%. Generalist and General Intuition are betting on human and gameplay video. Figure built Index, a crowdsourced video programme that had gathered 16 million videos by August 2026, and committed $3.5 billion to Nscale for GPUs to train Helix on it.
- Helix
- Figure AI
- Teleoperation
- 500
- π0
- Physical Intelligence
- Robot demonstrations
- 10,000
- π0.5
- Physical Intelligence
- Mobile manipulation (part of mix)
- 400
- GR00T N1
- NVIDIA
- Synthetic trajectories (hour equivalent)
- 6,500
- LBM
- Toyota Research Institute
- Robot demonstrations
- 1,700
- GEN-1
- Generalist AI
- Human wearable recordings, no robot data
- 500,000
Company posts and papers (TRI figure from arXiv preprint 2507.05331). Helix and π0.5 about; π0, GEN-1 more than. GR00T N1: 780,000 trajectories generated in 11 hours, which NVIDIA equates to about 6,500 demonstration hours, so about 590 demonstration hours per hour of generation (6,500 / 11).
As of Oct 1, 2026
The cost of moving a model to a new robot is falling as well. Google DeepMind said Gemini Robotics 2 adapts to a new two-armed robot in a few hours, typically with fewer than 200 examples. That is the figure a factory buyer should watch, because it decides whether a model can move from one cell to the next without a fresh data campaign, and it is what investors have been paying for.
Where the money has gone
Skild AI raised $1.4 billion in January 2026 at a valuation of more than $14 billion, Figure AI raised $1 billion at $39 billion post-money in September 2025, and Generalist AI raised $400 million in June 2026, according to the companies. Skild is the only one of the group to put revenue on record, saying on September 10, 2026 that it had passed $100 million in annual recurring revenue with more than 60 paying customers.
- Skild AI
- Series C
- 1400
- 14
- 2026-01-14
- Figure AI
- Series C
- 1000
- 39
- 2025-09-16
- Generalist AI
- Growth
- 400
- No data
- 2026-06-04
- General Intuition
- Venture
- 220
- No data
- 2026-09-29
- Dyna Robotics
- Series A
- 120
- No data
- 2025-09-15
- Genesis AI
- Launch round
- 105
- No data
- 2025-06-30
- Physical Intelligence
- Series B
- No data
- No data
- 2025-11-20
Latest round in the ROBOTNESS database, from company or investor releases. Skild AI valuation is stated as more than $14 bn. Skild step-up = $14 bn / $1.5 bn Series A post-money (July 2024) = 9.3x. null = not disclosed.
As of Oct 1, 2026
ROBOTNESS analysis
The VLA itself is becoming a shared layer; lasting value will sit with whoever owns the cheapest stream of relevant action data and the deployments that keep producing it.
The evidence is in the tables. Backbones are converging on a handful of public vision-language models, and capable open VLAs between 2 and 7 billion parameters can be downloaded free. The data behind them, by contrast, ranges from about 500 hours of teleoperation for Helix to more than 500,000 hours of human recordings for GEN-1. Skild's valuation rests on deployments more than on architecture: at more than $14 billion it is priced at about 140 times its stated recurring revenue ($14 bn / $100 m), after a 9.3-fold rise in 18 months from its $1.5 billion Series A valuation in July 2024.
The strongest counter-argument is that reliability, which is still a modelling problem, has not been solved by anyone. Gemini Robotics 2 succeeded in 68.4% of whole-body picks from a table and 45.7% from the floor on Apptronik's Apollo 2, by Google DeepMind's own count, numbers no plant manager would accept. A lab that closes that gap through a better architecture could win regardless of who holds more data, and world-action models from Runway or Generalist could make today's VLA design a transition step.
Bull case. Data costs keep falling through synthetic trajectories, wearables and in-context learning, open models spread to thousands of integrators, and the companies with fleets in the field compound their lead because every shift produces training data. Revenue like Skild's becomes common by 2028.
Bear case. Success rates stall in the 60% to 80% range on unstructured tasks, buyers stay with scripted cells plus narrow AI, and the valuations above $10 billion prove to have priced a research lead rather than a business. Consolidation follows.
What would change our view is a closed model showing above 90% task success at unseen customer sites without customer-specific data. We will track three dated signals.
- Runway's public release of Praxis-1 weights, promised in the months after September 30, 2026.
- LG's bipedal humanoid built on NVIDIA Jetson Thor and Isaac GR00T, due to be unveiled in the first quarter of 2027.
- Figure's first Vera Rubin GPU deployment with Nscale for Helix training, scheduled for the second half of 2027.
- One network mapping camera images plus a text instruction to robot motor commands
- OpenVLA (7 bn), π0 and π0.5 (openpi, Apache 2.0), GR00T N1 and N1.7
- π0: 3 bn PaliGemma plus 300 m action expert, 50-step chunks, up to 50 Hz
- NVIDIA: 780,000 trajectories in 11 hours, about 6,500 demonstration hours
- RT-2 splits 8 action dimensions into 256 bins each
- RT-2, Google DeepMind, arXiv preprint of July 28, 2023
- Skild AI, $1.4 bn Series C at more than $14 bn, January 2026
- Skild AI, $100 m ARR and 60+ customers, September 2026
- About 500 hours for Helix; more than 10,000 hours for π0
- Helix: 7 bn model at 7 to 9 Hz plus 80 m policy at 200 Hz
Why it matters
VLA models decide whether general-purpose robots become a software business or stay a custom-integration business. If one model can be adapted to a new robot in a few hours with fewer than 200 examples, as Google DeepMind now claims for Gemini Robotics 2, the cost of each new deployment falls from an engineering project to a data-collection task measured in days. That shift would move value from integrators and robot programmers towards model owners and whoever controls deployment data.
For investors, the explainer matters because nearly every large private round in robotics since 2024 is, at bottom, a bet on this model class. Skild AI, Figure AI, Physical Intelligence, Generalist AI and Dyna Robotics are all priced on the assumption that their models will generalise. Understanding the mechanics, from tokenised actions to flow-matching experts, is how one separates engineering progress from marketing.
Rival analysis
Google DeepMind has the strongest backbone and the broadest reasoning layer, with Gemini Robotics ER models offered through the Gemini API and action models kept with partners such as Apptronik, Boston Dynamics and Agile Robots. Its weakness is distribution: it does not build robots, and it depends on partners to generate field data. Physical Intelligence has the strongest open research footprint through openpi, which gives it mindshare among labs and small hardware makers, but it has not disclosed revenue.
NVIDIA plays a different game. GR00T is a means to sell Jetson Thor computers, Isaac simulation and Cosmos world models, so it can afford to give models away, and LG, FANUC and Hugging Face have all built on it. Figure and Skild are vertically integrated in different ways: Figure trains Helix only for its own humanoid and has committed $3.5 billion to compute, while Skild sells a cross-embodiment brain and is the only one to report $100 million in recurring revenue. In China AgiBot pairs its GO models with its own hardware volume, which it put above 10,000 robots produced by April 2026.
Valuation context
On disclosed numbers, Skild AI's valuation of more than $14 billion equals about 140 times its stated $100 million annual recurring revenue, and its valuation rose 9.3 times in the 18 months from its July 2024 Series A. Figure AI's $39 billion post-money from September 2025 carries no disclosed revenue, so no comparable multiple can be computed. Physical Intelligence's November 2025 Series B was announced without an amount or valuation in the investor post we hold.
The comparison that matters is not between these companies but between them and robot makers with revenue. Agility Robotics, which plans to list through a SPAC at a $2.5 billion pre-money value, reported $1.8 million of 2025 sales in its S-4. Investors are paying far more for the model layer than for hardware businesses with deployed fleets, which is the clearest sign the market has priced the thesis that intelligence, not hardware, captures margin.
Supply-chain implications
A VLA depends on three supply chains. Compute for training sits with NVIDIA GPUs and Google TPUs; Figure's Nscale commitment covers up to 100,000 Vera Rubin GPUs in Texas. Compute on the robot is increasingly NVIDIA's Jetson Thor, which LG, FANUC and NVIDIA's own reference humanoid use, leaving robot brains heavily exposed to one US supplier and to US export rules.
The third chain is data. Teleoperation rigs, motion-capture labs, wearable recorders and synthetic data pipelines are becoming a procurement category in their own right. LG has said its Yangjae data factory aims to produce 100,000 hours of data by the end of 2026, and China has ordered provinces and central state firms to open real-world training scenes. Countries with large manufacturing bases can turn factory floors into data sources, which favours China, Korea, Japan and Germany if their makers organise to share it.
Signals to watch
The first signal is reliability at unseen sites. Watch for any developer reporting task success above 90% at customer locations it did not train on; Figure's Helix 2.5 claim of 56% zero-shot success across 30 unseen homes is the current public reference.
The second is open-weight momentum: Runway's promised Praxis-1 weights, further GR00T releases and any open release of π-series successors. The third is revenue: a second company after Skild disclosing recurring revenue above $50 million would confirm that buyers pay for general models, while LG's GR00T-based bipedal humanoid in the first quarter of 2027 will show whether large manufacturers prefer to rent a model rather than build one.
Analyst view
Thesis: the model layer is commoditising faster than the data layer, so the winners will be the companies that own deployments and data flywheels, not those with the cleverest action head. Confidence: medium.
Why medium and not high: architectures still matter. The move from tokenised actions at 1 to 3 Hz to flow-matching experts at 50 Hz and dual systems at 200 Hz was an architectural gain, not a data gain, and similar steps could reset the race. Why not low: every disclosed data point shows data volume and adaptation cost, not parameter count, explaining the gap between leaders and followers, and open models have already closed much of the architecture gap.
Questions you should be asking
For Skild AI: what share of the $100 million recurring revenue comes from software licences rather than robots or services, and what is gross margin per deployment? For Figure: how much of Helix's improvement since 2025 comes from Index video versus teleoperation, and what does each hour cost?
For Physical Intelligence: what is the commercial model behind openpi, and which customers pay for closed successors? For Google DeepMind: when will Gemini Robotics action models be offered outside the trusted tester programme, and on what terms? For NVIDIA: will GR00T remain free once Jetson Thor volumes scale, and how is liability handled when an open model fails on a customer line?