NVIDIA previews GR00T N2, betting robots should predict video before they move
NVIDIA used its GTC keynote on March 16, 2026 to preview Isaac GR00T N2, a robot foundation model built on its DreamZero world action model research that it says succeeds at new tasks in new environments more than twice as often as leading VLA models. It also put the 3-billion-parameter GR00T N1.7 into commercial early access and announced Cosmos 3 and Isaac Lab 3.0, with N2 due by the end of 2026.

NVIDIA on March 16, 2026 previewed Isaac GR00T N2, the next generation of its robot foundation model, during chief executive Jensen Huang's keynote at the company's GTC conference. NVIDIA said the model uses a new world action model architecture taken from its DreamZero research and helps robots succeed at new tasks in new environments more than twice as often as leading vision language action (VLA) models. The company said GR00T N2 is slated to be available by the end of 2026.
The preview came with a wider set of releases. NVIDIA moved Isaac GR00T N1.7, a 3-billion-parameter open reasoning VLA model built for humanoids, into early access with commercial licensing, and said Humanoid, LG Electronics, NEURA Robotics and Noble Machines are adopting it. It also announced Cosmos 3, which it called the first world foundation model to combine synthetic world generation, vision reasoning and action simulation, and Isaac Lab 3.0 in early access, running on the new Newton 1.0 physics engine and the PhysX SDK. NVIDIA said GR00T N2 already ranks first on the MolmoSpaces and RoboArena leaderboards for generalist robot policies.
The research under N2 was posted on Feb. 17, 2026 as an arXiv preprint titled "World Action Models are Zero-shot Policies". Its authors include NVIDIA researchers Seonghyeon Ye and Joel Jang along with the company's robotics research leads Linxi "Jim" Fan and Yuke Zhu. DreamZero is a 14-billion-parameter autoregressive video diffusion model. An NVIDIA developer blog says it starts from the open Wan 2.1 image-to-video 14B model and that one transformer denoises video tokens and action tokens together, so the robot's next movements come out of the same pass that predicts what its cameras will see.
The preprint gives the numbers behind the claim. On AgiBot's G1 mobile two-armed robot, trained on about 500 hours of data from 22 environments, DreamZero reached 62.2% average task progress on tasks it had seen against 27.4% for pretrained VLA baselines, and 39.5% against 16.3% on unseen tasks. The baselines were NVIDIA's own GR00T N1.6 and Physical Intelligence's pi0.5. The authors cut inference latency from 5.7 seconds to 150 milliseconds, a 38-fold gain, on NVIDIA GB200 hardware through GPU parallelism, caching, CUDA graphs, NVFP4 quantisation and a faster "Flash" variant, which allowed closed-loop control at 7 Hz. They said weights, inference code and evaluation scripts are on GitHub.
- Average task progress, seen tasks (%)
- 62.2
- 27.4
- 2.27
- Average task progress, unseen tasks (%)
- 39.5
- 16.3
- 2.42
Figures from the DreamZero preprint (arXiv 2602.15922, Feb. 17, 2026). Ratio = DreamZero progress / baseline progress, computed by ROBOTNESS. Disclosed, not estimated.
As of Oct 1, 2026
Transfer between robots is the other headline result. According to the paper, adding 12 minutes of human egocentric video or 20 minutes of video from a YAM arm lifted performance on unseen tasks by more than 42% in relative terms, and the AgiBot-trained model adapted to the YAM robot with 30 minutes of unscripted play data while keeping its zero-shot ability.
GTC extended a programme NVIDIA laid out at CES on Jan. 5, 2026, when it released GR00T N1.6 as an open reasoning VLA with full-body control for humanoids, refreshed its Cosmos world models and began selling the Jetson T4000 module, a Blackwell part rated at 1,200 FP4 teraflops with 64 GB of memory, for $1,999 in 1,000-unit volumes. NVIDIA says 2 million developers use its robotics stack. The business logic is consistent across both events: the models are open or cheaply licensed, while training, simulation and onboard inference run on NVIDIA silicon.
The partner list at GTC showed how far that stack has spread. NVIDIA said FANUC, ABB Robotics, YASKAWA and KUKA, with a combined installed base of more than 2 million robots, are building its Omniverse libraries and Isaac simulation into virtual commissioning tools. AGIBOT is adopting GR00T N models for industrial humanoids, Skild AI and the Toyota Research Institute are among Cosmos 3 adopters, Hugging Face is folding Isaac and GR00T into its LeRobot framework, and Alibaba Cloud is integrating NVIDIA's full physical AI stack into its Platform for AI.
The bet is a technical one. VLA models such as pi0.5 and Google DeepMind's Gemini Robotics family adapt vision language models, trained mostly on images and text, to output motor commands. A world action model instead starts from a video generator that has absorbed how objects move and fall, and learns to predict actions alongside the next frames. Skild AI, which raised $1.4 billion at a valuation above $14 billion on Jan. 14, and Physical Intelligence are building their own general robot brains, and both compete with NVIDIA's models even as Skild uses NVIDIA tools.
- Skild AI
- US
- 2023
- Series C
- 1,400,000,000
- 14,000,000,000
- 2026-01-14
- Physical Intelligence
- US
- No data
- Series B
- No data
- No data
- 2025-11-20
- Generalist AI
- US
- No data
- Growth
- 400,000,000
- No data
- 2026-06-04
- Dyna Robotics
- US
- 2024
- Series A
- 120,000,000
- No data
- 2025-09-15
- Genesis AI
- FR
- No data
- Launch round
- 105,000,000
- No data
- 2025-06-30
Latest round each company disclosed on its own or its investor's page, per ROBOTNESS company database. Null = not disclosed. NVIDIA is public and not shown.
As of Oct 1, 2026
Cost is the main open question. A later NVIDIA developer blog measured world action model inference at 590 and 800 milliseconds per action chunk against 190 milliseconds for pi0.5 in its own comparison, a 3 to 4 times slowdown, and reaching 7 Hz in the paper required data-centre GB200 hardware. The leaderboard rankings are NVIDIA's own statements, the preprint has not been peer reviewed, and GR00T N2 itself had not shipped by the end of the quarter.
ROBOTNESS analysis
NVIDIA is steering robot foundation models from language-first VLAs toward video-first world action models, a shift that favours its business because those models need several times more compute per decision.
The evidence is in NVIDIA's own figures: a 14-billion-parameter policy that more than doubles task progress on unseen tasks over GR00T N1.6 and pi0.5, but needs a GB200 and a 38-fold optimisation effort to act at 7 Hz. The same blog that reports DreamZero at 1,750 Elo against 1,622 for pi0.5 on RoboArena's April 2026 snapshot also documents the latency penalty.
The strongest counter-argument is that factory buyers pay for cycle time and reliability, not benchmark progress scores. If robots must stream decisions from a data centre or carry far larger onboard computers, smaller VLAs that run at higher rates on a Jetson could win the deployments even with weaker generalisation.
Bull case. If GR00T N2 ships on time with a commercial licence and distils to Jetson Thor-class hardware, the 2 million robot installed base at FANUC, ABB, YASKAWA and KUKA gives NVIDIA a ready channel. Startups would then compete on data and hardware rather than on base models.
Bear case. If latency or cost stays high, N2 remains a lab result and rivals with leaner VLAs keep the commercial lead. NVIDIA's habit of grading its own models on leaderboards it cites would then weaken its standing with integrators.
Three dated signals to watch follow.
- By Dec. 31, 2026: whether GR00T N2 is released with weights and a commercial licence, as NVIDIA promised.
- Monthly RoboArena snapshots through 2026: whether DreamZero-based policies hold first place as rivals submit new models.
- GTC 2027 in spring 2027: whether NVIDIA shows a world action model running on a Jetson module inside a production humanoid.
- March 16, 2026, GTC keynote
- 3B parameters; early access with commercial licence
- 7 Hz after cutting latency from 5.7 s to 150 ms on GB200
- Cosmos 3, Isaac Lab 3.0, Newton 1.0
- 14B parameters, video diffusion world action model
- More than twice the success rate of leading VLAs on new tasks in new environments
- Preview; available by end of 2026
- FANUC, ABB, YASKAWA, KUKA (over 2 million installed robots)
- 39.5% vs 16.3% for pretrained VLAs
Why it matters
GR00T N2 is the first time a platform vendor of NVIDIA's reach has committed a commercial roadmap to world action models, where a policy predicts video and motor commands jointly. Until GTC 2026 the commercial default for general robot policies was the VLA recipe used by Physical Intelligence, Google DeepMind and NVIDIA itself in GR00T N1.6.
Because NVIDIA also supplies the simulation, synthetic data and onboard compute, its model choice sets defaults for thousands of smaller developers. A move to video-first policies raises the value of large video corpora and of GPU capacity at every stage of the robot lifecycle.
Rival analysis
Physical Intelligence is the reference VLA rival: its pi0.5 was the outside baseline in the DreamZero paper and scored 1,622 Elo on RoboArena's April 2026 snapshot against 1,750 for DreamZero, according to NVIDIA. Skild AI raised $1.4 billion at a valuation above $14 billion in January and sells its own robot brain, yet adopts Cosmos 3, which makes it both customer and competitor.
Google DeepMind's Gemini Robotics family remains closed and tied to Google's cloud, while NVIDIA publishes weights. That openness is NVIDIA's main lever with humanoid makers such as AGIBOT, LG Electronics and NEURA that do not want to depend on a rival model owner.
Valuation context
NVIDIA does not disclose robotics revenue separately, so the GTC releases cannot be tied to a reported figure. The relevant valuation signal sits with the model startups: Skild AI's $14 billion post-money in January 2026 and Generalist AI's $400 million raise in June 2026 price general robot brains as a standalone business.
If NVIDIA's open N2 matches or beats those companies' policies, the scarcity premium in their valuations comes under pressure, and their pitch shifts toward proprietary data and deployment revenue.
Supply-chain implications
DreamZero's 7 Hz control ran on GB200 data-centre hardware, while NVIDIA sells Jetson T4000 modules for robots at $1,999 in 1,000-unit volumes with 1,200 FP4 teraflops. The gap between those two classes of hardware is the practical supply-chain question for N2.
Training draws on large video corpora and robot data such as AgiBot's 500 hours across 22 environments, which raises demand for data-collection fleets, teleoperation and synthetic data from Cosmos and the Physical AI Data Factory blueprint now offered on Microsoft Azure and Nebius.
Signals to watch
The firm date is NVIDIA's promise of N2 availability by the end of 2026; a slip or a research-only licence would be telling. RoboArena and MolmoSpaces rankings, which NVIDIA cites, update as rivals submit models.
Integration news from FANUC, YASKAWA, KUKA and ABB will show whether world models move from virtual commissioning into robot controllers, and humanoid makers adopting GR00T N1.7 will indicate the upgrade path to N2.
Analyst view
Thesis: NVIDIA's shift to world action models is technically credible and commercially self-serving, and the outcome depends on whether N2 can be distilled to onboard hardware. Confidence: medium.
Reasons: the preprint gives detailed, reproducible numbers with open weights, which supports the capability claim. Against that, the headline leaderboard positions are self-reported, the model was a preview at GTC, and NVIDIA's own later blog documents a 3 to 4 times latency penalty against pi0.5.
Questions you should be asking
What onboard compute will GR00T N2 need for closed-loop control on a humanoid, and will NVIDIA publish a Jetson Thor benchmark?
Will N2 carry the same commercial licence terms as N1.7, and how will NVIDIA handle the licence of the Wan 2.1 backbone that DreamZero builds on?
How much of DreamZero's gain on unseen tasks survives in factory settings with strict cycle times?