NVIDIA opens Cosmos 3 Edge, a 4-billion-parameter world model that plans robot actions on a Jetson Thor
NVIDIA released Cosmos 3 Edge on July 20, the smallest member of its Cosmos 3 physical-AI model family, as open weights on Hugging Face. The 4-billion-parameter model runs on Jetson Thor and generates 32 robot actions per inference, and Doosan Robotics, Siemens, Agile Robots and Skild AI are evaluating it.
NVIDIA has pushed its world-model family down onto the robot itself. At SIGGRAPH on July 20, 2026 the company released Cosmos 3 Edge, a 4-billion-parameter open model that it says can reason about a scene and output robot actions in real time on a Jetson Thor computer, without a data-centre GPU.
Cosmos 3 Edge is the smallest of three Cosmos 3 models. NVIDIA launched the family at GTC Taipei on May 31 with the 64-billion-parameter Super and 16-billion-parameter Nano versions available immediately and Edge marked as coming soon. All three use what NVIDIA calls a mixture-of-transformers design that understands and generates text, images, video, ambient sound and actions. Edge pairs an autoregressive tower for reasoning with a diffusion tower for prediction, sharing attention layers across language, video, audio and action tokens. Its reasoning component is a 2-billion-parameter model based on Nemotron.
- Cosmos 3 Edge
- 4
- 2026-07-20
- Real-time inference on Jetson, RTX PRO, DGX, GeForce RTX
- Cosmos 3 Nano
- 16
- 2026-05-31
- Fast video and action reasoning
- Cosmos 3 Super
- 64
- 2026-05-31
- Highest physics accuracy for post-training robot and AV models
Parameter counts from NVIDIA's SIGGRAPH 2026 blog; dates from NVIDIA's May 31 launch release and July 20 Edge release. Disclosed figures only.
As of Oct 1, 2026
The weights are on Hugging Face, with inference and post-training code on GitHub. NVIDIA said Edge runs on Jetson modules, including the new Jetson T2000 and T3000, as well as RTX PRO, DGX and GeForce RTX GPUs. In BF16 precision the model occupies about 9 GB on a Jetson Thor. NVIDIA's technical blog reports that on a Jetson AGX Thor T5000 it generates a chunk of 32 future actions in about 1.53 seconds, and that each chunk covers roughly 2.13 seconds of robot motion at 15 Hz. The robot executes the start of a chunk and then replans, so a new plan is ready before the old one runs out.
For vision work, NVIDIA said Edge analyses 1080p video in as little as 22 milliseconds on RTX systems and about 30 milliseconds on L40 GPUs, and ranks first on the VANTAGE-Bench vision analytics benchmark in its parameter class. The company also released step-distilled versions of Cosmos 3 models that cut sampling from 50 denoising steps to 4, which it says makes inference 25 times faster with minimal quality loss.
The robot-control numbers are more modest, and NVIDIA publishes them. Post-trained on a dataset it calls Cosmos3-DROID, which contains 76,000 successful teleoperated trajectories, about 350 hours across 86 tasks and 564 scenes recorded on a Franka Panda arm with a Robotiq gripper, Edge reached 22.9% success across 120 language-conditioned manipulation tasks in closed-loop RoboLab evaluation. The baseline post-training run used 64 nodes of four GB200 chips each for 60,000 iterations, about 68 hours.
NVIDIA named Agile Robots, Doosan Robotics, Siemens and Skild AI as robotics companies evaluating Cosmos 3 Edge, and Centific, Vaidio and YUAN for smart-infrastructure vision agents. At the May launch it listed Agile Robots, Doosan Robotics, LG Electronics, Samsung Electronics and Skild AI as robotics adopters of Cosmos 3, and Li Auto in autonomous driving. Skild, which NVIDIA said in September uses Cosmos models and the Cosmos Curator tool in training its S1 model, is also a founding member of NVIDIA's Cosmos coalition alongside Agile Robots, Black Forest Labs, Generalist, LTX and Runway.
The target hardware matters to the strategy. NVIDIA's Jetson AGX Thor T5000, the module in its Isaac GR00T reference humanoid announced on June 1, delivers 2,070 FP4 teraflops with 128 GB of unified memory in a 40 to 130 watt envelope. A world model that fits in 9 GB leaves room on that module for perception, safety and a separate action policy. The reference humanoid itself is built on Unitree's H2 Plus and is expected from Unitree in late 2026.
The release sharpens a split in robot AI. Google DeepMind's Gemini Robotics 2, released ten days later, keeps its on-device action model with trusted testers. Skild AI offers S1 through an early-access sign-up. NVIDIA instead gives away weights for Cosmos 3 Edge and the 3-billion-parameter GR00T N1.7, which carries a commercial licence, and earns its return on the Jetson, RTX and data-centre hardware needed to train and run them.
- Cosmos 3 Edge
- NVIDIA
- 4
- Open weights (Hugging Face)
- 2026-07-20
- Isaac GR00T N1.7
- NVIDIA
- 3
- Open weights, commercial licence
- 2026-04-17
- Gemini Robotics On-Device 2
- Google DeepMind
- No data
- Trusted testers only
- 2026-07-30
- S1
- Skild AI
- No data
- Early access sign-up
- 2026-08-18
Parameters as disclosed on developer pages or model cards; null where not disclosed. GR00T N1.7 date is NVIDIA's early-access announcement.
As of Oct 1, 2026
In plain terms, a world model predicts what a scene will look like next. Cosmos 3 Edge uses that ability in two directions: it can generate or analyse video, and it can predict the motor commands that would move the scene towards a goal. Running both on the robot cuts the network delay of a cloud call and keeps camera footage on site, which matters for factories and for privacy rules in public spaces.
The limits are clear in NVIDIA's own figures. A 22.9% closed-loop success rate after 68 hours on 256 GB200 chips shows Edge is a research and prototyping base, not a deployable manipulation policy. A 1.53-second inference time is acceptable for slow tabletop tasks but not for fast dynamic motion. NVIDIA has not disclosed the licence terms for Cosmos 3 Edge in the materials reviewed, nor named any production deployment.
ROBOTNESS analysis
NVIDIA is commoditising the robot model layer so that every robot maker needs its chips, and Cosmos 3 Edge extends that strategy from the data centre to the robot's own computer.
The evidence is the packaging. Open weights, GitHub post-training recipes and a dataset recorded on a common Franka arm lower the cost of starting, while every published benchmark runs on NVIDIA silicon, from GB200 for training to Jetson Thor for inference. The evaluator list spans a Korean cobot maker, a German industrial group and the best-funded US robot-brain startup.
The strongest counter-argument is that a 22.9% success rate makes Edge a demonstration rather than a product, and that robot makers will deploy specialised action policies such as GR00T or their own models while using Cosmos only for synthetic data. In that case the edge model adds little to NVIDIA's revenue beyond what Jetson already earns.
Bull case: robot makers standardise on Cosmos for both simulation data and on-board reasoning, Jetson Thor becomes the default robot computer, and the distilled models close the gap in success rate within a year. Each new humanoid or cobot ships with NVIDIA silicon by default.
Bear case: Google, Skild and Chinese model developers out-perform Cosmos on manipulation, robot makers pick cheaper chips for inference, and Cosmos remains a training-time tool. Open weights also help rivals benchmark against NVIDIA and close the gap.
Signals to watch:
- Whether Doosan Robotics, Siemens or Agile Robots announce products that ship Cosmos 3 Edge on board by the end of 2026.
- Unitree's delivery of the Isaac GR00T reference humanoid in late 2026, the first broad hardware platform that pairs Jetson Thor with NVIDIA's open models.
- Updated RoboLab closed-loop results for Cosmos 3 Edge or distilled variants above the 22.9% baseline.
- July 20, 2026 (SIGGRAPH)
- 4 billion (2-billion-parameter Nemotron-based reasoner)
- Open weights on Hugging Face, code on GitHub
- 32 actions per inference, about 1.53 s on Jetson AGX Thor T5000
- 22.9% success on 120 RoboLab manipulation tasks
- Agile Robots, Doosan Robotics, Siemens, Skild AI
- About 9 GB in BF16
- Cosmos3-DROID, 76,000 trajectories, about 350 hours
- As low as 22 ms (1080p, RTX); about 30 ms on L40
Why it matters
NVIDIA positions Cosmos 3 Super and Nano for data-centre work such as post-training robot and autonomous-vehicle models. Edge puts a model of the same family on the robot, able both to analyse video and to emit action chunks, within about 9 GB of memory on a Jetson Thor.
That changes the procurement conversation for robot makers. A single NVIDIA stack now covers data generation (Cosmos Super and Nano), simulation (Isaac Sim and Isaac Lab), policy models (GR00T N1.7) and on-board reasoning (Cosmos 3 Edge). The more of that stack a robot maker adopts, the harder it becomes to move inference to another chip vendor.
Rival analysis
Google DeepMind is the closest rival on model capability but follows a cloud-first, closed approach: Gemini Robotics On-Device 2 remains limited to trusted testers. Skild AI builds its own proprietary brain while using NVIDIA tools to train it, which makes Skild both a customer and a potential competitor.
Physical Intelligence open-sourced its Ī0 policy in February 2025, and Hugging Face's LeRobot library integrates a growing list of open policies, including GR00T N1.7. Cosmos 3 Edge's distinguishing feature is that it is one open model for both vision analytics and action, sized for NVIDIA's own edge hardware.
Valuation context
Cosmos 3 Edge carries no direct price, so its value sits in NVIDIA hardware demand. The Jetson AGX Thor T5000 at the heart of NVIDIA's reference humanoid provides 2,070 FP4 teraflops and 128 GB of memory; every robot that runs Cosmos 3 Edge or GR00T on board needs a module in that class.
The partner list links the model to well-capitalised buyers. Skild AI raised $1.4 billion at a valuation of more than $14 billion in January 2026 and reports $100 million in ARR, and NVIDIA's venture arm NVentures took part in that round. Doosan Robotics is listed on the Korea Exchange. Their adoption decisions will show whether open models pull through hardware sales at scale.
Supply-chain implications
Training a Cosmos 3 Edge baseline took 64 nodes of four GB200 chips for about 68 hours, according to NVIDIA. Inference runs on Jetson Thor modules, with the newer T2000 and T3000 also supported. The post-training dataset was recorded on Franka Panda arms with Robotiq grippers, making those components the de facto reference hardware for anyone reproducing the results.
The Isaac GR00T reference humanoid ties the stack to Unitree's H2 Plus chassis and Sharpa Wave tactile hands, with deliveries from Unitree expected in late 2026. That gives a Chinese body maker a central role in NVIDIA's open robotics platform, a supply-chain dependency that Western buyers will weigh against export-control and procurement rules.
Signals to watch
First, product announcements from Doosan Robotics, Siemens, Agile Robots or Skild AI that ship Cosmos 3 Edge on board rather than only evaluate it. Second, new RoboLab results that improve on the 22.9% closed-loop baseline, especially from the step-distilled variants.
Third, the late-2026 delivery of the Unitree-built Isaac GR00T reference humanoid, whose named early adopters include Stanford, ETH Zurich, UC San Diego and Ai2. Fourth, any licence clarification for commercial on-device use of Cosmos 3 Edge, which will decide whether integrators can ship it in paid products.
Analyst view
Thesis: Cosmos 3 Edge is less a robot brain than a hardware pull-through tool that extends NVIDIA's open-model strategy to the robot's own computer. Confidence: high on the strategy, medium on adoption. The release materials themselves show every training and inference benchmark on NVIDIA silicon and an evaluator list drawn from NVIDIA's existing partner base.
Adoption confidence is lower because the 22.9% manipulation result is far from production quality and no evaluator has announced a shipping product. We expect Cosmos 3 Edge to be used first for on-robot vision analytics and data generation, with action control following only after further post-training.
Questions you should be asking
What licence governs commercial deployment of Cosmos 3 Edge on customer robots? How does the model perform on humanoid embodiments rather than a single Franka arm?
How much faster are the step-distilled models on Jetson Thor for action generation, not only video? Will Doosan Robotics or Siemens disclose measured latency and success rates from their evaluations?
