Toyota Research Institute's blind trials in Science Robotics find multitask robot models learn new tasks with a fraction of the data
An 82-author team from Toyota Research Institute, Cornell University and MIT reported in Science Robotics on April 15, 2026 that multitask large behavior models, built on diffusion policy, outperform single-task baselines in blind, randomized trials and need far less task-specific data. The preprint version describes about 1,700 hours of robot demonstrations across more than 500 tasks and 1,800 controlled real-world trials.
Toyota Research Institute (TRI) has published a peer-reviewed test of one of the central assumptions behind the money flowing into robot foundation models. In a paper that appeared in Science Robotics on April 15, 2026, an 82-author team reports that multitask robot manipulation policies, which TRI calls large behavior models, were more successful and more robust than single-task policies and learned complex new tasks with a fraction of the data.
The paper, titled A careful examination of large behavior models for multitask dexterous manipulation, appears in volume 11, issue 113 of the journal. Authors are affiliated with TRI in Cambridge, Massachusetts, Cornell University and the Massachusetts Institute of Technology, with J. Barreiros listed first and Russ Tedrake last. A preprint of the work was first posted on arXiv on July 7, 2025.
What sets the study apart is method rather than model. The team extended the diffusion policy approach, in which a network generates robot actions by iteratively refining noise, across a corpus of simulated and real robot data. It then compared the resulting models against single-task baselines in blind, randomized trials under controlled conditions, in both simulation and on hardware, and built an evaluation pipeline designed to give results with statistical confidence, the abstract says.
The scale is documented in the preprint version. TRI trained several models on roughly 1,700 hours of robot demonstrations covering more than 500 internally collected tasks, plus publicly available robot data, and evaluated them in 1,800 real-world trials. Each real-world task was run 50 times per policy and condition, and each simulation task 200 times. Tasks ranged from pick-and-place to long multi-step jobs such as coring an apple, setting a breakfast tray and installing a bicycle brake rotor.
The hardware was deliberately ordinary. The preprint describes bimanual tabletop stations built from two Franka FR3 arms with parallel grippers, nine stations in total, with policies running at 10 Hz. Simulation evaluation used lbm_eval, a benchmark TRI built on the Drake simulator, with four kitchen-style scenarios.
The headline findings in the Science Robotics abstract are three. Multitask pretraining made fine-tuned policies more successful and more robust than single-task ones. It allowed complex new tasks to be taught more quickly with a fraction of the data. And performance rose predictably as the scale and diversity of pretraining data grew. In the preprint, the authors estimate that in simulation a fine-tuned model needed less than 30% of the data a model trained from scratch required for similar performance, and that on one real task, setting a breakfast table, a model fine-tuned on 15% of the data already beat the baseline trained on all of it.
The work also argues for discipline in how the field reports progress. The authors deliberately tuned task difficulty toward success rates around 50%, used rubrics to score partial task completion, and applied statistical tests to decide whether differences between policies were real. They state that absolute success rates depend heavily on how hard a task is made, which is a caution against comparing headline percentages across company demos.
That matters commercially. Physical Intelligence, Skild AI, Figure AI, Google DeepMind and NVIDIA all market generalist robot models, and investors have priced them on the assumption that broad pretraining pays off. TRI's result supports that assumption with controlled evidence. Boston Dynamics, which has worked with TRI on large behavior models for Atlas, described separately in May 2026 how it trains whole-body skills in simulation.
The limits are stated by the authors. The preprint notes that confidence intervals do not capture randomness from training itself, that real-world noise could hide small effects, and that the models used modestly sized language encoders pretrained with CLIP rather than the large vision-language backbones in today's VLA models. All real-world tests were on parallel-gripper arms, not multi-fingered hands or humanoids.
- Authors
- 82
- Science Robotics
- Robot demonstration data, hours (approx.)
- 1700
- arXiv preprint
- Internally collected tasks (more than)
- 500
- arXiv preprint
- Real-world evaluation trials
- 1800
- arXiv preprint
- Real-world rollouts per task, policy and condition
- 50
- arXiv preprint
- Simulation rollouts per task, policy and condition
- 200
- arXiv preprint
- Robot stations
- 9
- arXiv preprint
- Policy rate, Hz
- 10
- arXiv preprint
Figures as disclosed in the Science Robotics record (DOI 10.1126/scirobotics.aea6201) and the July 2025 arXiv version (2507.05331). The published version may revise preprint figures.
As of Oct 1, 2026
- Skild AI
- US
- Series C
- 1,400,000,000
- 14,000,000,000
- 2026-01-14
- Figure AI
- US
- Series C
- 1,000,000,000
- 39,000,000,000
- 2025-09-16
- Apptronik
- US
- Series A extension
- 520,000,000
- No data
- 2026-02-11
- Walden Robotics
- US
- Seed
- 300,000,000
- 1,100,000,000
- 2026-07-15
- Physical Intelligence
- US
- Series B
- No data
- No data
- 2025-11-20
Amounts and valuations as disclosed by the companies or their investors; null where not disclosed. Toyota Research Institute is a Toyota research arm and has no separate funding round.
As of Oct 1, 2026
ROBOTNESS analysis
TRI's paper is the most rigorous public evidence yet that multitask pretraining pays off in robot manipulation, and its bigger contribution may be the evaluation standard it sets for an industry that mostly reports results through videos.
The evidence is the design. Blind A/B trials, 50 real rollouts per condition, simulation runs in the hundreds and statistical separation tests are rare in a field where many claims rest on curated demonstrations. The finding that a model fine-tuned on 15% of the data beat a baseline trained on all of it is a direct, measurable statement about data efficiency.
The strongest counter-argument is that the models tested are already a generation behind. Diffusion policies with CLIP language encoders on parallel grippers are not what Physical Intelligence or Google DeepMind ship today, so the paper may confirm a trend without telling buyers much about current products.
Bull case. Rigorous evaluation becomes a selling point, customers start demanding statistically sound acceptance tests, and TRI's open tooling on Drake becomes a reference. The data efficiency result strengthens the commercial case for pretrained models in factories, including at Toyota.
Bear case. The field keeps marketing through demonstrations, the evaluation cost of 50 rollouts per condition proves too high for start-ups, and scaling gains flatten once task diversity saturates, which the paper's own scaling curves cannot yet rule out.
Signals to watch:
- Whether TRI or its partners publish a follow-up that applies the same protocol to VLA-style models or humanoids, for example at CoRL 2026.
- Whether any commercial robot model developer adopts blind A/B evaluation with published statistics in its release materials before mid-2027.
- Toyota group announcements that move large behavior models from research into plant pilots, including through its work with Boston Dynamics.
- 82, from Toyota Research Institute, Cornell University and MIT
- Science Robotics, vol. 11, issue 113, eaea6201
- Bimanual Franka FR3 stations, 9 in total, policies at 10 Hz
- arXiv 2507.05331, July 7, 2025
- April 15, 2026
- Large behavior models based on diffusion policy
- Under 30% of from-scratch data needed in simulation; 15% enough to beat baseline on one real task
- About 1,700 hours of robot demonstrations, more than 500 tasks, plus public data
- 1,800, with 50 rollouts per task per policy per condition
Why it matters
Billions of dollars have gone into robot foundation models on the premise that pretraining on many tasks makes robots cheaper to teach and more reliable. Until this paper, most public support for that premise came from company demonstrations and benchmark tables without controlled comparisons. TRI's study supplies blind, randomized, statistically analysed evidence that the premise holds, at least for diffusion-based manipulation policies.
The second reason is the evaluation template. The paper shows what it costs to measure robot policy performance properly, with 50 real rollouts per task and condition, reproducible initial conditions and rubrics for partial completion. That standard can now be cited by customers and investors when they ask companies to back up their claims.
Rival analysis
Physical Intelligence, Skild AI, Google DeepMind, NVIDIA and Figure AI all develop generalist robot models, and most report results through internal benchmarks or videos. TRI's work is different in that it isolates the effect of pretraining by comparing against single-task baselines on the same hardware and tasks.
The closest industrial link is Boston Dynamics, which said in August 2025 that its large behavior model work for Atlas was part of a collaboration with TRI. That ties the Science Robotics result to a humanoid programme inside the Hyundai Motor Group, even though the paper itself tested only tabletop arms.
Valuation context
TRI is a research arm of Toyota and has no separate valuation. The relevance is to the private companies whose valuations rest on the same scaling thesis. Skild AI raised $1.4 billion at a $14 billion valuation in January 2026, Figure AI $1 billion at $39 billion in September 2025, Apptronik a $520 million Series A extension in February 2026 and Walden Robotics a $300 million seed at $1.1 billion in July 2026, according to their own announcements.
A rigorous confirmation that pretraining improves data efficiency supports those valuations in principle. It also raises the bar, because investors now have a public example of what proper evidence looks like.
Supply-chain implications
The study ran on commercially available components, Franka FR3 arms, WSG50 grippers, FRAMOS D415e scene cameras and FLIR Blackfly wrist cameras according to the preprint. That is a reminder that the binding constraint in robot learning is data and evaluation capacity rather than exotic hardware.
The data side is where supply matters. About 1,700 hours of demonstrations across more than 500 tasks required teleoperation stations, operators and curation over an extended period. Companies that can produce that volume at lower cost, through teleoperation services, wearable capture devices or simulation, hold a strategic position.
Signals to watch
Watch for TRI applying the same protocol to vision-language-action models and to humanoids, which would test whether the findings hold for the architectures companies are shipping now. Watch also whether Boston Dynamics reports Atlas results using comparable statistics.
On the commercial side, a useful signal is whether customer acceptance tests for robot deployments begin to specify rollout counts and statistical thresholds, which would show TRI's standard spreading into procurement.
Analyst view
Thesis: Multitask pretraining delivers measurable data-efficiency and robustness gains in robot manipulation, and TRI has provided the best controlled evidence for it. Confidence: high for the narrow claim, medium for its extension to current VLA products.
The narrow claim rests on blind trials, large sample sizes and peer review. The extension is less certain because the tested models used small CLIP language encoders, parallel grippers and tabletop tasks, while commercial systems now use large vision-language backbones, dexterous hands and mobile bodies.
Questions you should be asking
Do the same data-efficiency gains hold for VLA models with large language backbones, or do they shrink as models become more capable out of the box? How quickly do scaling gains flatten as task diversity grows beyond the 500-plus tasks in TRI's corpus?
Will Toyota move large behavior models into its plants, and will it disclose results with the same rigour? Can start-ups afford evaluation at this standard, or will rigorous testing remain the preserve of well-funded labs?
