RSS 2026's top paper, FlashSAC, cuts humanoid walking training on one GPU from hours to minutes
FlashSAC, an off-policy reinforcement learning method from a team spanning Holiday Robotics, KAIST, KRAFTON, TU Darmstadt, KTH and others, won the Outstanding Paper Award at Robotics: Science and Systems 2026 in Sydney. The authors report training a Unitree G1 walking policy in about 20 minutes on a single RTX 5090, against about three hours for the standard PPO method.
A reinforcement learning method called FlashSAC won the Outstanding Paper Award at Robotics: Science and Systems (RSS) 2026, the conference held from July 13 to 17 in Sydney, according to the RSS awards page. The paper's authors say it trains humanoid walking controllers roughly ten times faster than the method most robot companies use today, on a single consumer graphics card.
Who is behind it
The 13 authors list affiliations at Holiday Robotics, KAIST, KRAFTON, Turing Inc, TU Darmstadt, hessian.AI, KTH Royal Institute of Technology, the German Research Center for AI (DFKI) and the Robotics Institute Germany. Senior authors include Jan Peters of TU Darmstadt, Danica Kragic of KTH, Jaegul Choo of KAIST and Hojoon Lee. The work first appeared as an arXiv preprint (2604.04539) in April 2026, and the code is published on GitHub under the Holiday-Robot organisation.
The headline numbers
On a 29-degree-of-freedom Unitree G1 humanoid, the authors report that FlashSAC learned to walk on flat ground in about 20 minutes of training, compared with about three hours for PPO, the on-policy algorithm that dominates robot locomotion. For rough terrain and stair climbing, FlashSAC needed about four hours against about 20 hours for PPO. All wall-clock times were measured on a single NVIDIA RTX 5090. The policies were then transferred from simulation to the physical robot.
Across 25 state-based tasks in four GPU simulators (IsaacLab, MuJoCo Playground, ManiSkill and Genesis), the paper says FlashSAC matched PPO on 15 low-dimensional tasks and clearly beat it on 10 high-dimensional ones, including dexterous manipulation and humanoid locomotion, while using 50 million simulation steps against 200 million for PPO. On 40 tasks in CPU-based suites such as the DeepMind Control Suite, MyoSuite and HumanoidBench, the authors report that it consistently exceeded PPO, XQC, SimbaV2, TD-MPC2 and MR.Q without task-specific tuning. In total the project page lists more than 60 tasks across 10 simulators.
The method in plain terms
Robot learning splits into two families. On-policy methods such as PPO throw away experience after each update, which is wasteful but stable. Off-policy methods such as Soft Actor-Critic (SAC) store experience in a replay buffer and reuse it, which should be more efficient but tends to become unstable in robots with many joints. FlashSAC keeps SAC's reuse and fixes the instability.
It does this in two ways. First, it scales the way language models do: 1,024 parallel simulated robots, a replay buffer of 10 million transitions instead of the usual one million, a larger 2.5-million-parameter network with six layers, batches of 2,048 samples and very few gradient updates per unit of data, two for every 1,024 steps. Second, it constrains the size of weights, internal features and gradients through several normalisation layers, so training cannot drift into unstable regions. A fixed exploration target and repeated noise patterns help the simulated robot try longer, smoother movements.
Competitive context
PPO has been the default for sim-to-real locomotion at most humanoid companies because it is predictable. Faster off-policy alternatives such as TD-MPC2, SimbaV2 and MR.Q have shown gains in benchmarks but rarely on full humanoids. FlashSAC's claim is that it closes that gap on real hardware. NVIDIA's IsaacLab, one of the simulators used, is the most common training environment for humanoid makers, so the result plugs directly into existing pipelines.
What it means for companies
Training time is engineering time. If a walking controller that took a working day to train now takes 20 minutes, engineers can test many more designs, reward functions and hardware changes per week. Running on one RTX 5090 rather than a cluster also lowers the bar for startups and university spin-outs. The gains are largest on the high-dimensional problems, such as multi-fingered hands and full humanoids, which are exactly where the industry is spending.
Risks and open questions
The figures are the authors' own and are stated as approximate. Wall-clock comparisons depend on how well PPO was tuned, and the paper's PPO budget differs from FlashSAC's. The real-robot validation is on one platform, the Unitree G1, and on locomotion rather than manipulation. Whether the speed-up holds for contact-rich manipulation on real hardware has not yet been shown.
ROBOTNESS analysis
FlashSAC is the strongest evidence yet that off-policy reinforcement learning can replace PPO as the industry's default for training robot bodies, and it shifts competitive advantage from compute budgets to iteration speed.
The evidence: a roughly ninefold cut in flat-ground walking training time and a fivefold cut on rough terrain on real hardware, gains concentrated on high-dimensional tasks, and broad coverage of more than 60 tasks without per-task tuning, all recognised by the field's most selective conference award.
The strongest counter-argument is that PPO's appeal was never only speed but predictability, and an off-policy method with many interacting normalisation tricks may prove fragile when engineers change robots, rewards or simulators.
Bull case: humanoid and hand makers adopt FlashSAC as a drop-in replacement within months because the code is open, compressing development cycles and reducing GPU spending. Off-policy learning then extends to fine-tuning on real robots, where data reuse matters most.
Bear case: reproductions show smaller gains once PPO is properly tuned, and the method stays a research favourite while companies keep PPO for production controllers. Real-world manipulation results fail to match simulation.
Signals to watch:
- Independent reproductions on other humanoids, for example at CoRL 2026.
- Integration of FlashSAC into IsaacLab or MuJoCo Playground example libraries.
- Any humanoid maker naming FlashSAC in a technical report during 2026 or early 2027.
- Flat-terrain walking
- 20
- 180
- 9
- Rough terrain and stairs
- 240
- 1200
- 5
Approximate wall-clock times as reported by the authors (about 20 min vs about 3 h; about 4 h vs about 20 h). Ratio = PPO time / FlashSAC time. On GPU-simulator benchmarks FlashSAC used 50M steps vs 200M for PPO.
As of Oct 1, 2026
- Open on GitHub (Holiday-Robot/FlashSAC)
- RSS 2026 Outstanding Paper Award
- Unitree G1, 29 degrees of freedom
- Single NVIDIA RTX 5090
- arXiv 2604.04539 (April 2026)
- July 13 to 17, 2026, Sydney
- More than 60 tasks, 10 simulators
- About 20 min vs about 3 h for PPO
- About 4 h vs about 20 h for PPO
Why it matters
Most humanoid walking controllers in production today are trained with PPO in massively parallel simulation. That choice was made for stability, not efficiency, and it has shaped how companies budget GPUs and schedule engineering work. FlashSAC challenges the premise by showing a data-reusing method that is both faster and stable on a full humanoid, with open code.
The award matters as a signal. RSS named a single Outstanding Paper this year, and choosing an algorithmic training paper over hardware or system papers tells companies that the research community sees training efficiency as a central bottleneck for robot bodies.
Rival analysis
Within reinforcement learning, FlashSAC competes with SimbaV2, XQC, TD-MPC2 and MR.Q, which the authors say it beat on CPU-based benchmarks, and with PPO in GPU simulators. Model-based methods such as TD-MPC2 offer sample efficiency at higher compute per step, while FlashSAC's bet is on cheap, large-batch updates.
At the industry level, the alternative to faster RL is learning from demonstrations and large VLA models. The two are complementary: VLAs handle task understanding while RL-trained controllers handle balance and contact. Faster RL strengthens companies that combine both rather than those relying on teleoperation data alone.
Valuation context
FlashSAC is an academic and corporate research result without a direct price tag. Its economic value shows up in costs: if a controller that needed about 20 hours of GPU time on rough terrain now needs about four, a team can run roughly five times as many experiments on the same hardware.
For investors, the result slightly erodes the moat of companies whose advantage is compute scale for locomotion training, and favours smaller teams with strong engineering. KRAFTON's presence among the affiliations also shows non-robotics capital, here from a game publisher, funding physical AI research.
Supply-chain implications
The method runs on a single NVIDIA RTX 5090, a consumer-class GPU, and on widely used simulators: IsaacLab, MuJoCo Playground, ManiSkill and Genesis. That keeps demand anchored on NVIDIA hardware but shifts it from data-centre clusters toward workstations.
On the robot side, validation on the Unitree G1, a low-cost and widely available humanoid, reinforces Unitree's role as the default research platform. Faster training of high-dimensional hands may also accelerate demand for dexterous hand hardware, where control complexity has been a barrier.
Signals to watch
Watch for independent reproductions at CoRL 2026 and in arXiv preprints, especially on manipulation with real hands. Watch whether NVIDIA or the MuJoCo team adds FlashSAC to official example libraries, which would make it a default option.
Also watch Holiday Robotics, whose GitHub organisation hosts the code, for any product or funding announcement that builds on the method, and KAIST and KRAFTON for follow-up work.
Analyst view
Thesis: FlashSAC is likely to become a standard baseline in robot RL within a year and could replace PPO in some production pipelines for high-dimensional bodies. Confidence: medium.
The confidence is medium because the speed-ups are self-reported and approximate, and real-robot validation covers one humanoid and locomotion only. It is supported by open code, broad benchmark coverage across more than 60 tasks, and peer recognition from the RSS award committee.
Questions you should be asking
How sensitive is FlashSAC to its many normalisation choices when moved to a new robot? Does the speed advantage hold against a carefully tuned PPO with the same step budget?
Can the method fine-tune on physical robots where each sample is expensive, and how does it perform on real dexterous hands rather than in simulation? Will Holiday Robotics commercialise tooling around it?