ROBOTNESS
ExpertarXiv

Beyond Policy Alignment: Closing the Planning-Learning Loop for Robot Control with Learned World Models

Kowndinya Boyalakuntla, Yuhan Liu, Abdeslam Boularias
In 30 seconds

This preprint extends policy-constrained TD-MPC into PL-MPC by changing critic training targets, MPPI terminal values, and actor distillation without altering the world model or planner. On HumanoidBench, it lifts total average return from 98±18 to 387±255 on balance-hard and from 199±13 to 466±200 on hurdle; in zero-shot sim-to-real wrench–nut alignment on a KUKA IIWA14 it achieves 74.2% vs 61.3% success on the training object size. The result matters because it shows targeted interactions in the planning–learning loop can improve hard robot control tasks beyond the usual planner–policy alignment.

Research question

Can modifying critic supervision, planner terminal-value estimation, and planner-to-policy distillation improve robot control in a TD-MPC-style planning–learning loop beyond existing policy-alignment methods?

Problem

TD-MPC combines latent world models, MPPI, value critics, and an actor, but the planner and learner form a loop with three weak interfaces: one-step TD critics see little realized reward before bootstrapping; MPPI ranking uses possibly uncertain critic terminal values; and actor distillation from planner data ignores whether an episode actually succeeded. Existing policy-constrained variants address only the planner–actor mismatch, so hard tasks such as HumanoidBench balance-hard and hurdle remain weak.

Previous approach

Earlier methods use one-step TD-MPC targets, terminal values from the online critic ensemble without an uncertainty penalty, and actor updates constrained toward stored planner proposals in TD-M(PC)2. BMPC, BOOM, and PO-MPC also focus mainly on planner–policy alignment or value-weighted imitation. They do not alter critic multi-step supervision or penalize uncertain planner terminal values.

New approach

PL-MPC adds three mechanisms on top of TD-M(PC)2. MTD forms n-step critic targets with n=3 and default horizon H=3, using observed replay rewards where available and learned-model rollouts only beyond the sampled slice. ATE subtracts a normalized target-critic ensemble standard deviation from the MPPI terminal value. RAD adds return-weighted imitation of planner-executed actions, using realized episode returns normalized by a FIFO queue, with a warmup. The latent world model, critic ensemble, and MPPI optimizer are otherwise unchanged.

Results

On HumanoidBench locomotion, PL-MPC has the highest mean TAR on stand (933±3), maze (353±3), slide (910±7), hurdle (466±200), and balance-hard (387±255), and ties BOOM on pole; it improves mean over TD-M(PC)2 on 8 of 13 tasks and underperforms on 5, with large drops on run (658±355 vs 852±8) and sit-hard (635±233 vs 821±64). On DMControl it is competitive and highest on humanoid-walk (947±6). On balance-hard, removing MTD drops TAR from 387±255 to 186±51 and success from 19% to 0%; on hurdle, removing ATE drops to 296±21 and removing RAD to 231±72, while removing MTD yields 530±275. Targeted hurdle runs show MTD alone gives 288±116 vs 199±13 for TD-M(PC)2, and adding ATE+RAD gives 465±201. A 750K RAD warmup lifts crawl from 855±31 to 971±6 and sit-hard from 635±233 to 805±104 but lowers hurdle from 466±200 to 414±127. In zero-shot hardware tests, PL-MPC vs TD-M(PC)2 success rates are 74.2% vs 61.3% on training size 5 (31 trials each), 80.0% vs 73.3% on unseen size 3 (15 trials each), and 33.3% vs 20.0% on unseen size 1 (15 trials each); pooled unseen success is 56.7% vs 46.7%. PL-MPC also completes unseen-size successes in fewer steps, while final pose errors are comparable.

Limitations

Authors note the modifications are only tested on the TD-M(PC)2 backbone, performance is task-dependent with inconsistent gains, hardest tasks have high seed variability and survival rewards can mask success, and RAD requires a fixed warmup that is task-sensitive. Additional limitations: real-robot evidence is a single manipulation task on one KUKA IIWA14 with modest trial counts, no real-world fine-tuning, and code availability is only announced in the preprint. Training uses 3M environment steps, so compute may be substantial.

Industry impact

This could be adopted by robotics teams building humanoid locomotion or industrial manipulation controllers that already use TD-MPC-style learned world models, including developers of warehouse robots, industrial assembly, and humanoid platforms. As a training-time change, it is plausible within 1–3 years if the code is released, the results reproduce across backbones, and hardware validation is expanded. For production use, the task-dependent degradation and seed sensitivity would need to be resolved first.

The full text is not republished here because the paper's license does not allow it. Read the original on arXiv.