Beyond Policy Alignment: Closing the Planning-Learning Loop for Robot Control with Learned World Models
This preprint extends policy-constrained TD-MPC into PL-MPC by changing critic training targets, MPPI terminal values, and actor distillation without altering the world model or planner. On HumanoidBench, it lifts total average return from 98±18 to 387±255 on balance-hard and from 199±13 to 466±200 on hurdle; in zero-shot sim-to-real wrench–nut alignment on a KUKA IIWA14 it achieves 74.2% vs 61.3% success on the training object size. The result matters because it shows targeted interactions in the planning–learning loop can improve hard robot control tasks beyond the usual planner–policy alignment.