Text-to-3D Policy: Fine-Grained Language-Behavior Alignment for Unseen Specification Generalization
This preprint introduces T3DP, a text-to-3D policy framework that aligns instruction tokens with local behavior segments from demonstrations to improve generalization to unseen fine-grained behavior specifications. Across Meta-World, ManiSkill, and RoboTwin, T3DP improves average held-out-specification success over global language-behavior alignment by +11.0 to +14.2 points; on real-robot tasks it raises average success from 47.5% to 65.0% (+17.5 points). The result matters for robot systems that must follow precise language such as target positions, displacements, or articulated states without exhaustive demonstration coverage.