Text-to-3D Policy: Fine-Grained Language-Behavior Alignment for Unseen Specification Generalization
This preprint introduces T3DP, a text-to-3D policy framework that aligns instruction tokens with local behavior segments from demonstrations to improve generalization to unseen fine-grained behavior specifications. Across Meta-World, ManiSkill, and RoboTwin, T3DP improves average held-out-specification success over global language-behavior alignment by +11.0 to +14.2 points; on real-robot tasks it raises average success from 47.5% to 65.0% (+17.5 points). The result matters for robot systems that must follow precise language such as target positions, displacements, or articulated states without exhaustive demonstration coverage.
Can text-to-3D policies generalize to unseen fine-grained behavioral specifications—such as target position, displacement, or articulated state—within a known task family, when those specifications are absent from training demonstrations?
A robot instruction may share coarse task semantics with training tasks but require distinct behaviors, such as different target positions, distances, or desired states. Existing text-to-3D policies struggle because generic pretrained language embeddings and conventional global language-behavior alignment compress whole instructions and trajectories into single embeddings, blurring sparse but decisive local distinctions. Exhaustive demonstration coverage of the specification space is impractical, and the same observation can require different actions, so language must disambiguate behavior.
Prior 3D language-conditioned policies such as PerAct, Act3D, 3DDA, and 3D-LOTUS condition actions on language but are not designed to preserve fine-grained specification differences. Behavior grounding methods such as LIV, GRIF, DecisionNCE, and T2DA align globally pooled instruction and trajectory embeddings; this mostly enforces task-level or trajectory-level compatibility and can obscure local differences between closely related specifications.
T3DP uses three stages. First, a behavior encoder processes demonstration windows containing state, action, and reward tokens and is trained to predict the reward for held-out state-action pairs, forcing the embedding to capture specification-dependent behavior. Second, with the behavior encoder frozen, a pretrained text encoder is fine-tuned using a combined contrastive loss: global pooled alignment plus bidirectional late-interaction alignment between instruction tokens and local state-action hidden states. Third, the aligned text embedding conditions a point-cloud-based 3D diffusion policy; at inference only language and observation are needed.
On Meta-World, T3DP averages 50.3% success versus T2DA 36.1% and DP3 16.2% (+14.2 over global alignment); on ManiSkill 54.2% versus T2DA 40.7% and DP3 20.4% (+13.5); on RoboTwin the abstract reports +11.0 over global alignment. All 15 simulated task families improved. Meta-World per-task: Reach 38.3 vs 24.3, Push 64.7 vs 45.7, Pick Place 58.0 vs 46.0, Door Open 59.7 vs 41.3, Drawer Open 30.7 vs 23.3. ManiSkill per-task: Reach 38.0 vs 32.3, Push Cube 47.3 vs 15.0, Pull Cube 58.7 vs 48.3, Pick Cube 77.7 vs 72.0, Turn Valve 49.3 vs 36.0. Real robot: average 65.0% versus 47.5% for global alignment across two task families, 20 specifications each. Ablation on Meta-World: no-language 16.2, raw language 19.1, global-only 36.1, T3DP 50.3. Scaling with 10/20/40 training specifications: T3DP 32.4/50.3/76.4 versus T2DA 23.9/36.1/69.2. Joint multi-task Meta-World: T3DP 40.7 versus T2DA 35.5 and DP3 11.3. Action probe on Reach: MAE 0.1238 for T3DP versus 0.1423 for T2DA and 0.2567 state-only.
Authors state T3DP requires expert trajectories with reward signals and is evaluated only on held-out specifications within known task families under templated instructions. Real-robot evaluation uses a single PiPer X arm, a fixed overhead D435i camera, two task families, and 20 specifications per task. The paper does not report code availability.
Robotics companies building language-conditioned manipulation arms for logistics, assembly, or service tasks could use fine-grained alignment to make instruction following more precise without exhaustive demonstrations. Impact could be felt in 1–3 years where reward-labeled expert demonstrations and known task families are available; broader use requires free-form instructions, unseen task families, and less dependence on reward signals and per-task training.
The full text is not republished here because the paper's license does not allow it. Read the original on arXiv.