ROBOTNESS
ExpertarXiv

From Local Whole-Body VLA Behaviors to Scene-Scale Aerial Manipulation

Weixiang Guo, Rui Jin, Haotian Jin, Xinhang Xu, Ruiyang Liu, Haoran Zhao, Yi Wang, Weiqi Gai, Kun Cao, Lihua Xie
In 30 seconds

This preprint introduces a framework for scene-scale aerial manipulation on an articulated uncrewed aerial manipulator, combining synthetic whole-body VLA training, measured-progress-aligned trajectory realization, and scene-graph-guided mission composition. In simulation, local skills achieve 39/60 successes under oracle handoff, and full multi-site missions achieve 21/50 (42.0%), while a monolithic whole-task VLA baseline achieves 0/50. Physical trials validate representative tasks without platform demonstrations, reducing risk and data cost for aerial manipulation.

Research question

Can a vision-language-action policy trained only on synthetic whole-body demonstrations be reliably composed into scene-scale aerial manipulation missions on an articulated UAM despite asynchronous action execution and relational grounding requirements?

Problem

Extending VLA behavior to scene-scale aerial manipulation is difficult because collecting diverse whole-body demonstrations on physical articulated UAMs is costly and risky, asynchronous inference causes action-state misalignment, and relational language goals must be resolved to specific object instances and feasible whole-body handoff states. Existing systems address these aspects separately rather than providing an integrated path from local learned behavior to safe cross-site execution.

Previous approach

Prior VLA and visuomotor aerial work often relies on simulator teleoperation, handheld cross-embodiment demonstrations, Gaussian-splatting appearance augmentation, or navigation-centric trajectories; deployment methods such as AirVLA with RTC handle asynchronous execution without constrained whole-body trajectory realization. Scene-graph planners like SayPlan, DovSG, and USS-Nav resolve semantic instances or high-level skills but do not ground language goals to kinodynamically feasible whole-body interaction regions.

New approach

The framework uses a reconfigurable Gaussian-splatting-based supervision pipeline that samples task configurations, plans kinodynamically feasible whole-body expert trajectories, and renders synchronized base and wrist camera images without physical demonstrations. A policy fine-tuned from π0.5 via OpenPI predicts ten future actions at 0.25 s intervals from 224×224 multiview images and proprioception. MPAR projects estimated handoff states onto incoming VLA paths, builds a C² continuous takeover bridge, and jointly optimizes path progress and MINCO trajectory timing under platform constraints. A relational Scene Graph uses an LLM to compose a mission plan, grounds instructions to unique object instances, constructs skill-conditioned interaction regions, and performs topology-guided cross-site transfer with handoff verification.

Results

In simulation under oracle target and feasible handoff, local VLA skills succeed in 39/60 interactions (65.0%), with 14/30 object-acquisition and 25/30 post-acquisition successes. Runtime evaluation across 360 paired trials shows MPAR succeeds in 78/120 (65.0%) versus 71/120 (59.2%) for nominal-time alignment and 54/120 (45.0%) for RTC-Direct; under 500-ms added latency and stalls it reaches 65.0% versus 55.0% and 43.3%. MPAR reduces collision-bearing runs to 1/120 versus 4/120 nominal and 22/120 RTC-Direct, and lowers median takeover phase error by 0.212 s. Full scene-scale missions achieve 21/50 (42.0%) versus 0/50 for a monolithic whole-task VLA, and the scene-graph physical interface succeeds in 14/15 episodes with 9/15 full missions; removals of grounding, routing, or feasible handoff drop missions to 0/15, 4/15, and 2/15 respectively. Physical local trials achieve 2/5 bottle grasping, 1/5 block grasping, 3/5 watering, and 3/5 storage, while the full cross-site Water Plant mission succeeds in 2/5 trials.

Limitations

The authors state the system requires a preconstructed Scene Graph registered to a metric map and depends on remote VLA inference, and mission-level performance is limited by individual local skills. The physical evaluation is small, with 5 trials per local task and 5 full mission trials, and simulated full-mission success remains low at 42.0%. The approach uses one articulated UAM platform and a remote RTX 5090 for inference, lacks onboard scene-graph update or closed-loop failure processing, and no code release is reported.

Industry impact

This could inform companies building articulated aerial manipulators for infrastructure inspection, agriculture, logistics, or environmental interaction, especially where multi-site tasks require whole-body coordination. Product use is likely 3+ years away: current success rates are too low for unattended operation, inference runs remotely, and deployment requires a prebuilt metric scene graph. Onboard inference, online scene construction, stronger manipulation policies, and expanded physical validation would be prerequisites.

The full text is not republished here because the paper's license does not allow it. Read the original on arXiv.