Towards Agile Vision-Based Multi-UAV Flight: Revisiting State Estimation
This preprint introduces vision-based pose-aware state estimators for neighboring multirotor UAVs, integrating visual tilt measurements to infer thrust direction, including a new linear thrust-constrained Kalman filter. Across two real-world and one simulated dataset, pose-aware methods cut mean velocity and acceleration errors by 40% and 57% versus position-only filters, remove a ~300 ms acceleration delay, and enable simulated follower tracking of >2 g lateral maneuvers where position-only estimation fails. This matters for collision avoidance and coordination in agile multi-UAV operations.
Can adding visual tilt measurements to vision-based multi-UAV state estimation remove the structural delay in higher-order state estimates and enable agile close-proximity flight?
Agile multi-UAV flight requires accurate, low-latency onboard estimates of neighbor position, velocity, and acceleration for collision avoidance and motion coordination. Vision-only pipelines typically measure only position and derive velocity and acceleration indirectly from displacement, creating a fixed structural lag in those states and limiting achievable agility.
Previously, position-only Kalman filters and similar displacement-based estimators were the common vision-based approach. They infer higher-order states by differentiating position, causing roughly 300 ms of delay in acceleration step response that does not improve with more agile motion.
The authors add tilt measurements from a state-of-the-art visual detector, which encode the thrust direction of co-planar multirotors, and benchmark five pose-aware estimators, including a novel linear thrust-constraining Kalman filter. Tilt is observed before displacement accumulates, so the filters can respond near the camera frame-rate limit rather than waiting for positional change.
Across two real-world and one photorealistic simulated dataset, pose-aware estimation reduced average velocity error by 40% and acceleration error by 57%; the proposed thrust-constraining KF outperformed the other eight tested estimators. Position-only filters showed a ~300 ms delay in acceleration step response independent of agility, while tilt-constrained estimators operated near the camera frame-rate response limit. In closed-loop NMPC leader-follower simulation, position-only estimation could not maintain stable hovering, whereas the proposed estimator tracked lateral maneuvers exceeding 2 g of acceleration.
The paper has no explicit limitations section. Evident constraints include reliance on a visual detector that returns tilt only for co-planar multirotor UAVs, evaluation limited to two real-world datasets and one simulated dataset, and the leader-follower closed-loop result demonstrated only in simulation. It is not stated whether code or datasets are released.
Drone swarm operators, aerial inspection and delivery companies, and defense or light-show multi-UAV platforms could use this to improve relative state estimation and autonomous coordination. Adoption in products is likely 1–3 years away, provided the tilt detector runs at camera frame rate on embedded hardware, performs reliably under occlusions and outdoor lighting, and the approach is validated on larger real-world swarms. Advanced R&D groups could integrate it now for prototyping.
The full text is not republished here because the paper's license does not allow it. Read the original on arXiv.