Toward Real-Time VLAs: Stage-Aware Two-Step Flow Denoising and System-Level Evaluation
This preprint reports a two-stage non-uniform Flow Matching denoiser that cuts the π0.5 VLA's action-generation steps from 10 to 2 and model-inference time from 61.6 ms to 22.0 ms, a 2.8× speedup. The authors also built a distributed real-time VLA execution framework and benchmarked six action-scheduling methods on a 180-trial bimanual physical garment-folding task; Legato achieved 96.7% success as the best training-based method, while Temporal Smoothing led training-free methods at 76.7%. Combining the fast sampler with those methods gave 1.68–2.80× inference speedups at a 6.7–10.0 percentage-point success loss, highlighting a practical speed–quality tradeoff for real-time robot policies.