Toward Real-Time VLAs: Stage-Aware Two-Step Flow Denoising and System-Level Evaluation
This preprint reports a two-stage non-uniform Flow Matching denoiser that cuts the π0.5 VLA's action-generation steps from 10 to 2 and model-inference time from 61.6 ms to 22.0 ms, a 2.8× speedup. The authors also built a distributed real-time VLA execution framework and benchmarked six action-scheduling methods on a 180-trial bimanual physical garment-folding task; Legato achieved 96.7% success as the best training-based method, while Temporal Smoothing led training-free methods at 76.7%. Combining the fast sampler with those methods gave 1.68–2.80× inference speedups at a 6.7–10.0 percentage-point success loss, highlighting a practical speed–quality tradeoff for real-time robot policies.
How can model-side Flow Matching sampling cost and robot-system timing together be reduced and coordinated to close the gap between low-rate VLA inference and high-rate physical control?
VLA policies update actions too slowly for high-rate robot control, and repeated Flow Matching denoising accounts for much of the model cost; even with asynchronous execution, perception, state feedback, communication, and physical response add substantial delay. The paper measures primary-camera timestamp offset at 17.7 ms, image readout at 32.8 ms, proprioceptive feedback at 26.4 ms, and command-to-motion response at 15.3 ms, on top of 61.6 ms of model inference, so new action chunks are conditioned on stale observations and can create discontinuities.
Prior work either improves sampling speed via distillation or shortcut methods such as SnapFlow and Consistency Policy, improves chunk continuity via methods like Legato, VLASH, RTC, and Temporal Smoothing, or tunes system-level latency, but studies use different base models, platforms, tasks, and runtime settings. They lack a common measurement framework and often optimize one component without an end-to-end latency accounting.
The authors introduce two-stage non-uniform denoising: a standard Flow Matching step covers the interval [1, 0.3], then a SnapFlow-style shortcut step maps from t=0.3 to 0, with the same action expert trained on a 50/50 mixture of standard Flow and two-stage objectives and a terminal-refinement weight of 0.1. They also build a distributed, thread-decoupled VLA framework with separate inference, action-publication, and robot-control rates, pluggable execution strategies, joint sensor/actuation latency calibration, and action-provenance logging.
On an RTX 4090 D, two-stage non-uniform denoising reduced NFE from 10 to 2 and mean inference time from 61.557 ms to 21.956 ms (2.804×, 64.33% reduction); offline action errors remained on the same order, e.g., joint MAE 0.008525 vs 0.008396 rad on one sequence and 0.006503 vs 0.006797 on another. In 180 physical bimanual garment-folding trials, Legato achieved 29/30 success (96.7%) with 73.56 s mean time and 47.31 tasks/hour; VLASH achieved 28/30 (93.3%, 77.73 s, 43.22/h); Temporal Smoothing 23/30 (76.7%, 93.93 s, 29.38/h); Naive Asynchronous and Training-time RTC each 19/30 (63.3%); Inference-time RTC 18/30 (60.0%). Continuity on left_j3: Legato had Tail Gap 0.0193 rad and Switch Gap 0.0062 rad, while Naive Asynchronous reached 744.1 rad/s² maximum acceleration and gaps 0.1799/0.1538 rad; Temporal Smoothing had the lowest maximum acceleration at 25.6 rad/s² but higher gaps. Combining the 2-NFE sampler with Legato cut inference from 36.808 ms to 21.956 ms (1.676×) but lowered success from 96.7% to 86.7%; with Temporal Smoothing, inference dropped from 61.557 ms to 21.956 ms (2.804×) and success fell from 76.7% to 70.0%, with continuity metrics worsening by about 20%.
The authors state that aggressive NFE compression causes clear closed-loop losses: combined configurations lose 6.7–10.0 percentage points of success and increase continuity errors by about 20%. The evidence comes from one base policy (π0.5), one long-horizon bimanual garment-folding task, one hardware platform, 30 trials per method, and two offline validation sequences; the physical trials are unpaired and no between-method significance tests are performed. The two-stage schedule fixes τ=0.3 and has not been tested across other Flow Matching policies, tasks, or robot configurations; the work is a preprint.
Robot and VLA developers building dexterous manipulation, bimanual automation, humanoid, or fabric-handling systems could use the latency-calibration and execution framework now, while the trained two-stage sampler could reduce edge-GPU inference cost for Flow Matching policies within 1–3 years if validated across more models and tasks. Companies working on service robots, warehouse automation, or humanoid foundation models would benefit from faster action generation and explicit action continuity when deploying on physical hardware; adoption depends on broader multi-task validation, safety certification, and integration with production control stacks.
The full text is not republished here because the paper's license does not allow it. Read the original on arXiv.