ROBOTNESS
专业arXiv

GroundAnything: Reconciling Parallel Decoding with Precise Visual Grounding at Flash Speed

Qize Yu, Lianrui Fan, Bowen Ping, Xini Ding, Zetian Song, Junbo Niu, Kaixuan Wang, Tianxing Chen, Yue Chen, Minghua He, Yuran Wang, Jie Huang, Haojun Zhang, Min Chen, Hao Li, Wenxuan Song, Ruihai Wu, Xianming Liu, Shilong Liu, Shuchang Zhou
暂无中文版,显示英文原文。
30 秒速读

Preprint. The authors introduce GroundAnything, a 4B-parameter visual grounding model that uses bidirectional diffusion with blockwise denoising for parallel spatial decoding instead of sequential autoregressive token generation. Its autoregressive variant, GroundAnything-VLM, reports 72.42% across 30 grounding benchmarks, ahead of GPT-6 Astra at 71.35%. This matters for latency-sensitive robotics and interactive vision systems that need fast, precise localization.

研究问题

Can visual grounding be reformulated as parallel diffusion-based evidence extraction to remove autoregressive latency and fixed causal order while retaining precise localization?

问题

Autoregressive grounding models serialize spatial predictions, which adds sequential latency and forces a left-to-right causal order even though objects, locations, and spatial relations are jointly constrained by the image and query.

既有方法

Prior work used autoregressive (AR) grounding models that generate spatial outputs one token at a time. Their shortcoming is sequential decoding latency and an imposed causal order that does not reflect the joint nature of visual evidence.

新方法

GroundAnything is a 4B-parameter grounding foundation model that uses bidirectional diffusion with blockwise denoising so spatial hypotheses can emerge in parallel and be refined iteratively. Training includes grounding pretraining on public datasets and dedicated data engines, direct AR-to-diffusion conversion with joint AR and diffusion objectives, supervised fine-tuning, and GRPO-based reinforcement post-training; an entropy-guided decoding option is also used.

结果

Across 30 grounding benchmarks, the autoregressive variant GroundAnything-VLM reaches 72.42%, a new overall state of the art among similarly sized models, and is competitive with GPT-6 Astra at 71.35%. The abstract also notes that with entropy-guided decoding, GroundAnything surpasses an unspecified baseline, but the remaining comparison is cut off in the available abstract.

局限

Author-stated limitations are not reported in the abstract. Evident limitations include: preprint without peer review; no latency, throughput, or hardware details despite a 'flash speed' claim; no code or model release information in the abstract; results are limited to 30 grounding benchmarks with no robotic or real-world deployment data; and the truncated abstract omits the full entropy-guided decoding comparison.

产业影响

If the reported accuracy and parallel decoding hold up, this could benefit robotics companies that need fast visual grounding for manipulation, autonomous vehicle perception stacks, AR/VR platforms doing object localization, and VLM API providers. Adoption would likely take 1–3 years after peer review, open weight/code release, and validation on domain-specific latency and accuracy requirements.