Visual Grounding上級
GroundAnything: Reconciling Parallel Decoding with Precise Visual Grounding at Flash Speed
Preprint. The authors introduce GroundAnything, a 4B-parameter visual grounding model that uses bidirectional diffusion with blockwise denoising for parallel spatial decoding instead of sequential autoregressive token generation. Its autoregressive variant, GroundAnything-VLM, reports 72.42% across 30 grounding benchmarks, ahead of GPT-6 Astra at 71.35%. This matters for latency-sensitive robotics and interactive vision systems that need fast, precise localization.