GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives
This preprint introduces GroundingPI, a 4B grounding foundation model that predicts object points and boxes as quantized token coordinates rather than using a general-purpose vision-language backbone. Across 34 grounding benchmarks it averages 73.68%, ahead of a larger GPT-6 Astra baseline at 71.54%, and as a visual backbone it improves downstream manipulation (RoboTwin 2.0, RoboCasa-GR1) and autonomous driving (nuScenes L2 0.296 m). The result matters because it argues grounding is a distinct perceptual layer that can make embodied foundation models more precise.