RawVLA: Embodied Neural Image Signal Processor For Robotic Manipulation
From the abstract
Vision-language-action (VLA) models typically operate on RGB images produced by a fixed camera image signal processor (ISP), leaving the imaging pipeline outside the learning and evaluation loop. We systematically examine the consequences of this overlooked design choice across five fundamental ISP dimensions: gain, sensor noise, chromatic response, tonal response, and bit depth.
From the abstract. Our summary is in progress.