RawVLA: Embodied Neural Image Signal Processor For Robotic Manipulation
日本語版は未提供のため、英語原文で表示しています。
要旨より
Vision-language-action (VLA) models typically operate on RGB images produced by a fixed camera image signal processor (ISP), leaving the imaging pipeline outside the learning and evaluation loop. We systematically examine the consequences of this overlooked design choice across five fundamental ISP dimensions: gain, sensor noise, chromatic response, tonal response, and bit depth.
要旨より。当社による要約は作成中です。