STARS: From Spatiotemporal Dynamics to Social Representations in Human-Robot Interaction
日本語版は未提供のため、英語原文で表示しています。
要旨より
Robot navigation in dynamic, human-centered environments requires socially-compliant decisions grounded in robust scene understanding. Recent Vision-Language Models (VLMs) exhibit promising capabilities such as object recognition, common-sense reasoning, and contextual understanding, capabilities that align with the nuanced requirements of social robot navigation.
要旨より。当社による要約は作成中です。