STARS: From Spatiotemporal Dynamics to Social Representations in Human-Robot Interaction
暂无中文版,显示英文原文。
摘自论文摘要
Robot navigation in dynamic, human-centered environments requires socially-compliant decisions grounded in robust scene understanding. Recent Vision-Language Models (VLMs) exhibit promising capabilities such as object recognition, common-sense reasoning, and contextual understanding, capabilities that align with the nuanced requirements of social robot navigation.
摘自论文摘要,摘要生成中。