Does Local Video Understanding Transfer Across Encounters? The EgoGears Benchmark
暂无中文版,显示英文原文。
摘自论文摘要
Embodied systems must make knowledge acquired during one encounter usable in another despite changes in viewpoint, motion, and illumination. Yet aggregate cross-video accuracy conflates failures of local perception with failures to preserve observation identity, establish correspondence, and compose evidence, obscuring whether local video understanding actually transfers.
摘自论文摘要,摘要生成中。