Does Local Video Understanding Transfer Across Encounters? The EgoGears Benchmark
From the abstract
Embodied systems must make knowledge acquired during one encounter usable in another despite changes in viewpoint, motion, and illumination. Yet aggregate cross-video accuracy conflates failures of local perception with failures to preserve observation identity, establish correspondence, and compose evidence, obscuring whether local video understanding actually transfers.
From the abstract. Our summary is in progress.