NavHarness: Adaptive Goals for Agentic Vision-Language Navigation
Noch nicht auf Deutsch verfügbar: Anzeige im englischen Original.
Aus dem Abstract
Vision-Language Navigation (VLN) requires embodied agents to generate actions based on instructions and observations. General-purpose multimodal agents offer a promising basis for this task, but selecting plausible local actions does not ensure that execution remains consistent with the intended route, particularly in long-horizon tasks.
Aus dem Abstract. Unsere Zusammenfassung folgt.