Experience-Driven Continual Learning of Terrain Traversability for Quadruped Robots
This preprint presents a quadruped traversability system that predicts five foot-contact outcomes from pre-contact camera images using a frozen DINOv3 backbone and an evidential regressor, then continually updates the model through replay with a historical validation gate. On a sequential real-robot stream across three unseen surface types, the gate cut anchor negative log-likelihood degradation by 23.1% versus replay without the gate while keeping new-terrain adaptation nearly equal. It matters for legged robots that must operate safely on unfamiliar ground without forgetting earlier experience.
Can learned terrain-interaction predictions be updated continually from the robot's own experience so that new terrain types are learned without degrading performance on previously seen terrain?
The robot needs to know before contact how a surface will affect slip, foot loading, traction, and energy use, but visual geometry and appearance do not reveal these interaction outcomes; existing online adaptation methods either forget older terrains or produce overconfident predictions on unfamiliar ground.
Earlier systems used geometric and semantic traversability mapping, self-supervised visual-proprioceptive transfer, online learning with experience replay, and uncertainty-aware planners; they typically optimized on recent data or used a single traversability score without a held-out historical acceptance test, leaving retention and calibration incomplete.
The pipeline freezes a DINOv3 ViT-S/16 backbone and trains a compact multi-head evidential MLP on causal visual-proprioceptive samples. It produces five confidence-weighted indicators plus aleatoric and epistemic uncertainty. During deployment, a background candidate trains on recent contacts and a bounded replay set of up to 150 samples, drawn from a 1000-element memory that is quantile-sampled by slip. A candidate is promoted only if recent held-out negative log-likelihood improves and historical anchor negative log-likelihood does not worsen by more than 0.15% relative. Predictions are projected into a local 0.1 m grid map with configurable property weights and conservative uncertainty correction.
Offline held-out tests reported simulation slip MAE 0.0050 m, R2 0.3958, 80% coverage 82.11%; real-robot slip MAE 0.0047 m, R2 0.4110; real traction index MAE 0.0296, R2 0.5269; cost of transport MAE 0.235608, R2 0.3128; loading rate MAE 321.605 N/s, R2 0.5908. Spatial separation between favorable and unfavorable surfaces was delta tau 0.2991 in simulation and 0.4382 on hardware. On one hardware sequence of artificial grass to textured foam to foam with 951 accepted contacts, CL-P versus Replay gave adaptation 1.658% versus 1.664%, lab-test degradation 0.801% versus 0.924%, anchor NLL degradation 0.608% versus 0.790%, and final macro standardized MAE 1.03711 versus 1.03633 across five seeds. CL-P promoted 53 of 80 candidates and rejected 24 by the anchor gate; Replay promoted 70 of 80. In simulation navigation, both Offline and CL-P succeeded in 5 out of 5 trials; unfavorable-terrain exposure was 3.13% versus 2.76%, and path length was 23.75 m versus 24.53 m.
The authors state that the hardware comparison uses a single recorded stream with five optimization seeds and separately trained simulation and hardware models, so no overall accuracy advantage or isolated navigation benefit is established. They list longer terrain sequences, deformable terrain, varying locomotion conditions, and active uncertainty sampling as future work. Evident limitations include a small set of real surfaces, external PC compute rather than onboard deployment, no sim-to-real transfer, few closed-loop navigation trials, no reported code release, and the gate's downstream navigation impact was not isolated.
Legged robot manufacturers and operators using platforms such as the Unitree Go2 could apply this to inspection, construction, agriculture, and logistics robots that move over unfamiliar or variable ground. The system is a research prototype now; practical field deployment may be possible in 1–3 years if it is optimized for onboard compute, validated on more real-world sites and longer missions, and integrated with certified navigation stacks.
The full text is not republished here because the paper's license does not allow it. Read the original on arXiv.