ROBOTNESS
专业arXiv

MVP-SLAM:多相机视觉惯性SLAM用楼层平面图在线纠正施工场景漂移

Asier Bikandi-Noya, Miguel Fernandez-Cortizas, Muhammad Shaheer, Holger Voos, Jose Luis Sanchez-Lopez
30 秒速读

MVP-SLAM是一篇预印本,提出一种多相机视觉惯性SLAM系统,利用建筑楼层平面图在线校正漂移。在Hilti–Trimble SLAM Challenge 2026中,它在Localization任务以0.29米平均RMSE排名第2,SLAM任务以0.24米排名第5,是所有在线运行、整合平面图并定位到楼层平面图中的方法中成绩最好的。这项工作表明,仅靠相机和IMU也能在施工场景利用平面图先验在线持续抑制累积误差。

研究问题

能否在多相机视觉惯性SLAM系统中,仅依靠相机和IMU,在线利用预先获得的楼层平面图来持续纠正施工场景中的漂移并定位到平面图坐标系?

问题

室内建筑施工现场存在光线变化、低纹理和重复结构,视觉SLAM长期运行会产生漂移;现有利用平面图校正轨迹的方法大多离线进行,或依赖深度传感器,无法在手持相机设备上实时在线校正。

既有方法

此前方法分两类:离线方法如Z-FLoc、OmniRecon、CUFE在轨迹完成后将整体结果对齐到平面图;在线方法通常依赖LiDAR或RGB-D相机等主动深度传感器获取墙壁结构,缺乏纯视觉在线平面图整合方案。也有方法仅用平面图拒绝错误回环,不直接校正漂移。

新方法

MVP-SLAM扩展ORB-SLAM3为多相机视觉惯性前端,支持两台背对背鱼眼相机;通过跨相机度量初始化和IMU辅助的地图合并解决初始化与丢帧后地图断裂问题。墙检测使用预训练全景分割器EoMT对地图点分类,序列RANSAC拟合竖直平面;漂移感知关联根据行驶距离自适应放宽匹配阈值。平面图集成采用三阶段优化:局部对齐迷你图、传播位姿图、将校正持久化为各向异性先验,逐步将匹配墙壁纳入后端优化。

结果

在官方Hilti–Trimble SLAM Challenge 2026数据集上,24个Localization任务序列中,MVP-SLAM平均APE RMSE为0.29米,排名第2/22,仅落后Z-FLoc的0.24米;SLAM任务平均0.24米,排名第5/62,最高ACDC-VSLAM为0.09米。移除平面图集成后,Localization平均误差1.45米,SLAM平均误差0.73米,分别改善4.9倍和3.0倍。墙检测精度96.9%,召回38.4%;墙匹配精度、召回、F1均为98.9%。

局限

作者承认,系统仅在检测到墙壁并与平面图匹配时才校正漂移,在大型地下空间等以柱为主、缺少墙面场景中匹配不足。墙检测召回率仅38.4%,可见但点云稀疏或杂乱的墙壁无法稳定拟合。论文未报告在其他数据集或机器人平台上的验证,也未说明代码是否开源以及实时性能。

产业影响

该技术可应用于施工进度跟踪、建筑机器人巡检和基于BIM的定位导航,目标公司包括建筑施工技术、测量和室内定位方案商。短期(1至3年)内有望集成到手持或多相机机器人测绘系统,前提是具备可用的楼层平面图或BIM、粗初始位姿和足够的GPU算力。长期(3年以上)可扩展至更多结构元素(如柱)和动态施工环境。

论文全文

MVP-SLAM: Multi-Camera Visual-Inertial Floorplan-Prior SLAM

Asier Bikandi-Noya, Miguel Fernandez-Cortizas, Muhammad Shaheer, Holger Voos, Jose Luis Sanchez-Lopez

本文依据 CC BY 4.0 许可发布,经署名转载,原文见 arXiv:2609.39596(PDF)。

Abstract

Indoor building construction sites are demanding environments for visual SLAM, where variable lighting and repetitive, low-textured structures make the system drift over long trajectories, though structural elements such as walls remain distinguishable despite these conditions. These buildings are constructed according to their as-planned floor plans, available from the design phase, and although the actual as-built site can differ from this design, floor plans still provide a metric reference, both to localize the system in the building and to correct drift. Existing methods often use the floor plan to correct an already-built trajectory offline, and those that instead correct it online typically rely on depth sensors. We instead present MVP-SLAM, an online visual-inertial SLAM on two opposite-facing fisheye cameras that corrects drift from cameras alone by matching walls detected in its map to the floor plan, through a drift-aware policy. A multi-stage integration then turns each matched pair incrementally into a persistent correction, so the trajectory stays corrected and localized within the floor plan as it is built. MVP-SLAM was validated on the multi-floor construction sites of the Hilti–Trimble SLAM Challenge 2026, ranking 2nd of 22 teams in the Localization task (0.29 m mean RMSE) and 5th of 62 teams in the SLAM task (0.24 m), the top-ranked one in both tasks among those that operate online, integrate the floor plan, and localize within it.

I Introduction

Accurate and reliable localization in indoor building construction environments is essential for automating construction workflows, such as tracking a site’s progress or enabling autonomous robotic inspection [1]. Visual-inertial SLAM is a practical tool for this goal, but construction site environments pose a challenge for these systems due to variable lighting, moving workers, fast motions, and repetitive, low-textured structures [2], causing the estimated trajectory to accumulate drift. Nevertheless, construction environments present structural elements, such as walls and columns, that remain distinguishable despite these conditions.

Fig. 1: MVP-SLAM on a Hilti construction sequence, matching detected walls to the floor plan and correcting the estimated trajectory against ground truth.

Construction sites typically have floor plans available [3], 2D representations of these structural elements produced during the design phase that can serve as a reference to correct this drift. In more advanced cases, these are Building Information Models (BIMs), encoding this information in more detail, though they are not universally available and are costly to process [4]. Although discrepancies exist between the as-built site and this as-planned floor plan [5], they are also used to localize in the building. Visual localization algorithms, like Z-FLoc [6], localize within such a floor plan and correct drift against it, but only after a trajectory has been generated, in a post-processing step. Onsite operation, however, needs a camera position estimate while the site is still being traversed, not only once it is complete, so the correction must happen online, as the trajectory is built, rather than recovered afterwards.

Methods offering an online camera position estimation in construction sites integrate the floor plan directly into the SLAM back-end. These methods rely on semantic entities such as walls to establish correspondences between the as-built site and as-planned floor plan, incorporating each match into the back-end’s graph optimization [1]. However, these systems rely on active depth or range sensors, such as anchoring LiDAR point clouds to BIMs [5] or matching structural walls from RGB-D cameras [7], leaving visual-only settings largely unexplored. A particular challenge for these systems is to find reliable correspondences between detected structural elements, such as walls on the as-built site and their counterparts on the floor plan, since accumulated drift makes this association increasingly difficult as the environment grows.

The Hilti–Trimble SLAM Challenge 2026 [4] benchmarks such real-world construction sites and their challenging conditions, recorded with a hand-held device carrying two opposite-facing (back-to-back) fisheye cameras and an IMU, without depth sensor; the cameras cover 360∘, sharing too little overlap to act as a stereo pair. Floor plans are also available for these sites, and one of the challenge’s two tasks, Localization, recovers the camera trajectory within the floor-plan frame from an initial pose, while its SLAM task instead recovers it in an arbitrary frame. Both tasks are scored on trajectory accuracy alone, so neither requires online or real-time operation.

Fig. 2: System architecture of MVP-SLAM. Sensor and floor-plan inputs (left) feed a three-module pipeline whose optimized state is fed back into the SLAM system as persistent plan priors, producing the plan-aligned trajectory estimate (right).

To recover the trajectory in the plan’s frame while correcting drift, we present MVP-SLAM (Multi-Camera Visual-Inertial Floorplan-Prior SLAM), an online visual-inertial SLAM system that leverages the building’s floor plan to constrain its estimation, responding to the Hilti–Trimble SLAM Challenge 2026. Our contributions are: (i) an online multi-camera visual-inertial SLAM system that integrates the building’s floor plan to correct drift as the trajectory is built; (ii) a semantic wall detection and drift-aware association algorithm that establishes correspondences between the as-built map and the floor plan from cameras alone; and (iii) a floor-plan integration method that, through a multi-stage strategy, turns each match incrementally into a bounded, persistent correction, so the map and trajectory are jointly optimized online under visual, inertial, and plan constraints. MVP-SLAM ranked 2nd in the Localization task of the challenge and 5th in the SLAM task, the best-placed system in both among those that also operate online, integrate the floor plan, and localize within it.

II Related Work

Multi-camera visual-inertial SLAM

Monocular visual-inertial SLAM suits hand-held capture, but its limited field of view degrades tracking during poor lighting or occlusions. Multi-camera setups widen visual coverage to maintain feature continuity, but whereas MAVIS [8] relies on overlapping stereo pairs, purely visual non-overlapping setups [9] suffer from motion-dependent scale unobservability and drift in the absence of an IMU. On the Hilti–Trimble SLAM Challenge 2026 [4] rig, recent systems estimate the trajectory with feature-based visual-inertial SLAM [10, 11], while others recover the full trajectory offline through global factor-graph optimization [12]. These systems, however, leave the trajectory in an arbitrary frame, neither localizing within a floor plan nor exploiting the one available on these construction sites to correct it.

Localization with architectural priors

Several methods localize a camera within a pre-existing floor plan, aligning monocular detections or layout cues to it through semantic Monte-Carlo localization [3] or learned ray-based filtering [13], returning a global position but without using the floor plan to improve the trajectory or the map. In the challenge’s Localization task [4], Map-It Ralph [14] likewise makes no use of the floor-plan geometry, reaching the plan frame from the provided anchor pose alone. The top three teams, in contrast, do use the floor plan, aligning a completed estimate to it offline: Z-FLoc by a single global transform from a bird’s-eye reconstruction [6], OmniRecon by 2D ICP on an offline structure-from-motion cloud [15], and CUFE by refining completed trajectories for floor-plan consistency [16]. All of them, however, apply the floor plan only after the trajectory is complete, rather than integrating it into the estimation to correct drift as the trajectory is built.

Architectural priors inside SLAM

Architectural plans can be leveraged inside SLAM systems to localize estimates within a building structure while offering an external metric reference to correct drift. A recent system from the same challenge [17] relies entirely on the benchmark’s provided starting pose to localize within the building, whereas real-world inspection tasks normally start from an estimated position, such as that obtained by having an operator match a few detected walls to the floor plan [7]. Furthermore, it uses the floor plan only to reject false loop-closure candidates, without directly correcting the drift. Other works do integrate the floor plan into the estimation to correct drift, but sense the walls directly with a range or depth sensor. LiDAR pose-graph methods anchor scans to a BIM through multi-session alignment [5], while [1] matches an online scene graph of rooms and walls to one extracted from the floor plan, and ivS-Graphs [7] brings this to vision with a simpler, wall-only matching, from an RGB-D sensor. Such sensing is uncommon on hand-held site-capture rigs, and depth cameras reach only a few meters, short of a construction interior’s spans. A depth-free system [18] instead registers sparse map points to a digital twin, but its prior is a dense, photorealistic as-built mesh that must be acquired separately and updated as the site evolves, unlike an as-planned floor plan available from the design stage. Correcting drift online by integrating an as-planned floor plan into a camera-only visual-inertial estimation remains unexplored.

III Methodology

III-A System Overview

MVP-SLAM corrects the drift of a visual-inertial trajectory and localizes inside the building by introducing the as-planned floor plan into a graph-based back-end. We build it upon ORB-SLAM3 [19], a widely adopted and extensively benchmarked SLAM system of this kind.

MVP-SLAM takes four inputs: the two fisheye image streams, an IMU, a floor plan of the site, and one plan-frame initialization pose \mathbf{T}_{s}, needed only as a coarse estimate rather than ground truth, since later corrections absorb its error (Sec. III-D). It processes them through three modules, each supplying what the next requires (Fig. 2): a multi-camera visual-inertial SLAM (Sec. III-B) estimates a continuous, metric trajectory and reconstructs a map; a wall detection and matching module (Sec. III-C) detects walls W_{s} in the map and matches them to their floor-plan counterparts W_{\mathrm{plan}}; and an integration module (Sec. III-D) folds each matched wall into a multi-stage back-end optimization that jointly re-estimates the trajectory and map, aligning them with the floor plan (Fig. 1).

III-B Multi-camera visual-inertial SLAM

MVP-SLAM extends monocular-inertial SLAM to the sensor setup of the rig. Facing opposite directions, the two fisheye cameras share too little overlap for stereo matching; each stream is tracked monocularly against a single shared map, with features from both hemispheres and the high-frequency IMU constraining one trajectory. The extension spans initialization, tracking, and map merging when the track is lost.

Fast initialization. A first metric map anchors the rest of the trajectory. Monocular-inertial pipelines build it only once enough parallax is available [20, 19], which a sequence beginning at rest or turning in place does not provide, while stereo pipelines build it at once from a wide image overlap the rig lacks. In between, Li et al. [21] showed that, for multi-camera rigs with limited view overlap, the few features matched across cameras are enough to seed the map from a single frame. MVP-SLAM follows this idea, adapted to the rig and its IMU, trying two paths in order.

(a) Cross-camera metric seeding: Front-camera features falling in the narrow peripheral band shared by both cameras are searched in the back image along known cross-camera epipolar curves over a small set of depth hypotheses, and the median depth of the surviving matches sets the scale, every remaining feature placed at that depth along its ray. Unlike the vision-only method of [21] needing an accurately triangulated seed, this coarse map suffices: the 4 cm inter-camera baseline could not provide accurate depths, but the first inertial bundle adjustment recovers them together with the scale.

(b) Gyro-aided fallback: In visual odometry and outlier rejection, taking the inter-frame rotation from the gyroscope rather than estimating it has proven to make two-view geometry more robust [22, 23]: it reduces the problem from five degrees of freedom to the two of the translation direction, recovered by a 2-point RANSAC [23]. MVP-SLAM applies this to initialization: when too few cross-camera matches survive, it falls back to two views with the rotation from gyro preintegration. This also removes the choice between a homography and a fundamental matrix, unreliable where planar surfaces dominate and parallax is small, which improves the robustness of the initialization in these challenging scenes.

Multi-camera tracking and optimization. The back camera is rigidly attached to the front one, so, as in other multi-camera visual-inertial systems [24], its observations constrain the same visual-inertial state from the opposite viewing hemisphere. For a map point seen by the back camera, its reprojection residual composes the estimated front-camera pose with the fixed inter-camera extrinsic, known from the IMU-camera calibrations, before applying the Kannala–Brandt fisheye projection. Each camera populates the map on its own, back-camera points being triangulated temporally between keyframes with the rig motion as baseline. A point is then searched only in the camera that created it while it stays in that field of view: as the two hemispheres barely overlap, searching it in both cameras would only add computational cost and cross-camera mismatches between similar structures. When it leaves, during turns or corridor reversals, it is projected into the opposite camera, which re-observes it and keeps it tracked across the hemisphere change. When both cameras end up mapping the same structure, fusion merges the duplicated points and culling removes the redundant ones.

Inertial map merging. Tracking can still be lost, on a fast turn or in front of a textureless wall [2]. When the pose cannot be recovered in the current map, multi-map systems such as ORB-SLAM3 start a new one and connect it to the previous map only if place recognition later identifies an already mapped area [19]. Because the floor plan is anchored to the first map through the initialization pose, a new map opened at its own origin stays outside the plan frame until the operator walks through an already mapped area again, which may never happen in a walkthrough without loops, and no plan correction can be applied to it meanwhile.

MVP-SLAM instead keeps the interrupted map and merges it with the new one through the IMU: the last keyframe before the loss is propagated by dead reckoning over the frames the loss lasts, and comparing this prediction with the first keyframe of the new map gives the transform between the two map frames. Both maps are metric and gravity-aligned, so the transform reduces to yaw and position, the four degrees of freedom unobservable in visual-inertial estimation [20], and is applied once, leaving the internal geometry of each segment untouched. Both segments therefore remain in the anchored frame of the first map, so walls matched in either of them constrain the whole trajectory, and a structure seen on both sides of the loss refines the join itself (Sec. III-D).

III-C Semantic wall detection and floor-plan association

This module supplies the back-end with wall correspondences in three steps: wall detection recovers the as-built walls W_{s}, floor-plan processing extracts the plan walls W_{\mathrm{plan}}, and a drift-aware association matches the two.

Wall detection. Planar surfaces are established landmarks in structured-environment SLAM [25], but fitting them directly to a feature-based map is error-prone, since its points come from many structures and non-wall planar clutter such as cabinets or panels is easily mistaken for a wall. MVP-SLAM therefore detects walls semantically, classifying the map points and fitting wall planes only to those labelled as wall [7]. These labels come from a pretrained panoptic segmenter (EoMT [26]), which assigns every pixel of both fisheye images a COCO-panoptic class, a vocabulary that already contains the wall class, so no site-specific fine-tuning is required. Each local-map point is then assigned the semantic class of its projected pixel on a per-frame basis as observations arrive. Labels are stabilized with a saturating hysteresis counter, so a point’s label flips only under sustained contradictory evidence. Three filters then suppress spurious wall labels: a gravity filter demotes wall labels whose viewing normal is too close to vertical (floor/ceiling bleed-through), a depth band rejects unreliable projections, and newly created map points are withheld until their labels stabilize.

At every keyframe, these wall-labelled points are fitted into wall planes with a sequential RANSAC in a gravity-aligned frame. Verticality removes one rotational degree of freedom, so a wall is parameterized by azimuth and distance only and fitted from a minimal sample of two points; hypotheses are scored by a robust RANSAC objective (MSAC) and accepted only if their inliers form a single contiguous wall segment. A freshly fitted plane is tentative until confirmed by consistent re-fitting across several keyframes, after which it becomes a detected wall W_{s} eligible for matching with the floor plan.

Floor-plan processing. Since the wall detection module recovers only the larger, clearly-observed walls, not every thin partition or facade detail drawn in the CAD, the floor plan is reduced to a comparable, canonical set of wall faces the detected walls can be matched against. MVP-SLAM converts the CAD wall layers into 2D wall faces, registers them to the plan raster by a grid search maximizing wall-pixel overlap, and canonicalizes them by merging collinear spans, removing duplicate or ambiguous faces, assigning consistent outward normals, and completing missing opposite faces. Exterior/interior labels are inherited from the CAD layers, with exterior taking precedence when a face combines both.

Drift-aware association. Each accepted match becomes a persistent constraint on the trajectory, so the association commits conservatively to avoid corrupting the rest of the estimate. The initialization pose \mathbf{T}_{s} places the floor plan in the as-built frame, so this step can find the correct associations between each detected wall W_{s} and its floor-plan counterpart W_{\mathrm{plan}}. A wall constrains the estimate only along its normal, so the integration stage (Sec. III-D) applies each correction anisotropically, and how much a candidate W_{\mathrm{plan}} can be trusted then depends on how far the estimate has drifted along that direction. In principle, the estimator’s pose covariance could measure this, but a keyframe-windowed bundle adjustment holds out-of-window poses fixed [27], so this covariance reflects only the local window and not this accumulated drift.

We therefore estimate this drift from the trajectory instead. It grows the longer the trajectory runs without a correction in that direction, so for each detected wall we track it through \ell, the length traveled since the last accepted exterior association with a parallel wall normal (|\mathbf{n}_{i}^{\!\top}\mathbf{n}_{j}|\!\geq\!0.9); only exterior wall matches, being the more reliable, reset \ell. A small \ell means the perpendicular distance \Delta d to a plan face is still reliable; a large \ell means the trajectory may have drifted farther from the correct W_{\mathrm{plan}}. We use \ell in two ways. First, it widens a hard distance gate: a candidate W_{\mathrm{plan}} is admitted only if

\Delta d\;\leq\;\min\!\left(d_{0}+\alpha_{\ell}\,\ell,\;c\right),
(1)

so the tolerated perpendicular offset starts at d_{0}, grows with drift at rate \alpha_{\ell}, and is capped at a wall class dependent value c that limits how far the gate can open (Table II). This cap is tighter for interior walls than exterior ones (c_{\mathrm{int}}\!<\!c_{\mathrm{ext}}), because interior partitions are densely packed with near-parallel neighbours, where a looser gate risks matching the wrong twin, and are detected less reliably from shorter range, whereas well-separated facades can absorb more drift before a match becomes ambiguous. Second, it saturates the distance that enters the score, yielding the effective distance \widetilde{\Delta d},

\widetilde{\Delta d}\;=\;(1-\sigma)\,\Delta d\;+\;\sigma\,\min(\Delta d,\,d_{\mathrm{sat}}),
(2)

which equals the true distance \Delta d when drift is small (\sigma{=}0) and caps it at d_{\mathrm{sat}} when drift is large (\sigma{=}1). The blend factor \sigma rises linearly from 0 to 1 as \ell grows from \ell_{\min} to \ell_{\max} (Table II), holding at 0 below \ell_{\min} and 1 above \ell_{\max}. With large drift the distance therefore stops dominating, leaving the angle and overlap terms to decide among the remaining candidates. After this drift-aware distance handling, each surviving plan face is scored with

s=w_{a}\,\Delta\theta+w_{d}\,\widetilde{\Delta d}+w_{f}\,(1-f),
(3)

where s is the match cost, so the face minimizing s wins; \Delta\theta is the azimuth difference between the W_{s} and W_{\mathrm{plan}} normals, taken with sign so that a face whose normal points the opposite way is penalized rather than counted as aligned; f\!\in\![0,1] is the fraction of the detected extent contained in the face; and the weights w_{a},w_{d},w_{f} (Table II) bring the radian, metre, and unitless terms to a common scale. The angle term is weighted most strongly because small orientation errors make a wall match unreliable even when its centroid is close.

The final match of a detected wall W_{s} to its plan face W_{\mathrm{plan}} is committed only if it is unambiguous, which a rejection cascade enforces before the pair is handed to the integration stage. The lowest-cost candidate is discarded if its score exceeds a class-dependent ceiling, if it does not beat the runner-up by a sufficient margin (Table II), or if a same-class competitor lies at a comparable perpendicular distance.

III-D Multi-stage floor-plan integration

Each matched pair (W_{s},W_{\mathrm{plan}}) provides a plan correspondence, but it reduces drift only once integrated into the SLAM graph as factors that optimize the trajectory and map. Introducing every accepted correspondence into one joint plan-constrained optimization makes that problem grow with the trajectory and accumulated associations [28]. MVP-SLAM instead incorporates each correspondence incrementally through a multi-stage optimization run whenever a new association is committed. Because this re-estimation happens at every match, MVP-SLAM continually realigns and localizes within the floor plan, so errors such as an imperfect initial pose \mathbf{T}_{s} are progressively absorbed. It runs in three steps: a local alignment over a bounded window (Step 1), a pose propagation of that correction (Step 2), and its persistence as a prior in later optimizations (Step 3), illustrated in Fig. 3.

Fig. 3: Incremental multi-stage integration of a matched wall, preserving earlier corrections.
III-D1 Step 1: Local alignment mini-graph

The first stage builds a local 4-DoF mini-graph for the newly matched pair (W_{s},W_{\mathrm{plan}}), optimizing only yaw and translation because the wall constraints are vertical and roll/pitch are fixed by gravity. In it, the plan wall W_{\mathrm{plan}} is a fixed vertex and the detected wall W_{s} an optimizable one, joined by a high-information alignment edge; every keyframe observing W_{s} is a pose vertex, tied to W_{s} by a plane-projection edge and to its covisibility neighbours by relative 4-DoF edges. Aligning W_{s} with W_{\mathrm{plan}} therefore shifts the observing keyframes and removes their accumulated drift. Besides these observers, only the keyframes in the temporal gap between them and their covisibility neighbours are free to correct. If a wall was observed over a long span, this set is capped at a maximum window size (Table II), keeping the most recent keyframes. Each match is corrected on its own window, keeping the update incremental and per-association. While physical walls are finite surfaces represented by W_{s} and W_{\mathrm{plan}}, their supporting infinite planes, \pi_{s} and \pi_{\mathrm{plan}}, are used in this mini-graph formulation. The local alignment minimizes the joint cost:

\min_{\{\mathbf{T}_{k}\},\,\pi_{s}}\;\left\lVert\pi_{s}\ominus\pi_{\mathrm{plan}}\right\rVert^{2}_{\mathbf{\Lambda}_{B}}+\sum_{k\in\mathcal{O}}\left\lVert\mathbf{e}^{\pi}_{k}\right\rVert^{2}_{\mathbf{\Lambda}_{\pi}}+\sum_{(i,j)\in\mathcal{E}_{\mathrm{loc}}}\left\lVert\mathbf{e}_{ij}\right\rVert^{2}_{\mathbf{\Lambda}_{ij}},
(4)

whose three terms weight these edges: the alignment edge (\ominus the difference between infinite plane parameters [25]) pulls the detected wall \pi_{s} onto the fixed plan face \pi_{\mathrm{plan}}, penalizing their azimuth and signed-distance mismatch; plane-projection residuals \mathbf{e}^{\pi}_{k} tie each observing keyframe k\in\mathcal{O} to the wall measurement stored when it observed the wall, so consistency is enforced against the local observation, not a global drifted pose; and relative 4-DoF covisibility edges \mathcal{E}_{\mathrm{loc}} preserve the local trajectory shape while yaw and translation adapt to the wall. The first keyframe is a gauge when no earlier plan correction exists. Rather than modifying SLAM poses directly, the mini-graph outputs corrected keyframe targets storing their associated wall normals to guide subsequent directional priors

III-D2 Step 2: Correction propagation to the system

Applied on their own, the mini-graph targets would end at the window boundary, leaving a discontinuity against the rest of the trajectory. A second 4-DoF pose graph therefore propagates them to the surrounding keyframes. Its adaptive window is built around the newly corrected keyframes and an anchor, chosen when possible as the latest previous correction whose stored wall normal is parallel to the current one. The window bridges from this anchor to the newly corrected span and extends backward from it, farther if too few same-axis old priors fall inside, so the previous correction supports the update over multiple keyframes rather than a single fixed anchor. The graph combines relative edges (covisibility and temporal-chain) with unary priors (fresh targets from Step 1 and previous corrections) by minimizing the joint cost:

\min_{\{\mathbf{T}_{k}\}}\sum_{(i,j)\in\mathcal{E}}\left\lVert\mathbf{e}_{ij}\right\rVert^{2}_{\mathbf{\Lambda}_{ij}}+\sum_{k\in\mathcal{P}}\left\lVert\mathbf{e}_{k}\right\rVert^{2}_{\mathbf{\Lambda}_{k}},
(5)

where \mathcal{E} contains relative 4-DoF edges, \mathcal{P} is the set of keyframes carrying targets from Step 1, \mathbf{e}_{ij} is the relative residual, and \mathbf{e}_{k} is the unary prior residual pulling keyframe k toward its target pose, weighted by the block-diagonal information matrix \mathbf{\Lambda}_{k} (Fig. 3, Step 2 inset):

\displaystyle\mathbf{\Lambda}_{k} \\ \displaystyle=\;\begin{bmatrix}\mathbf{\Lambda}_{t}&\mathbf{0}\\
\mathbf{0}&\lambda_{r}\,\mathbf{I}_{3}\end{bmatrix}\in\mathbb{R}^{6\times 6}, \\ \displaystyle\mathbf{\Lambda}_{t} \\ \displaystyle=\;\tau\,\mathbf{I}_{3}\;+\;s\!\!\sum_{\mathbf{n}\in\mathcal{N}_{k}}\mathbf{n}\mathbf{n}^{\!\top}\;\in\mathbb{R}^{3\times 3},

The translation block \mathbf{\Lambda}_{t} builds an anisotropic constraint based on environment geometry. For each observed wall normal \mathbf{n}\in\mathcal{N}_{k}, the outer-product term s\,\mathbf{n}\mathbf{n}^{\!\top} injects high information s perpendicular to the wall, while \tau\mathbf{I}_{3} provides a small isotropic information floor in every direction. The prior therefore pins the keyframe stiffly across each wall it observed while leaving it free to slide along the wall surface and vertically. Keyframes matched to multiple non-parallel walls are naturally pinned along all their respective normals. The rotation block \lambda_{r}\mathbf{I}_{3} provides an isotropic constraint on heading (yaw). Because roll and pitch remain fixed by gravity alignment in this 4-DoF formulation, yaw is the only rotational degree of freedom optimized, making a single scalar weight \lambda_{r} sufficient.

TABLE I: Absolute Pose Error (APE) RMSE (\mathrm{cm}) on Hilti–Trimble SLAM Challenge 2026 sequences, as reported on the official challenge leaderboard on the day of the challenge. Bold and underline denote best and second-best results per column per task; dashes (–) mark unavailable entries for Localization.
  • —
    无数据
    Floors and Sequences
    Ground Floors
    无数据
    无数据
    无数据
    无数据
    无数据
    无数据
    Upper Floors
    无数据
    无数据
    无数据
    无数据
    无数据
    无数据
    无数据
    无数据
    无数据
    无数据
    无数据
    Underground Levels
    无数据
    无数据
    无数据
    无数据
    无数据
    无数据
  • —
    无数据
    Floors and Sequences
    Floor 1
    无数据
    Floor 2
    无数据
    Floor EG
    无数据
    无数据
    Floor 3
    无数据
    Floor 4
    无数据
    Floor 5
    Floor 6
    无数据
    无数据
    无数据
    Floor 7
    无数据
    无数据
    Floor UG1
    无数据
    无数据
    无数据
    无数据
    Floor UG2
    无数据
  • —
    无数据
    Floors and Sequences
    07-07
    12-02
    12-02
    12-03
    10-16
    12-02a
    12-02b
    05-19
    12-02
    05-19
    12-02
    12-02
    06-18
    07-07
    12-02a
    12-02b
    12-02a
    12-02b
    12-03
    05-19
    06-18
    12-02a
    12-02b
    12-03
    12-02
    Average
  • Duration (\mathrm{sec.})
    无数据
    Floors and Sequences
    133.7
    276.3
    143.7
    149.9
    242.9
    126.4
    162.5
    115.5
    132.7
    91.9
    193.4
    160.5
    69.3
    73.0
    170.7
    124.6
    117.9
    155.1
    198.3
    205.5
    165.8
    254.8
    219.7
    134.4
    223.1
    161.7
  • Length (\mathrm{m})
    无数据
    Floors and Sequences
    157.8
    321.8
    154.8
    138.8
    240.4
    114.9
    158.6
    128.2
    148.8
    97.8
    214.1
    174.9
    79.4
    72.2
    210.6
    145.8
    145.4
    197.1
    152.6
    262.5
    222.6
    351.7
    322.5
    119.1
    292.1
    185.0
  • Localization
    Map-It Ralph [14]
    Floors and Sequences
    71.92
    39.68
    577.06
    233.11
    41.86
    87.98
    320.64
    103.89
    182.71
    38.67
    49.77
    82.38
    49.74
    41.31
    50.59
    109.38
    301.10
    105.34
    69.70
    95.13
    147.77
    264.79
    88.25
    127.67
    –
    136.69
  • CUFE [16]
    66.98
    Floors and Sequences
    41.20
    60.07
    42.64
    25.54
    30.83
    18.22
    84.93
    106.07
    75.80
    45.97
    65.83
    20.80
    28.94
    41.87
    78.11
    63.01
    88.80
    76.51
    38.88
    66.50
    48.09
    83.49
    142.69
    –
    60.07
    无数据
  • OmniRecon [15]
    38.05
    Floors and Sequences
    30.51
    38.79
    45.59
    33.32
    19.88
    41.99
    40.12
    64.74
    44.20
    17.58
    53.57
    30.20
    40.37
    52.20
    27.42
    14.26
    74.43
    27.60
    46.85
    55.00
    58.41
    53.50
    58.08
    –
    41.94
    无数据
  • Z-FLoc [6]
    28.87
    Floors and Sequences
    17.78
    32.55
    29.98
    15.07
    12.48
    18.37
    17.54
    36.45
    24.22
    12.60
    14.31
    27.69
    33.57
    16.10
    19.08
    13.30
    21.04
    22.61
    32.17
    33.56
    24.25
    42.62
    23.77
    –
    23.75
    无数据
  • Ours (MVP-SLAM)
    17.69
    Floors and Sequences
    25.48
    33.57
    18.54
    22.10
    13.48
    37.79
    13.65
    31.77
    17.37
    21.30
    31.65
    21.56
    18.85
    34.09
    18.58
    26.07
    31.57
    21.30
    27.35
    79.10
    46.90
    62.10
    31.19
    –
    29.29
    无数据
  • Ours (w/o plan)
    131.48
    Floors and Sequences
    251.39
    151.83
    95.04
    243.66
    26.28
    100.84
    55.17
    50.33
    100.81
    244.43
    54.88
    58.60
    51.21
    75.51
    35.91
    163.64
    113.40
    98.81
    220.39
    206.57
    350.86
    483.25
    109.04
    –
    144.72
    无数据
  • SLAM
    ACDC-VSLAM [10]
    Floors and Sequences
    10.07
    9.82
    7.69
    11.03
    7.84
    5.53
    6.70
    12.60
    6.50
    8.01
    7.12
    5.47
    6.36
    5.82
    6.02
    5.37
    5.95
    5.82
    10.27
    15.36
    19.88
    11.99
    13.17
    6.24
    12.86
    8.94
  • \sqrt{\text{VINS}} [12]
    14.65
    Floors and Sequences
    10.60
    13.54
    18.83
    9.87
    7.70
    6.67
    9.07
    8.66
    7.83
    9.18
    8.37
    5.88
    6.93
    8.83
    8.82
    8.89
    7.49
    11.34
    13.18
    18.97
    16.16
    16.43
    10.93
    13.67
    10.90
    无数据
  • Undisclosed
    17.59
    Floors and Sequences
    11.76
    36.04
    12.55
    9.46
    8.99
    12.75
    12.86
    36.01
    12.97
    6.74
    5.43
    16.71
    14.72
    5.41
    7.25
    8.92
    7.46
    16.32
    27.20
    33.24
    49.54
    50.03
    6.13
    40.83
    18.68
    无数据
  • QQ [11]
    29.42
    Floors and Sequences
    28.59
    23.23
    15.14
    10.97
    20.88
    20.48
    12.29
    38.15
    6.78
    16.40
    12.61
    5.32
    9.14
    12.36
    9.03
    13.18
    19.41
    23.93
    24.08
    39.23
    28.89
    39.67
    16.66
    28.34
    20.17
    无数据
  • Ours (MVP-SLAM)
    18.35
    Floors and Sequences
    17.93
    19.19
    15.88
    18.15
    8.93
    20.20
    11.95
    25.05
    13.92
    15.72
    33.58
    17.05
    10.32
    29.84
    10.50
    18.31
    11.49
    12.35
    30.98
    56.87
    56.47
    83.06
    20.52
    29.17
    24.23
    无数据
  • Ours (w/o plan)
    104.96
    Floors and Sequences
    128.91
    50.79
    78.35
    116.47
    10.40
    22.32
    35.78
    93.05
    55.04
    100.69
    21.00
    29.62
    14.69
    49.29
    18.78
    37.41
    46.37
    32.66
    155.01
    68.01
    177.93
    187.25
    74.11
    121.42
    73.21
    无数据

After propagation, the optimized poses are written back to the SLAM state. Map points move rigidly with their most recent corrected observer, velocities rotate with their keyframes, and the IMU preintegration terms are re-evaluated at the new states. Without this patch, the next inertial local BA would see inconsistent inertial residuals and pull the trajectory back toward the pre-correction state.

III-D3 Step 3: Persistence in subsequent optimization

A correction is useful only if it persists in the optimizations that follow. We thus store each corrected target with its wall-normal set \mathcal{N}_{k} and reintroduce it as a unary anisotropic prior whenever the affected keyframe appears in a later optimization.

In the local inertial BA that follows each correction, the priors keep the continuous SLAM estimate coherent with them, using the same anisotropic form but a softer information ratio (Table II) so that reprojection and inertial terms still refine the map along the wall while the floor plan prevents drift back across its normal. Each local BA jointly optimizes map points, keyframe poses, velocities, and IMU biases under these plan priors, so the final map and trajectory jointly satisfy visual, inertial, and structural constraints.

To enforce global alignment without noise, floor-plan priors enter the essential-graph optimization [19] at loop closures and map merges. These constraints are applied only to keyframes associated with exterior walls, more reliable than interior ones.

IV Experimental Results

TABLE II: Key parameters of MVP-SLAM.
  • Drift-aware association (Sec. III-C)
    Symbol
    无数据
    Value
    无数据
    Description
    无数据
  • score weights
    Symbol
    w_{a,d,f}
    Value
    6 rad-1, 1 m-1, 1
    Description
    angle, dist., overlap
  • gate ramp
    Symbol
    d_{0},\alpha_{\ell}
    Value
    1 m, 0.1
    Description
    offset, drift growth
  • gate cap
    Symbol
    c_{\mathrm{ext/int}}
    Value
    5, 2.5 m
    Description
    ext./int. distance cap
  • dist. saturation
    Symbol
    d_{\mathrm{sat}}
    Value
    1 m
    Description
    score distance cap
  • saturation range
    Symbol
    \ell_{\min/\max}
    Value
    15, 40 m
    Description
    score ramp
  • runner-up margin
    Symbol
    –
    Value
    1.6\times
    Description
    match uniqueness
  • Floor-plan integration (Sec. III-D)
    Symbol
    无数据
    Value
    无数据
    Description
    无数据
  • mini-graph window
    Symbol
    –
    Value
    200 KF
    Description
    max. KFs in Step1
  • along-normal info
    Symbol
    s
    Value
    10^{4}
    Description
    stiffness across wall
  • tangent floor
    Symbol
    \tau
    Value
    1 / 50
    Description
    propagation / local BA
  • rotation info
    Symbol
    \lambda_{r}
    Value
    10^{3} / 10^{4}
    Description
    propagation / local BA

IV-A Validation Methodology

Datasets. We evaluate on the Hilti–Trimble SLAM Challenge 2026 dataset [4], recorded in active construction sites with a hand-held rig carrying two back-to-back \sim200∘ fisheye cameras operating at 30 Hz and a 1 kHz IMU. We address both official tasks with the same system. Localization is scored in the floor-plan frame and SLAM is scored after rigid alignment to the reference. Localization contains 24 scored sequences, excluding Floor UG2 because no floor plan is released.

We group the sequences into ground floors, upper floors, and underground levels. Across all of them, the site is under construction and can differ from the floor-plan. The ground floors range from open spaces to the more built-out entrance level EG; the upper floors include interior areas and one outdoor terrace run; and the underground levels are large, sparsely-walled spaces dominated by structural columns.

Baselines. We compare MVP-SLAM against the other top-five competitors of each task (Table I), with per-sequence results taken from the official challenge release [4].

Implementation Details. MVP-SLAM runs on a workstation equipped with an Intel Core i9-11950H CPU and an NVIDIA T600 GPU (4 GB), 32 GB RAM. Key parameters are listed in Table II.

Fig. 4: Trajectory comparison in the Localization task, in the floor-plan frame, on three representative sequences.

Trajectory estimation Performance. The challenge provides a LiDAR-inertial reference trajectory, used only for evaluation. Both tasks report per-sequence RMSE of the Absolute Pose Error (APE) against it: 2D in the floor-plan plane for Localization, 3D after rigid alignment for SLAM. A trajectory must cover at least 99\% of the reference poses to be scored. To isolate the impact of map priors, we also evaluate an ablated variant, Ours (w/o plan). It uses the same visual-inertial SLAM setup but disables wall associations and floor-plan integration.

SLAM initialization and robustness ablation. We assess the fast metric initialization and the map preservation across tracking losses of the multi-camera front-end (Sec. III-B) with an ablation. The reduced configuration disables them, so it initializes and recovers as a standard visual-inertial SLAM would, while the full configuration keeps them enabled. We measure initialization success, time to metric scale, trajectory coverage, and whether the run satisfies the challenge coverage requirement on six sequences, two per environment type.

Wall detection and matching Performance. We evaluate wall detection and wall-to-plan matching on the same six sequences, using precision, recall, and F1. The reference set is hand-labeled from the floor plan, camera images, and reconstructed map, and contains only walls that are visible to the cameras and supported by reconstructed map points. Unbuilt planned walls and transient clutter are ignored. For detection, a true positive is a detected plane on the same physical wall as a reference wall; a false positive is a plane fitted to non-wall structure, clutter, or empty space; and a false negative a reference wall not recovered. For matching, we manually label each detected wall’s correct plan face; a match is correct only when the committed association links to it.

IV-B Results and Discussion

IV-B1 Trajectory Estimation Performance

Table I reports the APE RMSE for both official challenge tasks. In the Localization task, MVP-SLAM achieves a mean APE of 0.29 m and ranks second among the 22 Localization teams, close to the winning Z-FLoc submission (0.24 m). MVP-SLAM is best or second-best on most sequences, showing that the plan constraints keep the trajectory well aligned with the building frame across floors and environment types. Without plan integration, Ours (w/o plan) reaches 1.45 m mean APE; while tracking remains stable, accumulated drift is no longer bounded. Adding the plan integration reduces the Localization error by 4.9\times, from 1.45 m to 0.29 m. Fig. 4 illustrates this effect on representative ground-floor, upper-floor, and underground sequences.

In the SLAM task, MVP-SLAM obtains 0.24 m mean APE and ranks fifth among 62 teams. As this metric only accounts for the trajectory shape (Sec. IV-A), the top-ranked systems, led by ACDC-VSLAM [10] (0.09 m) and \sqrt{\text{VINS}} [12] (0.11 m), reach lower APE through mechanisms such as enhanced loop closure, effective since the challenge sequences often revisit previously seen areas, and richer visual features. MVP-SLAM instead uses the floor plan to both localize the trajectory and correct drift, reaching the top five and demonstrating that the floor-plan integration reduces the aligned error 3.0\times from 0.73 m (Ours w/o plan) to 0.24 m.

As established in Sec. II, the other Localization teams apply the floor plan offline, once the trajectory is built [6, 15, 16]; MVP-SLAM is instead the only one to integrate it online, folding each match into the estimation as it is detected. The same distinction holds in the SLAM task, where MVP-SLAM is again the top-ranked team among those that integrate the floor plan online, and the only one to also localize within it, showing the building’s floor plan is a promising, distinct prior for improving a SLAM system.

TABLE III: Initialization and robustness ablation for Ours (w/o plan): 3 repeats per sequence, both configurations initialized every run. t_{\mathrm{metric}}: median time to metric scale (s); Cov.: mean coverage (%); APE: median Localization error (m).
  • Ground
    Sequence
    Floor 1 (07-07)
    Reduced t_{\mathrm{metric}}
    3.1
    Reduced Cov.
    100
    Reduced APE
    1.38
    Full t_{\mathrm{metric}}
    2.5
    Full Cov.
    100
    Full APE
    1.31
  • Floor EG (10-16)
    Sequence
    2.5
    Reduced t_{\mathrm{metric}}
    100
    Reduced Cov.
    3.02
    Reduced APE
    2.2
    Full t_{\mathrm{metric}}
    100
    Full Cov.
    2.44
    Full APE
    无数据
  • Upper
    Sequence
    Floor 3 (05-19)
    Reduced t_{\mathrm{metric}}
    2.8
    Reduced Cov.
    100
    Reduced APE
    0.58
    Full t_{\mathrm{metric}}
    2.5
    Full Cov.
    100
    Full APE
    0.55
  • Floor 6 (12-02a)
    Sequence
    2.5
    Reduced t_{\mathrm{metric}}
    100
    Reduced Cov.
    1.01
    Reduced APE
    2.5
    Full t_{\mathrm{metric}}
    100
    Full Cov.
    0.76
    Full APE
    无数据
  • Under
    Sequence
    Floor UG1 (06-18)
    Reduced t_{\mathrm{metric}}
    6.6
    Reduced Cov.
    96.7
    Reduced APE
    28.0
    Full t_{\mathrm{metric}}
    2.5
    Full Cov.
    100
    Full APE
    2.07
  • Floor UG1 (12-02a)
    Sequence
    8.1
    Reduced t_{\mathrm{metric}}
    100
    Reduced Cov.
    5.10
    Reduced APE
    2.4
    Full t_{\mathrm{metric}}
    100
    Full Cov.
    3.51
    Full APE
    无数据
IV-B2 SLAM Initialization and Robustness Ablation

Table III shows that both configurations initialize reliably and reach metric scale in \sim2.5s on upper floors, and diverge only in underground environments. Here, the reduced configuration takes up to 3\times longer to reach metric scale (6.6–8.1s) and fails the 99\% coverage threshold on Floor UG1 (06-18). This degradation occurs when a hard initialization briefly drops tracking, creating a second map frame. Without map preservation, the two maps remain unmerged. As a result, abandoned keyframes drop out of the trajectory, causing task failures and misaligning the floor plan reference. By maintaining continuity across tracking losses, map preservation merges these segments, recovering full coverage and reducing Floor UG1 (06-18) Localization error from 28m to 2.07m.

TABLE IV: Wall-detection and wall-matching performance (precision P, recall R, F1, in percent). Overall pools counts.
  • Floor
    Sequence
    Detection
    P
    R
    F1
    Matching
    P
    R
    F1
  • Ground
    Floor 1 (07-07)
    Detection
    94.4
    50.0
    65.4
    Matching
    100.0
    94.1
    97.0
  • Floor EG (10-16)
    100.0
    Detection
    28.6
    44.4
    100.0
    Matching
    100.0
    100.0
    无数据
  • Upper
    Floor 3 (05-19)
    Detection
    100.0
    57.9
    73.3
    Matching
    100.0
    100.0
    100.0
  • Floor 6 (12-02a)
    84.6
    Detection
    35.5
    50.0
    90.9
    Matching
    100.0
    95.2
    无数据
  • Under
    Floor UG1 (06-18)
    Detection
    100.0
    37.8
    54.9
    Matching
    100.0
    100.0
    100.0
  • Floor UG1 (12-02a)
    100.0
    Detection
    40.4
    57.6
    100.0
    Matching
    100.0
    100.0
    无数据
  • —
    Overall
    Detection
    96.9
    38.4
    55.0
    Matching
    98.9
    98.9
    98.9
IV-B3 Wall Detection and Matching Performance

Table IV reports wall detection and wall-to-plan matching on six representative sequences. Detection is conservative by design: it reaches high precision (96.9\%) but moderate recall (38.4\%). This choice follows from the integration design (Sec. III-D), where each accepted wall becomes a persistent plan prior; a wrong wall is therefore more harmful than a missed one. Most missed walls are visible in the images but have too few or too noisy reconstructed map points to support a stable plane fit, especially in cluttered or texture-scarce areas.

Once a wall is detected, association is highly reliable: matching precision, recall, and F1 are all 98.9\% pooled. The detected wall candidates are usually close to their true plan faces and separated from alternatives, so the score (Eq. 3) and rejection cascade resolve most matches cleanly. The single wrong match occurs on Floor 6, where two nearly collinear plan faces fall within the gate, and the single missed match is a detected wall that receives no association. Thus, in these sequences, the plan prior is limited mainly by which walls can be detected from the sparse map, not by the association step. This is sufficient for localization when the accepted walls are distributed: Floor EG, for example, attains a 22 cm Localization APE despite only 28.6\% wall-detection recall.

Limitations. MVP-SLAM corrects drift only where it detects a wall and matches it to the floor plan. Its wall module fits planar walls, so in open areas with few walls, notably large underground spaces dominated by structural columns, few matches are available and the trajectory drifts uncorrected.

V Conclusions and Future Work

We presented MVP-SLAM, a multi-camera visual-inertial SLAM system that uses an as-planned floor plan as a reference to reduce drift accumulated on construction sites. It tracks its position on two opposite-facing fisheye cameras and an IMU, and corrects the drift online by mapping walls, matching them to the floor plan, and integrating it into the SLAM back-end through each match as the trajectory is built. On the Hilti–Trimble SLAM Challenge 2026 it ranked second in Localization (0.29 m mean RMSE) and fifth in SLAM (0.24 m). It is the top team in both tasks among those that also operate online, integrate the floor plan, and localize within it.

Future work will explore lightweight map densification to make the detection of walls and additional structural primitives such as columns easier, integrating them into the optimization for drift correction. It will also target the SLAM’s internal consistency directly, strengthening loop closure and adopting the richer-feature mechanisms of higher-ranked SLAM teams.

References

  1. [1] M. Shaheer, J. A. Millan-Romera, H. Bavle, J. L. Sanchez-Lopez, J. Civera, and H. Voos (2023) Graph-based global robot localization informing situational graphs with architectural graphs. In 2023 IEEE/RSJ Int. Conf. Intelligent Robots and Systems (IROS), pp. 9155–9162.
  2. [2] M. Helmberger, K. Morin, B. Berner, N. Kumar, G. Cioffi, and D. Scaramuzza (2022) The Hilti SLAM challenge dataset. IEEE Robot. Autom. Lett. 7 (3), pp. 7518–7525.
  3. [3] O. Mendez, S. Hadfield, N. Pugeault, and R. Bowden (2018) SeDAR: semantic detection and ranging — humans can localise without LiDAR, can robots?. In 2018 IEEE Int. Conf. Robotics and Automation (ICRA), pp. 6053–6060.
  4. [4] S. Centanni, Y. Zhang, Y. Tao, J. Kindle, F. Neuhaus, T. Koß, A. Patel, M. Helmberger, E. Szymańska, T. Gräber, et al. (2026) Hilti-trimble-oxford dataset: 360 visual-inertial benchmark with floor plan priors for slam and localization. arXiv preprint arXiv:2607.06464.
  5. [5] M. A. V. Torres, A. Braun, and A. Borrmann (2023) BIM-SLAM: integrating BIM models in multi-session SLAM for lifelong mapping using 3D LiDAR. In Proc. Int. Symp. Automation and Robotics in Construction (ISARC), Vol. 40, pp. 521–528.
  6. [6] A. Umemura, T. Kuwahara, M. Pollefeys, and D. Barath (2026) Z-FLoc: zero-shot floorplan localization via geometric primitives. arXiv preprint arXiv:2606.04788.
  7. [7] A. Bikandi-Noya, M. Fernandez-Cortizas, M. Shaheer, A. Tourani, H. Voos, and J. L. Sanchez-Lopez (2025) BIM-informed visual SLAM for construction monitoring. arXiv preprint arXiv:2509.13972.
  8. [8] Y. Wang, Y. Ng, I. Sa, Á. Parra, C. Rodriguez-Opazo, T. Lin, and H. Li (2024) MAVIS: multi-camera augmented visual-inertial SLAM using SE2(3)-based exact IMU pre-integration. In 2024 IEEE Int. Conf. Robotics and Automation (ICRA), pp. 1694–1700.
  9. [9] M. J. Tribou, A. Harmat, D. W. L. Wang, I. Sharf, and S. L. Waslander (2015) Multi-camera parallel tracking and mapping with non-overlapping fields of view. Int. J. Robotics Research 34 (12), pp. 1480–1500.
  10. [10] J. Jeon, D. Seo, J. Choi, S. Lee, J. Nam, H. Lim, and H. Myung (2026) ACDC-VSLAM: adaptive constraints for dual-fisheye-camera-based visual-inertial SLAM with point and line features. Note: Hilti–Trimble SLAM Challenge 2026, SLAM track report
  11. [11] J. Jiang (2026) Technical report for the Hilti \times Trimble SLAM challenge 2026. Note: Hilti–Trimble SLAM Challenge 2026, SLAM track report
  12. [12] J. Lee, H. Kim, J. Choi, J. Jeong, and Y. Cho (2026) \sqrt{\text{VINS}} with factor graph optimization-based visual-inertial SLAM in the Hilti–Trimble SLAM challenge 2026. Note: Hilti–Trimble SLAM Challenge 2026, SLAM track report
  13. [13] C. Chen, R. Wang, C. Vogel, and M. Pollefeys (2024) F3Loc: fusion and filtering for floorplan localization. In Proc. IEEE/CVF Conf. Computer Vision and Pattern Recognition (CVPR), pp. 18029–18038.
  14. [14] A. C. Demirtaş, B. Şekeroglu, and A. Topaloglu (2026) Two-stage visual-inertial SLAM with sliding-window filtering and global bundle adjustment for the Hilti \times Trimble SLAM challenge 2026. Note: Hilti–Trimble SLAM Challenge 2026, Localization track report
  15. [15] G. Tanner, G. Evangelou, J. Lechner, Z. Pataki, X. Jiang, P. Sarlin, and S. Liu (2026) Establishing the gold standard for 360{}^{\circ} visual-inertial reconstruction. Note: Hilti–Trimble SLAM Challenge 2026, Localization track report
  16. [16] CUFE Team, Cairo University (2026) A fully automated multi-stage localization pipeline for indoor floorplan-constrained trajectory estimation. Note: Hilti–Trimble SLAM Challenge 2026, Localization track report
  17. [17] Y. Zang (2026) Non-overlapping dual-fisheye visual-inertial SLAM based on ORB-SLAM3. Note: Hilti–Trimble SLAM Challenge 2026, organizer baseline
  18. [18] R. Merat, G. Cioffi, L. Bauersfeld, and D. Scaramuzza (2025) Drift-free visual slam using digital twins. IEEE Robot. Autom. Lett. 10 (2), pp. 1633–1640. External Links: Document
  19. [19] C. Campos, R. Elvira, J. J. G. Rodríguez, J. M. M. Montiel, and J. D. Tardós (2021) ORB-SLAM3: an accurate open-source library for visual, visual–inertial, and multimap SLAM. IEEE Trans. Robotics 37 (6), pp. 1874–1890.
  20. [20] T. Qin, P. Li, and S. Shen (2018) VINS-Mono: a robust and versatile monocular visual-inertial state estimator. IEEE Trans. Robotics 34 (4), pp. 1004–1020.
  21. [21] A. Li, D. Zou, and W. Yu (2021) Robust initialization of multi-camera SLAM with limited view overlaps and inaccurate extrinsic calibration. In 2021 IEEE/RSJ Int. Conf. Intelligent Robots and Systems (IROS), pp. 3361–3367. External Links: Document
  22. [22] L. Kneip, M. Chli, and R. Siegwart (2011) Robust real-time visual odometry with a single camera and an IMU. In Proc. British Machine Vision Conf. (BMVC),
  23. [23] C. Troiani, A. Martinelli, C. Laugier, and D. Scaramuzza (2014) 2-point-based outlier rejection for camera-IMU systems with applications to micro aerial vehicles. In 2014 IEEE Int. Conf. Robotics and Automation (ICRA), pp. 5530–5536.
  24. [24] Y. He, H. Yu, W. Yang, and S. Scherer (2022) Towards robust visual-inertial odometry with multiple non-overlapping monocular cameras. In 2022 IEEE/RSJ Int. Conf. Intelligent Robots and Systems (IROS), pp. 9452–9458.
  25. [25] M. Kaess (2015) Simultaneous localization and mapping with infinite planes. In 2015 IEEE Int. Conf. Robotics and Automation (ICRA),
  26. [26] T. Kerssies, N. Cavagnero, A. Hermans, N. Norouzi, G. Averta, B. Leibe, G. Dubbelman, and D. De Geus (2025) Your ViT is secretly an image segmentation model. In Proc. Computer Vision and Pattern Recognition Conf. (CVPR), pp. 25303–25313.
  27. [27] R. Mur-Artal and J. D. Tardós (2017) Visual-inertial monocular slam with map reuse. IEEE Robot. Autom. Lett. 2 (2), pp. 796–803. External Links: Document
  28. [28] M. Kaess, H. Johannsson, R. Roberts, V. Ila, J. J. Leonard, and F. Dellaert (2012) ISAM2: incremental smoothing and mapping using the bayes tree. Int. J. Robotics Research 31 (2), pp. 216–235.