ROBOTNESS
上級arXiv

MVP-SLAM、床プラン事前情報でドリフト補正する複数カメラ視覚慣性SLAM

Asier Bikandi-Noya, Miguel Fernandez-Cortizas, Muhammad Shaheer, Holger Voos, Jose Luis Sanchez-Lopez
30秒で読む

ルクセンブルク大学の研究チームが、建設現場向けに2台の魚眼カメラとIMUのみで動作し、設計図の床プランを逐次統合してドリフトを補正するオンラインSLAM「MVP-SLAM」を発表した(査読前プレプリント)。Hilti–Trimble SLAM Challenge 2026のLocalizationタスクで22チーム中2位(平均RMSE 0.29m)、SLAMタスクで62チーム中5位(0.24m)となり、オンラインで床プランを統合し位置特定する手法としては両タスクで首位だった。床プランをカメラのみの視覚慣性SLAMにオンラインで組み込み、軌跡構築中に補正できる点が重要である。

研究課題

建設現場のオンライン視覚慣性SLAMにおいて、設計時の床プランをカメラのみの観測から検出した壁と対応付けて逐次統合し、ドリフトを補正しながら床プラン座標系で位置推定できるか。

問題

建設現場は照明変化、周期的で低テクスチャの構造、作業員の移動などにより視覚慣性SLAMの軌跡にドリフトが蓄積する。既存の床プラン利用手法は軌跡完成後のオフライン補正か、LiDARやRGB-Dなどの深度センサに依存しており、魚眼カメラとIMUのみのオンライン補正は未開拓だった。設計図と実際の建物には差異があり、ドリフトが大きくなるとカメラのみで壁の対応付けを取るのが難しくなる。

従来の手法

従来は、Z-FLocやOmniRecon、CUFEのように完成した軌跡を床プランにオフラインで位置合わせする手法、あるいはLiDAR点群のBIMアンカリングやRGB-Dカメラによる壁マッチング(ivS-Graphsなど)など深度・測距センサを統合する手法があった。オフライン手法は現場巡回中に位置が必要な用途に使えず、深度センサは手持ち装置では一般的でなく、RGB-Dは到達距離が数メートルに限られる。深度不要の手法でも精密なas-builtメッシュを別途取得する必要があった。

新しい手法

MVP-SLAMはORB-SLAM3を基盤とし、2台の背向き魚眼カメラを単眼で同一マップに追跡し、視野重複が少ないための交差カメラ初期化とジャイロ支援フォールバック、追跡喪失時の慣性マップ統合を導入する。事前学習済みのEoMTパノプティック分割で画素を壁クラスに分類し、地図点から壁平面を逐次RANSACで検出。床プランはCAD壁レイヤを2D壁面に変換して正準化し、走行距離に応じたドリフト対応型のゲートとスコアで対応付けを決定する。各対応は、局所ミニグラフ、異方性事前情報による補正伝播、以降の最適化での持続的制約という多段階統合でSLAMグラフに逐次組み込まれる。

結果

Hilti–Trimble SLAM Challenge 2026の実データ(24のLocalizationシーケンス)で評価。Localizationタスクは平均APE RMSE 0.29mで22チーム中2位、トップのZ-FLocは0.24m。床プラン統合を無効にした自社アブレーションは1.45mで、4.9倍の低減。SLAMタスクは0.24mで62チーム中5位、1位のACDC-VSLAMは0.09m、2位の√VINSは0.11m。アブレーションは0.73mで3.0倍の低減。初期化アブレーションでは、完全構成が約2.5秒でメトリックスケールに到達し、UG1ではカバレッジ100%、エラー2.07mに対し、縮小構成は最大8.1秒、カバレッジ96.7%、エラー28.0m。壁検出は精密さ96.9%、再現率38.4%、壁対応付けは精密さ・再現率・F1すべて98.9%だった。評価はLiDAR慣性参照軌跡との比較による。

限界

著者は、壁を検出して床プランと対応付けられた場所でのみドリフトを補正できるため、柱が支配的で壁が少ない地下空間などでは補正が働かないと述べている。壁検出は保守的で再現率が38.4%と低く、検出されない壁は補正に使えない。また、論文ではオンライン動作としているが処理時間やリアルタイム性の評価は報告されておらず、コードやデータセットの公開についても記載がない。単一のデータセットと単一のハンドヘルド装置での検証であり、実環境の設計図差異への頑健性は限定的にしか示されていない。

産業への影響

建設現場の進捗管理、自律ロボット検査、手持ち型スキャン装置などに応用できる。設計段階で入手可能な床プランを利用するため、追加の深度センサや精密メッシュ取得が不要な点は導入コストを下げる。ただし現時点ではHilti–Trimbleチャレンジのデータセットでの評価にとどまり、壁検出の再現率向上、処理時間のリアルタイム保証、as-builtとas-plannedの差異への頑健性確認が製品化前に必要である。これらが満たされれば、1〜3年以内に建設現場向けの自律走行ロボットや現場記録装置へ組み込める可能性がある。

論文全文

MVP-SLAM: Multi-Camera Visual-Inertial Floorplan-Prior SLAM

Asier Bikandi-Noya, Miguel Fernandez-Cortizas, Muhammad Shaheer, Holger Voos, Jose Luis Sanchez-Lopez

CC BY 4.0 のもとで公開された論文です。出典を明記して転載しています。原文は arXiv:2609.39596(PDF)。

Abstract

Indoor building construction sites are demanding environments for visual SLAM, where variable lighting and repetitive, low-textured structures make the system drift over long trajectories, though structural elements such as walls remain distinguishable despite these conditions. These buildings are constructed according to their as-planned floor plans, available from the design phase, and although the actual as-built site can differ from this design, floor plans still provide a metric reference, both to localize the system in the building and to correct drift. Existing methods often use the floor plan to correct an already-built trajectory offline, and those that instead correct it online typically rely on depth sensors. We instead present MVP-SLAM, an online visual-inertial SLAM on two opposite-facing fisheye cameras that corrects drift from cameras alone by matching walls detected in its map to the floor plan, through a drift-aware policy. A multi-stage integration then turns each matched pair incrementally into a persistent correction, so the trajectory stays corrected and localized within the floor plan as it is built. MVP-SLAM was validated on the multi-floor construction sites of the Hilti–Trimble SLAM Challenge 2026, ranking 2nd of 22 teams in the Localization task (0.29 m mean RMSE) and 5th of 62 teams in the SLAM task (0.24 m), the top-ranked one in both tasks among those that operate online, integrate the floor plan, and localize within it.

I Introduction

Accurate and reliable localization in indoor building construction environments is essential for automating construction workflows, such as tracking a site’s progress or enabling autonomous robotic inspection [1]. Visual-inertial SLAM is a practical tool for this goal, but construction site environments pose a challenge for these systems due to variable lighting, moving workers, fast motions, and repetitive, low-textured structures [2], causing the estimated trajectory to accumulate drift. Nevertheless, construction environments present structural elements, such as walls and columns, that remain distinguishable despite these conditions.

Fig. 1: MVP-SLAM on a Hilti construction sequence, matching detected walls to the floor plan and correcting the estimated trajectory against ground truth.

Construction sites typically have floor plans available [3], 2D representations of these structural elements produced during the design phase that can serve as a reference to correct this drift. In more advanced cases, these are Building Information Models (BIMs), encoding this information in more detail, though they are not universally available and are costly to process [4]. Although discrepancies exist between the as-built site and this as-planned floor plan [5], they are also used to localize in the building. Visual localization algorithms, like Z-FLoc [6], localize within such a floor plan and correct drift against it, but only after a trajectory has been generated, in a post-processing step. Onsite operation, however, needs a camera position estimate while the site is still being traversed, not only once it is complete, so the correction must happen online, as the trajectory is built, rather than recovered afterwards.

Methods offering an online camera position estimation in construction sites integrate the floor plan directly into the SLAM back-end. These methods rely on semantic entities such as walls to establish correspondences between the as-built site and as-planned floor plan, incorporating each match into the back-end’s graph optimization [1]. However, these systems rely on active depth or range sensors, such as anchoring LiDAR point clouds to BIMs [5] or matching structural walls from RGB-D cameras [7], leaving visual-only settings largely unexplored. A particular challenge for these systems is to find reliable correspondences between detected structural elements, such as walls on the as-built site and their counterparts on the floor plan, since accumulated drift makes this association increasingly difficult as the environment grows.

The Hilti–Trimble SLAM Challenge 2026 [4] benchmarks such real-world construction sites and their challenging conditions, recorded with a hand-held device carrying two opposite-facing (back-to-back) fisheye cameras and an IMU, without depth sensor; the cameras cover 360∘, sharing too little overlap to act as a stereo pair. Floor plans are also available for these sites, and one of the challenge’s two tasks, Localization, recovers the camera trajectory within the floor-plan frame from an initial pose, while its SLAM task instead recovers it in an arbitrary frame. Both tasks are scored on trajectory accuracy alone, so neither requires online or real-time operation.

Fig. 2: System architecture of MVP-SLAM. Sensor and floor-plan inputs (left) feed a three-module pipeline whose optimized state is fed back into the SLAM system as persistent plan priors, producing the plan-aligned trajectory estimate (right).

To recover the trajectory in the plan’s frame while correcting drift, we present MVP-SLAM (Multi-Camera Visual-Inertial Floorplan-Prior SLAM), an online visual-inertial SLAM system that leverages the building’s floor plan to constrain its estimation, responding to the Hilti–Trimble SLAM Challenge 2026. Our contributions are: (i) an online multi-camera visual-inertial SLAM system that integrates the building’s floor plan to correct drift as the trajectory is built; (ii) a semantic wall detection and drift-aware association algorithm that establishes correspondences between the as-built map and the floor plan from cameras alone; and (iii) a floor-plan integration method that, through a multi-stage strategy, turns each match incrementally into a bounded, persistent correction, so the map and trajectory are jointly optimized online under visual, inertial, and plan constraints. MVP-SLAM ranked 2nd in the Localization task of the challenge and 5th in the SLAM task, the best-placed system in both among those that also operate online, integrate the floor plan, and localize within it.

II Related Work

Multi-camera visual-inertial SLAM

Monocular visual-inertial SLAM suits hand-held capture, but its limited field of view degrades tracking during poor lighting or occlusions. Multi-camera setups widen visual coverage to maintain feature continuity, but whereas MAVIS [8] relies on overlapping stereo pairs, purely visual non-overlapping setups [9] suffer from motion-dependent scale unobservability and drift in the absence of an IMU. On the Hilti–Trimble SLAM Challenge 2026 [4] rig, recent systems estimate the trajectory with feature-based visual-inertial SLAM [10, 11], while others recover the full trajectory offline through global factor-graph optimization [12]. These systems, however, leave the trajectory in an arbitrary frame, neither localizing within a floor plan nor exploiting the one available on these construction sites to correct it.

Localization with architectural priors

Several methods localize a camera within a pre-existing floor plan, aligning monocular detections or layout cues to it through semantic Monte-Carlo localization [3] or learned ray-based filtering [13], returning a global position but without using the floor plan to improve the trajectory or the map. In the challenge’s Localization task [4], Map-It Ralph [14] likewise makes no use of the floor-plan geometry, reaching the plan frame from the provided anchor pose alone. The top three teams, in contrast, do use the floor plan, aligning a completed estimate to it offline: Z-FLoc by a single global transform from a bird’s-eye reconstruction [6], OmniRecon by 2D ICP on an offline structure-from-motion cloud [15], and CUFE by refining completed trajectories for floor-plan consistency [16]. All of them, however, apply the floor plan only after the trajectory is complete, rather than integrating it into the estimation to correct drift as the trajectory is built.

Architectural priors inside SLAM

Architectural plans can be leveraged inside SLAM systems to localize estimates within a building structure while offering an external metric reference to correct drift. A recent system from the same challenge [17] relies entirely on the benchmark’s provided starting pose to localize within the building, whereas real-world inspection tasks normally start from an estimated position, such as that obtained by having an operator match a few detected walls to the floor plan [7]. Furthermore, it uses the floor plan only to reject false loop-closure candidates, without directly correcting the drift. Other works do integrate the floor plan into the estimation to correct drift, but sense the walls directly with a range or depth sensor. LiDAR pose-graph methods anchor scans to a BIM through multi-session alignment [5], while [1] matches an online scene graph of rooms and walls to one extracted from the floor plan, and ivS-Graphs [7] brings this to vision with a simpler, wall-only matching, from an RGB-D sensor. Such sensing is uncommon on hand-held site-capture rigs, and depth cameras reach only a few meters, short of a construction interior’s spans. A depth-free system [18] instead registers sparse map points to a digital twin, but its prior is a dense, photorealistic as-built mesh that must be acquired separately and updated as the site evolves, unlike an as-planned floor plan available from the design stage. Correcting drift online by integrating an as-planned floor plan into a camera-only visual-inertial estimation remains unexplored.

III Methodology

III-A System Overview

MVP-SLAM corrects the drift of a visual-inertial trajectory and localizes inside the building by introducing the as-planned floor plan into a graph-based back-end. We build it upon ORB-SLAM3 [19], a widely adopted and extensively benchmarked SLAM system of this kind.

MVP-SLAM takes four inputs: the two fisheye image streams, an IMU, a floor plan of the site, and one plan-frame initialization pose \mathbf{T}_{s}, needed only as a coarse estimate rather than ground truth, since later corrections absorb its error (Sec. III-D). It processes them through three modules, each supplying what the next requires (Fig. 2): a multi-camera visual-inertial SLAM (Sec. III-B) estimates a continuous, metric trajectory and reconstructs a map; a wall detection and matching module (Sec. III-C) detects walls W_{s} in the map and matches them to their floor-plan counterparts W_{\mathrm{plan}}; and an integration module (Sec. III-D) folds each matched wall into a multi-stage back-end optimization that jointly re-estimates the trajectory and map, aligning them with the floor plan (Fig. 1).

III-B Multi-camera visual-inertial SLAM

MVP-SLAM extends monocular-inertial SLAM to the sensor setup of the rig. Facing opposite directions, the two fisheye cameras share too little overlap for stereo matching; each stream is tracked monocularly against a single shared map, with features from both hemispheres and the high-frequency IMU constraining one trajectory. The extension spans initialization, tracking, and map merging when the track is lost.

Fast initialization. A first metric map anchors the rest of the trajectory. Monocular-inertial pipelines build it only once enough parallax is available [20, 19], which a sequence beginning at rest or turning in place does not provide, while stereo pipelines build it at once from a wide image overlap the rig lacks. In between, Li et al. [21] showed that, for multi-camera rigs with limited view overlap, the few features matched across cameras are enough to seed the map from a single frame. MVP-SLAM follows this idea, adapted to the rig and its IMU, trying two paths in order.

(a) Cross-camera metric seeding: Front-camera features falling in the narrow peripheral band shared by both cameras are searched in the back image along known cross-camera epipolar curves over a small set of depth hypotheses, and the median depth of the surviving matches sets the scale, every remaining feature placed at that depth along its ray. Unlike the vision-only method of [21] needing an accurately triangulated seed, this coarse map suffices: the 4 cm inter-camera baseline could not provide accurate depths, but the first inertial bundle adjustment recovers them together with the scale.

(b) Gyro-aided fallback: In visual odometry and outlier rejection, taking the inter-frame rotation from the gyroscope rather than estimating it has proven to make two-view geometry more robust [22, 23]: it reduces the problem from five degrees of freedom to the two of the translation direction, recovered by a 2-point RANSAC [23]. MVP-SLAM applies this to initialization: when too few cross-camera matches survive, it falls back to two views with the rotation from gyro preintegration. This also removes the choice between a homography and a fundamental matrix, unreliable where planar surfaces dominate and parallax is small, which improves the robustness of the initialization in these challenging scenes.

Multi-camera tracking and optimization. The back camera is rigidly attached to the front one, so, as in other multi-camera visual-inertial systems [24], its observations constrain the same visual-inertial state from the opposite viewing hemisphere. For a map point seen by the back camera, its reprojection residual composes the estimated front-camera pose with the fixed inter-camera extrinsic, known from the IMU-camera calibrations, before applying the Kannala–Brandt fisheye projection. Each camera populates the map on its own, back-camera points being triangulated temporally between keyframes with the rig motion as baseline. A point is then searched only in the camera that created it while it stays in that field of view: as the two hemispheres barely overlap, searching it in both cameras would only add computational cost and cross-camera mismatches between similar structures. When it leaves, during turns or corridor reversals, it is projected into the opposite camera, which re-observes it and keeps it tracked across the hemisphere change. When both cameras end up mapping the same structure, fusion merges the duplicated points and culling removes the redundant ones.

Inertial map merging. Tracking can still be lost, on a fast turn or in front of a textureless wall [2]. When the pose cannot be recovered in the current map, multi-map systems such as ORB-SLAM3 start a new one and connect it to the previous map only if place recognition later identifies an already mapped area [19]. Because the floor plan is anchored to the first map through the initialization pose, a new map opened at its own origin stays outside the plan frame until the operator walks through an already mapped area again, which may never happen in a walkthrough without loops, and no plan correction can be applied to it meanwhile.

MVP-SLAM instead keeps the interrupted map and merges it with the new one through the IMU: the last keyframe before the loss is propagated by dead reckoning over the frames the loss lasts, and comparing this prediction with the first keyframe of the new map gives the transform between the two map frames. Both maps are metric and gravity-aligned, so the transform reduces to yaw and position, the four degrees of freedom unobservable in visual-inertial estimation [20], and is applied once, leaving the internal geometry of each segment untouched. Both segments therefore remain in the anchored frame of the first map, so walls matched in either of them constrain the whole trajectory, and a structure seen on both sides of the loss refines the join itself (Sec. III-D).

III-C Semantic wall detection and floor-plan association

This module supplies the back-end with wall correspondences in three steps: wall detection recovers the as-built walls W_{s}, floor-plan processing extracts the plan walls W_{\mathrm{plan}}, and a drift-aware association matches the two.

Wall detection. Planar surfaces are established landmarks in structured-environment SLAM [25], but fitting them directly to a feature-based map is error-prone, since its points come from many structures and non-wall planar clutter such as cabinets or panels is easily mistaken for a wall. MVP-SLAM therefore detects walls semantically, classifying the map points and fitting wall planes only to those labelled as wall [7]. These labels come from a pretrained panoptic segmenter (EoMT [26]), which assigns every pixel of both fisheye images a COCO-panoptic class, a vocabulary that already contains the wall class, so no site-specific fine-tuning is required. Each local-map point is then assigned the semantic class of its projected pixel on a per-frame basis as observations arrive. Labels are stabilized with a saturating hysteresis counter, so a point’s label flips only under sustained contradictory evidence. Three filters then suppress spurious wall labels: a gravity filter demotes wall labels whose viewing normal is too close to vertical (floor/ceiling bleed-through), a depth band rejects unreliable projections, and newly created map points are withheld until their labels stabilize.

At every keyframe, these wall-labelled points are fitted into wall planes with a sequential RANSAC in a gravity-aligned frame. Verticality removes one rotational degree of freedom, so a wall is parameterized by azimuth and distance only and fitted from a minimal sample of two points; hypotheses are scored by a robust RANSAC objective (MSAC) and accepted only if their inliers form a single contiguous wall segment. A freshly fitted plane is tentative until confirmed by consistent re-fitting across several keyframes, after which it becomes a detected wall W_{s} eligible for matching with the floor plan.

Floor-plan processing. Since the wall detection module recovers only the larger, clearly-observed walls, not every thin partition or facade detail drawn in the CAD, the floor plan is reduced to a comparable, canonical set of wall faces the detected walls can be matched against. MVP-SLAM converts the CAD wall layers into 2D wall faces, registers them to the plan raster by a grid search maximizing wall-pixel overlap, and canonicalizes them by merging collinear spans, removing duplicate or ambiguous faces, assigning consistent outward normals, and completing missing opposite faces. Exterior/interior labels are inherited from the CAD layers, with exterior taking precedence when a face combines both.

Drift-aware association. Each accepted match becomes a persistent constraint on the trajectory, so the association commits conservatively to avoid corrupting the rest of the estimate. The initialization pose \mathbf{T}_{s} places the floor plan in the as-built frame, so this step can find the correct associations between each detected wall W_{s} and its floor-plan counterpart W_{\mathrm{plan}}. A wall constrains the estimate only along its normal, so the integration stage (Sec. III-D) applies each correction anisotropically, and how much a candidate W_{\mathrm{plan}} can be trusted then depends on how far the estimate has drifted along that direction. In principle, the estimator’s pose covariance could measure this, but a keyframe-windowed bundle adjustment holds out-of-window poses fixed [27], so this covariance reflects only the local window and not this accumulated drift.

We therefore estimate this drift from the trajectory instead. It grows the longer the trajectory runs without a correction in that direction, so for each detected wall we track it through \ell, the length traveled since the last accepted exterior association with a parallel wall normal (|\mathbf{n}_{i}^{\!\top}\mathbf{n}_{j}|\!\geq\!0.9); only exterior wall matches, being the more reliable, reset \ell. A small \ell means the perpendicular distance \Delta d to a plan face is still reliable; a large \ell means the trajectory may have drifted farther from the correct W_{\mathrm{plan}}. We use \ell in two ways. First, it widens a hard distance gate: a candidate W_{\mathrm{plan}} is admitted only if

\Delta d\;\leq\;\min\!\left(d_{0}+\alpha_{\ell}\,\ell,\;c\right),
(1)

so the tolerated perpendicular offset starts at d_{0}, grows with drift at rate \alpha_{\ell}, and is capped at a wall class dependent value c that limits how far the gate can open (Table II). This cap is tighter for interior walls than exterior ones (c_{\mathrm{int}}\!<\!c_{\mathrm{ext}}), because interior partitions are densely packed with near-parallel neighbours, where a looser gate risks matching the wrong twin, and are detected less reliably from shorter range, whereas well-separated facades can absorb more drift before a match becomes ambiguous. Second, it saturates the distance that enters the score, yielding the effective distance \widetilde{\Delta d},

\widetilde{\Delta d}\;=\;(1-\sigma)\,\Delta d\;+\;\sigma\,\min(\Delta d,\,d_{\mathrm{sat}}),
(2)

which equals the true distance \Delta d when drift is small (\sigma{=}0) and caps it at d_{\mathrm{sat}} when drift is large (\sigma{=}1). The blend factor \sigma rises linearly from 0 to 1 as \ell grows from \ell_{\min} to \ell_{\max} (Table II), holding at 0 below \ell_{\min} and 1 above \ell_{\max}. With large drift the distance therefore stops dominating, leaving the angle and overlap terms to decide among the remaining candidates. After this drift-aware distance handling, each surviving plan face is scored with

s=w_{a}\,\Delta\theta+w_{d}\,\widetilde{\Delta d}+w_{f}\,(1-f),
(3)

where s is the match cost, so the face minimizing s wins; \Delta\theta is the azimuth difference between the W_{s} and W_{\mathrm{plan}} normals, taken with sign so that a face whose normal points the opposite way is penalized rather than counted as aligned; f\!\in\![0,1] is the fraction of the detected extent contained in the face; and the weights w_{a},w_{d},w_{f} (Table II) bring the radian, metre, and unitless terms to a common scale. The angle term is weighted most strongly because small orientation errors make a wall match unreliable even when its centroid is close.

The final match of a detected wall W_{s} to its plan face W_{\mathrm{plan}} is committed only if it is unambiguous, which a rejection cascade enforces before the pair is handed to the integration stage. The lowest-cost candidate is discarded if its score exceeds a class-dependent ceiling, if it does not beat the runner-up by a sufficient margin (Table II), or if a same-class competitor lies at a comparable perpendicular distance.

III-D Multi-stage floor-plan integration

Each matched pair (W_{s},W_{\mathrm{plan}}) provides a plan correspondence, but it reduces drift only once integrated into the SLAM graph as factors that optimize the trajectory and map. Introducing every accepted correspondence into one joint plan-constrained optimization makes that problem grow with the trajectory and accumulated associations [28]. MVP-SLAM instead incorporates each correspondence incrementally through a multi-stage optimization run whenever a new association is committed. Because this re-estimation happens at every match, MVP-SLAM continually realigns and localizes within the floor plan, so errors such as an imperfect initial pose \mathbf{T}_{s} are progressively absorbed. It runs in three steps: a local alignment over a bounded window (Step 1), a pose propagation of that correction (Step 2), and its persistence as a prior in later optimizations (Step 3), illustrated in Fig. 3.

Fig. 3: Incremental multi-stage integration of a matched wall, preserving earlier corrections.
III-D1 Step 1: Local alignment mini-graph

The first stage builds a local 4-DoF mini-graph for the newly matched pair (W_{s},W_{\mathrm{plan}}), optimizing only yaw and translation because the wall constraints are vertical and roll/pitch are fixed by gravity. In it, the plan wall W_{\mathrm{plan}} is a fixed vertex and the detected wall W_{s} an optimizable one, joined by a high-information alignment edge; every keyframe observing W_{s} is a pose vertex, tied to W_{s} by a plane-projection edge and to its covisibility neighbours by relative 4-DoF edges. Aligning W_{s} with W_{\mathrm{plan}} therefore shifts the observing keyframes and removes their accumulated drift. Besides these observers, only the keyframes in the temporal gap between them and their covisibility neighbours are free to correct. If a wall was observed over a long span, this set is capped at a maximum window size (Table II), keeping the most recent keyframes. Each match is corrected on its own window, keeping the update incremental and per-association. While physical walls are finite surfaces represented by W_{s} and W_{\mathrm{plan}}, their supporting infinite planes, \pi_{s} and \pi_{\mathrm{plan}}, are used in this mini-graph formulation. The local alignment minimizes the joint cost:

\min_{\{\mathbf{T}_{k}\},\,\pi_{s}}\;\left\lVert\pi_{s}\ominus\pi_{\mathrm{plan}}\right\rVert^{2}_{\mathbf{\Lambda}_{B}}+\sum_{k\in\mathcal{O}}\left\lVert\mathbf{e}^{\pi}_{k}\right\rVert^{2}_{\mathbf{\Lambda}_{\pi}}+\sum_{(i,j)\in\mathcal{E}_{\mathrm{loc}}}\left\lVert\mathbf{e}_{ij}\right\rVert^{2}_{\mathbf{\Lambda}_{ij}},
(4)

whose three terms weight these edges: the alignment edge (\ominus the difference between infinite plane parameters [25]) pulls the detected wall \pi_{s} onto the fixed plan face \pi_{\mathrm{plan}}, penalizing their azimuth and signed-distance mismatch; plane-projection residuals \mathbf{e}^{\pi}_{k} tie each observing keyframe k\in\mathcal{O} to the wall measurement stored when it observed the wall, so consistency is enforced against the local observation, not a global drifted pose; and relative 4-DoF covisibility edges \mathcal{E}_{\mathrm{loc}} preserve the local trajectory shape while yaw and translation adapt to the wall. The first keyframe is a gauge when no earlier plan correction exists. Rather than modifying SLAM poses directly, the mini-graph outputs corrected keyframe targets storing their associated wall normals to guide subsequent directional priors

III-D2 Step 2: Correction propagation to the system

Applied on their own, the mini-graph targets would end at the window boundary, leaving a discontinuity against the rest of the trajectory. A second 4-DoF pose graph therefore propagates them to the surrounding keyframes. Its adaptive window is built around the newly corrected keyframes and an anchor, chosen when possible as the latest previous correction whose stored wall normal is parallel to the current one. The window bridges from this anchor to the newly corrected span and extends backward from it, farther if too few same-axis old priors fall inside, so the previous correction supports the update over multiple keyframes rather than a single fixed anchor. The graph combines relative edges (covisibility and temporal-chain) with unary priors (fresh targets from Step 1 and previous corrections) by minimizing the joint cost:

\min_{\{\mathbf{T}_{k}\}}\sum_{(i,j)\in\mathcal{E}}\left\lVert\mathbf{e}_{ij}\right\rVert^{2}_{\mathbf{\Lambda}_{ij}}+\sum_{k\in\mathcal{P}}\left\lVert\mathbf{e}_{k}\right\rVert^{2}_{\mathbf{\Lambda}_{k}},
(5)

where \mathcal{E} contains relative 4-DoF edges, \mathcal{P} is the set of keyframes carrying targets from Step 1, \mathbf{e}_{ij} is the relative residual, and \mathbf{e}_{k} is the unary prior residual pulling keyframe k toward its target pose, weighted by the block-diagonal information matrix \mathbf{\Lambda}_{k} (Fig. 3, Step 2 inset):

\displaystyle\mathbf{\Lambda}_{k} \\ \displaystyle=\;\begin{bmatrix}\mathbf{\Lambda}_{t}&\mathbf{0}\\
\mathbf{0}&\lambda_{r}\,\mathbf{I}_{3}\end{bmatrix}\in\mathbb{R}^{6\times 6}, \\ \displaystyle\mathbf{\Lambda}_{t} \\ \displaystyle=\;\tau\,\mathbf{I}_{3}\;+\;s\!\!\sum_{\mathbf{n}\in\mathcal{N}_{k}}\mathbf{n}\mathbf{n}^{\!\top}\;\in\mathbb{R}^{3\times 3},

The translation block \mathbf{\Lambda}_{t} builds an anisotropic constraint based on environment geometry. For each observed wall normal \mathbf{n}\in\mathcal{N}_{k}, the outer-product term s\,\mathbf{n}\mathbf{n}^{\!\top} injects high information s perpendicular to the wall, while \tau\mathbf{I}_{3} provides a small isotropic information floor in every direction. The prior therefore pins the keyframe stiffly across each wall it observed while leaving it free to slide along the wall surface and vertically. Keyframes matched to multiple non-parallel walls are naturally pinned along all their respective normals. The rotation block \lambda_{r}\mathbf{I}_{3} provides an isotropic constraint on heading (yaw). Because roll and pitch remain fixed by gravity alignment in this 4-DoF formulation, yaw is the only rotational degree of freedom optimized, making a single scalar weight \lambda_{r} sufficient.

TABLE I: Absolute Pose Error (APE) RMSE (\mathrm{cm}) on Hilti–Trimble SLAM Challenge 2026 sequences, as reported on the official challenge leaderboard on the day of the challenge. Bold and underline denote best and second-best results per column per task; dashes (–) mark unavailable entries for Localization.
  • —
    データなし
    Floors and Sequences
    Ground Floors
    データなし
    データなし
    データなし
    データなし
    データなし
    データなし
    Upper Floors
    データなし
    データなし
    データなし
    データなし
    データなし
    データなし
    データなし
    データなし
    データなし
    データなし
    データなし
    Underground Levels
    データなし
    データなし
    データなし
    データなし
    データなし
    データなし
  • —
    データなし
    Floors and Sequences
    Floor 1
    データなし
    Floor 2
    データなし
    Floor EG
    データなし
    データなし
    Floor 3
    データなし
    Floor 4
    データなし
    Floor 5
    Floor 6
    データなし
    データなし
    データなし
    Floor 7
    データなし
    データなし
    Floor UG1
    データなし
    データなし
    データなし
    データなし
    Floor UG2
    データなし
  • —
    データなし
    Floors and Sequences
    07-07
    12-02
    12-02
    12-03
    10-16
    12-02a
    12-02b
    05-19
    12-02
    05-19
    12-02
    12-02
    06-18
    07-07
    12-02a
    12-02b
    12-02a
    12-02b
    12-03
    05-19
    06-18
    12-02a
    12-02b
    12-03
    12-02
    Average
  • Duration (\mathrm{sec.})
    データなし
    Floors and Sequences
    133.7
    276.3
    143.7
    149.9
    242.9
    126.4
    162.5
    115.5
    132.7
    91.9
    193.4
    160.5
    69.3
    73.0
    170.7
    124.6
    117.9
    155.1
    198.3
    205.5
    165.8
    254.8
    219.7
    134.4
    223.1
    161.7
  • Length (\mathrm{m})
    データなし
    Floors and Sequences
    157.8
    321.8
    154.8
    138.8
    240.4
    114.9
    158.6
    128.2
    148.8
    97.8
    214.1
    174.9
    79.4
    72.2
    210.6
    145.8
    145.4
    197.1
    152.6
    262.5
    222.6
    351.7
    322.5
    119.1
    292.1
    185.0
  • Localization
    Map-It Ralph [14]
    Floors and Sequences
    71.92
    39.68
    577.06
    233.11
    41.86
    87.98
    320.64
    103.89
    182.71
    38.67
    49.77
    82.38
    49.74
    41.31
    50.59
    109.38
    301.10
    105.34
    69.70
    95.13
    147.77
    264.79
    88.25
    127.67
    –
    136.69
  • CUFE [16]
    66.98
    Floors and Sequences
    41.20
    60.07
    42.64
    25.54
    30.83
    18.22
    84.93
    106.07
    75.80
    45.97
    65.83
    20.80
    28.94
    41.87
    78.11
    63.01
    88.80
    76.51
    38.88
    66.50
    48.09
    83.49
    142.69
    –
    60.07
    データなし
  • OmniRecon [15]
    38.05
    Floors and Sequences
    30.51
    38.79
    45.59
    33.32
    19.88
    41.99
    40.12
    64.74
    44.20
    17.58
    53.57
    30.20
    40.37
    52.20
    27.42
    14.26
    74.43
    27.60
    46.85
    55.00
    58.41
    53.50
    58.08
    –
    41.94
    データなし
  • Z-FLoc [6]
    28.87
    Floors and Sequences
    17.78
    32.55
    29.98
    15.07
    12.48
    18.37
    17.54
    36.45
    24.22
    12.60
    14.31
    27.69
    33.57
    16.10
    19.08
    13.30
    21.04
    22.61
    32.17
    33.56
    24.25
    42.62
    23.77
    –
    23.75
    データなし
  • Ours (MVP-SLAM)
    17.69
    Floors and Sequences
    25.48
    33.57
    18.54
    22.10
    13.48
    37.79
    13.65
    31.77
    17.37
    21.30
    31.65
    21.56
    18.85
    34.09
    18.58
    26.07
    31.57
    21.30
    27.35
    79.10
    46.90
    62.10
    31.19
    –
    29.29
    データなし
  • Ours (w/o plan)
    131.48
    Floors and Sequences
    251.39
    151.83
    95.04
    243.66
    26.28
    100.84
    55.17
    50.33
    100.81
    244.43
    54.88
    58.60
    51.21
    75.51
    35.91
    163.64
    113.40
    98.81
    220.39
    206.57
    350.86
    483.25
    109.04
    –
    144.72
    データなし
  • SLAM
    ACDC-VSLAM [10]
    Floors and Sequences
    10.07
    9.82
    7.69
    11.03
    7.84
    5.53
    6.70
    12.60
    6.50
    8.01
    7.12
    5.47
    6.36
    5.82
    6.02
    5.37
    5.95
    5.82
    10.27
    15.36
    19.88
    11.99
    13.17
    6.24
    12.86
    8.94
  • \sqrt{\text{VINS}} [12]
    14.65
    Floors and Sequences
    10.60
    13.54
    18.83
    9.87
    7.70
    6.67
    9.07
    8.66
    7.83
    9.18
    8.37
    5.88
    6.93
    8.83
    8.82
    8.89
    7.49
    11.34
    13.18
    18.97
    16.16
    16.43
    10.93
    13.67
    10.90
    データなし
  • Undisclosed
    17.59
    Floors and Sequences
    11.76
    36.04
    12.55
    9.46
    8.99
    12.75
    12.86
    36.01
    12.97
    6.74
    5.43
    16.71
    14.72
    5.41
    7.25
    8.92
    7.46
    16.32
    27.20
    33.24
    49.54
    50.03
    6.13
    40.83
    18.68
    データなし
  • QQ [11]
    29.42
    Floors and Sequences
    28.59
    23.23
    15.14
    10.97
    20.88
    20.48
    12.29
    38.15
    6.78
    16.40
    12.61
    5.32
    9.14
    12.36
    9.03
    13.18
    19.41
    23.93
    24.08
    39.23
    28.89
    39.67
    16.66
    28.34
    20.17
    データなし
  • Ours (MVP-SLAM)
    18.35
    Floors and Sequences
    17.93
    19.19
    15.88
    18.15
    8.93
    20.20
    11.95
    25.05
    13.92
    15.72
    33.58
    17.05
    10.32
    29.84
    10.50
    18.31
    11.49
    12.35
    30.98
    56.87
    56.47
    83.06
    20.52
    29.17
    24.23
    データなし
  • Ours (w/o plan)
    104.96
    Floors and Sequences
    128.91
    50.79
    78.35
    116.47
    10.40
    22.32
    35.78
    93.05
    55.04
    100.69
    21.00
    29.62
    14.69
    49.29
    18.78
    37.41
    46.37
    32.66
    155.01
    68.01
    177.93
    187.25
    74.11
    121.42
    73.21
    データなし

After propagation, the optimized poses are written back to the SLAM state. Map points move rigidly with their most recent corrected observer, velocities rotate with their keyframes, and the IMU preintegration terms are re-evaluated at the new states. Without this patch, the next inertial local BA would see inconsistent inertial residuals and pull the trajectory back toward the pre-correction state.

III-D3 Step 3: Persistence in subsequent optimization

A correction is useful only if it persists in the optimizations that follow. We thus store each corrected target with its wall-normal set \mathcal{N}_{k} and reintroduce it as a unary anisotropic prior whenever the affected keyframe appears in a later optimization.

In the local inertial BA that follows each correction, the priors keep the continuous SLAM estimate coherent with them, using the same anisotropic form but a softer information ratio (Table II) so that reprojection and inertial terms still refine the map along the wall while the floor plan prevents drift back across its normal. Each local BA jointly optimizes map points, keyframe poses, velocities, and IMU biases under these plan priors, so the final map and trajectory jointly satisfy visual, inertial, and structural constraints.

To enforce global alignment without noise, floor-plan priors enter the essential-graph optimization [19] at loop closures and map merges. These constraints are applied only to keyframes associated with exterior walls, more reliable than interior ones.

IV Experimental Results

TABLE II: Key parameters of MVP-SLAM.
  • Drift-aware association (Sec. III-C)
    Symbol
    データなし
    Value
    データなし
    Description
    データなし
  • score weights
    Symbol
    w_{a,d,f}
    Value
    6 rad-1, 1 m-1, 1
    Description
    angle, dist., overlap
  • gate ramp
    Symbol
    d_{0},\alpha_{\ell}
    Value
    1 m, 0.1
    Description
    offset, drift growth
  • gate cap
    Symbol
    c_{\mathrm{ext/int}}
    Value
    5, 2.5 m
    Description
    ext./int. distance cap
  • dist. saturation
    Symbol
    d_{\mathrm{sat}}
    Value
    1 m
    Description
    score distance cap
  • saturation range
    Symbol
    \ell_{\min/\max}
    Value
    15, 40 m
    Description
    score ramp
  • runner-up margin
    Symbol
    –
    Value
    1.6\times
    Description
    match uniqueness
  • Floor-plan integration (Sec. III-D)
    Symbol
    データなし
    Value
    データなし
    Description
    データなし
  • mini-graph window
    Symbol
    –
    Value
    200 KF
    Description
    max. KFs in Step1
  • along-normal info
    Symbol
    s
    Value
    10^{4}
    Description
    stiffness across wall
  • tangent floor
    Symbol
    \tau
    Value
    1 / 50
    Description
    propagation / local BA
  • rotation info
    Symbol
    \lambda_{r}
    Value
    10^{3} / 10^{4}
    Description
    propagation / local BA

IV-A Validation Methodology

Datasets. We evaluate on the Hilti–Trimble SLAM Challenge 2026 dataset [4], recorded in active construction sites with a hand-held rig carrying two back-to-back \sim200∘ fisheye cameras operating at 30 Hz and a 1 kHz IMU. We address both official tasks with the same system. Localization is scored in the floor-plan frame and SLAM is scored after rigid alignment to the reference. Localization contains 24 scored sequences, excluding Floor UG2 because no floor plan is released.

We group the sequences into ground floors, upper floors, and underground levels. Across all of them, the site is under construction and can differ from the floor-plan. The ground floors range from open spaces to the more built-out entrance level EG; the upper floors include interior areas and one outdoor terrace run; and the underground levels are large, sparsely-walled spaces dominated by structural columns.

Baselines. We compare MVP-SLAM against the other top-five competitors of each task (Table I), with per-sequence results taken from the official challenge release [4].

Implementation Details. MVP-SLAM runs on a workstation equipped with an Intel Core i9-11950H CPU and an NVIDIA T600 GPU (4 GB), 32 GB RAM. Key parameters are listed in Table II.

Fig. 4: Trajectory comparison in the Localization task, in the floor-plan frame, on three representative sequences.

Trajectory estimation Performance. The challenge provides a LiDAR-inertial reference trajectory, used only for evaluation. Both tasks report per-sequence RMSE of the Absolute Pose Error (APE) against it: 2D in the floor-plan plane for Localization, 3D after rigid alignment for SLAM. A trajectory must cover at least 99\% of the reference poses to be scored. To isolate the impact of map priors, we also evaluate an ablated variant, Ours (w/o plan). It uses the same visual-inertial SLAM setup but disables wall associations and floor-plan integration.

SLAM initialization and robustness ablation. We assess the fast metric initialization and the map preservation across tracking losses of the multi-camera front-end (Sec. III-B) with an ablation. The reduced configuration disables them, so it initializes and recovers as a standard visual-inertial SLAM would, while the full configuration keeps them enabled. We measure initialization success, time to metric scale, trajectory coverage, and whether the run satisfies the challenge coverage requirement on six sequences, two per environment type.

Wall detection and matching Performance. We evaluate wall detection and wall-to-plan matching on the same six sequences, using precision, recall, and F1. The reference set is hand-labeled from the floor plan, camera images, and reconstructed map, and contains only walls that are visible to the cameras and supported by reconstructed map points. Unbuilt planned walls and transient clutter are ignored. For detection, a true positive is a detected plane on the same physical wall as a reference wall; a false positive is a plane fitted to non-wall structure, clutter, or empty space; and a false negative a reference wall not recovered. For matching, we manually label each detected wall’s correct plan face; a match is correct only when the committed association links to it.

IV-B Results and Discussion

IV-B1 Trajectory Estimation Performance

Table I reports the APE RMSE for both official challenge tasks. In the Localization task, MVP-SLAM achieves a mean APE of 0.29 m and ranks second among the 22 Localization teams, close to the winning Z-FLoc submission (0.24 m). MVP-SLAM is best or second-best on most sequences, showing that the plan constraints keep the trajectory well aligned with the building frame across floors and environment types. Without plan integration, Ours (w/o plan) reaches 1.45 m mean APE; while tracking remains stable, accumulated drift is no longer bounded. Adding the plan integration reduces the Localization error by 4.9\times, from 1.45 m to 0.29 m. Fig. 4 illustrates this effect on representative ground-floor, upper-floor, and underground sequences.

In the SLAM task, MVP-SLAM obtains 0.24 m mean APE and ranks fifth among 62 teams. As this metric only accounts for the trajectory shape (Sec. IV-A), the top-ranked systems, led by ACDC-VSLAM [10] (0.09 m) and \sqrt{\text{VINS}} [12] (0.11 m), reach lower APE through mechanisms such as enhanced loop closure, effective since the challenge sequences often revisit previously seen areas, and richer visual features. MVP-SLAM instead uses the floor plan to both localize the trajectory and correct drift, reaching the top five and demonstrating that the floor-plan integration reduces the aligned error 3.0\times from 0.73 m (Ours w/o plan) to 0.24 m.

As established in Sec. II, the other Localization teams apply the floor plan offline, once the trajectory is built [6, 15, 16]; MVP-SLAM is instead the only one to integrate it online, folding each match into the estimation as it is detected. The same distinction holds in the SLAM task, where MVP-SLAM is again the top-ranked team among those that integrate the floor plan online, and the only one to also localize within it, showing the building’s floor plan is a promising, distinct prior for improving a SLAM system.

TABLE III: Initialization and robustness ablation for Ours (w/o plan): 3 repeats per sequence, both configurations initialized every run. t_{\mathrm{metric}}: median time to metric scale (s); Cov.: mean coverage (%); APE: median Localization error (m).
  • Ground
    Sequence
    Floor 1 (07-07)
    Reduced t_{\mathrm{metric}}
    3.1
    Reduced Cov.
    100
    Reduced APE
    1.38
    Full t_{\mathrm{metric}}
    2.5
    Full Cov.
    100
    Full APE
    1.31
  • Floor EG (10-16)
    Sequence
    2.5
    Reduced t_{\mathrm{metric}}
    100
    Reduced Cov.
    3.02
    Reduced APE
    2.2
    Full t_{\mathrm{metric}}
    100
    Full Cov.
    2.44
    Full APE
    データなし
  • Upper
    Sequence
    Floor 3 (05-19)
    Reduced t_{\mathrm{metric}}
    2.8
    Reduced Cov.
    100
    Reduced APE
    0.58
    Full t_{\mathrm{metric}}
    2.5
    Full Cov.
    100
    Full APE
    0.55
  • Floor 6 (12-02a)
    Sequence
    2.5
    Reduced t_{\mathrm{metric}}
    100
    Reduced Cov.
    1.01
    Reduced APE
    2.5
    Full t_{\mathrm{metric}}
    100
    Full Cov.
    0.76
    Full APE
    データなし
  • Under
    Sequence
    Floor UG1 (06-18)
    Reduced t_{\mathrm{metric}}
    6.6
    Reduced Cov.
    96.7
    Reduced APE
    28.0
    Full t_{\mathrm{metric}}
    2.5
    Full Cov.
    100
    Full APE
    2.07
  • Floor UG1 (12-02a)
    Sequence
    8.1
    Reduced t_{\mathrm{metric}}
    100
    Reduced Cov.
    5.10
    Reduced APE
    2.4
    Full t_{\mathrm{metric}}
    100
    Full Cov.
    3.51
    Full APE
    データなし
IV-B2 SLAM Initialization and Robustness Ablation

Table III shows that both configurations initialize reliably and reach metric scale in \sim2.5s on upper floors, and diverge only in underground environments. Here, the reduced configuration takes up to 3\times longer to reach metric scale (6.6–8.1s) and fails the 99\% coverage threshold on Floor UG1 (06-18). This degradation occurs when a hard initialization briefly drops tracking, creating a second map frame. Without map preservation, the two maps remain unmerged. As a result, abandoned keyframes drop out of the trajectory, causing task failures and misaligning the floor plan reference. By maintaining continuity across tracking losses, map preservation merges these segments, recovering full coverage and reducing Floor UG1 (06-18) Localization error from 28m to 2.07m.

TABLE IV: Wall-detection and wall-matching performance (precision P, recall R, F1, in percent). Overall pools counts.
  • Floor
    Sequence
    Detection
    P
    R
    F1
    Matching
    P
    R
    F1
  • Ground
    Floor 1 (07-07)
    Detection
    94.4
    50.0
    65.4
    Matching
    100.0
    94.1
    97.0
  • Floor EG (10-16)
    100.0
    Detection
    28.6
    44.4
    100.0
    Matching
    100.0
    100.0
    データなし
  • Upper
    Floor 3 (05-19)
    Detection
    100.0
    57.9
    73.3
    Matching
    100.0
    100.0
    100.0
  • Floor 6 (12-02a)
    84.6
    Detection
    35.5
    50.0
    90.9
    Matching
    100.0
    95.2
    データなし
  • Under
    Floor UG1 (06-18)
    Detection
    100.0
    37.8
    54.9
    Matching
    100.0
    100.0
    100.0
  • Floor UG1 (12-02a)
    100.0
    Detection
    40.4
    57.6
    100.0
    Matching
    100.0
    100.0
    データなし
  • —
    Overall
    Detection
    96.9
    38.4
    55.0
    Matching
    98.9
    98.9
    98.9
IV-B3 Wall Detection and Matching Performance

Table IV reports wall detection and wall-to-plan matching on six representative sequences. Detection is conservative by design: it reaches high precision (96.9\%) but moderate recall (38.4\%). This choice follows from the integration design (Sec. III-D), where each accepted wall becomes a persistent plan prior; a wrong wall is therefore more harmful than a missed one. Most missed walls are visible in the images but have too few or too noisy reconstructed map points to support a stable plane fit, especially in cluttered or texture-scarce areas.

Once a wall is detected, association is highly reliable: matching precision, recall, and F1 are all 98.9\% pooled. The detected wall candidates are usually close to their true plan faces and separated from alternatives, so the score (Eq. 3) and rejection cascade resolve most matches cleanly. The single wrong match occurs on Floor 6, where two nearly collinear plan faces fall within the gate, and the single missed match is a detected wall that receives no association. Thus, in these sequences, the plan prior is limited mainly by which walls can be detected from the sparse map, not by the association step. This is sufficient for localization when the accepted walls are distributed: Floor EG, for example, attains a 22 cm Localization APE despite only 28.6\% wall-detection recall.

Limitations. MVP-SLAM corrects drift only where it detects a wall and matches it to the floor plan. Its wall module fits planar walls, so in open areas with few walls, notably large underground spaces dominated by structural columns, few matches are available and the trajectory drifts uncorrected.

V Conclusions and Future Work

We presented MVP-SLAM, a multi-camera visual-inertial SLAM system that uses an as-planned floor plan as a reference to reduce drift accumulated on construction sites. It tracks its position on two opposite-facing fisheye cameras and an IMU, and corrects the drift online by mapping walls, matching them to the floor plan, and integrating it into the SLAM back-end through each match as the trajectory is built. On the Hilti–Trimble SLAM Challenge 2026 it ranked second in Localization (0.29 m mean RMSE) and fifth in SLAM (0.24 m). It is the top team in both tasks among those that also operate online, integrate the floor plan, and localize within it.

Future work will explore lightweight map densification to make the detection of walls and additional structural primitives such as columns easier, integrating them into the optimization for drift correction. It will also target the SLAM’s internal consistency directly, strengthening loop closure and adopting the richer-feature mechanisms of higher-ranked SLAM teams.

References

  1. [1] M. Shaheer, J. A. Millan-Romera, H. Bavle, J. L. Sanchez-Lopez, J. Civera, and H. Voos (2023) Graph-based global robot localization informing situational graphs with architectural graphs. In 2023 IEEE/RSJ Int. Conf. Intelligent Robots and Systems (IROS), pp. 9155–9162.
  2. [2] M. Helmberger, K. Morin, B. Berner, N. Kumar, G. Cioffi, and D. Scaramuzza (2022) The Hilti SLAM challenge dataset. IEEE Robot. Autom. Lett. 7 (3), pp. 7518–7525.
  3. [3] O. Mendez, S. Hadfield, N. Pugeault, and R. Bowden (2018) SeDAR: semantic detection and ranging — humans can localise without LiDAR, can robots?. In 2018 IEEE Int. Conf. Robotics and Automation (ICRA), pp. 6053–6060.
  4. [4] S. Centanni, Y. Zhang, Y. Tao, J. Kindle, F. Neuhaus, T. Koß, A. Patel, M. Helmberger, E. Szymańska, T. Gräber, et al. (2026) Hilti-trimble-oxford dataset: 360 visual-inertial benchmark with floor plan priors for slam and localization. arXiv preprint arXiv:2607.06464.
  5. [5] M. A. V. Torres, A. Braun, and A. Borrmann (2023) BIM-SLAM: integrating BIM models in multi-session SLAM for lifelong mapping using 3D LiDAR. In Proc. Int. Symp. Automation and Robotics in Construction (ISARC), Vol. 40, pp. 521–528.
  6. [6] A. Umemura, T. Kuwahara, M. Pollefeys, and D. Barath (2026) Z-FLoc: zero-shot floorplan localization via geometric primitives. arXiv preprint arXiv:2606.04788.
  7. [7] A. Bikandi-Noya, M. Fernandez-Cortizas, M. Shaheer, A. Tourani, H. Voos, and J. L. Sanchez-Lopez (2025) BIM-informed visual SLAM for construction monitoring. arXiv preprint arXiv:2509.13972.
  8. [8] Y. Wang, Y. Ng, I. Sa, Á. Parra, C. Rodriguez-Opazo, T. Lin, and H. Li (2024) MAVIS: multi-camera augmented visual-inertial SLAM using SE2(3)-based exact IMU pre-integration. In 2024 IEEE Int. Conf. Robotics and Automation (ICRA), pp. 1694–1700.
  9. [9] M. J. Tribou, A. Harmat, D. W. L. Wang, I. Sharf, and S. L. Waslander (2015) Multi-camera parallel tracking and mapping with non-overlapping fields of view. Int. J. Robotics Research 34 (12), pp. 1480–1500.
  10. [10] J. Jeon, D. Seo, J. Choi, S. Lee, J. Nam, H. Lim, and H. Myung (2026) ACDC-VSLAM: adaptive constraints for dual-fisheye-camera-based visual-inertial SLAM with point and line features. Note: Hilti–Trimble SLAM Challenge 2026, SLAM track report
  11. [11] J. Jiang (2026) Technical report for the Hilti \times Trimble SLAM challenge 2026. Note: Hilti–Trimble SLAM Challenge 2026, SLAM track report
  12. [12] J. Lee, H. Kim, J. Choi, J. Jeong, and Y. Cho (2026) \sqrt{\text{VINS}} with factor graph optimization-based visual-inertial SLAM in the Hilti–Trimble SLAM challenge 2026. Note: Hilti–Trimble SLAM Challenge 2026, SLAM track report
  13. [13] C. Chen, R. Wang, C. Vogel, and M. Pollefeys (2024) F3Loc: fusion and filtering for floorplan localization. In Proc. IEEE/CVF Conf. Computer Vision and Pattern Recognition (CVPR), pp. 18029–18038.
  14. [14] A. C. Demirtaş, B. Şekeroglu, and A. Topaloglu (2026) Two-stage visual-inertial SLAM with sliding-window filtering and global bundle adjustment for the Hilti \times Trimble SLAM challenge 2026. Note: Hilti–Trimble SLAM Challenge 2026, Localization track report
  15. [15] G. Tanner, G. Evangelou, J. Lechner, Z. Pataki, X. Jiang, P. Sarlin, and S. Liu (2026) Establishing the gold standard for 360{}^{\circ} visual-inertial reconstruction. Note: Hilti–Trimble SLAM Challenge 2026, Localization track report
  16. [16] CUFE Team, Cairo University (2026) A fully automated multi-stage localization pipeline for indoor floorplan-constrained trajectory estimation. Note: Hilti–Trimble SLAM Challenge 2026, Localization track report
  17. [17] Y. Zang (2026) Non-overlapping dual-fisheye visual-inertial SLAM based on ORB-SLAM3. Note: Hilti–Trimble SLAM Challenge 2026, organizer baseline
  18. [18] R. Merat, G. Cioffi, L. Bauersfeld, and D. Scaramuzza (2025) Drift-free visual slam using digital twins. IEEE Robot. Autom. Lett. 10 (2), pp. 1633–1640. External Links: Document
  19. [19] C. Campos, R. Elvira, J. J. G. Rodríguez, J. M. M. Montiel, and J. D. Tardós (2021) ORB-SLAM3: an accurate open-source library for visual, visual–inertial, and multimap SLAM. IEEE Trans. Robotics 37 (6), pp. 1874–1890.
  20. [20] T. Qin, P. Li, and S. Shen (2018) VINS-Mono: a robust and versatile monocular visual-inertial state estimator. IEEE Trans. Robotics 34 (4), pp. 1004–1020.
  21. [21] A. Li, D. Zou, and W. Yu (2021) Robust initialization of multi-camera SLAM with limited view overlaps and inaccurate extrinsic calibration. In 2021 IEEE/RSJ Int. Conf. Intelligent Robots and Systems (IROS), pp. 3361–3367. External Links: Document
  22. [22] L. Kneip, M. Chli, and R. Siegwart (2011) Robust real-time visual odometry with a single camera and an IMU. In Proc. British Machine Vision Conf. (BMVC),
  23. [23] C. Troiani, A. Martinelli, C. Laugier, and D. Scaramuzza (2014) 2-point-based outlier rejection for camera-IMU systems with applications to micro aerial vehicles. In 2014 IEEE Int. Conf. Robotics and Automation (ICRA), pp. 5530–5536.
  24. [24] Y. He, H. Yu, W. Yang, and S. Scherer (2022) Towards robust visual-inertial odometry with multiple non-overlapping monocular cameras. In 2022 IEEE/RSJ Int. Conf. Intelligent Robots and Systems (IROS), pp. 9452–9458.
  25. [25] M. Kaess (2015) Simultaneous localization and mapping with infinite planes. In 2015 IEEE Int. Conf. Robotics and Automation (ICRA),
  26. [26] T. Kerssies, N. Cavagnero, A. Hermans, N. Norouzi, G. Averta, B. Leibe, G. Dubbelman, and D. De Geus (2025) Your ViT is secretly an image segmentation model. In Proc. Computer Vision and Pattern Recognition Conf. (CVPR), pp. 25303–25313.
  27. [27] R. Mur-Artal and J. D. Tardós (2017) Visual-inertial monocular slam with map reuse. IEEE Robot. Autom. Lett. 2 (2), pp. 796–803. External Links: Document
  28. [28] M. Kaess, H. Johannsson, R. Roberts, V. Ila, J. J. Leonard, and F. Dellaert (2012) ISAM2: incremental smoothing and mapping using the bayes tree. Int. J. Robotics Research 31 (2), pp. 216–235.