EDBT 2026 Demo / reviewers in the wild / expert
Hao Wei 0008
dblp:96/133-8
· DBLP profile ↗
13ranked-venue papers
1as first author
13since 2021 · last 2026
0000-0002-7708-0243ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 1 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Sparse3DPR: Training-Free 3D Hierarchical Scene Parsing and Task-Adaptive Subgraph Reasoning from Sparse RGB ViewsabstractRecently, large language models (LLMs) have been explored widely for 3D scene understanding. Among them, training-free approaches are gaining attention for their flexibility and generalization over training-based methods. However, they typically struggle with accuracy and efficiency in practical deployment. To address the problems, we propose Sparse3DPR, a novel training-free framework for open-ended scene understanding, which leverages the reasoning capabilities of pre-trained LLMs and requires only sparse-view RGB inputs. Specifically, we introduce a hierarchical plane-enhanced scene graph that supports open vocabulary and adopts dominant planar structures as spatial anchors, which enables clearer reasoning chains and more reliable high-level inferences. Furthermore, we design a task-adaptive subgraph extraction method to filter query-irrelevant information dynamically, reducing contextual noise and improving 3D scene reasoning efficiency and accuracy. Experimental results demonstrate the superiority of Sparse3DPR, which achieves a 28.7% EM@1 improvement and a 78.2% speedup compared with ConceptGraphs on the Space3D-Bench. Moreover, Sparse3DPR obtains comparable performance to training-based methods on ScanQA, with additional real-world experiments confirming its robustness and generalization capability. Haida Feng, Hao Wei 0008, Zewen Xu, Haolin Wang 0005, Chade Li, Yihong Wu 0002 |
AAAI | 2 |
| 2026 | SLNMapping: Super Lightweight Neural Mapping in Large-Scale Scenes
Chenhui Shi 0001, Fulin Tang, Hao Wei 0008, Yihong Wu 0002 |
Int. J. Comput. Vis. | 3 |
| 2026 | CGFMamba-PCR: Color-Geometric Fusion Mamba-Based Color Point Cloud RegistrationabstractRecently, color point cloud registration has begun to receive attention. Unlike geometric-only point clouds, color point clouds incorporate additional color information. Therefore, color point cloud registration can achieve higher accuracy than geometric-only point cloud registration. Despite the success, existing methods are computationally intensive due to the high resource demands of Transformer. In this paper, we propose a Mamba architecture based registration algorithm, CGFMamba-PCR, with color-geometric fusion. Specifically, we propose CGFMamba, a novel color-geometric fusion Mamba network for color point cloud registration, which enhances feature representation through color-geometric guided point ordering and positional encoding. For the input color point cloud pair, they are passed through CGFMamba based feature extraction module to obtain their corresponding features. Then, these features are passed through feature matching and outlier rejection modules to obtain final registration result. Furthermore, an ordering method for the Mamba architecture is proposed that clusters color hue and sorts spatial coordinates. Experiments on Color3DMatch and Color3DLoMatch datasets demonstrate that the proposed algorithm outperforms the state-of-the-art (SOTA) methods. The code of the proposed algorithm will be open-sourced upon acceptance of this paper. Shiyi Guo, Tong Jia 0001, Bi Yang, Yihong Wu 0002, Hao Wei 0008, Nannan Liu, Ning An 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | A Rotation-Translation Decoupled Solution for Visual-Inertial Initialization and Online Spatial-Temporal CalibrationabstractWe propose a novel initialization and online spatial-temporal calibration method for visual-inertial odometry (VIO), which decouples rotation and translation estimation to achieve higher accuracy and better robustness. Existing initialization methods suffer from limited accuracy or robustness (e.g., in scenarios with small translational motion) and rarely integrate simultaneous spatial-temporal calibration during initialization, despite its considerable practical value. Our proposed method leverages rotation-translation decoupling constraints to enable simultaneous estimation of gyroscope bias, extrinsic rotation, and camera-IMU time offset-even under pure rotational motion. Moreover, we are the first to conduct observability analysis on rotational constraints in rotation-translation decoupling methods, experimentally identifying the unobservable state-space directions under three degenerate motions within our approach. We also perform extensive experiments to delineate practical parameter solution boundaries for our method, with both efforts substantially enhancing the overall practical applicability of decoupling-based methods. Extensive experiments on simulated and real-world datasets demonstrate that our method outperforms state-of-the-art approaches in accuracy and robustness while maintaining computational efficiency. Furthermore, experiments verify that it significantly improves convergence in VIO systems. Bo Xu 0022, Zewen Xu, Yijia He, Zhanpeng Ouyang, Hao Wei 0008, Yihong Wu 0002, Jiancheng Li, Hongdong Li |
IEEE Trans. Robotics | 5 |
| 2025 | FEAST-Mamba: FEAture and SpaTial Aware Mamba Network with Bidirectional Orthogonal Fusion for Cross-Modal Point Cloud SegmentationabstractPoint cloud segmentation has a wide range of applications in autonomous driving, augmented reality and virtual reality. Multi-modal fusion strategies have received increasing attention in point cloud segmentation recently. Despite the success, existing methods usually generate unnecessary information loss or redundancy. In this paper, we propose FEAST-Mamba, a novel FEAture and SpaTial aware Mamba network to tackle multi-modal point cloud segmentation. To exploit the complementarity between different modals, we propose a bidirectional orthogonal attention module, where features are first bidirectionally interacted with each other through cross-modal attention, and then orthogonal fusion is used to reduce feature redundancy. Furthermore, a reordering strategy is proposed for the Mamba architecture that takes into account both spatial and semantic information during cross-modal feature ordering. Experiments on indoor datasets, S3DIS and ScanNet, and outdoor datasets, nuScenes and SemanticKITTI, show that the proposed method achieves state-of-the-art performances. Chade Li, Hao Wei 0008, Yihong Wu 0002 |
AAAI | 4 |
| 2025 | Hi-Gaussian: Hierarchical Gaussians Under Normalized Spherical Projection for Single-View 3D Reconstruction
Binjian Xie, Hao Wei 0008, Yihong Wu 0002 |
ICCV | 3 |
| 2025 | DOGE: An Extrinsic Orientation and Gyroscope Bias Estimation for Visual-Inertial Odometry InitializationabstractMost existing visual-inertial odometry (VIO) initialization methods rely on accurate pre-calibrated extrinsic parameters. However, during long-term use, irreversible structural deformation caused by temperature changes, mechanical squeezing, etc. will cause changes in extrinsic parameters, especially in the rotational part. Existing initialization methods that simultaneously estimate extrinsic parameters suffer from poor robustness, low precision, and long initialization latency due to the need for sufficient translational motion. To address these problems, we propose a novel VIO initialization method, which jointly considers extrinsic orientation and gyroscope bias within the normal epipolar constraints, achieving higher precision and better robustness without delayed rotational calibration. First, a rotation-only constraint is designed for extrinsic orientation and gyroscope bias estimation, which tightly couples gyroscope measurements and visual observations and can be solved in pure-rotation cases. Second, we propose a weighting strategy together with a failure detection strategy to enhance the precision and robustness of the estimator. Finally, we leverage Maximum A Posteriori to refine the results before enough translation parallax comes. Extensive experiments have demonstrated that our method outperforms the state-of-the-art methods in both accuracy and robustness while maintaining competitive efficiency. Zewen Xu, Yijia He, Hao Wei 0008, Yihong Wu 0002 |
ICRA | 3 |
| 2025 | Floorplan-SLAM: A Real-Time, High-Accuracy, and Long-Term Multi-Session Point-Plane SLAM for Efficient Floorplan ReconstructionabstractFloorplan reconstruction provides structural priors essential for reliable indoor robot navigation and high-level scene understanding. However, existing approaches either require time-consuming offline processing with a complete map, or rely on expensive sensors and substantial computational resources. To address the problems, we propose FloorplanSLAM, which incorporates floorplan reconstruction tightly into a multi-session SLAM system by seamlessly interacting with plane extraction, pose estimation, back-end optimization, and loop & map merging, achieving real-time, high-accuracy, and long-term floorplan reconstruction using only a stereo camera. Specifically, we present a robust plane extraction algorithm that operates in a compact plane parameter space and leverages spatially complementary features to accurately detect planar structures, even in weakly textured scenes. Furthermore, we propose a floorplan reconstruction module tightly coupled with the SLAM system, which uses continuously optimized plane landmarks and poses to formulate and solve a novel optimization problem, thereby enabling real-time and high-accuracy floorplan reconstruction. Note that by leveraging the map merging capability of multi-session SLAM, our method supports long-term floorplan reconstruction across multiple sessions without redundant data collection. Experiments on the VECtor and the self-collected datasets indicate that Floorplan-SLAM significantly outperforms state-of-the-art methods in terms of plane extraction robustness, pose estimation accuracy, and floorplan reconstruction fidelity and speed, achieving real-time performance at 25–45 FPS without GPU acceleration, which reduces the floorplan reconstruction time for a 1000 m2scene from 16 hours and 44 minutes to just 9.4 minutes. Haolin Wang 0005, Zeren Lv, Hao Wei 0008, Haijiang Zhu, Yihong Wu 0002 |
IROS | 3 |
| 2025 | Maximum Clique-Based Floorplan Association for Robust Multi-Session Stereo SLAM in Challenging Indoor EnvironmentsabstractExisting multi-session visual simultaneous localization and mapping (SLAM) systems struggle severely to achieve robust localization and map merging under extreme viewpoint and illumination variations, particularly when handling completely opposite viewpoints and drastic day-night lighting changes. These challenges stem largely from the limited viewpoint/illumination invariance of conventional low-level visual features and their inability to capture a global structural context. In this paper, we make the critical observation that a life-long floorplan not only encodes rich geometric and semantic information—serving as a robust high-level structural representation—but is also inherently more robust to severe viewpoint and illumination variations than purely visual data. Building on this insight, we propose a novel hierarchical framework for multi-session SLAM that integrates a floorplan-based map as a global feature to achieve robust indoor localization and map merging under drastic viewpoint and illumination shifts. In particular, we innovatively formulate floorplan association as a maximum clique problem augmented with trajectory data to achieve robust floorplan-level global localization. We further introduce a novel coarse-to-fine localization and map merging strategy that seamlessly integrates floorplan alignment, multistage point cloud registration, and feature matching, fully leveraging the macro-level stability of global features and the micro-level precision of local features to achieve keyframe-level fine localization. Extensive experiments on both public and self-collected datasets demonstrate that our method consistently outperforms state-of-the-art (SOTA) approaches reliant solely on low-level visual or geometric features. Crucially, it delivers superior accuracy and robustness even in the face of completely opposite viewpoints and extreme day–night illumination changes. This work underscores the promise of fusing macro-level floorplan representations with conventional SLAM frameworks to advance long-term, robust indoor localization and map merging under the most challenging conditions. Haolin Wang 0005, Hao Wei 0008, Zeren Lv, Haijiang Zhu, Yihong Wu 0002 |
IROS | 2 |
| 2024 | Novel camera self-calibration method with clustering prior and nonlinear optimization from an image sequence
Xiaohui Jiang, Haijiang Zhu, Ning An 0002, Binjian Xie, Hao Wei 0008, Fulin Tang, Yihong Wu 0002 |
Multim. Tools Appl. | 5 |
| 2024 | An Efficient Outlier Rejection Algorithm for Point Cloud RegistrationabstractPoint cloud registration plays an important role in many applications of computer vision and robotics. Outlier rejection is an essential step in this task. In this letter, we propose an efficient hand-crafted outlier rejection algorithm for point cloud registration. The proposed method includes four modules: seed selection, consensus set construction and sampling, transformation matrix calculation, and hypothesis selection. Particularly, we propose a novel seed selection module based on the addition of elements in the Spatial Consistency (SC) matrix and a new noise suppression strategy in constructing consensus sets. Experimental results on various datasets show that the proposed algorithm presents superior performance over the previous hand-crafted methods. Siyi Xiang, Shiyi Guo, Hao Wei 0008, Bingxi Liu 0001, Dabo Zhang |
IEEE Signal Process. Lett. | 3 |
| 2023 | PLPL-VIO: A Novel Probabilistic Line Measurement Model for Point-Line-Based Visual-Inertial OdometryabstractPoint and line features are complementary in Visual-Inertial Odometry (VIO) or Visual-Inertial Simultaneous Localization And Mapping (VI-SLAM) systems. The advantage of combining these two types of features relies on their proper weighting in the cost function, usually set by their uncertainty. Compared with point features, setting line segment endpoints' uncertainty with isotropic distribution is unreasonable. But the uncertainty of line feature observation, especially for the endpoints' uncertainty along the line, is difficult to set due to occlusion and fragmentation problems. In this article, we use infinite lines as the line feature observations and prove that the uncertainty of these observations is only related to the vertical uncertainty of the endpoints, thus avoiding setting the parallel uncertainty of the endpoints. Besides, we introduce a novel consistent measurement model for line features. Furthermore, for long-time constraints, we add 3D line segments into the state vector and derive how to update them properly. Finally, we construct a point-line-based VIO system that takes into account the uncertainty of line feature observations and the consistency of line feature measurements. The proposed VIO system is validated on two public datasets. The results show that the proposed method obtains the best accuracy compared with the state-of-the-art point-based VIO systems (OpenVINS, VINS-Mono), a point-line-based VIO system (PL-VINS), and a structural line-based system (StructVIO). Zewen Xu, Hao Wei 0008, Fulin Tang, Yihong Wu 0002, Gang Ma 0007, Shuzhe Wu, Xin Jin 0004 |
IROS | 2 |
| 2021 | Highly Efficient Line Segment Tracking with an IMU-KLT Prediction and a Convex Geometric Distance MinimizationabstractLine segment features become popular in SLAM community. Usually, line-based SLAM systems utilize local appearance descriptors for line segment tracking. However, traditional descriptor-based line segment tracking algorithms suffer from the problem that accuracy and speed cannot be possessed simultaneously, which affects the performance of line-based SLAM systems negatively. We propose a novel line segment tracking method with an IMU-KLT line segment prediction and a convex geometric distance minimization to boost line segment tracking performance in both accuracy and speed. Particularly, the proposed convex geometric distance minimization uses a ℓ1-norm model to minimize geometric constraints between predicted line segments and extracted line segments efficiently. Furthermore, the line segment tracking is embedded into a VIO system and we adapt it to obtain more reliable point tracking. Experimental results on public datasets show that the proposed line segment tracking method achieves much higher accuracy and much less time cost than state-of-the-art level, where not only the number of correct matches increases but also the inlier ratios are increased by at least 35.1% along with a 3 times faster speed. Besides, the VIO system combining the proposed line segment tracking is improved in terms of accuracy. Hao Wei 0008, Fulin Tang, Chaofan Zhang, Yihong Wu 0002 |
ICRA | 1 |