EDBT 2026 Demo / reviewers in the wild / expert
Hongming Shen
dblp:140/3990
· DBLP profile ↗
9ranked-venue papers
3as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | UniLGL: Learning Uniform Place Recognition for FOV-Limited/Panoramic LiDAR Global LocalizationabstractLiDAR-based Global Localization (LGL) is an essential ingredient for autonomous robots. However, existing LGL methods typically consider only partial information (e.g., geometric features) from LiDAR observations or are designed for homogeneous LiDAR sensors, overlooking the uniformity in LGL. In this work, a uniform LGL method is proposed, termed UniLGL, which simultaneously achieves spatial and material uniformity, as well as sensor-type uniformity. The key idea of the proposed method is to encode the complete point cloud, which contains both geometric and material information, into a pair of Bird's Eye View (BEV) images (i.e., a spatial BEV image and an intensity BEV image), thereby transforming the LGL problem into a cascaded LiDAR Place Recognition (LPR) and pose estimation problem from the perspective of image fusion. An end-to-end multi-BEV fusion network is designed to extract uniform features, equipping UniLGL with spatial and material uniformity. To ensure robust LGL across heterogeneous LiDAR sensors, a viewpoint invariance hypothesis is introduced, which replaces the conventional translation equivariance assumption commonly used in existing LPR networks and supervises UniLGL to achieve sensor type uniformity in both global descriptors and local feature representations. Moreover, UniLGL introduces a pipeline that leverages a pre-trained single-image Vision Foundation Model (VFM) for feature extraction to enhance the multi-BEV fusion LPR network, enabling strong generalization with only a few LiDAR data for fine-tuning. Finally, based on the mapping between local features on the 2D BEV image and the point cloud, a robust global pose estimator is derived that determines the global minimum of the global pose on <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"><tex-math notation="LaTeX">$\text{SE}(3)$</tex-math></inline-formula> without requiring additional registration. To validate the effectiveness of the proposed uniform LGL, extensive benchmarks are conducted in real-world environments, and the results show that the proposed UniLGL is demonstratively competitive compared to other State-of-the-Art (SOTA) LGL methods. Furthermore, UniLGL has been deployed on diverse platforms, including full-size trucks and agile Micro Aerial Vehicles (MAVs), to enable high-precision localization and mapping as well as multi-MAV collaborative exploration in port and forest environments, demonstrating the applicability of UniLGL in industrial and field scenarios. The code will be released at <uri xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">https://github.com/shenhm516/UniLGL</uri>. Hongming Shen, Yulin Hui, Zhenyu Wu 0001, Qiyang Lyu, Tianchen Deng, Danwei Wang |
IEEE Trans. Robotics | 1 |
| 2025 | MNE-SLAM: Multi-Agent Neural SLAM for Mobile RobotsabstractNeural implicit scene representations have recently shown promising results in dense visual SLAM. However, existing implicit SLAM algorithms are constrained to single-agent scenarios, and fall difficulty in large indoor scenes and long sequences. Existing multi-agent SLAM frameworks cannot meet the constraints of communication bandwidth. To this end, we propose the first distributed multi-agent collaborative SLAM framework with distributed mapping and camera tracking, joint scene representation, intra-to-inter loop closure, and multi-submap fusion. Specifically, our proposed distributed neural mapping and tracking framework only needs peer-to-peer communication, which can greatly improve multi-agent cooperation and communication efficiency. A novel intra-to-inter loop closure method is designed to achieve local (single-agent) and global (multi-agent) consistency. Furthermore, to the best of our knowledge, there is no real-world dataset for NeRF-based/GS-based SLAM that provides both continuous-time trajectories groundtruth and high-accuracy 3D meshes groundtruth. To this end, we propose the first real-world indoor neural slam (INS) dataset covering both single-agent and multi-agent scenarios, ranging from small room to large-scale scenes, with high-accuracy ground truth for both 3D mesh and continuous-time camera trajectory. This dataset can advance the development of the community. Experiments on various datasets demonstrate the superiority of the proposed method in both mapping, tracking, and communication. The dataset and code will be open-source on https://github.com/dtc111111/MNESLAM. Tianchen Deng, Guole Shen, Chen Xun, Shenghai Yuan 0001, Tongxin Jin, Hongming Shen, Jingchuan Wang, Hesheng Wang 0001, Danwei Wang, Weidong Chen 0001 |
CVPR | 6 |
| 2024 | CT-MLO: Voxel-Based Multi-LiDAR Odometry Using Continuous-Time Kalman FilterabstractIn recent years, LiDAR-based localization and mapping methods have achieved significant progress thanks to their reliable and real-time localization capability. However, single LiDAR odometry often faces hardware failures and degradation in practical scenarios, and the continuous-time measurement characteristic is constantly neglected by existing LiDAR odometry. This motivates us to develop a continuous-time Multi-LiDAR Odometry (MLO) method, namely CT-MLO, which can realize accurate and real-time state estimation using multi-LiDAR measurements through a continuous-time perspective. Due to the advantageous continuous-time formulation, each LiDAR point in a point stream can query the corresponding continuous-time trajectory within its time instants. Additionally, a decentralized multi-LiDAR synchronization scheme is devised to combine points from separate LiDARs into a single point cloud without the need for primary LiDAR assignment. With the detailed derivation of the analytic Jacobians for continuous-time LiDAR observation, the proposed method integrates synchronization, continuous-time estimation, and voxel map management within a Kalman filter framework, which can achieve real-time state estimation with only a few linear iterations. The effectiveness of the proposed method is demonstrated through various scenarios, including public datasets and real-world autonomous driving experiments. The results demonstrate that the proposed CT-MLO can achieve high-accuracy continuous-time state estimations in real-time and is demonstratively competitive compared to other State-of-the-Art (SOTA) methods. Hongming Shen, Zhenyu Wu 0001, Qiyang Lyu, Huiqin Zhou, Yeqing Zhu |
ICARCV | 1 |
| 2024 | PLP-SLAM: Point-Line-Plane Simultaneous Localization and MappingabstractFor indoor environments, prior point-based visual SLAM cannot be processed in real time under low texture and illumination. To address this issue, this work proposes PLP-SLAM (Point-Line-Plane-SLAM) with RGB-D camera. Firstly, point and line features are detected in RGB images. For line features, establish length suppression and near line merge strategy to improve the line extraction quality. Secondly, plane features are extracted based on agglomerative hierarchical clustering method in point cloud obtained by RGB-D camera. Point clouds are divided into several nodes, unlike prior methods spend a lot of time to estimate the normal vector for each individual point, this work assumes that points within each node sharing the same plane normal vector, which can significantly improve the computational efficiency. Thirdly, sparse maps including points, lines and planes are established, meanwhile the scenes are reconstructed by creating the dense maps to show plan features directly. Finally, the performance of proposed method is compared against the state-of-the-art SLAM on public datasets to evaluate the pose estimation. All modules are run in real-time on a CPU, experiments clarify that PLP-SLAM can significantly enhance the robustness of 6DoF pose of the camera and simultaneously creating more detailed maps of the environment. Yeqing Zhu, Liangyu Zhao, Qingjie Zhao, Zhenyu Wu 0001, Hongming Shen, Danwei Wang |
ICARCV | 5 |
| 2024 | MM4MM: Map Matching Framework for Multi-Session Mapping in Ambiguous and Perceptually-Degraded EnvironmentsabstractMulti-session mapping serves as the pre-requisite for autonomous robots to fulfill various long-term tasks (e.g., map updating, navigation, collaboration). However, it is challenging to implement multi-session mapping in enclosed or partially enclosed ambiguous environments (e.g., long corridors, industrial warehouses). Existing solutions either depend heavily on the matching of elementary geometric features (e.g., points, lines, and planes), which tends to fail in environments with ambiguous geometric features; or depend on the given guess of the initial transformation matrix of multiple single-session maps, which is not always obtainable and accurate enough. The ambient magnetic field has exhibited ubiquity and high distinctiveness at different location, which makes it suitable for estimating the initial transformation matrix. Thus, this paper proposes a novel probabilistic magnetic-aware Map Matching framework for Multi-session Mapping, namely MM4MM, to estimate the relative transformation of multiple single-session maps and to build the globally consistent maps in ambiguous and perceptually-degraded environments. The key novelties of this work are the designing of the hierarchical probabilistic map matching framework and the Particle Swarm Optimization strategy to associate the magnetic data of multiple sessions. Evaluations on both simulated and real world experiments demonstrate the greatly improved utility, accuracy, and robustness of multi-session mapping over the comparative methods. Zhenyu Wu 0001, Yufeng Yue, Jun Zhang 0042, Hongming Shen, Danwei Wang |
ICRA | 6 |
| 2024 | S-GPR: Sliding Gaussian Process Regression-based Magnetic Mapping and Evaluation of Different Magnetic Mapping MethodsabstractThe localization of autonomous robots in modern enclosed or semi-enclosed environments, such as office/hotel/hospital, supermarket, and indoor car park environments where GPS signals are severely challenged, remains a bottleneck for the deployment of fully autonomous mobile systems. Existing infrastructure-based (e.g., QR codes, RFID) localization methods are troubled by high maintenance cost and inflexibility issues, while onboard sensors-based solutions (e.g., LiDAR/camera-based) suffer from the ambiguous geometric features and view obstructions from crowded dynamic obstacles (e.g., pedestrians). Magnetic field (MF)-based localization has been gradually utilized in recent years due to its independence from positioning infrastructures and geometric features, thus making it ideal for applications such as service robots and security robots. Magnetic map building serves as the basis and prerequisite component for MF-based localization tasks. The well-acknowledged Gaussian Process Regression (GPR) method can be implemented to build magnetic maps but with heavy computational burdens. Thus in this paper, we propose an efficient and accurate magnetic mapping system based on a novel Sliding-GPR (i.e., S-GPR) method, and evaluate different magnetic mapping methods. A unique region-of-interest (ROI) selection technique and a down/up-sampling method are proposed for the S-GPR to dramatically decrease the computational time while maintaining the mapping accuracy. Extensive experiments in a high-fidelity simulated warehouse and real-world car park environments show that our proposed S-GPR mapping method has exhibited the highest accuracy and relatively low computational time compared with the SOTA magnetic mapping methods. Qiyang Lyu, Zhenyu Wu 0001, Hongming Shen, Jun Zhang 0042, Huiqin Zhou, Danwei Wang |
IECON | 3 |
| 2024 | IDF-MFL: Infrastructure-free and Drift-free Magnetic Field Localization for Mobile RobotabstractIn recent years, infrastructure-based localization methods have achieved significant progress thanks to their reliable and drift-free localization capability. However, the preinstalled infrastructures suffer from inflexibilities and high maintenance costs. This poses an interesting problem of how to develop a drift-free localization system without using the preinstalled infrastructures. In this paper, an infrastructure-free and drift-free localization system is proposed using the ambient magnetic field (MF) information, namely IDF-MFL. IDF-MFL is infrastructure-free thanks to the high distinctiveness of the ambient MF information produced by inherent ferromagnetic objects in the environment, such as steel and reinforced concrete structures of buildings, and underground pipelines. The MF-based localization problem is defined as a stochastic optimization problem with the consideration of the non-Gaussian heavy-tailed noise introduced by MF measurement outliers (caused by dynamic ferromagnetic objects), and an outlier-robust state estimation algorithm is derived to find the optimal distribution of robot state that makes the expectation of MF matching cost achieves its lower bound. The proposed method is evaluated in multiple scenarios1, including experiments on high-fidelity simulation, and real-world environments. The results demonstrate that the proposed method can achieve high-accuracy, reliable, and real-time localization without any pre-installed infrastructures. Hongming Shen, Zhenyu Wu 0001, Qiyang Lyu, Huiqin Zhou, Danwei Wang |
IROS | 1 |
| 2024 | Towards Kbps-level Vehicle Teleoperation via Persistent-Transient Environment ModellingabstractTraditional teleoperation technologies based on video streaming are facing several challenges in practical applications, including limited bandwidth, constrained spatial awareness, and sensitivity to illumination. Existing studies have not adequately addressed these issues. This paper presents a novel non-video based teleoperation framework for autonomous vehicles operating in bandwidth-limited environments. To reduce the amount of data being transmitted, a persistent-transient environment model is proposed for telepresence. Initially, a digital twin of the environment is preconstructed, containing only persistent environmental information. Subsequently, transient information captured by onboard sensors, such as vehicle state and dynamic objects, necessitate real-time transmission. Based on this model, a 3D virtual scene is rendered in front of the teleoperator, offering any desired virtual viewpoint to enhance spatial awareness. This telepresence model only requires real-time transmission of minimal data, i.e., vehicle state and detected objects, and remains unaffected by illumination conditions, enabling teleoperation even in applications with Kbps-level bandwidth constraints. Experimental results showcase the substantial potential of the proposed framework in bandwidth-limited settings. Dogan Kircali, Guoyi Chi, Hongming Shen, Yuanzhe Wang, Danwei Wang |
IROS | 6 |
| 2024 | MRS-Net: Brain tumour segmentation network based on feature fusion and attention mechanismabstractAbstract Accurate segmentation of brain tumor magnetic resonance imaging (MRI) is crucial for treatment planning. Addressing the challenges of complex tumor structures and inadequate cross‐channel information utilization in Unet‐based segmentation, this paper proposes the multi‐scale residual brain tumor MRI segmentation network (MRS‐Net) incorporating an attention mechanism to enhance segmentation accuracy. First, the double residual feature fusion module is utilized to enhance the fusion of feature information between different levels. Second, the Atrous Spatial Pyramid Pooling is introduced as a bridging module of the network to capture the features at different scales of the image, so as to enhance the extraction capability of the network for detailed features. Finally, the inverted residual coordinate attention module replaces the direct splicing in Unet to fuse the large feature information at each level and scale, thus enhancing the model's ability to recognize the spatial location information of brain tumors. The Dice coefficients, positive predictive values (PPVs), sensitivities (Sensitivity) and Hausdorff distance (HD), which are the four evaluation indexes, reach 84.54%, 87.43%, 88.37% and 2.248, respectively, which are improved by 1.85%, 2.11%, 2.88% and 6.0%, respectively, compared with Unet. The experimental results show that MRS‐Net achieves better brain tumor image segmentation. Xiaoyan Shen, Jiakai Zhang, Hongming Shen |
IET Image Process. | 7 |