EDBT 2026 Demo / reviewers in the wild / expert
Hengwang Zhao
dblp:282/8738
· DBLP profile ↗
6ranked-venue papers
3as first author
5since 2021 · last 2025
0000-0003-3556-8029ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021Systems, architecture and hardware · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CrossGLoc: Cross-Modal Global Localization Leveraging Pretrained Diffusion Models and Semantic Cues for Intelligent VehiclesabstractCross-modal global localization matches visual information with pre-built LiDAR maps, which has attracted more and more attention for its low cost and potential robustness. However, the inherent modality difference between images and point clouds makes it challenging. This paper proposes a novel cross-modal global localization system, named CrossGLoc, which leverages pre-trained diffusion models and semantic cues to address this challenge. The main idea is leveraging the semantic cues shared between different modalities to bridge the modality gap, and utilizing pre-trained diffusion models to extract modality-consistent high-dimensional features guided by these semantic cues. To achieve this, ControlNet is used to generate intermediate feature maps from semantic images and semantic map projections, and a semantic categories-based feature aggregation algorithm is proposed to aggregate these feature maps into global descriptors. Furthermore, a semantic edge key points-based pose estimation algorithm is proposed to estimate the pose of retrieved image and point cloud pairs. Extensive experiments on the KITTI dataset, the KITTI360 dataset and the self-collected dataset demonstrate that the proposed method achieves state-of-the-art performance in cross-modal global localization. Hengwang Zhao, Qiyuan Shen, Hanyang Zhuang, Tong Qin 0001, Ming Yang 0002 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2024 | MOSFormer: A Transformer-based Multi-Modal Fusion Network for Moving Object Segmentationabstract3D moving object segmentation (MOS) is vital for autonomous systems, providing essential information for downstream tasks like mapping and localization. However, current MOS methods face challenges due to the limitation of existing datasets, which are sparse in moving objects and limited in scene diversity. Meanwhile, the prevalent methods are projection-based, struggling with the challenge of blurred boundaries. To tackle the dataset issue, we introduce a nuScenes-based MOS dataset, which provides richer scenes and more dynamic instances. To alleviate the boundary blur-ring issue and further improve accuracy and generalizability, we propose a dual-branch multimodal fusion MOS network, MOSFormer. The Transformer structure is incorporated to extract spatio-temporal information better, while image semantic information is utilized to refine the boundaries of moving objects. Finally, experiments on two datasets show that our method achieves state-of-the-art performance, and a mapping experiment with our method confirms its effectiveness in downstream tasks such as mapping and localization. Zike Cheng, Hengwang Zhao, Qiyuan Shen, Weihao Yan 0001, Ming Yang 0002 |
IROS | 2 |
| 2024 | Cross-Modal Visual Relocalization in Prior LiDAR Maps Utilizing Intensity TexturesabstractCross-modal localization has drawn increasing attention in recent years, while the visual relocalization in prior LiDAR maps is less studied. Related methods usually suffer from inconsistency between the 2D texture and 3D geometry, neglecting the intensity features in the LiDAR point cloud. In this paper, we propose a cross-modal visual relocalization system in prior LiDAR maps utilizing intensity textures, which consists of three main modules: map projection, coarse retrieval, and fine relocalization. In the map projection module, we construct the database of intensity channel map images leveraging the dense characteristic of panoramic projection. The coarse retrieval module retrieves the top-K most similar map images to the query image from the database, and retains the top-K’ results by covisibility clustering. The fine relocalization module applies a two-stage 2D-3D association and a covisibility inlier selection method to obtain robust correspondences for 6DoF pose estimation. The experimental results on our self-collected datasets demonstrate the effectiveness in both place recognition and pose estimation tasks. Qiyuan Shen, Hengwang Zhao, Weihao Yan 0001, Tong Qin 0001, Ming Yang 0002 |
IROS | 2 |
| 2023 | Cross-Modal Monocular Localization in Prior LiDAR Maps Utilizing Semantic ConsistencyabstractVisual localization for mobile robots and intelligent vehicles in prior LiDAR maps can achieve high accuracy and low cost. However, algorithms for finding the cross-modal correspondences between images and LiDAR map points are not yet stable. In this paper, we propose a monocular visual localization system in prior LiDAR maps, which is based on the cross-modal registration to optimize the camera pose. To align the point clouds from vision and LiDAR map, a point-to-plane Iterative Closest Point algorithm utilizing semantic consistency is designed, and a decoupling optimization strategy is proposed to compute the affine transformation for the monocular scale ambiguity. Experiments on KITTI dataset show that utilizing the semantic consistency and geometric information of the map makes our system competitive with other methods. On the self-collected dataset, experiments on different light intensities demonstrate the robustness of the system in long-term localization tasks, and the ablation study demonstrates the effectiveness of the proposed algorithms. Hengwang Zhao, Xuanlai Tang, Ming Yang 0002 |
ICRA | 2 |
| 2023 | Cy-CNN: cylinder convolution based rotation-invariant neural network for point cloud registration
Hengwang Zhao, Zhidong Liang, Yuesheng He, Ming Yang 0002 |
Sci. China Inf. Sci. | 1 |
| 2020 | Time-of-Flight Camera Based Indoor Parking Localization Leveraging Manhattan World RegulationabstractLocalization is a key problem for autonomous driving in indoor parking. There have been some previously proposed methods based on UWB, LiDAR, fisheye cameras, etc. However, most of these methods have some drawbacks such as high cost or dependency on light conditions. To address these challenges, this paper proposes a novel Time-of-Flight (ToF) camera based mapping and localization system for indoor parking lots leveraging Manhattan World Regulation. ToF cameras are low-cost and can actively generate dense point clouds of the environment without external light sources. To overcome the shortcoming of ToF camera small field of view, the proposed system utilizes the structural information of the ceiling of indoor parking lots and Manhattan World Regulation. We track the surface normals on the unit sphere for drift-free rotation estimation. Based on this drift-free rotation, we can effectively calculate 6-DOF pose with decoupled rotation and translation estimation during mapping or global localization. This new system runs in real-time on limited computation resources and is demonstrated on two different challenging indoor parking, achieving real-time performance at 10 Hz and localization error less than 0.1 meter. Hengwang Zhao, Ming Yang 0002, Yuesheng He |
IV | 1 |