Zhijian Qiao

dblp:268/5386 · DBLP profile ↗
← Back
14ranked-venue papers
5as first author
13since 2021 · last 2025
0000-0001-6639-1110ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 3 first-author · 7 since 2021Systems, architecture and hardware · 8 · 3 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
YearPublicationVenuePosition
2025 SLABIM: A SLAM-BIM Coupled Dataset in HKUST Main Building
abstract
Existing indoor SLAM datasets primarily focus on robot sensing, often lacking building architectures. To address this gap, we design and construct the first dataset to couple the SLAM and BIM, named SLABIM. This dataset provides BIM and SLAM -oriented sensor data, both modeling a university building at HKUST. The as-designed BIM is decomposed and converted for ease of use. We employ a multi-sensor suite for multi-session data collection and mapping to obtain the as-built model. All the related data are timestamped and organized, enabling users to deploy and test effectively. Furthermore, we deploy advanced methods and report the experimental results on three tasks: registration, localization and semantic mapping, demonstrating the effectiveness and practicality of SLAB 1M. We make our dataset open-source at https://github.com/HKUST-Aerial-Robotics/SLABIM.
Haoming Huang, Zhijian Qiao, Zehuan Yu, Chuhao Liu, Shaojie Shen, Fumin Zhang 0001, Huan Yin
ICRA2
2025 Speak the Same Language: Global LiDAR Registration on BIM Using Pose Hough Transform
abstract
Light detection and ranging (LiDAR) point clouds and building information modeling (BIM) represent two distinct data modalities in the fields of robot perception and construction. These modalities originate from different sources and are associated with unique reference frames. The primary goal of this study is to align these modalities within a shared reference frame using a global registration approach, effectively enabling them to “speak the same language”. To achieve this, we propose a cross-modality registration method, spanning from the front end to the back end. At the front end, we extract triangle descriptors by identifying walls and intersected corners, enabling the matching of corner triplets with a complexity independent of the BIM’s size. For the back-end transformation estimation, we utilize the Hough transform to map the matched triplets to the transformation space and introduce a hierarchical voting mechanism to hypothesize multiple pose candidates. The final transformation is then verified using our designed occupancy-aware scoring method. To assess the effectiveness of our approach, we conducted real-world multi-session experiments in a large-scale university building, employing two different types of LiDAR sensors. We make the collected datasets and codes publicly available to benefit the community. Note to Practitioners—Our proposed registration method leverages walls and corners as shared features between LiDAR and BIM data, making it particularly well-suited for scenarios with well-defined structural layouts. Accumulating a larger LiDAR submap provides richer structural information, which further aids in achieving accurate alignment. To optimize computational efficiency, we recommend constructing the descriptor database offline and loading it during runtime, enabling a theoretical retrieval complexity of$O(1)$. Despite its advantages, our approach has certain limitations. First, it primarily focuses on planar structures, which limits its effectiveness in utilizing non-planar features. Second, the method may underperform in cases where significant deviations exist between the as-designed BIM and as-is LiDAR data. Lastly, in ambiguous scenarios, such as long corridors or similar layouts within the same or across different floors, our method may struggle to verify the correct transformation among candidates. To address these challenges, incorporating additional information, particularly semantic cues such as floor numbers, room numbers, and room types, could enhance its robustness and reliability.
Zhijian Qiao, Haoming Huang, Chuhao Liu, Zehuan Yu, Shaojie Shen, Fumin Zhang 0001, Huan Yin
IEEE Trans Autom. Sci. Eng.1
2025 G3Reg: Pyramid Graph-Based Global Registration Using Gaussian Ellipsoid Model
abstract
This study introduces a novel framework, G3Reg, for fast and robust global registration of LiDAR point clouds. In contrast to conventional complex keypoints and descriptors, we extract fundamental geometric primitives, including planes, clusters, and lines (PCL) from the raw point cloud to obtain low-level semantic segments. Each segment is represented as a unified Gaussian Ellipsoid Model (GEM), using a probability ellipsoid to ensure the ground truth centers are encompassed with a certain degree of probability. Utilizing these GEMs, we present a distrust-and-verify scheme based on a Pyramid Compatibility Graph for Global Registration (PAGOR). Specifically, we establish an upper bound, which can be traversed based on the confidence level for compatibility testing to construct the pyramid graph. Then, we solve multiple maximum cliques (MAC) for each level of the pyramid graph, thus generating the corresponding transformation candidates. In the verification phase, we adopt a precise and efficient metric for point cloud alignment quality, founded on geometric primitives, to identify the optimal candidate. The algorithm’s performance is validated on three publicly available datasets and a self-collected multi-session dataset. Parameter settings remained unchanged during the experiment evaluations. The results exhibit superior robustness and real-time performance of the G3Reg framework compared to state-of-the-art methods. Furthermore, we demonstrate the potential for integrating individual GEM and PAGOR components into other registration frameworks to enhance their efficacy.Note to Practitioners—Our proposed method aims to perform global registration for outdoor LiDAR point clouds. Our methodology, which extracts point cloud segments and utilizes their centers for registration, differs from conventional approaches that rely on keypoints and descriptors. We further propose GEM to model the uncertainty of the centers and embed it into our distrust-and-verify framework. In theory, our method can be applied to any registration task that involves primitives representable as sets of Gaussians or points. Additionally, practitioners should consider the following to enhance applicability. First, practitioners can fine-tune the parameters of the segmentation algorithm to generate more repeatable segmentation results. Second, although our default setting uses four compatibility test thresholds, fewer may suffice, especially when translations between point clouds are minor. Finally, for geometrically uninformative segments such as vegetation, consider extracting descriptors within these segments to increase correspondences.
Zhijian Qiao, Zehuan Yu, Binqian Jiang, Huan Yin, Shaojie Shen
IEEE Trans Autom. Sci. Eng.1
2025 SG-Reg: Generalizable and Efficient Scene Graph Registration
abstract
This paper addresses the challenges of registering two rigid semantic scene graphs, an essential capability when an autonomous agent needs to register its map against a remote agent, or against a prior map. The hand-crafted descriptors in classical semantic-aided registration, or the ground-truth annotation reliance in learning-based scene graph registration, impede their application in practical real-world environments. To address the challenges, we design a scene graph network to encode multiple modalities of semantic nodes: open-set semantic feature, local topology with spatial awareness, and shape feature. These modalities are fused to create compact semantic node features. The matching layers then search for correspondences in a coarse-to-fine manner. In the back-end, we employ a robust pose estimator to decide transformation according to the correspondences. We manage to maintain a sparse and hierarchical scene representation. Our approach demands fewer GPU resources and fewer communication bandwidth in multi-agent tasks. Moreover, we design a new data generation approach using vision foundation models and a semantic mapping module to reconstruct semantic scene graphs. It differs significantly from previous works, which rely on ground-truth semantic annotations to generate data. We validate our method in a two-agent SLAM benchmark. It significantly outperforms the hand-crafted baseline in terms of registration success rate. Compared to visual loop closure networks, our method achieves a slightly higher registration recall while requiring only 52 KB of communication bandwidth for each query frame. Code available at:http://github.com/HKUST-Aerial-Robotics/SG-Reg
Chuhao Liu, Zhijian Qiao, Jieqi Shi, Ke Wang 0058, Peize Liu, Shaojie Shen
IEEE Trans. Robotics2
2025 SLIM: Scalable and Lightweight LiDAR Mapping in Urban Environments
abstract
Light detection and ranging (LiDAR) point cloud maps are extensively utilized on roads for robot navigation due to their high consistency. However, dense point clouds face challenges of high memory consumption and reduced maintainability for long-term operations. In this study, we introduce scalable and lightweight LiDAR mapping (SLIM), a scalable and lightweight mapping system for long-term LiDAR mapping in urban environments. The system begins by parameterizing structural point clouds into lines and planes. These lightweight and structural representations meet the requirements of map merging, pose graph optimization, and bundle adjustment, ensuring incremental management and local consistency. For long-term operations, a map-centric nonlinear factor recovery method is designed to sparsify poses while preserving mapping accuracy. We validate the SLIM system with multisession real-world LiDAR data from classical LiDAR mapping datasets, including KITTI, NCLT, HeLiPR, and M2DGR. The experiments demonstrate its capabilities in mapping accuracy, lightweightness, and scalability. Map reuse is also verified through map-based robot localization. Finally, with multisession LiDAR data, the SLIM system provides a globally consistent map with low memory consumption ($\sim$130 KB/km on KITTI).
Zehuan Yu, Zhijian Qiao, Huan Yin, Shaojie Shen
IEEE Trans. Robotics2
2024 Less is More: Physical-Enhanced Radar-Inertial Odometry
abstract
Radar offers the advantage of providing additional physical properties related to observed objects. In this study, we design a physical-enhanced radar-inertial odometry system that capitalizes on the Doppler velocities and radar cross-section information. The filter for static radar points, correspondence estimation, and residual functions are all strengthened by integrating the physical properties. We conduct experiments on both public datasets and our self-collected data, with different mobile platforms and sensor types. Our quantitative results demonstrate that the proposed radar-inertial odometry system outperforms alternative methods using the physical-enhanced components. Our findings also reveal that using the physical properties results in fewer radar points for odometry estimation, but the performance is still guaranteed and even improved, thus aligning with the "less is more" principle.
Qiucan Huang, Zhijian Qiao, Shaojie Shen, Huan Yin
ICRA3
2024 VCounselor: a psychological intervention chat agent based on a knowledge-enhanced large language model
Hanzhong Zhang, Zhijian Qiao, Jibin Yin
Multim. Syst.2
2023 SeasonDepth: Cross-Season Monocular Depth Prediction Dataset and Benchmark Under Multiple Environments
abstract
Different environments pose a great challenge to the outdoor robust visual perception for long-term autonomous driving, and the generalization of learning-based algorithms on different environments is still an open problem. Although monocular depth prediction has been well studied recently, few works focus on the robustness of learning-based depth prediction across different environments, e.g. changing illumination and seasons, owing to the lack of such a multi-environment real-world dataset and benchmark. To this end, the cross-season monocular depth prediction dataset and benchmark, SeasonDepth, is introduced to benchmark the depth estimation performance under different environments. We investigate several state-of-the-art representative open-source supervised and self-supervised depth prediction methods using newly-formulated metrics. Through extensive experimental evaluation on the proposed dataset and cross-dataset evaluation with current autonomous driving datasets, the performance and robustness against the influence of multiple environments are analyzed qualitatively and quantitatively. We show that long-term monocular depth prediction is still challenging and believe our work can boost further research on the long-term robustness and generalization for outdoor visual perception. The dataset is available on https://seasondepth.github.io.
Hanjiang Hu, Baoquan Yang, Zhijian Qiao, Shiqi Liu 0005, Zuxin Liu, Wenhao Ding, Ding Zhao, Hesheng Wang 0001
IROS3
2023 Online Monocular Lane Mapping Using Catmull-Rom Spline
abstract
In this study, we introduce an online monocular lane mapping approach that solely relies on a single camera and odometry for generating spline-based maps. Our proposed technique models the lane association process as an assignment issue utilizing a bipartite graph, and assigns weights to the edges by incorporating Chamfer distance, pose uncertainty, and lateral sequence consistency. Furthermore, we meticulously design control point initialization, spline parameterization, and optimization to progressively create, expand, and refine splines. In contrast to prior research that assessed performance using self-constructed datasets, our experiments are conducted on the openly accessible OpenLane dataset. The experimental outcomes reveal that our suggested approach enhances lane association and odometry precision, as well as overall lane map quality. We have open-sourced out code11https://github.com/HKUST-Aerial-Robotics/MonoLaneMapping for this project.
Zhijian Qiao, Zehuan Yu, Huan Yin, Shaojie Shen
IROS1
2023 Pyramid Semantic Graph-Based Global Point Cloud Registration with Low Overlap
abstract
Global point cloud registration is essential in many robotics tasks like loop closing and relocalization. Unfortunately, the registration often suffers from the low overlap between point clouds, a frequent occurrence in practical applications due to occlusion and viewpoint change. In this paper, we propose a graph-theoretic framework to address the problem of global point cloud registration with low overlap. To this end, we construct a consistency graph to facilitate robust data association and employ graduated non-convexity (GNC) for reliable pose estimation, following the state-of-the-art (SoTA) methods. Unlike previous approaches, we use semantic cues to scale down the dense point clouds, thus reducing the problem size. Moreover, we address the ambiguity arising from the consistency threshold by constructing a pyramid graph with multi-level consistency thresholds. Then we propose a cascaded gradient ascend method to solve the resulting densest clique problem and obtain multiple pose candidates for every consistency threshold. Finally, fast geometric verification is employed to select the optimal estimation from multiple pose candidates. Our experiments, conducted on a self-collected indoor dataset and the public KITTI dataset, demonstrate that our method achieves the highest success rate despite the low overlap of point clouds and low semantic quality. We have open-sourced our code1for this project.
Zhijian Qiao, Zehuan Yu, Huan Yin, Shaojie Shen
IROS1
2023 Multi-Session, Localization-Oriented and Lightweight LiDAR Mapping Using Semantic Lines and Planes
abstract
In this paper, we present a centralized framework for multi-session LiDAR mapping in urban environments, by utilizing lightweight line and plane map representations instead of widely used point clouds. The proposed framework achieves consistent mapping in a coarse-to-fine manner. Global place recognition is achieved by associating lines and planes on the Grassmannian manifold, followed by an outlier rejection-aided pose graph optimization for map merging. Then a novel bundle adjustment is also designed to improve the local consistency of lines and planes. In the experimental section, both public and self-collected datasets are used to demonstrate efficiency and effectiveness. Extensive results validate that our LiDAR mapping framework could merge multi-session maps globally, optimize maps incrementally, and is applicable for lightweight robot localization.
Zehuan Yu, Zhijian Qiao, Liuyang Qiu, Huan Yin, Shaojie Shen
IROS2
2021 A Registration-aided Domain Adaptation Network for 3D Point Cloud Based Place Recognition
abstract
In the field of large-scale SLAM for autonomous driving and mobile robotics, 3D point cloud based place recognition has aroused significant research interest due to its robustness to changing environments with drastic daytime and weather variance. However, it is time-consuming and effort-costly to obtain high-quality point cloud data for place recognition model training and ground truth for registration in the real world. To this end, a novel registration-aided 3D domain adaptation network for point cloud based place recognition is proposed. A structure-aware registration network is introduced to help to learn features with geometric information and a 6-DoFs pose between two point clouds with partial overlap can be estimated. The model is trained through a synthetic virtual LiDAR dataset through GTA-V with diverse weather and daytime conditions and domain adaptation is implemented to the real-world domain by aligning the global features. Our results outperform state-of-the-art 3D place recognition baselines or achieve comparable on the real-world Oxford RobotCar dataset with the visualization of registration on the virtual dataset.
Zhijian Qiao, Hanjiang Hu, Weiang Shi, Zhe Liu 0022, Hesheng Wang 0001
IROS1
2021 DASGIL: Domain Adaptation for Semantic and Geometric-Aware Image-Based Localization
abstract
Long-Term visual localization under changing environments is a challenging problem in autonomous driving and mobile robotics due to season, illumination variance, etc. Image retrieval for localization is an efficient and effective solution to the problem. In this paper, we propose a novel multi-task architecture to fuse the geometric and semantic information into the multi-scale latent embedding representation for visual place recognition. To use the high-quality ground truths without any human effort, the effective multi-scale feature discriminator is proposed for adversarial training to achieve the domain adaptation from synthetic virtual KITTI dataset to real-world KITTI dataset. The proposed approach is validated on the Extended CMU-Seasons dataset and Oxford RobotCar dataset through a series of crucial comparison experiments, where our performance outperforms state-of-the-art baselines for retrieval-based localization and large-scale place recognition under the challenging environment.
Hanjiang Hu, Zhijian Qiao, Ming Cheng 0004, Zhe Liu 0022, Hesheng Wang 0001
IEEE Trans. Image Process.2
2020 End-to-End 3D Point Cloud Learning for Registration Task Using Virtual Correspondences
abstract
3D Point cloud registration is still a very challenging topic due to the difficulty in finding the rigid transformation between two point clouds with partial correspondences, and it's even harder in the absence of any initial estimation information. In this paper, we present an end-to-end deep-learning based approach to resolve the point cloud registration problem. Firstly, the revised LPD-Net is introduced to extract features and aggregate them with the graph network. Secondly, the self-attention mechanism is utilized to enhance the structure information in the point cloud and the cross-attention mechanism is designed to enhance the corresponding information between the two input point clouds. Based on which, the virtual corresponding points can be generated by a soft pointer based method, and finally, the point cloud registration problem can be solved by implementing the SVD method. Comparison results in ModelNet40 dataset validate that the proposed approach reaches the state-of-the-art in point cloud registration tasks and experiment resutls in KITTI dataset validate the effectiveness of the proposed approach in real applications.
Huanshu Wei, Zhijian Qiao, Zhe Liu 0022, Chuanzhe Suo, Peng Yin 0001, Yueling Shen, Haoang Li, Hesheng Wang 0001
IROS2