Jianhao Jiao

dblp:207/2281 · DBLP profile ↗
← Back
35ranked-venue papers
10as first author
26since 2021 · last 2026
0000-0001-7372-266XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 22 · 7 first-author · 16 since 2021Systems, architecture and hardware · 20 · 6 first-author · 16 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 2 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-authorComputer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Short Paper: The Starlink Robot: A Platform and Dataset for Mobile Satellite Communication
abstract
The integration of satellite communication into mobile devices represents a paradigm shift in connectivity, yet the performance characteristics under motion and environmental occlusion remain poorly understood. We present the Starlink Robot, the first mobile robotic platform equipped with Starlink satellite internet, comprehensive sensor suite including upward-facing camera, LiDAR, and IMU, designed to systematically study satellite communication performance during movement. Our multi-modal dataset captures synchronized communication metrics, motion dynamics, sky visibility, and 3D environmental context across diverse scenarios including steady-state motion, variable speeds, and different occlusion conditions. This platform and dataset enable researchers to develop motion-aware communication protocols, predict connectivity disruptions, and optimize satellite communication for emerging mobile applications from smartphones to autonomous vehicles. In this work, we use LEOViz for real-time data collection and visualization. The project is available at https://starlinkrobot.github.io.
Boyi Liu 0003, Qianyi Zhang, Qiang Yang 0018, Jianhao Jiao, Jagmohan Chauhan, Dimitrios Kanoulas
SenSys4
2026 ROVER: Robust Loop Closure Verification With Trajectory Prior in Repetitive Environments
abstract
Loop closure detection is important for simultaneous localization and mapping (SLAM), which associates current observations with historical keyframes, achieving drift correction and global relocalization. However, a falsely detected loop can be fatal, and this is especially difficult in repetitive environments where appearance-based features fail due to the high similarity. Therefore, verifying a loop closure is a critical step to avoid false-positive detections. Existing works in loop closure verification predominantly focus on learning invariant appearance features, neglecting the prior knowledge of the robot’s spatial-temporal motion cue, i.e., trajectory. In this article, we propose ROVER, a loop closure verification method that leverages the historical trajectory as a prior constraint to reject false loops in challenging repetitive environments. For each loop candidate, it is first used to estimate the robot trajectory with pose-graph optimization. This trajectory is then submitted to a scoring scheme that assesses its compliance with the trajectory without the loop, which we refer to as the trajectory prior constraint (TPC), to determine if the loop candidate should be accepted. Benchmark comparisons and real-world experiments demonstrate the effectiveness of the proposed method. Furthermore, we integrate ROVER into state-of-the-art SLAM systems to verify its robustness and efficiency.
Jingwen Yu, Jianhao Jiao, Anjun Hu, Zhonghang Liu, Jiankun Wang 0001, Ping Tan 0002, Hong Zhang 0013
IEEE Trans Autom. Sci. Eng.3
2026 Co-Attention Guided Multimodal Fusion for Robust 6-D Pose Estimation in Occluded Scenarios
Hui Zhang 0074, Jianhao Jiao, Fengqin He
IEEE Trans Autom. Sci. Eng.3
2026 MobileROS: A Wireless-Native Robot Operating System for Mobile Robotics
abstract
The increasing deployment of mobile robots in dynamic outdoor environments necessitates robotic systems capable of maintaining reliability amidst fluctuating wireless connectivity. While the Robot Operating System (ROS) has established itself as the de facto standard for such networked robotics, its abstraction of communication as an opaque, besteffort utility creates a critical bottleneck: it fails to leverage physical layer (PHY) information, resulting in degraded performance and unreliable execution in fluctuating networks. To address this, this paper presents MobileROS, a wireless-native robot operating system that transforms wireless communication from an external service into a core system resource. Grounded in the Symbiotic Paradigm, MobileROS establishes a bidirectional exchange where network conditions inform robotic decisions and mission requirements guide network resource allocation. Based on service mesh principles and domain-driven design, our architecture implements a Hub-Engines-Cells (HEC) model. It features a central Hub for global optimization, three specialized engines (the Radio Information Engine, the Cross Domain Engine, and the Physical Adaptive Engine) for crosslayer intelligence, and distributed Cells as functional units. A key mechanism, Application-Driven Bidirectional Dynamic Slicing, allows robots to actively reconfigure network resources based on semantic urgency, transforming the robot from a passive observer into an active network controller. We systematically evaluate MobileROS across three cities (London, Hong Kong, and Shenzhen) in five scenarios: distributed visual SLAM, cross-domain LiDAR perception, V2X autonomous driving, hybrid multi-robot collaboration against WebRTC baselines, and partition recovery validating CAP-theorem-aware failsafe mechanisms. Results demonstrate that MobileROS maintains significantly more stable performance than standard ROS in mobile wireless deployments.We provide implementation details athttps://github.com/MobileROS.
Boyi Liu 0003, Qianyi Zhang, Yongguang Lu, Jianhao Jiao, Jagmohan Chauhan, Wen Wu 0003, Jun Zhang 0004, Dimitrios Kanoulas
IEEE Trans. Robotics4
2025 Watch Your STEPP: Semantic Traversability Estimation Using Pose Projected Features
abstract
Understanding the traversability of terrain is essential for autonomous robot navigation, particularly in unstructured environments such as natural landscapes. Although traditional methods, such as occupancy mapping, provide a basic framework, they often fail to account for the complex mobility capabilities of some platforms such as legged robots. In this work, we propose a method for estimating terrain traversability by learning from demonstrations of human walking. Our approach leverages dense, pixel-wise feature embeddings generated using the DINOv2 vision Transformer model, which are processed through an encoder-decoder MLP architecture to analyze terrain segments. The averaged feature vectors, extracted from the masked regions of interest, are used to train the model in a reconstruction-based framework. By minimizing reconstruction loss, the network distinguishes between familiar terrain with a low reconstruction error and unfamiliar or hazardous terrain with a higher reconstruction error. This approach facilitates the detection of anomalies, allowing a legged robot to navigate more effectively through challenging terrain. We run real-world experiments on the ANYmal legged robot both indoor and outdoor to prove our proposed method. The code is open-source, while video demonstrations can be found on our website: https://rpl-cs-ucl.github.io/STEPP/
Sebastian Aegidius, Denis Hadjivelichkov, Jianhao Jiao, Jonathan Embley-Riches, Dimitrios Kanoulas
ICRA3
2025 LoGS: Visual Localization via Gaussian Splatting with Fewer Training Images
abstract
Visual localization involves estimating a query image's 6-DoF (degrees of freedom) camera pose, which is a fundamental component in various computer vision and robotic tasks. This paper presents LoGS, a vision-based localization pipeline utilizing the 3D Gaussian Splatting (GS) technique as scene representation. This novel representation allows high-quality novel view synthesis. During the mapping phase, structure-from-motion (SfM) is applied first, followed by the generation of a GS map. During localization, the initial position is obtained through image retrieval, local feature matching coupled with a PnP solver, and then a high-precision pose is achieved through the analysis-bysynthesis manner on the GS map. Experimental results on four large-scale datasets demonstrate the proposed approach's SoTA accuracy in estimating camera poses and robustness under challenging few-shot conditions. Codes can be found at: https://github.com/RPL-CS-UCL/gs_localization.
Yuzhou Cheng, Jianhao Jiao, Yue Wang 0020, Dimitrios Kanoulas
ICRA2
2025 LiteVLoc: Map-Lite Visual Localization for Image Goal Navigation
abstract
This paper presents Lite VLoc, a hierarchical vi-sual localization framework that uses a lightweight topo-metric map to represent the environment. The method consists of three sequential modules that estimate camera poses in a coarse-to-fine manner. Unlike dense 3D mapping methods, LiteVLoc reduces storage by avoiding geometric reconstruction. It uses a learning-based feature matcher to establish dense correspondences between sparse keyframes and observations, and then refines poses with a geometric solver, enabling robustness to viewpoint changes. The system assumes depth sensors or stereo camera for deployment. A novel dataset for the map-free relocalization task is also introduced. Extensive experiments including localization and navigation in both simulated and real-world scenarios have validate the system's performance and demonstrated its precision and efficiency for large-scale deployment. Code and data will be made publicly available at the webpage:https://rpl-cs-ucl.github.io/LiteVLoc.
Jianhao Jiao, Jinhao He, Changkun Liu 0001, Sebastian Aegidius, Xiangcheng Hu, Tristan Braud, Dimitrios Kanoulas
ICRA1
2025 AIR-HLoc: Adaptive Retrieved Images Selection for Efficient Visual Localisation
abstract
State-of-the-art hierarchical localisation pipelines (HLoc) employ image retrieval (IR) to establish 2D-3D correspondences by selecting the top-k most similar images from a reference database. While increasing$k$improves localisation robustness, it also linearly increases computational cost and runtime, creating a significant bottleneck. This paper investigates the relationship between global and local descriptors, showing that greater similarity between the global descriptors of query and database images increases the proportion of feature matches. Low similarity queries significantly benefit from increasing k, while high similarity queries rapidly experience diminishing returns. Building on these observations, we propose an adaptive strategy that adjusts$k$based on the similarity between the query's global descriptor and those in the database, effectively mitigating the feature-matching bottleneck. Our approach reduces computational costs and processing time without sacrificing accuracy. Experiments on three indoor and outdoor datasets show that AIR-HLoc reduces feature matching time by up to 30% while preserving state-of-the-art accuracy. The results demonstrate that AIR-HLoc facilitates a latency-sensitive localisation system.
Changkun Liu 0001, Jianhao Jiao, Huajian Huang, Zhengyang Ma, Dimitrios Kanoulas, Tristan Braud
ICRA2
2025 From Satellite to Street: Semantic and Depth Information for Enhanced Geo-Localization
abstract
Accurate positioning is essential for autonomous driving, but localization using 2D maps is challenging due to the domain gap between perspective view and 2D map. While GNSS accuracy is often limited by atmospheric effects, multipath, and signal blockages. We propose a novel positioning method that combines perspective view images with satellite images retrieved based on rough GNSS positions to achieve precise three-degree-of-freedom (3-DoF) pose estimation. Our method leverages the Swin Transformer for satellite image processing and semantic completion for monocular image analysis. By extracting depth and semantic information from monocular images, we convert these to overhead projections, effectively bridging the gap between different viewpoints. This cross-view transformation allows for precise alignment of features from monocular images onto semantically enriched satellite images. Additionally, we integrate a robust global position estimator using the semantic information from satellite images to further enhance accuracy and robustness. The experimental results demonstrate that our method excels in various complex scenarios; we successfully improved the positioning accuracy within 1 m to 80.67% and the heading in 1° to 33.78%. However, longitudinal localization remains more challenging, with higher errors than lateral positioning.
Yilong Zhu, Jianhao Jiao, Hexiang Wei, Jin Wu 0002, Bohuan Xue, Shaojie Shen
IROS2
2025 Event-Driven Dynamic Scene Depth Completion
abstract
Depth completion in dynamic scenes poses significant challenges due to rapid ego-motion and object motion, which can severely degrade the quality of input modalities such as RGB images and LiDAR measurements. Conventional RGB-D sensors often struggle to align precisely and capture reliable depth under such conditions. In contrast, event cameras with their high temporal resolution and sensitivity to motion at the pixel level provide complementary cues that are beneficial in dynamic environments. To this end, we propose EventDC, the first event-driven depth completion framework. It consists of two key components: Event-Modulated Alignment (EMA) and Local Depth Filtering (LDF). Both modules adaptively learn the two fundamental components of convolution operations: offsets and weights conditioned on motion-sensitive event streams. In the encoder, EMA leverages events to modulate the sampling positions of RGB-D features to achieve pixel redistribution for improved alignment and fusion. In the decoder, LDF refines depth estimations around moving objects by learning motion-aware masks from events. Additionally, EventDC incorporates two loss terms to further benefit global alignment and enhance local depth recovery. Moreover, we establish the first benchmark for event-based depth completion comprising one real-world and two synthetic datasets to facilitate future research. Extensive experiments on this benchmark demonstrate the superiority of our EventDC. [Project page](https://yanzq95.github.io/projectpage/EventDC/index.html).
Zhiqiang Yan 0001, Jianhao Jiao, Zhengxue Wang, Gim Hee Lee
NeurIPS2
2025 Real-Time Metric-Semantic Mapping for Autonomous Navigation in Outdoor Environments
abstract
The creation of a metric-semantic map, which encodes human-prior knowledge, represents a high-level abstraction of environments. However, constructing such a map poses challenges related to the fusion of multi-modal sensor data, the attainment of real-time mapping performance, and the preservation of structural and semantic information consistency. In this paper, we introduce an online metric-semantic mapping system that utilizes LiDAR-Visual-Inertial sensing to generate a global metric-semantic mesh map of large-scale outdoor environments. Leveraging GPU acceleration, our mapping process achieves exceptional speed, with frame processing taking less than$7ms$, regardless of scenario scale. Furthermore, we seamlessly integrate the resultant map into a real-world navigation system, enabling metric-semantic-based terrain assessment and autonomous point-to-point navigation within a campus environment. Through extensive experiments conducted on both publicly available and self-collected datasets comprising 24 sequences, we demonstrate the effectiveness of our mapping and navigation methodologies. Note to Practitioners—This paper tackles the challenge of autonomous navigation for mobile robots in complex, unstructured environments with rich semantic elements. Traditional navigation relies on geometric analysis and manual annotations, struggling to differentiate similar structures like roads and sidewalks. We propose an online mapping system that creates a global metric-semantic mesh map for large-scale outdoor environments, utilizing GPU acceleration for speed and overcoming the limitations of existing real-time semantic mapping methods, which are generally confined to indoor settings. Our map integrates into a real-world navigation system, proven effective in localization and terrain assessment through experiments with both public and proprietary datasets. Future work will focus on integrating kernel-based methods to improve the map’s semantic accuracy.
Jianhao Jiao, Ruoyu Geng, Yuanhang Li, Ren Xin, Jin Wu 0002, Lujia Wang 0001, Ming Liu 0001, Rui Fan 0001, Dimitrios Kanoulas
IEEE Trans Autom. Sci. Eng.1
2025 General Place Recognition Survey: Toward Real-World Autonomy
abstract
In the realm of robotics, the quest for achieving real-world autonomy, capable of executing large-scale and long-term operations, has positioned place recognition (PR) as a cornerstone technology. Despite the PR community's remarkable strides over the past two decades, garnering attention from fields like computer vision and robotics, the development of PR methods that sufficiently support real-world robotic systems remains a challenge. This article aims to bridge this gap by highlighting the crucial role of PR within the framework of simultaneous localization and mapping 2.0. This new phase in robotic navigation calls for scalable, adaptable, and efficient PR solutions by integrating advanced artificial intelligence technologies. For this goal, we provide a comprehensive review of the current state-of-the-art advancements in PR, alongside the remaining challenges, and underscore its broad applications in robotics. This article begins with an exploration of PR's formulation and key research challenges. We extensively review literature, focusing on related methods on place representation and solutions to various PR challenges. Applications showcasing PR's potential in robotics, key PR datasets, and open-source libraries are discussed.
Peng Yin 0001, Jianhao Jiao, Guoquan Huang 0001, Howie Choset, Sebastian A. Scherer, Jianda Han
IEEE Trans. Robotics2
2024 Accurate Prior-centric Monocular Positioning with Offline LiDAR Fusion
abstract
Unmanned vehicles usually rely on Global Positioning System (GPS) and Light Detection and Ranging (LiDAR) sensors to achieve high-precision localization results for navigation purpose. However, this combination with their associated costs and infrastructure demands, poses challenges for widespread adoption in mass-market applications. In this paper, we aim to use only a monocular camera to achieve comparable onboard localization performance by tracking deep-learning visual features on a LiDAR-enhanced visual prior map. Experiments show that the proposed algorithm can provide centimeter-level global positioning results with scale, which is effortlessly integrated and favorable for low-cost robot system deployment in real-world applications.
Jinhao He, Huaiyang Huang, Jianhao Jiao, Ming Liu 0001
ICRA4
2024 OmniColor: A Global Camera Pose Optimization Approach of LiDAR-360Camera Fusion for Colorizing Point Clouds
abstract
A Colored point cloud, as a simple and efficient 3D representation, has many advantages in various fields, including robotic navigation and scene reconstruction. This representation is now commonly used in 3D reconstruction tasks relying on cameras and LiDARs. However, fusing data from these two types of sensors is poorly performed in many existing frameworks, leading to unsatisfactory mapping results, mainly due to inaccurate camera poses. This paper presents Omni-Color, a novel and efficient algorithm to colorize point clouds using an independent 360-degree camera. Given a LiDAR-based point cloud and a sequence of panorama images with initial coarse camera poses, our objective is to jointly optimize the poses of all frames for mapping images onto geometric reconstructions. Our pipeline works in an off-the-shelf manner that does not require any feature extraction or matching process. Instead, we find optimal poses by directly maximizing the photometric consistency of LiDAR maps. In experiments, we show that our method can overcome the severe visual distortion of omnidirectional images and greatly benefit from the wide field of view (FOV) of 360-degree cameras to reconstruct various scenarios with accuracy and stability. The code will be released at https://github.com/liubonan123/OmniColor/.
Bonan Liu, Guoyang Zhao, Jianhao Jiao, Guang Cai, Handi Yin, Ming Liu 0001
ICRA3
2024 An Image Acquisition Scheme for Visual Odometry based on Image Bracketing and Online Attribute Control
abstract
Visual odometry (VO) system is challenged by complex illumination environments. Image quality and its consistency in the time domain directly determine feature detection and tracking performance, which further affect the robustness and accuracy of the entire system. In this paper, an image acquisition scheme with image bracketing patterns is proposed. Images with different exposure levels are continuously captured to sufficiently explore the scene under varying illumination. An attribute control method is designed to adjust image exposures within the brackets online. Gaussian process regression fits the relationship between image quality metric and exposure via image synthesis technique. The optimal exposures for the next bracket are obtained directly without attempts to ensure a quick response. Experiments show our acquisition system’s effectiveness and performance improvement for VO tasks in complex illumination scenes.
Jinhao He, Bohuan Xue, Jin Wu 0002, Pengyu Yin, Jianhao Jiao, Ming Liu 0001
ICRA6
2024 DHP-Mapping: A Dense Panoptic Mapping System with Hierarchical World Representation and Label Optimization Techniques
abstract
Maps provide robots with crucial environmental knowledge, thereby enabling them to perform interactive tasks effectively. Easily accessing accurate abstract-to-detailed geometric and semantic concepts from maps is crucial for robots to make informed and efficient decisions. To comprehensively model the environment and effectively manage the map data structure, we propose DHP-Mapping, a dense mapping system that utilizes multiple Truncated Signed Distance Field (TSDF) submaps and panoptic labels to hierarchically model the environment. The output map is able to maintain both voxel- and submap-level metric and semantic information. Two modules are presented to enhance the mapping efficiency and label consistency: (1) an inter-submaps label fusion strategy to eliminate duplicate points across submaps and (2) a conditional random field (CRF) based approach to enhance panoptic labels. We conducted experiments with two public datasets including indoor and outdoor scenarios. Our system performs comparably to state-of-the-art (SOTA) methods across geometry and label accuracy evaluation metrics. The experiment results highlight the effectiveness and scalability of our system, as it is capable of constructing precise geometry and maintaining consistent panoptic labels. Our code is publicly available at https://github.com/hutslib/DHP-Mapping.
Tianshuai Hu, Jianhao Jiao, Hongji Liu, Sheng Wang 0017, Ming Liu 0001
IROS2
2024 GV-Bench: Benchmarking Local Feature Matching for Geometric Verification of Long-term Loop Closure Detection
abstract
Visual loop closure detection is an important module in visual simultaneous localization and mapping (SLAM), which associates current camera observation with previously visited places. Loop closures correct drifts in trajectory estimation to build a globally consistent map. However, a false loop closure can be fatal, so verification is required as an additional step to ensure robustness by rejecting the false positive loops. Geometric verification has been a well-acknowledged solution that leverages spatial clues provided by local feature matching to find true positives. Existing feature matching methods focus on homography and pose estimation in long-term visual localization, lacking references for geometric verification. To fill the gap, this paper proposes a unified benchmark targeting geometric verification of loop closure detection under long-term conditional variations. Furthermore, we evaluate six representative local feature matching methods (handcrafted and learning-based) under the benchmark, with in-depth analysis for limitations and future directions.
Jingwen Yu, Hanjing Ye, Jianhao Jiao, Ping Tan 0002, Hong Zhang 0013
IROS3
2023 Completely Rational $\text{SO}(n)$ Orthonormalization
abstract
The rotation orthonormalization on the special orthogonal group$\text{SO}(n)$, also known as the high dimensional nearest rotation problem, has been revisited. A new generalized simple iterative formula has been proposed that solves this problem in a completely rational manner. Rational operations allow for efficient implementation on various platforms and also significantly simplify the synthesis of large-scale circuitization. The developed scheme is also capable of designing efficient fundamental rational algorithms, for example, quaternion normalization, which outperforms long-exisiting solvers. Furthermore, an$\text{SO}(n)$neural network has been developed for further learning purpose on the rotation group. Simulation results verify the effectiveness of the proposed scheme and show the superiority against existing representatives. Applications show that the proposed orthonormalizer is of potential in robotic pose estimation problems, e.g., hand-eye calibration.
Jin Wu 0002, Soheil Sarabandi, Jianhao Jiao, Huaiyang Huang, Bohuan Xue, Ruoyu Geng, Lujia Wang 0001, Ming Liu 0001
ICRA3
2023 EmPointMovSeg: Sparse Tensor-Based Moving-Object Segmentation in 3-D LiDAR Point Clouds for Autonomous Driving-Embedded System
abstract
Object segmentation is a per-pixel label prediction task that targets at providing context analysis for autonomous driving. Moving-object segmentation (MOS) serves as a subbranch of object segmentation, targeting to separating the surrounding objects into binary options: dynamic and static. MOS is vital for the safety-critical task in autonomous driving because dynamic objects are often a true potential threat to self-driving cars compared to static ones. Current methods typically address the MOS problem as a category feature to label the mapping task, which is not rational in reality. For example, a parking car should be considered as static instead of a moving-object category. There is a little systematic theory to differentiate object moving characteristics from nonmoving characteristics in MOS. Furthermore, restricted by limited resources in the embedded system, MOS is often in an offline manner due to huge computational requirements. An online and low computational cost MOS is an urgent demand for the practical safety-critical mission which takes immediate reaction as compulsory. In this article, we propose EmPointMovSeg, an efficient and practical 3-D LiDAR MOS solution for autonomous driving. Leveraging the power of the well-adapted autoregressive system identification (AR-SI) theory, EmPointMovSeg theoretically explains the moving-object feature in large-scale 3-D LiDAR semantic segmentation. An end-to-end sparse tensor-based CNN which balances segmentation accuracy and online process ability is proposed. We construct our experiment on both representative dataset benchmarks and practical embedded systems. The evaluation result shows the effectiveness and accuracy of our proposed solution, conquering the bottleneck in the online large-scale 3-D LiDAR semantic segmentation.
Zhijian He, Xueli Fan, Zhaoyan Shen, Jianhao Jiao, Ming Liu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2023 A VT-HMM-Based Framework for Countdown Timer Traffic Light State Estimation
abstract
Traffic lights are important components of traffic systems, and perceptual tasks on traffic lights are crucial for intelligent agents on the road. Auxiliary countdown timers, providing the remaining time of the current traffic phase, improve the safety and smoothness of the entire traffic system. This work proposes a state estimation framework for countdown timer traffic lights. Time-domain information is adequately integrated into a variable transition Hidden Markov Model (VT-HMM), and our system provides optimal estimates of traffic light colors and countdown numbers based on noisy detection inputs. A dynamic state transition matrix is designed based on a 1-step transition logic and a probability of the number of transitions related to the current state sojourn duration. A recursive decoding method based on the Viterbi algorithm is proposed to update all the state candidates and select the optimal state chain. Extensive experiments evaluate the robustness and effectiveness of the proposed work. The performance boundaries of this system are also found under various input noise levels. The source code is available here:https://github.com/ShuyangUni/countdown-timer-traffic-light-estimation
Qingwen Zhang, Feiyi Chen, Jin Wu 0002, Jianhao Jiao, Lujia Wang 0001
IEEE Trans. Intell. Transp. Syst.5
2022 R-PCC: A Baseline for Range Image-based Point Cloud Compression
abstract
In autonomous vehicles or robots, point clouds from LiDAR can provide accurate depth information of objects compared with 2D images, but they also suffer a large volume of data, which is inconvenient for data storage or transmission. In this paper, we propose a Range image-based Point Cloud Compression method, R-PCC, which can reconstruct the point cloud with uniform or non-uniform accuracy loss. We segment the original large-scale point cloud into small and compact regions for spatial redundancy and salient region classification. Our range image-based method can keep and align all points from the original point cloud in the reconstructed point cloud, and the setting of the quantization module restricts the maximum reconstruction error. In the experiments, we prove that our easier FPS-based segmentation method can achieve better performance than instance-based segmentation methods such as DBSCAN, and our non-uniform compression framework shows a great improvement on the downstream tasks compared with the state-of-the-art large-scale point cloud compression methods. Our real-time method can achieve 40 × compression ratio without affecting downstream tasks, to act as a baseline for range image-based point cloud compression. The code is available on https://github.com/StevenWang30/R-PCC.git.
Sukai Wang, Jianhao Jiao, Peide Cai, Lujia Wang 0001
ICRA2
2022 FusionPortable: A Multi-Sensor Campus-Scene Dataset for Evaluation of Localization and Mapping Accuracy on Diverse Platforms
abstract
Combining multiple sensors enables a robot to maximize its perceptual awareness of environments and enhance its robustness to external disturbance, crucial to robotic navigation. This paper proposes the FusionPortable benchmark, a complete multi-sensor dataset with a diverse set of sequences for mobile robots. This paper presents three contributions. We first advance a portable and versatile multi-sensor suite that offers rich sensory measurements: 10Hz LiDAR point clouds, 20Hz stereo frame images, high-rate and asynchronous events from stereo event cameras, 200Hz inertial readings from an IMU, and 10Hz GPS signal. Sensors are already temporally synchronized in hardware. This device is lightweight, self-contained, and has plug-and-play support for mobile robots. Second, we construct a dataset by collecting 17 sequences that cover a variety of environments on the campus by exploiting multiple robot platforms for data collection. Some sequences are challenging to existing SLAM algorithms. Third, we provide ground truth for the decouple localization and mapping performance evaluation. We additionally evaluate state-of-the-art SLAM approaches and identify their limitations. The dataset, consisting of raw sensor measurements, ground truth, calibration data, and evaluated algorithms, will be released.
Jianhao Jiao, Hexiang Wei, Tianshuai Hu, Xiangcheng Hu, Yilong Zhu, Zhijian He, Jin Wu 0002, Jingwen Yu, Xupeng Xie, Huaiyang Huang, Ruoyu Geng, Lujia Wang 0001, Ming Liu 0001
IROS1
2022 An Online Interactive Approach for Crowd Navigation of Quadrupedal Robots
abstract
Robot navigation in human crowds remains the challenge of understanding human behaviors in different scenarios. We present an approach for interactive and human-friendly crowd navigation in complex static environments. The planner models the online interactions among the robot, humans, and the static environment based on game theory. It recurrently expands and optimizes the estimated trajectories for the robot and neighboring agents and provides human-friendly navigation commands. We use various indicators to evaluate the social awareness of the planners and show that our method outperforms existing approaches in success rate to reach the goals and compatibility with humans while maintaining low navigation times. The planner is successfully deployed on a real-world quadrupedal robot, demonstrating safe and interactive crowd navigation with real-time performance.
Jianhao Jiao, Lujia Wang 0001, Ming Liu 0001
IROS2
2022 Robust Odometry and Mapping for Multi-LiDAR Systems With Online Extrinsic Calibration
abstract
Combining multiple LiDARs enables a robot to maximize its perceptual awareness of environments and obtain sufficient measurements, which is promising for simultaneous localization and mapping (SLAM). This article proposes a system to achieve robust and simultaneous extrinsic calibration, odometry, and mapping for multiple LiDARs. Our approach starts with measurement preprocessing to extract edge and planar features from raw measurements. After a motion and extrinsic initialization procedure, a sliding window-based multi-LiDAR odometry runs onboard to estimate poses with an online calibration refinement and convergence identification. We further develop a mapping algorithm to construct a global map and optimize poses with sufficient features together with a method to capture and reduce data uncertainty. We validate our approach’s performance with extensive experiments on 10 sequences (4.60-km total length) for the calibration and SLAM and compare it against the state of the art. We demonstrate that the proposed work is a complete, robust, and extensible system for various multi-LiDAR setups. The source code, datasets, and demonstrations are available at:https://ram-lab.com/file/site/m-loam.
Jianhao Jiao, Haoyang Ye, Yilong Zhu, Ming Liu 0001
IEEE Trans. Robotics1
2022 Quadratic Pose Estimation Problems: Globally Optimal Solutions, Solvability/Observability Analysis, and Uncertainty Description
abstract
Pose estimation problems are fundamental in robotics. Most of these problems are challenging due to the nonconvex nature. This also sets up an obstacle for uncertainty description that is essential for pose integration and quality control. In this article, we show that a large class of related problems can be categorized as the quadratic pose estimation problems (QPEPs) and we propose a general quaternion-based mathematical model to unify these problems. To solve the nonconvex QPEPs, a Gröbner-basis method is investigated to derive their globally optimal and robust solutions. Furthermore, we develop the rules for characterizing the solvability and observability of these solutions. In addition, the uncertainty description, i.e., covariance matrix, as an important piece of information in robotic state estimation frameworks, is analyzed in detail. Theoretical results show that the covariance can be estimated via online optimization, in an efficient and unbiased manner. In this way, both the solution and covariance are guaranteed to be globally optimal. Through simulations and experiments, we show that the proposed QPEP-based solver is not only accurate, robust, and efficient but outperforms the representatives for covariance estimation. The designed algorithms are also assembled as a C++/MATLAB/Octave/ROS library, while these developed interfaces are built for main stream platforms and simultaneous localization and mapping schemes.
Jin Wu 0002, Yu Zheng 0001, Zhi Gao 0005, Yi Jiang 0007, Xiangcheng Hu, Yilong Zhu, Jianhao Jiao, Ming Liu 0001
IEEE Trans. Robotics7
2021 Greedy-Based Feature Selection for Efficient LiDAR SLAM
abstract
Modern LiDAR-SLAM (L-SLAM) systems have shown excellent results in large-scale, real-world scenarios. However, they commonly have a high latency due to the expensive data association and nonlinear optimization. This paper demonstrates that actively selecting a subset of features significantly improves both the accuracy and efficiency of an L-SLAM system. We formulate the feature selection as a combinatorial optimization problem under a cardinality constraint to preserve the information matrix's spectral attributes. The stochastic-greedy algorithm is applied to approximate the optimal results in real-time. To avoid ill-conditioned estimation, we also propose a general strategy to evaluate the environment's degeneracy and modify the feature number online. The proposed feature selector is integrated into a multi-LiDAR SLAM system. We validate this enhanced system with extensive experiments covering various scenarios on two sensor setups and computation platforms. We show that our approach exhibits low localization error and speedup compared to the state-of-the-art L-SLAM systems. To benefit the community, we have released the source code: https://ram-lab.com/file/site/m-loam.
Jianhao Jiao, Yilong Zhu, Haoyang Ye, Huaiyang Huang, Peng Yun, Lingxin Jiang, Lujia Wang 0001, Ming Liu 0001
ICRA1
2020 Smart-Inspect: Micro Scale Localization and Classification of Smartphone Glass Defects for Industrial Automation
abstract
The presence of any type of defect on the glass screen of smart devices has a great impact on their quality. We present a robust semi-supervised learning framework for intelligent micro-scaled localization and classification of defects on a 16K pixel image of smartphone glass. Our model features the efficient recognition and labeling of three types of defects: scratches, light leakage due to cracks, and pits. Our method also differentiates between the defects and light reflections due to dust particles and sensor regions, which are classified as non-defect areas. We use a partially labeled dataset to achieve high robustness and excellent classification of defect and non-defect areas as compared to principal components analysis (PCA), multi-resolution and information-fusion-based algorithms. In addition, we incorporated two classifiers at different stages of our inspection framework for labeling and refining the unlabeled defects. We successfully enhanced the inspection depth-limit up to 5 microns. The experimental results show that our method outperforms manual inspection in testing the quality of glass screen samples by identifying defects on samples that have been marked as good by human inspection.
M. Usman Maqbool Bhutta, Shoaib Aslam, Peng Yun, Jianhao Jiao, Ming Liu 0001
IROS4
2020 MLOD: Awareness of Extrinsic Perturbation in Multi-LiDAR 3D Object Detection for Autonomous Driving
abstract
Extrinsic perturbation always exists in multiple sensors. In this paper, we focus on the extrinsic uncertainty in multi-LiDAR systems for 3D object detection. We first analyze the influence of extrinsic perturbation on geometric tasks with two basic examples. To minimize the detrimental effect of extrinsic perturbation, we propagate an uncertainty prior on each point of input point clouds, and use this information to boost an approach for 3D geometric tasks. Then we extend our findings to propose a multi-LiDAR 3D object detector called MLOD. MLOD is a two-stage network where the multi-LiDAR information is fused through various schemes in stage one, and the extrinsic perturbation is handled in stage two. We conduct extensive experiments on a real-world dataset, and demonstrate both the accuracy and robustness improvement of MLOD. The code, data and supplementary materials are available at: https://ram-lab.com/file/site/mlod.
Jianhao Jiao, Peng Yun, Lei Tai, Ming Liu 0001
IROS1
2019 Using DP Towards A Shortest Path Problem-Related Application
abstract
The detection of curved lanes is still challenging for autonomous driving systems. Although current cutting-edge approaches have performed well in real applications, most of them are based on strict model assumptions. Similar to other visual recognition tasks, lane detection can be formulated as a two-dimensional graph searching problem, which can be solved by finding several optimal paths along with line segments and boundaries. In this paper, we present a directed graph model, in which dynamic programming is used to deal with a specific shortest path problem. This model is particularly suitable to represent objects with long continuous shape structure, e.g., lanes and roads. We apply the designed model and proposed an algorithm for detecting lanes by formulating it as the shortest path problem. To evaluate the performance of our proposed algorithm, we tested five sequences (including 1573 frames) from the KITTI database. The results showed that our method achieves an average successful detection precision of 97.5%.
Jianhao Jiao, Rui Fan 0001, Ming Liu 0001
ICRA1
2019 Real-Time Binocular Vision Implementation on an SoC TMS320C6678 DSP
Rui Fan 0001, Sicheng Duanmu, Yilong Zhu, Jianhao Jiao, Mohammud Junaid Bocus, Yang Yu 0028, Lujia Wang 0001, Ming Liu 0001
ICVS5
2019 Automatic Calibration of Multiple 3D LiDARs in Urban Environments
abstract
Multiple LiDARs have progressively emerged on autonomous vehicles for rendering a rich view and dense measurements. However, the lack of precise calibration negatively affects their potential applications. In this paper, we propose a novel system that enables automatic multi-LiDAR calibration method without any calibration target, prior environment information, and manual initialization. Our approach starts with a hand-eye calibration by aligning the motion of each sensor. The initial results are then refined by an appearance-based method by minimizing a cost function constructed by point-plane distance. Experimental results on simulated and real-world data demonstrate the reliability and accuracy of our calibration approach. The proposed approach can calibrate a multi-LiDAR system with the rotation and translation errors less than 0. 04rad and 0. 1m respectively for a mobile platform.
Jianhao Jiao, Yang Yu 0028, Qinghai Liao, Haoyang Ye, Rui Fan 0001, Ming Liu 0001
IROS1
2019 Road Crack Detection Using Deep Convolutional Neural Network and Adaptive Thresholding
abstract
Crack is one of the most common road distresses which may pose road safety hazards. Generally, crack detection is performed by either certified inspectors or structural engineers. This task is, however, time-consuming, subjective and labor-intensive. In this paper, a novel road crack detection algorithm which is based on deep learning and adaptive image segmentation is proposed. Firstly, a deep convolutional neural network is trained to determine whether an image contains cracks or not. The images containing cracks are then smoothed using bilateral filtering, which greatly minimizes the number of noisy pixels. Finally, cracks are extracted from the road surface using an adaptive thresholding method. The experimental results illustrate that our network can classify images with an accuracy of 99.92%, and the cracks can be successfully extracted from the images using our proposed thresholding algorithm.
Rui Fan 0001, Mohammud Junaid Bocus, Yilong Zhu, Jianhao Jiao, Fulong Ma, Ming Liu 0001
IV4
2019 A Novel Dual-Lidar Calibration Algorithm Using Planar Surfaces
abstract
Multiple lidars are used on mobile vehicles for rendering a broad view to enhance the performance of perception systems. However, precise calibration of multiple lidars is challenging since the feature correspondences in scan points are sparse for providing enough constraints. To address this problem, existing methods require fixed calibration targets in scenes or rely exclusively on additional sensors. In this paper, we present a novel method that enables automatic lidar calibration without these restrictions. Three linearly independent planar surfaces appearing in surroundings is utilized to find correspondences. Two components are developed to ensure the extrinsic parameters to be found: a closed-form solver for initialization and an optimizer for refinement by minimizing a nonlinear cost function. Simulation and experimental results demonstrate the accuracy of our calibration approach with the rotation and translation errors smaller than 0.05rad and 0.1m respectively.
Jianhao Jiao, Qinghai Liao, Yilong Zhu, Tianyu Liu 0008, Yang Yu 0028, Rui Fan 0001, Lujia Wang 0001, Ming Liu 0001
IV1
2017 A Cloud-Based Visual SLAM Framework for Low-Cost Agents
Jianhao Jiao, Peng Yun, Ming Liu 0001
ICVS1
2017 Towards a Cloud Robotics Platform for Distributed Visual SLAM
Peng Yun, Jianhao Jiao, Ming Liu 0001
ICVS2