VLDB 2026 Research / reviewers in the wild / expert
Jun Zhang 0042
dblp:29/4190-42
· DBLP profile ↗
29ranked-venue papers
6as first author
22since 2021 · last 2025
0000-0001-7406-5243ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 24 · 6 first-author · 20 since 2021Systems, architecture and hardware · 18 · 4 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Overlapping Free: Anchorless UWB-Assisted Relative Pose Estimation for Multi-Robot SystemsabstractAccurate Relative Pose Estimation (RPE) is critical for effective collaboration of multi-robot systems. Traditional methods using cameras or LiDARs heavily rely on overlapping Fields of View (FoV) between robots, which is highly demanding in practical applications and may hinder collaboration efficiency. To accommodate this issue, we propose Anchorless UWB-Assisted Relative Pose Estimation (AURPE), a novel approach that leverages ultra-wideband (UWB) technology in an anchorless setup to achieve multi-robot RPE without requiring overlapping FoVs or external infrastructure. AURPE first estimates the initial relative poses between robots using inter-robot UWB ranging combined with a Bayesian framework and constrained optimization. During robot operation, AURPE continuously refines the relative poses by integrating UWB measurements with LiDAR-inertial odometry (LIO) and employs a consensus voting mechanism to identify the most reliable pose estimates. Additionally, a pose graph-based backend optimization is incorporated to enhance the accuracy of both initial and real-time relative pose. Extensive simulations and real-world experiments demonstrate that AURPE achieves accurate RPE even in non-overlapping scenarios where traditional methods fail. Compared to state-of-the-art point cloud registration methods, AURPE shows superior performance in both accuracy and robustness, highlighting its potential to significantly enhance cooperative tasks in multi-robot systems operating in complex environments. Yanpu Yun, Guohao Peng, Jun Zhang 0042, Yiyao Liu, Kaimin Mao, Danwei Wang |
ICRA | 4 |
| 2025 | LCSPose: Efficient, Accurate and Scalable Markerless 6-DoF Pose Estimation of a Quay Crane Spreader Based on LiDAR and CameraabstractAccurate Six Degrees of Freedom (6-DoF) pose estimation of Ship-To-Shore (STS) quay crane spreaders is crucial for ensuring safe and efficient container handling in port automation. However, existing pose estimation techniques face significant challenges, as camera-based systems either rely on markers, which are prone to damage, or struggle with depth estimation inaccuracies. Additionally, 3D sensor-based approaches, particularly point cloud registration (PCR), face challenges such as initial pose errors, high-latency inference, and difficulties in object identification based purely on geometric features. To address these limitations, we propose LCSPose, a LiDAR-camera fusion-based 6-DoF pose estimation method that is marker-free, accurate, efficient, and scalable. Our approach integrates three key modules: (1) a semantic-geometric segmentation module for spreader segmentation and outlier removal, (2) a spatial consistency template sampling module based on Spatial Consistency Score (SC-Score) for reliable template selection across varying distances, and (3) a multi-view coarse-to-fine pose refinement module which incorporates multi-view PCA alignment for robust initial posture prior estimation and iterative pose refinement strategy for long-range registration. Our method demonstrates a 60% improvement in registration recall over state-of-the-art (SOTA) PCR methods, achieving up to 6 cm in translation error and 0.19 degrees in rotation error, while maintaining real-time processing at 20Hz. Jun Zhang 0042, Guohao Peng, Yanpu Yun, Yiyao Liu, Yuanzhe Wang, Danwei Wang |
ICRA | 2 |
| 2024 | TransLoc4D: Transformer-Based 4D Radar Place RecognitionabstractPlace recognition is crucial for unmanned vehicles in terms of localization and mapping. Recent years have witnessed numerous explorations in the field, where 2D cameras and 3D LiDARs are mostly employed. Despite their admirable performance, they may encounter challenges in adverse weather such as rain and fog. Hopefully, 4D millimeter-wave radar emerges as a promising alternative, as its longer wavelength makes it virtually immune to interference from tiny particles of fog and rain. Therefore, in this work, we propose a novel 4D radar place recognition model, TransLoc4D, based on sparse convolutions and Transformer structures. Specifically, a MinkLoc4D back-bone is first proposed to leverage the multimodal information from 4D radar scans. Rather than merely capturing geometric structures of point clouds, MinkLoc4D additionally explores their intensity and velocity properties. After feature extraction, a Transformer layer is introduced to enhance local features before aggregation, where linear self-attention captures the long-range dependencies of the point cloud, alleviating its sparsity and noise. To validate TransLoc4D, we construct two datasets and set up benchmarks for 4D radar place recognition. Experiments vali-date the feasibility of TransLoc4D and demonstrate it can robustly deal with dynamic and adverse environments. Guohao Peng, Heshan Li, Jun Zhang 0042, Zhenyu Wu 0001, Pengyu Zheng, Danwei Wang |
CVPR | 4 |
| 2024 | Cross-View Detection of Crowded Objects Based on Multi-Sensor FusionabstractTraditional object detection methods are limited by single-sensor constraints, high computational requirements, and poor real-time performance. In addition, occlusion often occurs under the condition of restricted single-view. In this paper, we introduce a camera and LiDAR fusion-based object detection method, which achieves excellent detection performance under limited computational resources. We also explores a fusion detection method deployed with multi-view, which can effectively solve the occlusion issue encountered by single view. The proposed method is valuable for single view as well as multi-view in various application scenarios. Our fusion method significantly improves detection accuracy and reliability, and solves the problems of data discrepancy, interference between sensors, and occlusion due to restricted view. Simulations and extensive experiments show that our proposed object detection method exhibited high accuracy and relatively low computational time. Zhipeng Gu, Guohao Peng, Yanpu Yun, Yiyao Liu, Zhenyu Wu 0001, Jun Zhang 0042, Xudong Suo, Danwei Wang |
ICARCV | 6 |
| 2024 | ACS-MM-Explore: Adaptive Circular Search Strategy for Multi-Modal Robot Exploration in Large-Scale Urban EnvironmentsabstractAutonomous exploration has become a crucial technology for mobile robots, and numerous broadly applicable algorithms have emerged. However, few exploration methods effectively utilize the features of specified types of areas to enhance the efficiency of autonomous exploration in a complex environment. In this paper, we propose ACS-MM-Explore, an adaptive-circular-search-based exploration framework for large-scale urban road environments, focusing on extracting and utilizing the boundaries of roads to enhance exploration efficiency. Our approach integrates a multi-modal traversabil-ity analysis module to distinguish between road and non-traversable areas on a 2D costmap. A novel mechanism for gen-erating exploration viewpoints is introduced, efficiently creating exploration viewpoints with a circular search process with an adaptive radius. An optimized viewpoint selection mechanism is included, taking into account the geographical and geomet-ric information of each viewpoint. The framework extends the move base and TEB local planner as a viewpoint-based navigation module. A comprehensive evaluation is concluded in a simulation environment, demonstrating the framework's effectiveness and robustness. Kaimin Mao, Mingxing Wen, Jun Zhang 0042, Guohao Peng, Zhenyu Wu 0001, Danwei Wang |
ICARCV | 4 |
| 2024 | MM4MM: Map Matching Framework for Multi-Session Mapping in Ambiguous and Perceptually-Degraded EnvironmentsabstractMulti-session mapping serves as the pre-requisite for autonomous robots to fulfill various long-term tasks (e.g., map updating, navigation, collaboration). However, it is challenging to implement multi-session mapping in enclosed or partially enclosed ambiguous environments (e.g., long corridors, industrial warehouses). Existing solutions either depend heavily on the matching of elementary geometric features (e.g., points, lines, and planes), which tends to fail in environments with ambiguous geometric features; or depend on the given guess of the initial transformation matrix of multiple single-session maps, which is not always obtainable and accurate enough. The ambient magnetic field has exhibited ubiquity and high distinctiveness at different location, which makes it suitable for estimating the initial transformation matrix. Thus, this paper proposes a novel probabilistic magnetic-aware Map Matching framework for Multi-session Mapping, namely MM4MM, to estimate the relative transformation of multiple single-session maps and to build the globally consistent maps in ambiguous and perceptually-degraded environments. The key novelties of this work are the designing of the hierarchical probabilistic map matching framework and the Particle Swarm Optimization strategy to associate the magnetic data of multiple sessions. Evaluations on both simulated and real world experiments demonstrate the greatly improved utility, accuracy, and robustness of multi-session mapping over the comparative methods. Zhenyu Wu 0001, Yufeng Yue, Jun Zhang 0042, Hongming Shen, Danwei Wang |
ICRA | 5 |
| 2024 | LB-R2R-Calib: Accurate and Robust Extrinsic Calibration of Multiple Long Baseline 4D Imaging Radars for V2XabstractAs a new sensor, 4D radar (x, y, z, velocity) has great potential for V2X, due to its 3D point cloud, direct doppler velocity output, long distance ranging, low-cost, and more importantly, robust perception in all weathers. However, the extrinsic calibration of multiple long baseline 4D radars is rarely researched in V2X, which is the key to fuse multi-radars. The main reasons are three-folds: (1) New sensor. Thus, it is not surprising that little related work can be found. (2) Long baseline and large viewpoint-difference. Current works are mainly focused on unmanned vehicles, which is short baseline and small viewpoint-difference. (3) Sparse, noisy, and very cluttered 4D radar point cloud. Thus, it is challenging to rapidly and accurately locate the target and extract the feature. In this paper, LB-R2R-Calib (Long Baseline Radar to Radar extrinsic Calibration) is proposed to address these problems. The novelties are: (1) A new target is introduced: an eight-quadrant corner reflector enclosed by a foam sphere. The benefit is the target center is a viewpoint-invariant feature. Thus, it is ideal for large viewpoint-difference calibration. (2) A new feature extraction algorithm is proposed to rapidly locate the target and extract the target center from a very cluttered point cloud, as we observed some important characteristics of 4D radar. Experiments with two 4D radars in real environments with four configurations demonstrate our method is highly accurate and robust. Jun Zhang 0042, Fangwei Zhang, Zhenyu Wu 0001, Guohao Peng, Yiyao Liu, Qiyang Lyu, Mingxing Wen, Danwei Wang |
ICRA | 1 |
| 2024 | S-GPR: Sliding Gaussian Process Regression-based Magnetic Mapping and Evaluation of Different Magnetic Mapping MethodsabstractThe localization of autonomous robots in modern enclosed or semi-enclosed environments, such as office/hotel/hospital, supermarket, and indoor car park environments where GPS signals are severely challenged, remains a bottleneck for the deployment of fully autonomous mobile systems. Existing infrastructure-based (e.g., QR codes, RFID) localization methods are troubled by high maintenance cost and inflexibility issues, while onboard sensors-based solutions (e.g., LiDAR/camera-based) suffer from the ambiguous geometric features and view obstructions from crowded dynamic obstacles (e.g., pedestrians). Magnetic field (MF)-based localization has been gradually utilized in recent years due to its independence from positioning infrastructures and geometric features, thus making it ideal for applications such as service robots and security robots. Magnetic map building serves as the basis and prerequisite component for MF-based localization tasks. The well-acknowledged Gaussian Process Regression (GPR) method can be implemented to build magnetic maps but with heavy computational burdens. Thus in this paper, we propose an efficient and accurate magnetic mapping system based on a novel Sliding-GPR (i.e., S-GPR) method, and evaluate different magnetic mapping methods. A unique region-of-interest (ROI) selection technique and a down/up-sampling method are proposed for the S-GPR to dramatically decrease the computational time while maintaining the mapping accuracy. Extensive experiments in a high-fidelity simulated warehouse and real-world car park environments show that our proposed S-GPR mapping method has exhibited the highest accuracy and relatively low computational time compared with the SOTA magnetic mapping methods. Qiyang Lyu, Zhenyu Wu 0001, Hongming Shen, Jun Zhang 0042, Huiqin Zhou, Danwei Wang |
IECON | 5 |
| 2024 | Secure Object Detection of Autonomous Vehicles Against Adversarial AttacksabstractThis paper addresses the critical challenge of reliable object detection in autonomous vehicles operating in dynamic urban environments, particularly when facing adversarial attacks on perception systems. It presents a novel dualvalidation methodology leveraging the synergy of image based and point cloud based object detection systems. The approach comprises sensor calibration to align data from both sources, an attack detection algorithm utilizing cross-validation technique to identify inconsistencies, and a self-restoration strategy to ensure correct detection despite malicious manipulation. The methodology has been tested in a complex urban environment with adversarial scenarios including patch and random removal attacks. Experimental results demonstrate the robustness and accuracy of the proposed method in maintaining reliable object detection under adversarial conditions. Haoyi Wang, Jun Zhang 0042, Yuanzhe Wang, Danwei Wang |
IECON | 2 |
| 2023 | CAHIR: Co-Attentive Hierarchical Image Representations for Visual Place RecognitionabstractRobust visual place recognition (VPR) against significant appearance changes is crucial for the life-long operation of mobile robots. Focusing on this task, we propose a Co-Attentive Hierarchical Image Representations (CAHIR) framework for VPR, which unifies attention-sharing global and local descriptor generation into one encoding pipeline. The hierarchical descriptors are applied to a coarse-to-fine VPR system with global retrieval and local geometric verification. To explore high-quality local matches between task-relevant visual elements, a cross-attention mutual enhancement layer is introduced to strengthen the information interaction between the local descriptors. Through the proposed selective matching distillation, the mutual enhancement layer can learn from state-of-the-art local matchers in a distillation manner. After weighted cross-matching of the enhanced local descriptors, geometric verification is applied to evaluate the spatial consistency of the compared image pair. Experiments show CAHIR outperforms the existing global and local representations for VPR in terms of performance and efficiency. Quantitatively, it achieves state-of-the-art results on three city-scale benchmark datasets. Qualitatively, CAHIR proves to attach great importance to task-relevant visual elements and excels at finding local correspondences that are discriminative to the VPR task. Guohao Peng, Heshan Li, Jun Zhang 0042, Mingxing Wen, Singh Rahul, Danwei Wang |
ICRA | 4 |
| 2023 | Global Localization in Repetitive and Ambiguous EnvironmentsabstractAccurate global localization is an essential ingredient for autonomous mobile robots (AMRs) operating in enclosed or partially enclosed repetitive environments (e.g., office corridors, industrial warehouses, transportation centers). In such environments, the Global Navigation Satellite System (GNSS) signals are unreliable or severely degraded. The highly ambiguous structures in such challenging scenarios would also lead the ordinary geometric feature-based LiDAR/visual localization methods to fail. The ambient magnetic field (MF) has exhibited high distinctiveness at different location, which makes it a viable alternative for infrastructure-free AMR localization. However, few of the previous research has been focused on the orientation-dependency and similar-sequential-route limitations of MF-based localization. Thus, this paper proposes a novel probabilistic global localization system with 2-D LiDAR and rotation-invariant magnetic field for AMRs operating in challenging repetitive and ambiguous environments. The proposed localization system mainly consists of: 1) Two-step Initialization: laser distance and MF sequence based matching, and 2) MF-based Pose Tracking: recursive multi-dimensional MF sequence based matching. Extensive experimental results demonstrate the advantageous localization performances of the proposed localization system over the existing methods. Zhenyu Wu 0001, Jun Zhang 0042, Qiyang Lyu, Danwei Wang |
ICRA | 3 |
| 2023 | 4DRadarSLAM: A 4D Imaging Radar SLAM System for Large-scale Environments based on Pose Graph OptimizationabstractLiDAR-based SLAM may easily fail in adverse weathers (e.g., rain, snow, smoke, fog), while mmWave Radar remains unaffected. However, current researches are primarily focused on 2D$(x,y)$or 3D ($x, y$, doppler) Radar and 3D LiDAR, while limited work can be found for 4D Radar ($x, y, z$, doppler). As a new entrant to the market with unique characteristics, 4D Radar outputs 3D point cloud with added elevation information, rather than 2D point cloud; compared with 3D LiDAR, 4D Radar has noisier and sparser point cloud, making it more challenging to extract geometric features (edge and plane). In this paper, we propose a full system for 4D Radar SLAM consisting of three modules: 1) Front-end module performs scan-to-scan matching to calculate the odometry based on GICP, considering the probability distribution of each point; 2) Loop detection utilizes multiple rule-based loop pre-filtering steps, followed by an intensity scan context step to identify loop candidates, and odometry check to reject false loop; 3) Back-end builds a pose graph using front-end odometry, loop closure, and optional GPS data. Optimal pose is achieved through$\mathrm{g}2\mathrm{o}$. We conducted real experiments on two platforms and five datasets (ranging from 240m to 4.8km) and will make the code open-source to promote further research at: https://github.com/zhuge2333/4DRadarSLAM Jun Zhang 0042, Huayang Zhuge, Zhenyu Wu 0001, Guohao Peng, Mingxing Wen, Yiyao Liu, Danwei Wang |
ICRA | 1 |
| 2023 | LB-L2L-Calib 2.0: A Novel Online Extrinsic Calibration Method for Multiple Long Baseline 3D LiDARs Using ObjectsabstractIn V2X (Vehicle-to-Everything), one important work is to extrinsically calibrate multiple 3D LiDARs, which are mounted with a long baseline and large viewpoint-difference at the road-side. Current solutions either require a specific target being set up (e.g., a sphere), or require specific features existing in the environment (e.g., mutually orthogonal planes). However, it is time-consuming, sometimes even inconvenient, to set up specific targets, e.g., at busy intersections and highways. Furthermore, specific features do not always exist in the traffic scenario. Thus, the current solutions are not feasible. To address this problem, a novel extrinsic calibration method is proposed in this paper, namely LB-L2L-Calib 2.0. It is the 2.0 version of our previous work. The novelties are: 1) We propose to use the easily accessible objects on the road as features for calibration (i.e., the vehicles). Thus, it is not necessary to set up any specific targets and we do not need to worry whether specific features exist or not. The key point is we observed that the 3D bounding box centers of the vehicles are viewpoint-invariant from different viewpoints, which makes them ideal features for long baseline and large viewpoint-difference calibration. 2) To establish correct correspondence between the bounding box centers detected from different LiDARs, we propose an exhaustive searching strategy. It can robustly output correct correspondence. Extensive experiments are performed in three scenarios (simulation: intersection, real: carpark and highway), with two types of LiDAR (Velodyne and Livox), demonstrating that LB-L2L-Calib 2.0 is robust, effective, and accurate. Jun Zhang 0042, Qiao Yan, Mingxing Wen, Qiyang Lyu, Guohao Peng, Zhenyu Wu 0001, Danwei Wang |
IROS | 1 |
| 2023 | L2V2T2Calib: Automatic and Unified Extrinsic Calibration Toolbox for Different 3D LiDAR, Visual Camera and Thermal CameraabstractExtrinsic calibration between LiDAR-Camera and LiDAR-LiDAR has been researched extensively, because it is the foundation for sensor fusion. Meanwhile, many projects are open-sourced and significantly promote related research. However, limited solutions can unify the calibration between repetitive scanning and non-repetitive scanning 3D LiDAR, sparse and dense 3D LiDAR, visual and thermal camera. Currently, to achieve that, we normally need to use different targets and extract different features for different sensor combinations. Sometimes, human intervention is required to locate the target. It is inconvenient and time-consuming. In this paper, L2V2T2Calib is introduced and open-sourced as a trial to unify the calibration. 1). A four-circular-holes board is adopted for all sensors. The four circle centers can be detected by all the sensors, thus are ideal common features. Previous works also use this target, but the algorithms don’t consider non-repetitive scanning LiDARs, thus cannot be directly applied. 2). To unify the process, an important step is to automatically and robustly detect the target from different types of LiDARs. However, this does not receive enough attention. We propose a method based on template matching. It is simple, but effective and general to different depth sensors. 3). We provide two types of output, minimizing 2D re-projection error (Min2D) and minimizing 3D matching error (Min3D), for different users. And their performance is compared. Extensive experiments conducted in both simulation and real environment demonstrate L2V2T2Calib is accurate, robust, more importantly, unified. The code will be open-sourced to promote related research at: https://github.com/Clothooo/lvt2calib Jun Zhang 0042, Yiyao Liu, Mingxing Wen, Yufeng Yue, Danwei Wang |
IV | 1 |
| 2022 | C-TM: Topo-metric Mapping and Localization based on Place Categorization and Place Recognition for a Delivery Robot on FootpathabstractIn this work, C-TM is presented: a method to build a topo-metric map for delivery robot navigation in largescale city environments. This system automatically generates a compact map by only saving expensive LIDAR information at key locations. These locations form the nodes of a topological map. Nodes are identified using a Place-Categorization (PC) neural network which output the place category from RGB cameras. Inside nodes, we generate and save high quality LIDAR submaps. Global localization within the map is done with a Visual-Place-Recognition (VPR) neural network. The topo-metric map can be used for navigation on footpath. We deploy C-TM on a four-wheeled autonomous delivery robot and test the effectiveness in two environments, both day and night. Timothy Chia, Jun Zhang 0042, Heshan Li, Guohao Peng, Mingxing Wen, Dawei Kee, P. G. C. N. Senarathne |
ICARCV | 2 |
| 2022 | Vision Based Sidewalk Navigation for Last-mile Delivery RobotabstractNavigating delivery robot along the sidewalk safely and robustly in a campus environment is extremely challenging due to the narrow motion space, appearance changes and unstable GPS localization signal under canopies of trees, etc. To that end, we have completed a systematic implementation for delivery robot sidewalk navigation, where a robust vision based navigation algorithm has been proposed. And it consists of three main modules: sidewalk segmentation, costmap generation and motion planning. More Specifically, the first module is to find the drivable area of the surrounding environment, where an image-based segmentation neural network has been developed to extract where the robot can traverse. Since it only takes as input immediate and local sensory data, thus releasing the high dependence on a prior map. Then, an inverse perspective mapping follows to generate a bird-eye-view of the drivable area and constructs the local occupancy grid map intuitively. Next, two different motion planners, control-based primitives (Dynamic Window Approach) and state-based primitives (state lattice planner), have been adopted to generate a trajectory candidate for navigating the robot along the sidewalk. Both simulation and real-world sidewalk navigation experiments have been conducted to test and evaluate their performance. The results show that our algorithm can precisely extract the sidewalk area for traversing, and the state-based primitive planner demonstrates superior performance in terms of trajectory length and time cost, achieving 14.3% and 18.7% improvement compared with control-based primitive planner. Mingxing Wen, Jun Zhang 0042, Tairan Chen, Guohao Peng, Timothy Chia, Yingchong Ma |
ICARCV | 2 |
| 2022 | LB-L2L-Calib: Accurate and Robust Extrinsic Calibration for Multiple 3D LiDARs with Long Baseline and Large Viewpoint DifferenceabstractMulti-LiDAR system is an important part of V2X (Vehicle to Everything) to enhance the perception information for unmanned vehicles. To fuse the information from multiple 3D LiDARs, accurate extrinsic calibration between the LiDARs is essential. However, the existing multi-LiDAR calibration methods mainly focus on short baseline scenarios, where multiple LiDARs are closely mounted on a single platform (e.g., an unmanned vehicle). Besides, most methods typically use a planar target for calibration. Some of the methods require the motion of the multi-LiDAR system. The above conditions severely limit the application of these methods to V2X, where LiDARs are non-movable, the baseline and viewpoint difference between the LiDARs can be very large. In order to meet these challenges, we propose an accurate and robust extrinsic calibration method for long baseline multi-LiDAR systems, named LB-L2L-Calib (Large Baseline LiDAR to LiDAR extrinsic Calibration). (1) We use a sphere as the calibration target for multiple LiDARs with large viewpoint difference, leveraging the viewpoint-invariance of the sphere. (2) A improved sphere detection and sphere center estimation strategy is introduced to detect and extract the sphere center from a cluttered point cloud in large-scale outdoor scenario. (3) A extrinsic parameter regression scheme is introduced. Both simulation and real experiments demonstrate that LB-L2L-Calib is highly accurate and robust. Quantitative results show that the rotation and translation error is less than 0.01m and 0.01° (in simulation, Gauss noise 0.03m, the distance and viewpoint difference between two LiDARs is more than 30m and 90°). Jun Zhang 0042, Qiyang Lyu, Guohao Peng, Zhenyu Wu 0001, Qiao Yan, Danwei Wang |
ICRA | 1 |
| 2022 | SectionKey: 3-D Semantic Point Cloud Descriptor for Place RecognitionabstractPlace recognition is seen as a crucial factor to correct cumulative errors in Simultaneous Localization and Mapping (SLAM) applications. Most existing studies focus on visual place recognition, which is inherently sensitive to environmental changes such as illumination, weather and seasons. Considering these facts, more recent attention has been attracted to use 3-D Light Detection and Ranging (LiDAR) scans for place recognition, which demonstrates more credibility by exerting accurate geometric information. Different from pure geometric-based studies, this paper proposes a novel global descriptor, named SectionKey, which leverages both semantic and geometric information to tackle the problem of place recognition in large-scale urban environments. The proposed descriptor is robust and invariant to viewpoint changes. Specifically, the encoded three-layers key serves as a pre-selection step and a ‘candidate center’ selection strategy is deployed before calculating the similarity score, thus improving the accuracy and efficiency significantly. Then, a two-step semantic iterative closest point (ICP) algorithm is applied to acquire the 3-D pose (x, y, θ) that is used to align the candidate point clouds with the query frame and calculate the similarity score. Extensive experiments have been conducted on public Semantic KITTI dataset to demonstrate the superior performance of our proposed system over state-of-the-art baselines. Shutong Jin, Zhenyu Wu 0001, Jun Zhang 0042, Guohao Peng, Danwei Wang |
IROS | 4 |
| 2022 | A Robust Sidewalk Navigation Method for Mobile Robots Based on Sparse Semantic Point CloudabstractLast-mile delivery robots are usually required to navigate on the sidewalk through a fixed route. The current solutions heavily rely on the image-based perception and GPS localization to successfully complete delivery tasks. However, it is prone to fail and become unreliable when the robot runs in challenging conditions, such as operating in different illuminations, or under canopies of trees or buildings. To address these issues, this paper proposes a novel robust sidewalk navigation method for the last-mile delivery robots with an affordable sparse LiDAR, which consists of two main modules: Semantic Point Cloud Network (SegPCn) and Reactive Nav-igation Network (RNn), as shown in Fig. 1. More specifically, SegPCn takes the raw 3D point cloud as input and predicts the point-wise segmentation labels, presenting a robust perception capability even in the night. Then, the semantic point clouds are fed to RNn to generate an angular velocity to navigate the robot along the sidewalk, where the localization of the robot is not required. Moreover, an autolabeling mechanism is developed to reduce the labor involved in data preparation as well. And the LSTM neural network is explored to effectively leverage the historical context and derive correct decisions. Extensive experiments have been carried out to verify the efficacy of this method, and the results show that this method enables the robot to navigate on the sidewalk robustly during day and night. We open source the code and the data set on https://github.com/lukewenMX/Robust-Navigation-Method. Mingxing Wen, Yunxiang Dai, Tairan Chen, Jun Zhang 0042, Danwei Wang |
IROS | 5 |
| 2021 | Attentional Pyramid Pooling of Salient Visual Residuals for Place RecognitionabstractThe core of visual place recognition (VPR) lies in how to identify task-relevant visual cues and embed them into dis- criminative representations. Focusing on these two points, we propose a novel encoding strategy named Attentional Pyramid Pooling of Salient Visual Residuals (APPSVR). It incorporates three types of attention modules to model the saliency of local features in individual, spatial and cluster dimensions respectively. (1) To inhibit task-irrelevant local features, a semantic-reinforced local weighting scheme is employed for local feature refinement; (2) To leverage the spatial context, an attentional pyramid structure is constructed to adaptively encode regional features according to their relative spatial saliency; (3) To distinguish the different importance of visual clusters to the task, a parametric normalization is proposed to adjust their contribution to image descriptor generation. Experiments demonstrate APPSVR outperforms the existing techniques and achieves a new state-of-the-art performance on VPR benchmark datasets. The visualization shows the saliency map learned in a weakly supervised manner is largely consistent with human cognition. Guohao Peng, Jun Zhang 0042, Heshan Li, Danwei Wang |
ICCV | 2 |
| 2021 | Semantic Reinforced Attention Learning for Visual Place RecognitionabstractLarge-scale visual place recognition (VPR) is inherently challenging because not all visual cues in the image are beneficial to the task. In order to highlight the task-relevant visual cues in the feature embedding, the existing attention mechanisms are either based on artificial rules or trained in a thorough data-driven manner. To fill the gap between the two types, we propose a novel Semantic Reinforced Attention Learning Network (SRALNet), in which the inferred attention can benefit from both semantic priors and data-driven fine-tuning. The contribution lies in two-folds. (1) To suppress misleading local features, an interpretable local weighting scheme is proposed based on hierarchical feature distribution. (2) By exploiting the interpretability of the local weighting scheme, a semantic constrained initialization is proposed so that the local attention can be reinforced by semantic priors. Experiments demonstrate that our method outperforms state-of-the-art techniques on city-scale VPR benchmark datasets. Guohao Peng, Yufeng Yue, Jun Zhang 0042, Zhenyu Wu 0001, Danwei Wang |
ICRA | 3 |
| 2021 | MSTSL: Multi-Sensor Based Two-Step Localization in Geometrically Symmetric EnvironmentsabstractSymmetric environment is one of the most intractable and challenging scenarios for mobile robots to accomplish global localization tasks, due to the highly similar geometrical structures and insufficient distinctive features. Existing localization solutions in such scenarios either depend on pre-deployed infrastructures which are expensive, inflexible, and hard to maintain; or rely on single sensor-based methods whose initialization module is incapable to provide enough unique information. Thus, this paper proposes a novel Multi-Sensor based Two-Step Localization framework named MSTSL, which addresses the problem of mobile robot global localization in geometrically symmetric environments by utilizing the measured magnetic field, 2-D LiDAR, and wheel odometry information. The proposed system mainly consists of two steps: 1) Magnetic Field-based Initialization, and 2) LiDAR-based Localization. Based on the pre-built magnetic field database, multiple initial hypotheses poses can firstly be determined by the proposed two-stage initialization algorithm. Then, utilizing the obtained multiple initial hypotheses, the robot can be localized more accurately by LiDAR-based localization. Extensive experiments demonstrate the practical utility and accuracy of the proposed system over the alternative approaches in real-world scenarios. Zhenyu Wu 0001, Yufeng Yue, Mingxing Wen, Jun Zhang 0042, Guohao Peng, Danwei Wang |
ICRA | 4 |
| 2020 | Human-Robot Teaming and Coordination in Day and Night EnvironmentsabstractAs robots are sharing work spaces with human, human-robot teamwork is becoming increasingly important. It is foreseeable that the daily work team will be composed of human and robots. The integration of the appropriate decision-making process is an essential part to design and develop the team. If robots can understand the activities and intents of human, it is convenient for a person to cooperate with robots in a natural manner. This paper proposes a system that enables robots to understand human pose and execute given command. The system provides two options for different hardware systems: the first one is suitable for powerful computational units; the second model is compact and efficient on a normal robot platform. In order to enrich application scenarios, we propose a method to extract human pose from thermal images so that our system can be used in all-weather scenario. In addition, we collected extensive training data and trained a MLP neural network to classify several human poses. The experimental results show the accuracy and efficiency of the proposed MLP neural network in day and night environments. Yufeng Yue, Yuanzhe Wang, Jun Zhang 0042, Danwei Wang |
ICARCV | 4 |
| 2020 | Day and Night Collaborative Dynamic Mapping in Unstructured Environment Based on Multimodal SensorsabstractEnabling long-term operation during day and night for collaborative robots requires a comprehensive understanding of the unstructured environment. Besides, in the dynamic environment, robots must be able to recognize dynamic objects and collaboratively build a global map. This paper proposes a novel approach for dynamic collaborative mapping based on multimodal environmental perception. For each mission, robots first apply heterogeneous sensor fusion model to detect humans and separate them to acquire static observations. Then, the collaborative mapping is performed to estimate the relative position between robots and local 3D maps are integrated into a globally consistent 3D map. The experiment is conducted in the day and night rainforest with moving people. The results show the accuracy, robustness, and versatility in 3D map fusion missions. Yufeng Yue, Chule Yang, Jun Zhang 0042, Mingxing Wen, Zhenyu Wu 0001, Danwei Wang |
ICRA | 3 |
| 2020 | A Hierarchical Framework for Collaborative Probabilistic Semantic MappingabstractPerforming collaborative semantic mapping is a critical challenge for cooperative robots to maintain a comprehensive contextual understanding of the surroundings. Most of the existing work either focus on single robot semantic mapping or collaborative geometry mapping. In this paper, a novel hierarchical collaborative probabilistic semantic mapping framework is proposed, where the problem is formulated in a distributed setting. The key novelty of this work is the mathematical modeling of the overall collaborative semantic mapping problem and the derivation of its probability decomposition. In the single robot level, the semantic point cloud is obtained based on heterogeneous sensor fusion model and is used to generate local semantic maps. Since the voxel correspondence is unknown in collaborative robots level, an Expectation-Maximization approach is proposed to estimate the hidden data association, where Bayesian rule is applied to perform semantic and occupancy probability update. The experimental results show the high quality global semantic map, demonstrating the accuracy and utility of 3D semantic map fusion algorithm in real missions. Yufeng Yue, Chule Yang, Jun Zhang 0042, Mingxing Wen, Yuanzhe Wang, Danwei Wang |
ICRA | 5 |
| 2020 | Infrastructure-Free Global Localization in Repetitive Environments: An OverviewabstractRepetitive environment is a challenging scenario for mobile robot global localization due to its highly similar structures and lack of distinctive features. Existing solutions in such environments rely heavily on pre-installed infrastructures, which are neither flexible nor cost-effective. Besides, few of the previous research have been focused on the implementation of infrastructure-free localization approaches in repetitive scenarios. Thus, this paper serves as a survey to investigate the problem of infrastructure-free mobile robot global localization with low-cost and efficient sensors in repetitive environments. Three of the most popular infrastructure-free localization methods, namely LiDAR-based localization (LBL), vision-based localization (VBL), and magnetic field-based localization (MFL), are analyzed and evaluated. Extensive global localization experiments are conducted in real-world repetitive scenarios and the results demonstrate that VBL methods perform slightly better than LBL and MFL methods. The overall evaluations indicate that infrastructure-free global localization in repetitive environment is still a challenging problem which deserves more research efforts to develop new solutions. Zhenyu Wu 0001, Jun Zhang 0042, Yufeng Yue, Mingxing Wen, Zichen Jiang, Danwei Wang |
IECON | 2 |
| 2019 | Probabilistic Reasoning for Unique Role Recognition Based on the Fusion of Semantic-Interaction and Spatio-Temporal FeaturesabstractThis paper deals with the problem of recognizing the unique role in dynamic environments. Different from social roles, the unique role refers to those who are unusual in their carrying items or movements in the scene. In this paper, we propose a hierarchical probabilistic reasoning method that relates spatial relationships between interested objects and humans with their temporal changes to recognize the unique individual. Two observation models, Object Existence Model (OEM) and Human Action Model (HAM), are established to support role inference by analyzing the corresponding semantic-interaction features and spatio-temporal features. Then, OEM and HAM results of each person are compared with the overall distribution in the scene, respectively. Finally, we can determine the role through the fusion of two observation models. Experiments are conducted in both indoor and outdoor environments concerning different settings, degrees of clutter, and occlusions. The results show that the proposed method can adapt to a variety of scenarios and outperforms other methods on accuracy and robustness, moreover, exhibiting stable performance even in complex scenes. Chule Yang, Yufeng Yue, Jun Zhang 0042, Mingxing Wen, Danwei Wang |
IEEE Trans. Multim. | 3 |
| 2018 | Probabilistic Fusion Framework for Collaborative Robots 3D MappingabstractFusion of local 3D maps generated by individual robots to a globally consistent 3D map is one of the fundamental challenges in multi-robot mapping missions. In this paper, we propose a probabilistic mathematical formulation to address the integrated map fusion problem. More specifically, the problem of estimating fused map posterior can be factorized into a product of relative transformation posterior and the global map posterior, which enables us to solve map matching and map merging problems efficiently. In addition, a distributed communication strategy is employed to share map information among robots. The proposed approach is evaluated in indoor and mixed environments, which shows its utility in 3D map fusion for multi-robot mapping missions. Yufeng Yue, P. G. C. N. Senarathne, Chule Yang, Jun Zhang 0042, Mingxing Wen, Danwei Wang |
FUSION | 4 |
| 2018 | A Two-step Method for Extrinsic Calibration between a Sparse 3D LiDAR and a Thermal CameraabstractTo obtain the 6 DOF extrinsic parameters (rotation and translation matrix) between a 3D ranging sensor and a thermal camera, previous methods require a high-resolution 3D ranging sensor to reliably detect features. Although sparse 3D LiDARs are widely used on autonomous robots, to the best of our knowledge, the extrinsic calibration between a sparse 3D LiDAR (particularly Velodyne VLP-16) and a thermal camera has not been considered in the literature. In this paper, we present a two-step method to address the problem, where a monocular visual camera is used to assist the process. The proposed method decomposes the problem into two steps: extrinsic calibration between a sparse 3D LiDAR and a visual camera; extrinsic calibration between a visual camera and a thermal camera. Experiments are conducted to demonstrate the effectiveness of the proposed two-step method. Jun Zhang 0042, Prarinya Siritanawan, Yufeng Yue, Chule Yang, Mingxing Wen, Danwei Wang |
ICARCV | 1 |