VLDB 2026 Research / reviewers in the wild / expert
Zhenyu Wu 0001
dblp:87/6581-1
· DBLP profile ↗
23ranked-venue papers
4as first author
19since 2021 · last 2026
0000-0002-9094-4982ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 3 first-author · 16 since 2021Systems, architecture and hardware · 15 · 4 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | UniLGL: Learning Uniform Place Recognition for FOV-Limited/Panoramic LiDAR Global LocalizationabstractLiDAR-based Global Localization (LGL) is an essential ingredient for autonomous robots. However, existing LGL methods typically consider only partial information (e.g., geometric features) from LiDAR observations or are designed for homogeneous LiDAR sensors, overlooking the uniformity in LGL. In this work, a uniform LGL method is proposed, termed UniLGL, which simultaneously achieves spatial and material uniformity, as well as sensor-type uniformity. The key idea of the proposed method is to encode the complete point cloud, which contains both geometric and material information, into a pair of Bird's Eye View (BEV) images (i.e., a spatial BEV image and an intensity BEV image), thereby transforming the LGL problem into a cascaded LiDAR Place Recognition (LPR) and pose estimation problem from the perspective of image fusion. An end-to-end multi-BEV fusion network is designed to extract uniform features, equipping UniLGL with spatial and material uniformity. To ensure robust LGL across heterogeneous LiDAR sensors, a viewpoint invariance hypothesis is introduced, which replaces the conventional translation equivariance assumption commonly used in existing LPR networks and supervises UniLGL to achieve sensor type uniformity in both global descriptors and local feature representations. Moreover, UniLGL introduces a pipeline that leverages a pre-trained single-image Vision Foundation Model (VFM) for feature extraction to enhance the multi-BEV fusion LPR network, enabling strong generalization with only a few LiDAR data for fine-tuning. Finally, based on the mapping between local features on the 2D BEV image and the point cloud, a robust global pose estimator is derived that determines the global minimum of the global pose on <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"><tex-math notation="LaTeX">$\text{SE}(3)$</tex-math></inline-formula> without requiring additional registration. To validate the effectiveness of the proposed uniform LGL, extensive benchmarks are conducted in real-world environments, and the results show that the proposed UniLGL is demonstratively competitive compared to other State-of-the-Art (SOTA) LGL methods. Furthermore, UniLGL has been deployed on diverse platforms, including full-size trucks and agile Micro Aerial Vehicles (MAVs), to enable high-precision localization and mapping as well as multi-MAV collaborative exploration in port and forest environments, demonstrating the applicability of UniLGL in industrial and field scenarios. The code will be released at <uri xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">https://github.com/shenhm516/UniLGL</uri>. Hongming Shen, Yulin Hui, Zhenyu Wu 0001, Qiyang Lyu, Tianchen Deng, Danwei Wang |
IEEE Trans. Robotics | 4 |
| 2025 | Curb-Tracker: An Integrated Curb Following System for Autonomous Vehicles
Yuanzhe Wang, Guohao Peng, Zhenyu Wu 0001, Danwei Wang |
IEEE Trans. Robotics | 4 |
| 2024 | TransLoc4D: Transformer-Based 4D Radar Place RecognitionabstractPlace recognition is crucial for unmanned vehicles in terms of localization and mapping. Recent years have witnessed numerous explorations in the field, where 2D cameras and 3D LiDARs are mostly employed. Despite their admirable performance, they may encounter challenges in adverse weather such as rain and fog. Hopefully, 4D millimeter-wave radar emerges as a promising alternative, as its longer wavelength makes it virtually immune to interference from tiny particles of fog and rain. Therefore, in this work, we propose a novel 4D radar place recognition model, TransLoc4D, based on sparse convolutions and Transformer structures. Specifically, a MinkLoc4D back-bone is first proposed to leverage the multimodal information from 4D radar scans. Rather than merely capturing geometric structures of point clouds, MinkLoc4D additionally explores their intensity and velocity properties. After feature extraction, a Transformer layer is introduced to enhance local features before aggregation, where linear self-attention captures the long-range dependencies of the point cloud, alleviating its sparsity and noise. To validate TransLoc4D, we construct two datasets and set up benchmarks for 4D radar place recognition. Experiments vali-date the feasibility of TransLoc4D and demonstrate it can robustly deal with dynamic and adverse environments. Guohao Peng, Heshan Li, Jun Zhang 0042, Zhenyu Wu 0001, Pengyu Zheng, Danwei Wang |
CVPR | 5 |
| 2024 | Cross-View Detection of Crowded Objects Based on Multi-Sensor FusionabstractTraditional object detection methods are limited by single-sensor constraints, high computational requirements, and poor real-time performance. In addition, occlusion often occurs under the condition of restricted single-view. In this paper, we introduce a camera and LiDAR fusion-based object detection method, which achieves excellent detection performance under limited computational resources. We also explores a fusion detection method deployed with multi-view, which can effectively solve the occlusion issue encountered by single view. The proposed method is valuable for single view as well as multi-view in various application scenarios. Our fusion method significantly improves detection accuracy and reliability, and solves the problems of data discrepancy, interference between sensors, and occlusion due to restricted view. Simulations and extensive experiments show that our proposed object detection method exhibited high accuracy and relatively low computational time. Zhipeng Gu, Guohao Peng, Yanpu Yun, Yiyao Liu, Zhenyu Wu 0001, Jun Zhang 0042, Xudong Suo, Danwei Wang |
ICARCV | 5 |
| 2024 | ACS-MM-Explore: Adaptive Circular Search Strategy for Multi-Modal Robot Exploration in Large-Scale Urban EnvironmentsabstractAutonomous exploration has become a crucial technology for mobile robots, and numerous broadly applicable algorithms have emerged. However, few exploration methods effectively utilize the features of specified types of areas to enhance the efficiency of autonomous exploration in a complex environment. In this paper, we propose ACS-MM-Explore, an adaptive-circular-search-based exploration framework for large-scale urban road environments, focusing on extracting and utilizing the boundaries of roads to enhance exploration efficiency. Our approach integrates a multi-modal traversabil-ity analysis module to distinguish between road and non-traversable areas on a 2D costmap. A novel mechanism for gen-erating exploration viewpoints is introduced, efficiently creating exploration viewpoints with a circular search process with an adaptive radius. An optimized viewpoint selection mechanism is included, taking into account the geographical and geomet-ric information of each viewpoint. The framework extends the move base and TEB local planner as a viewpoint-based navigation module. A comprehensive evaluation is concluded in a simulation environment, demonstrating the framework's effectiveness and robustness. Kaimin Mao, Mingxing Wen, Jun Zhang 0042, Guohao Peng, Zhenyu Wu 0001, Danwei Wang |
ICARCV | 6 |
| 2024 | CT-MLO: Voxel-Based Multi-LiDAR Odometry Using Continuous-Time Kalman FilterabstractIn recent years, LiDAR-based localization and mapping methods have achieved significant progress thanks to their reliable and real-time localization capability. However, single LiDAR odometry often faces hardware failures and degradation in practical scenarios, and the continuous-time measurement characteristic is constantly neglected by existing LiDAR odometry. This motivates us to develop a continuous-time Multi-LiDAR Odometry (MLO) method, namely CT-MLO, which can realize accurate and real-time state estimation using multi-LiDAR measurements through a continuous-time perspective. Due to the advantageous continuous-time formulation, each LiDAR point in a point stream can query the corresponding continuous-time trajectory within its time instants. Additionally, a decentralized multi-LiDAR synchronization scheme is devised to combine points from separate LiDARs into a single point cloud without the need for primary LiDAR assignment. With the detailed derivation of the analytic Jacobians for continuous-time LiDAR observation, the proposed method integrates synchronization, continuous-time estimation, and voxel map management within a Kalman filter framework, which can achieve real-time state estimation with only a few linear iterations. The effectiveness of the proposed method is demonstrated through various scenarios, including public datasets and real-world autonomous driving experiments. The results demonstrate that the proposed CT-MLO can achieve high-accuracy continuous-time state estimations in real-time and is demonstratively competitive compared to other State-of-the-Art (SOTA) methods. Hongming Shen, Zhenyu Wu 0001, Qiyang Lyu, Huiqin Zhou, Yeqing Zhu |
ICARCV | 2 |
| 2024 | PLP-SLAM: Point-Line-Plane Simultaneous Localization and MappingabstractFor indoor environments, prior point-based visual SLAM cannot be processed in real time under low texture and illumination. To address this issue, this work proposes PLP-SLAM (Point-Line-Plane-SLAM) with RGB-D camera. Firstly, point and line features are detected in RGB images. For line features, establish length suppression and near line merge strategy to improve the line extraction quality. Secondly, plane features are extracted based on agglomerative hierarchical clustering method in point cloud obtained by RGB-D camera. Point clouds are divided into several nodes, unlike prior methods spend a lot of time to estimate the normal vector for each individual point, this work assumes that points within each node sharing the same plane normal vector, which can significantly improve the computational efficiency. Thirdly, sparse maps including points, lines and planes are established, meanwhile the scenes are reconstructed by creating the dense maps to show plan features directly. Finally, the performance of proposed method is compared against the state-of-the-art SLAM on public datasets to evaluate the pose estimation. All modules are run in real-time on a CPU, experiments clarify that PLP-SLAM can significantly enhance the robustness of 6DoF pose of the camera and simultaneously creating more detailed maps of the environment. Yeqing Zhu, Liangyu Zhao, Qingjie Zhao, Zhenyu Wu 0001, Hongming Shen, Danwei Wang |
ICARCV | 4 |
| 2024 | MM4MM: Map Matching Framework for Multi-Session Mapping in Ambiguous and Perceptually-Degraded EnvironmentsabstractMulti-session mapping serves as the pre-requisite for autonomous robots to fulfill various long-term tasks (e.g., map updating, navigation, collaboration). However, it is challenging to implement multi-session mapping in enclosed or partially enclosed ambiguous environments (e.g., long corridors, industrial warehouses). Existing solutions either depend heavily on the matching of elementary geometric features (e.g., points, lines, and planes), which tends to fail in environments with ambiguous geometric features; or depend on the given guess of the initial transformation matrix of multiple single-session maps, which is not always obtainable and accurate enough. The ambient magnetic field has exhibited ubiquity and high distinctiveness at different location, which makes it suitable for estimating the initial transformation matrix. Thus, this paper proposes a novel probabilistic magnetic-aware Map Matching framework for Multi-session Mapping, namely MM4MM, to estimate the relative transformation of multiple single-session maps and to build the globally consistent maps in ambiguous and perceptually-degraded environments. The key novelties of this work are the designing of the hierarchical probabilistic map matching framework and the Particle Swarm Optimization strategy to associate the magnetic data of multiple sessions. Evaluations on both simulated and real world experiments demonstrate the greatly improved utility, accuracy, and robustness of multi-session mapping over the comparative methods. Zhenyu Wu 0001, Yufeng Yue, Jun Zhang 0042, Hongming Shen, Danwei Wang |
ICRA | 1 |
| 2024 | LB-R2R-Calib: Accurate and Robust Extrinsic Calibration of Multiple Long Baseline 4D Imaging Radars for V2XabstractAs a new sensor, 4D radar (x, y, z, velocity) has great potential for V2X, due to its 3D point cloud, direct doppler velocity output, long distance ranging, low-cost, and more importantly, robust perception in all weathers. However, the extrinsic calibration of multiple long baseline 4D radars is rarely researched in V2X, which is the key to fuse multi-radars. The main reasons are three-folds: (1) New sensor. Thus, it is not surprising that little related work can be found. (2) Long baseline and large viewpoint-difference. Current works are mainly focused on unmanned vehicles, which is short baseline and small viewpoint-difference. (3) Sparse, noisy, and very cluttered 4D radar point cloud. Thus, it is challenging to rapidly and accurately locate the target and extract the feature. In this paper, LB-R2R-Calib (Long Baseline Radar to Radar extrinsic Calibration) is proposed to address these problems. The novelties are: (1) A new target is introduced: an eight-quadrant corner reflector enclosed by a foam sphere. The benefit is the target center is a viewpoint-invariant feature. Thus, it is ideal for large viewpoint-difference calibration. (2) A new feature extraction algorithm is proposed to rapidly locate the target and extract the target center from a very cluttered point cloud, as we observed some important characteristics of 4D radar. Experiments with two 4D radars in real environments with four configurations demonstrate our method is highly accurate and robust. Jun Zhang 0042, Fangwei Zhang, Zhenyu Wu 0001, Guohao Peng, Yiyao Liu, Qiyang Lyu, Mingxing Wen, Danwei Wang |
ICRA | 4 |
| 2024 | S-GPR: Sliding Gaussian Process Regression-based Magnetic Mapping and Evaluation of Different Magnetic Mapping MethodsabstractThe localization of autonomous robots in modern enclosed or semi-enclosed environments, such as office/hotel/hospital, supermarket, and indoor car park environments where GPS signals are severely challenged, remains a bottleneck for the deployment of fully autonomous mobile systems. Existing infrastructure-based (e.g., QR codes, RFID) localization methods are troubled by high maintenance cost and inflexibility issues, while onboard sensors-based solutions (e.g., LiDAR/camera-based) suffer from the ambiguous geometric features and view obstructions from crowded dynamic obstacles (e.g., pedestrians). Magnetic field (MF)-based localization has been gradually utilized in recent years due to its independence from positioning infrastructures and geometric features, thus making it ideal for applications such as service robots and security robots. Magnetic map building serves as the basis and prerequisite component for MF-based localization tasks. The well-acknowledged Gaussian Process Regression (GPR) method can be implemented to build magnetic maps but with heavy computational burdens. Thus in this paper, we propose an efficient and accurate magnetic mapping system based on a novel Sliding-GPR (i.e., S-GPR) method, and evaluate different magnetic mapping methods. A unique region-of-interest (ROI) selection technique and a down/up-sampling method are proposed for the S-GPR to dramatically decrease the computational time while maintaining the mapping accuracy. Extensive experiments in a high-fidelity simulated warehouse and real-world car park environments show that our proposed S-GPR mapping method has exhibited the highest accuracy and relatively low computational time compared with the SOTA magnetic mapping methods. Qiyang Lyu, Zhenyu Wu 0001, Hongming Shen, Jun Zhang 0042, Huiqin Zhou, Danwei Wang |
IECON | 2 |
| 2024 | IDF-MFL: Infrastructure-free and Drift-free Magnetic Field Localization for Mobile RobotabstractIn recent years, infrastructure-based localization methods have achieved significant progress thanks to their reliable and drift-free localization capability. However, the preinstalled infrastructures suffer from inflexibilities and high maintenance costs. This poses an interesting problem of how to develop a drift-free localization system without using the preinstalled infrastructures. In this paper, an infrastructure-free and drift-free localization system is proposed using the ambient magnetic field (MF) information, namely IDF-MFL. IDF-MFL is infrastructure-free thanks to the high distinctiveness of the ambient MF information produced by inherent ferromagnetic objects in the environment, such as steel and reinforced concrete structures of buildings, and underground pipelines. The MF-based localization problem is defined as a stochastic optimization problem with the consideration of the non-Gaussian heavy-tailed noise introduced by MF measurement outliers (caused by dynamic ferromagnetic objects), and an outlier-robust state estimation algorithm is derived to find the optimal distribution of robot state that makes the expectation of MF matching cost achieves its lower bound. The proposed method is evaluated in multiple scenarios1, including experiments on high-fidelity simulation, and real-world environments. The results demonstrate that the proposed method can achieve high-accuracy, reliable, and real-time localization without any pre-installed infrastructures. Hongming Shen, Zhenyu Wu 0001, Qiyang Lyu, Huiqin Zhou, Danwei Wang |
IROS | 2 |
| 2023 | Global Localization in Repetitive and Ambiguous EnvironmentsabstractAccurate global localization is an essential ingredient for autonomous mobile robots (AMRs) operating in enclosed or partially enclosed repetitive environments (e.g., office corridors, industrial warehouses, transportation centers). In such environments, the Global Navigation Satellite System (GNSS) signals are unreliable or severely degraded. The highly ambiguous structures in such challenging scenarios would also lead the ordinary geometric feature-based LiDAR/visual localization methods to fail. The ambient magnetic field (MF) has exhibited high distinctiveness at different location, which makes it a viable alternative for infrastructure-free AMR localization. However, few of the previous research has been focused on the orientation-dependency and similar-sequential-route limitations of MF-based localization. Thus, this paper proposes a novel probabilistic global localization system with 2-D LiDAR and rotation-invariant magnetic field for AMRs operating in challenging repetitive and ambiguous environments. The proposed localization system mainly consists of: 1) Two-step Initialization: laser distance and MF sequence based matching, and 2) MF-based Pose Tracking: recursive multi-dimensional MF sequence based matching. Extensive experimental results demonstrate the advantageous localization performances of the proposed localization system over the existing methods. Zhenyu Wu 0001, Jun Zhang 0042, Qiyang Lyu, Danwei Wang |
ICRA | 1 |
| 2023 | 4DRadarSLAM: A 4D Imaging Radar SLAM System for Large-scale Environments based on Pose Graph OptimizationabstractLiDAR-based SLAM may easily fail in adverse weathers (e.g., rain, snow, smoke, fog), while mmWave Radar remains unaffected. However, current researches are primarily focused on 2D$(x,y)$or 3D ($x, y$, doppler) Radar and 3D LiDAR, while limited work can be found for 4D Radar ($x, y, z$, doppler). As a new entrant to the market with unique characteristics, 4D Radar outputs 3D point cloud with added elevation information, rather than 2D point cloud; compared with 3D LiDAR, 4D Radar has noisier and sparser point cloud, making it more challenging to extract geometric features (edge and plane). In this paper, we propose a full system for 4D Radar SLAM consisting of three modules: 1) Front-end module performs scan-to-scan matching to calculate the odometry based on GICP, considering the probability distribution of each point; 2) Loop detection utilizes multiple rule-based loop pre-filtering steps, followed by an intensity scan context step to identify loop candidates, and odometry check to reject false loop; 3) Back-end builds a pose graph using front-end odometry, loop closure, and optional GPS data. Optimal pose is achieved through$\mathrm{g}2\mathrm{o}$. We conducted real experiments on two platforms and five datasets (ranging from 240m to 4.8km) and will make the code open-source to promote further research at: https://github.com/zhuge2333/4DRadarSLAM Jun Zhang 0042, Huayang Zhuge, Zhenyu Wu 0001, Guohao Peng, Mingxing Wen, Yiyao Liu, Danwei Wang |
ICRA | 3 |
| 2023 | LB-L2L-Calib 2.0: A Novel Online Extrinsic Calibration Method for Multiple Long Baseline 3D LiDARs Using ObjectsabstractIn V2X (Vehicle-to-Everything), one important work is to extrinsically calibrate multiple 3D LiDARs, which are mounted with a long baseline and large viewpoint-difference at the road-side. Current solutions either require a specific target being set up (e.g., a sphere), or require specific features existing in the environment (e.g., mutually orthogonal planes). However, it is time-consuming, sometimes even inconvenient, to set up specific targets, e.g., at busy intersections and highways. Furthermore, specific features do not always exist in the traffic scenario. Thus, the current solutions are not feasible. To address this problem, a novel extrinsic calibration method is proposed in this paper, namely LB-L2L-Calib 2.0. It is the 2.0 version of our previous work. The novelties are: 1) We propose to use the easily accessible objects on the road as features for calibration (i.e., the vehicles). Thus, it is not necessary to set up any specific targets and we do not need to worry whether specific features exist or not. The key point is we observed that the 3D bounding box centers of the vehicles are viewpoint-invariant from different viewpoints, which makes them ideal features for long baseline and large viewpoint-difference calibration. 2) To establish correct correspondence between the bounding box centers detected from different LiDARs, we propose an exhaustive searching strategy. It can robustly output correct correspondence. Extensive experiments are performed in three scenarios (simulation: intersection, real: carpark and highway), with two types of LiDAR (Velodyne and Livox), demonstrating that LB-L2L-Calib 2.0 is robust, effective, and accurate. Jun Zhang 0042, Qiao Yan, Mingxing Wen, Qiyang Lyu, Guohao Peng, Zhenyu Wu 0001, Danwei Wang |
IROS | 6 |
| 2022 | LB-L2L-Calib: Accurate and Robust Extrinsic Calibration for Multiple 3D LiDARs with Long Baseline and Large Viewpoint DifferenceabstractMulti-LiDAR system is an important part of V2X (Vehicle to Everything) to enhance the perception information for unmanned vehicles. To fuse the information from multiple 3D LiDARs, accurate extrinsic calibration between the LiDARs is essential. However, the existing multi-LiDAR calibration methods mainly focus on short baseline scenarios, where multiple LiDARs are closely mounted on a single platform (e.g., an unmanned vehicle). Besides, most methods typically use a planar target for calibration. Some of the methods require the motion of the multi-LiDAR system. The above conditions severely limit the application of these methods to V2X, where LiDARs are non-movable, the baseline and viewpoint difference between the LiDARs can be very large. In order to meet these challenges, we propose an accurate and robust extrinsic calibration method for long baseline multi-LiDAR systems, named LB-L2L-Calib (Large Baseline LiDAR to LiDAR extrinsic Calibration). (1) We use a sphere as the calibration target for multiple LiDARs with large viewpoint difference, leveraging the viewpoint-invariance of the sphere. (2) A improved sphere detection and sphere center estimation strategy is introduced to detect and extract the sphere center from a cluttered point cloud in large-scale outdoor scenario. (3) A extrinsic parameter regression scheme is introduced. Both simulation and real experiments demonstrate that LB-L2L-Calib is highly accurate and robust. Quantitative results show that the rotation and translation error is less than 0.01m and 0.01° (in simulation, Gauss noise 0.03m, the distance and viewpoint difference between two LiDARs is more than 30m and 90°). Jun Zhang 0042, Qiyang Lyu, Guohao Peng, Zhenyu Wu 0001, Qiao Yan, Danwei Wang |
ICRA | 4 |
| 2022 | SectionKey: 3-D Semantic Point Cloud Descriptor for Place RecognitionabstractPlace recognition is seen as a crucial factor to correct cumulative errors in Simultaneous Localization and Mapping (SLAM) applications. Most existing studies focus on visual place recognition, which is inherently sensitive to environmental changes such as illumination, weather and seasons. Considering these facts, more recent attention has been attracted to use 3-D Light Detection and Ranging (LiDAR) scans for place recognition, which demonstrates more credibility by exerting accurate geometric information. Different from pure geometric-based studies, this paper proposes a novel global descriptor, named SectionKey, which leverages both semantic and geometric information to tackle the problem of place recognition in large-scale urban environments. The proposed descriptor is robust and invariant to viewpoint changes. Specifically, the encoded three-layers key serves as a pre-selection step and a ‘candidate center’ selection strategy is deployed before calculating the similarity score, thus improving the accuracy and efficiency significantly. Then, a two-step semantic iterative closest point (ICP) algorithm is applied to acquire the 3-D pose (x, y, θ) that is used to align the candidate point clouds with the query frame and calculate the similarity score. Extensive experiments have been conducted on public Semantic KITTI dataset to demonstrate the superior performance of our proposed system over state-of-the-art baselines. Shutong Jin, Zhenyu Wu 0001, Jun Zhang 0042, Guohao Peng, Danwei Wang |
IROS | 2 |
| 2022 | LSDNet: A Lightweight Self-Attentional Distillation Network for Visual Place RecognitionabstractVisual Place Recognition (VPR) has become an indispensable capacity for mobile robots to operate in large-scale environments. Existing methods in this field mostly focus on exploring high-performance encoding strategies, while few attempts are devoted to lightweight models that balance per-formance and computational cost. In this work, we propose a Lightweight Self-attentional Distillation Network (LSDNet) aiming to obtain advantages of both performance and efficiency. (1) From a performance perspective, an attentional encoding strategy is proposed to integrate crucial information in the scene. It extends the NetVlad architecture with a self-attention module to facilitate non-local information interaction between local features. Through further visual word vector rescaling, the final image representation can benefit from both non-local spatial integration and cluster-wise weighting. (2) From an efficiency perspective, LSDNet is built upon a lightweight back-bone. To maintain comparable performance to large backbone models, a dual distillation strategy is introduced. It prompts LSDNet to learn both encoding patterns in the hidden space and feature distributions in the encoding space from the teacher model. Through distillation-augmented training, LSDNet is able to rival the teacher model and outperform SOTA global representations with the same lightweight backbone. Guohao Peng, Heshan Li, Zhenyu Wu 0001, Danwei Wang |
IROS | 4 |
| 2021 | Semantic Reinforced Attention Learning for Visual Place RecognitionabstractLarge-scale visual place recognition (VPR) is inherently challenging because not all visual cues in the image are beneficial to the task. In order to highlight the task-relevant visual cues in the feature embedding, the existing attention mechanisms are either based on artificial rules or trained in a thorough data-driven manner. To fill the gap between the two types, we propose a novel Semantic Reinforced Attention Learning Network (SRALNet), in which the inferred attention can benefit from both semantic priors and data-driven fine-tuning. The contribution lies in two-folds. (1) To suppress misleading local features, an interpretable local weighting scheme is proposed based on hierarchical feature distribution. (2) By exploiting the interpretability of the local weighting scheme, a semantic constrained initialization is proposed so that the local attention can be reinforced by semantic priors. Experiments demonstrate that our method outperforms state-of-the-art techniques on city-scale VPR benchmark datasets. Guohao Peng, Yufeng Yue, Jun Zhang 0042, Zhenyu Wu 0001, Danwei Wang |
ICRA | 4 |
| 2021 | MSTSL: Multi-Sensor Based Two-Step Localization in Geometrically Symmetric EnvironmentsabstractSymmetric environment is one of the most intractable and challenging scenarios for mobile robots to accomplish global localization tasks, due to the highly similar geometrical structures and insufficient distinctive features. Existing localization solutions in such scenarios either depend on pre-deployed infrastructures which are expensive, inflexible, and hard to maintain; or rely on single sensor-based methods whose initialization module is incapable to provide enough unique information. Thus, this paper proposes a novel Multi-Sensor based Two-Step Localization framework named MSTSL, which addresses the problem of mobile robot global localization in geometrically symmetric environments by utilizing the measured magnetic field, 2-D LiDAR, and wheel odometry information. The proposed system mainly consists of two steps: 1) Magnetic Field-based Initialization, and 2) LiDAR-based Localization. Based on the pre-built magnetic field database, multiple initial hypotheses poses can firstly be determined by the proposed two-stage initialization algorithm. Then, utilizing the obtained multiple initial hypotheses, the robot can be localized more accurately by LiDAR-based localization. Extensive experiments demonstrate the practical utility and accuracy of the proposed system over the alternative approaches in real-world scenarios. Zhenyu Wu 0001, Yufeng Yue, Mingxing Wen, Jun Zhang 0042, Guohao Peng, Danwei Wang |
ICRA | 1 |
| 2020 | HILPS: Human-in-Loop Policy Search for Mobile Robot NavigationabstractReinforcement learning has obtained increasing attention in mobile robot mapless navigation in recent years. However, there are still some obvious challenges including the sample efficiency, safety due to dilemma of exploration and exploitation. These problems are addressed in this paper by proposing the Human-in-Loop Policy Search (HILPS) framework, where learning from demonstration, learning from human intervention and Near Optimal Policy strategies are integrated together. Firstly, the former two make sure that expert experience grant mobile robot a more informative and correct decision for accomplishing the task and also maintaining the safety of the mobile robot due to the priority of human control. Then the Near Optimal Policy (NOP) provides a way to selectively store the similar experience with respect to the preexisting human demonstration, in which case the sample efficiency can be improved by eliminating exclusively exploratory behaviors. To verify the performance of the algorithm, the mobile robot navigation experiments are extensively conducted in simulation and real world. Results show that HILPS can improve sample efficiency and safety in comparison to state-of-art reinforcement learning. Mingxing Wen, Yufeng Yue, Zhenyu Wu 0001, Ehsan Mihankhah, Danwei Wang |
ICARCV | 3 |
| 2020 | Day and Night Collaborative Dynamic Mapping in Unstructured Environment Based on Multimodal SensorsabstractEnabling long-term operation during day and night for collaborative robots requires a comprehensive understanding of the unstructured environment. Besides, in the dynamic environment, robots must be able to recognize dynamic objects and collaboratively build a global map. This paper proposes a novel approach for dynamic collaborative mapping based on multimodal environmental perception. For each mission, robots first apply heterogeneous sensor fusion model to detect humans and separate them to acquire static observations. Then, the collaborative mapping is performed to estimate the relative position between robots and local 3D maps are integrated into a globally consistent 3D map. The experiment is conducted in the day and night rainforest with moving people. The results show the accuracy, robustness, and versatility in 3D map fusion missions. Yufeng Yue, Chule Yang, Jun Zhang 0042, Mingxing Wen, Zhenyu Wu 0001, Danwei Wang |
ICRA | 5 |
| 2020 | Infrastructure-Free Global Localization in Repetitive Environments: An OverviewabstractRepetitive environment is a challenging scenario for mobile robot global localization due to its highly similar structures and lack of distinctive features. Existing solutions in such environments rely heavily on pre-installed infrastructures, which are neither flexible nor cost-effective. Besides, few of the previous research have been focused on the implementation of infrastructure-free localization approaches in repetitive scenarios. Thus, this paper serves as a survey to investigate the problem of infrastructure-free mobile robot global localization with low-cost and efficient sensors in repetitive environments. Three of the most popular infrastructure-free localization methods, namely LiDAR-based localization (LBL), vision-based localization (VBL), and magnetic field-based localization (MFL), are analyzed and evaluated. Extensive global localization experiments are conducted in real-world repetitive scenarios and the results demonstrate that VBL methods perform slightly better than LBL and MFL methods. The overall evaluations indicate that infrastructure-free global localization in repetitive environment is still a challenging problem which deserves more research efforts to develop new solutions. Zhenyu Wu 0001, Jun Zhang 0042, Yufeng Yue, Mingxing Wen, Zichen Jiang, Danwei Wang |
IECON | 1 |
| 2020 | Collaborative Semantic Perception and Relative Localization Based on Map MatchingabstractIn order to enable a team of robots to operate successfully, retrieving accurate relative transformation between robots is the fundamental requirement. So far, most research on relative localization mainly focus on geometry features such as points, lines and planes. To address this problem, collaborative semantic map matching is proposed to perform semantic perception and relative localization. This paper performs semantic perception, probabilistic data association and nonlinear optimization within an integrated framework. Since the voxel correspondence between partial maps is a hidden variable, a probabilistic semantic data association algorithm is proposed based on Expectation-Maximization. Instead of specifying hard geometry data association, semantic and geometry association are jointly updated and estimated. The experimental verification on Semantic KITTI benchmarks demonstrate the improved robustness and accuracy. Yufeng Yue, Mingxing Wen, Zhenyu Wu 0001, Danwei Wang |
IROS | 4 |