Peng Yin 0001

dblp:23/4378-1 · DBLP profile ↗
← Back
21ranked-venue papers
12as first author
14since 2021 · last 2025
0000-0001-8398-0988ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 4 first-author · 3 since 2021Systems, architecture and hardware · 9 · 4 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 6 first-author · 6 since 2021Security and privacy · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2025 DTSG-Net: Dynamic Time Series Graph Neural Network and Its Application in Modulation Recognition
abstract
Modulation recognition of communication signals is of great importance in the context of the Internet of Everything (IoE), as wireless communication technology is a key foundation for implementing the IoE. Recently, graph neural networks (GNNs) have been successfully applied to modulation recognition tasks due to their ability to merge messages transmitted between adjacent nodes in the graph. However, GNN-based models are more computationally intensive when processing long signals, potentially reducing their practicality. In this article, we explore a novel signal representation from a graph perspective and propose a graph-powered modulation recognition framework. We first propose the dynamic time series graph (DTSG) algorithm, which segments the signals and maps each segment into a patch graph, with corresponding patches from different signals sharing connected edges. By integrating DTSG with both GNNs and recurrent neural networks (RNNs), we have designed an end-to-end signal classification framework, DTSG-Net, for modulation recognition. Experimental results on four datasets: 1) RML2016.10a; 2) RML2018.01a; 3) Sig2019-12; and 4) HKDD_AMC36—demonstrate that our DTSG-Net can achieve high signal modulation classification accuracy (Acc) with minimal computational resources, outperforming existing methods based on signal graph representation in terms of computational resource savings and higher accuracy.
Peng Yin 0001, Jinchao Zhou, Yizheng Ge, Zhuangzhi Chen
IEEE Internet Things J.1
2025 General Place Recognition Survey: Toward Real-World Autonomy
abstract
In the realm of robotics, the quest for achieving real-world autonomy, capable of executing large-scale and long-term operations, has positioned place recognition (PR) as a cornerstone technology. Despite the PR community's remarkable strides over the past two decades, garnering attention from fields like computer vision and robotics, the development of PR methods that sufficiently support real-world robotic systems remains a challenge. This article aims to bridge this gap by highlighting the crucial role of PR within the framework of simultaneous localization and mapping 2.0. This new phase in robotic navigation calls for scalable, adaptable, and efficient PR solutions by integrating advanced artificial intelligence technologies. For this goal, we provide a comprehensive review of the current state-of-the-art advancements in PR, alongside the remaining challenges, and underscore its broad applications in robotics. This article begins with an exploration of PR's formulation and key research challenges. We extensively review literature, focusing on related methods on place representation and solutions to various PR challenges. Applications showcasing PR's potential in robotics, key PR datasets, and open-source libraries are discussed.
Peng Yin 0001, Jianhao Jiao, Guoquan Huang 0001, Howie Choset, Sebastian A. Scherer, Jianda Han
IEEE Trans. Robotics1
2025 iLoc: An Adaptive, Efficient, and Robust Visual Localization System
abstract
In this article, we introduceiLoc, an innovative visual localization system designed to enhance the autonomy and adaptability of robotic agents in long-term and large-scale applications.iLocspecializes in: 1) extracting stable and consistent descriptors for place recognition, unaffected by changes in viewpoint and illumination; 2) performing swift and precise global relocalization to establish a robot's position within a large and complex environment; and 3) generating real-time tracking trajectories aligned with reference maps, ensuring continual orientation within known spaces. Distinctively,iLocincorporates a transformer-based learning module and an attention-enhanced recognition approach, enabling it to adapt to diverse environmental and viewpoint conditions.iLocleverages a coarse-to-fine global feature matching technique for enhanced localization and integrates robust state estimation combining visual odometry and loop closures through local refinement and pose graph optimization.iLocdemonstrates remarkable proficiency in place recognition, achieving localization over distances of up to 2 km within 0.5 s with average accuracy at 1 m. It maintains stable localization accuracy, even under variable conditions. Its versatile design allows integration across various environments, significantly broadening the scope of universal localization capabilities in robotics.iLocrepresents a substantial step forward in visual-based localization systems, delivering unparalleled speed and accuracy in place recognition. Its ability to adapt and respond to diverse environmental stimuli marks it as a crucial tool in advancing the field of robotic localization.
Peng Yin 0001, Jing Wang 0193, Ruohai Ge, Jianmin Ji, Yeping Hu, Huaping Liu 0001, Jianda Han
IEEE Trans. Robotics1
2024 CAFE: Robust Detection of Malicious Macro based on Cross-modal Feature Extraction
abstract
The detection of malicious macros has been a prominent focus of research. Previous approaches exhibit two notable shortcomings. Firstly, methods centered on document and macro code features often fall short in effectively countering targeted adversarial strategies. Secondly, detection techniques relying on deceptive information, such as visual and textual cues, although alleviating certain challenges, introduce a new vulnerability to adversarial machine learning techniques. In this paper, we present Collaborative Adaptive Feature Extraction method (CAFE), designed for robust detection based on deceptive information. The core of CAFE is a feature fusion network architecture, where modality-shared associations and modalityprivate information are modeled from feature of different modalities, resulting in independently valid and comprehensive feature representations. An adaptive feature sampling module is introduced to address partial feature absence, enhancing detection robustness. Experimental results, conducted on two datasets, demonstrate that CAFE adeptly captures shared and complementary information from two modalities, showcasing its capability for robust malicious macro detection in the presence of input noise and adversarial samples. Index Terms—Malicious Macro Detection, Multi-modal Features, Model Robustness, Security Wen Wang is corresponding author.
Huaifeng Bao, Xingyu Wang 0003, Wenhao Li 0005, Jinpeng Xu, Peng Yin 0001, Wen Wang 0008, Feng Liu 0001
CSCWD5
2024 LF-3PM: a LiDAR-based Framework for Perception-aware Planning with Perturbation-induced Metric
abstract
Just as humans can become disoriented in featureless deserts or thick fogs, not all environments are conducive to the Localization Accuracy and Stability (LAS) of autonomous robots. This paper introduces an efficient framework designed to enhance LiDAR-based LAS through strategic trajectory generation, known as Perception-aware Planning. Unlike vision-based frameworks, the LiDAR-based requires different considerations due to unique sensor attributes. Our approach focuses on two main aspects: firstly, assessing the impact of LiDAR observations on LAS. We introduce a perturbation-induced metric to provide a comprehensive and reliable evaluation of LiDAR observations. Secondly, we aim to improve motion planning efficiency. By creating a Static Observation Loss Map (SOLM) as an intermediary, we logically separate the time-intensive evaluation and motion planning phases, significantly boosting the planning process. In the experimental section, we demonstrate the effectiveness of the proposed metrics across various scenes and the feature of trajectories guided by different metrics. Ultimately, our framework is tested in a real-world scenario, enabling the robot to actively choose topologies and orientations preferable for localization. The source code is accessible at https://github.com/ZJU-FAST-Lab/LF-3PM.
Kaixin Chai, Long Xu 0002, Qianhao Wang, Chao Xu 0001, Peng Yin 0001, Fei Gao 0011
IROS5
2024 A LLM-based agent for the automatic generation and generalization of IDS rules
abstract
Cyberattacks on digital services and Internet of Things (IoT) are rising, employing complex tactics. Using intrusion detection systems (IDS) to detect and counter threats at key network points is vital for strong cybersecurity. Traditional rule-based network IDS rely on predefined rules, which may not effectively recognize the myriad complex variants of potential attacks. AI-driven methods for detecting malicious traffic offer enhanced capabilities but can fall short in terms of interpretability and performance under high-throughput network conditions. To address these challenges, we propose a LLM-based (Large Language Model) agent that utilizes multiple sources inputs to generate and generalize rules. The generated rules are designed to detect a variety of corresponding malicious threats, while the generalized rules are crafted to identify similar variant attacks. We have amassed an extensive dataset, comprising vulnerability security reports, malicious traffic, and original IDS rules from authoritative sources, which serve as input for the LLM-based agent. Subsequently, comparative experiments were conducted to assess the performance of the new rules in detecting malicious traffic. The experimental results demonstrate the superior performance of these new rules across various metrics for malicious traffic detection.
Haoning Chen, Huaifeng Bao, Wen Wang 0008, Feng Liu 0001, Guoqiao Zhou, Peng Yin 0001
TrustCom7
2024 Analysis on dendritic deep learning model for AMR task
abstract
Abstract This study introduces a novel hybrid deep learning model featuring a dendritic layer for enhancing the performance of automatic modulation recognition (AMR). By replacing the fully connected layer, the proposed model demonstrates superior classification accuracy in AMR tasks. Comparative experiments with nine state-of-the-art deep learning models on the RadioML2016.10a dataset reveal its consistent superiority. Statistical analyses, including the Friedman test and Wilcoxon signed-rank test, confirm the significant advantage of the HDM-D model.
Peng Yin 0001, Sanli Zhu, Zhuangzhi Chen
Cybersecur.1
2023 360FusionNeRF: Panoramic Neural Radiance Fields with Joint Guidance
abstract
Based on the neural radiance fields (NeRF), we present a pipeline for generating novel views from a single 360° panoramic image. Prior research relied on the neighborhood interpolation capability of multi-layer perceptions to complete missing regions caused by occlusion. This resulted in artifacts in their predictions. We propose 360FusionNeRF, a semi-supervised learning framework that employs geometric supervision and semantic consistency to guide the progressive training process. Firstly, the input image is reprojected to 360° images, and depth maps are extracted at different camera positions. In addition to the NeRF color guidance, the depth supervision enhances the geometry of the synthesized views. Furthermore, we include a semantic consistency loss that encourages realistic renderings of novel views. We extract these semantic features using a pre-trained visual encoder CLIP, a Vision Transformer (ViT) trained on hundreds of millions of diverse 2D photographs mined from the web with natural language supervision. Experiments indicate that our proposed method is capable of producing realistic completions of unobserved regions while preserving the features of the scene. 360FusionNeRF consistently delivers state-of-the-art performance when transferring to synthetic Structured3D dataset (PSNR ~ 5%, SSIM ~3% LPIPS ~13%), real-world Matterport3D dataset (PSNR ~3%, SSIM ~3% LPIPS ~9%) and Replica360 dataset (PSNR ~8%, SSIM ~2% LPIPS ~18%). We provide the source code at https://github.com/MetaSLAM/360FusionNeRF.
Peng Yin 0001, Sebastian A. Scherer
IROS2
2023 BioSLAM: A Bioinspired Lifelong Memory System for General Place Recognition
abstract
We present BioSLAM, a lifelong (lifelong simultaneous localization and mapping) SLAM framework for learning various new appearances incrementally and maintaining accurate place recognition for previously visited areas. Unlike humans, artificial neural networks suffer from catastrophic forgetting and may forget the previously visited areas when trained with new arrivals. For humans, researchers discover that there exists a memory replay mechanism in the brain to keep the neuron active for previous events. Inspired by this discovery, BioSLAM designs a gated generative replay to control the robot's learning behavior based on the feedback rewards. Specifically, BioSLAM provides a novel dual-memory mechanism for the maintenance of: 1) a dynamic memory to efficiently learn new observations; and 2) a static memory to balance new–old knowledge. When the agent is encountered with different appearances under new domains, the complete processing pipeline can help to incrementally update the place recognition ability, robust to the increasing complexity of long-term place recognition. We demonstrate BioSLAM in three incremental SLAM scenarios as follows. 1) A 120 km city-scale trajectories with LiDAR-based inputs. 2) A multivisited 4.5 km campus-scale trajectories with LiDAR-vision inputs. 3) An official Oxford dataset with 10 km visual inputs under different environmental conditions. We show that BioSLAM can incrementally update the agent's place recognition ability and outperform the state-of-the-art incremental approach, generative replay, by 24% in terms of place recognition accuracy. To the best of our knowledge, BioSLAM is the first memory-enhanced lifelong SLAM system to help incremental place recognition in long-term navigation tasks.
Peng Yin 0001, Abulikemu Abuduweili, Changliu Liu, Sebastian A. Scherer
IEEE Trans. Robotics1
2023 iSimLoc: Visual Global Localization for Previously Unseen Environments With Simulated Images
abstract
The camera is an attractive device for use in beyond visual line of sight drone operation since cameras are low in size, weight, power, and cost. However, state-of-the-art visual localization algorithms have trouble matching visual data that have significantly different appearances due to changes in illumination or viewpoint. This article presents iSimLoc, a learning-based global relocalization approach that is robust to appearance and viewpoint differences. The features learned by iSimLoc's place recognition network can be utilized to match query images to reference images of a different stylistic domain and viewpoint. In addition, our hierarchical global relocalization module searches in a coarse-to-fine manner, allowing iSimLoc to perform fast and accurate pose estimation. We evaluate our method on a dataset with appearance variations and a dataset that focuses on demonstrating large-scale matching over a long flight over complex terrain. iSimLoc achieves 88.7% and 83.8% successful retrieval rates on our two datasets, with 1.5 s inference time, compared to 45.8% and 39.7% using the next best method. These results demonstrate robust localization in a range of environments and conditions.
Peng Yin 0001, Ivan Cisneros, Ji Zhang 0003, Howie Choset, Sebastian A. Scherer
IEEE Trans. Robotics1
2023 AutoMerge: A Framework for Map Assembling and Smoothing in City-Scale Environments
abstract
In the era of advancing autonomous driving and increasing reliance on geospatial information, high-precision mapping not only demands accuracy but also flexible construction. Current approaches mainly rely on expensive mapping devices, which are time consuming for city-scale map construction and vulnerable to erroneous data associations without accurate GPS assistance. In this article, we present AutoMerge, a novel framework for merging large-scale maps that surpasses these limitations, which: 1) provides robust place recognition performance despite differences in both translation and viewpoint; 2) is capable of identifying and discarding incorrect loop closures caused by perceptual aliasing; and 3) effectively associates and optimizes large-scale and numerous map segments in the real-world scenario. AutoMerge utilizes multiperspective fusion and adaptive loop closure detection for accurate data associations, and it uses incremental merging to assemble large maps from individual trajectory segments given in random order and with no initial estimations. Furthermore, AutoMerge performs pose graph optimization after assembling the segments to smooth the merged map globally. We demonstrate AutoMerge on both city-scale merging (120 km) and campus-scale repeated merging (4.5 km × 8). The experiments show that AutoMerge: 1) surpasses the second- and third-best methods by 0.9% and 6.5% recall in segment retrieval; 2) achieves comparable 3-D mapping accuracy for 120-km large-scale map assembly; and 3) and is robust to temporally spaced revisits. To our knowledge, AutoMerge is the first mapping approach to merge hundreds of kilometers of individual segments without using GPS.
Peng Yin 0001, Haowen Lai, Ruohai Ge, Ji Zhang 0003, Howie Choset, Sebastian A. Scherer
IEEE Trans. Robotics1
2022 Modeling and Prediction of User Stability and Comfortability on Autonomous Wheelchairs With 3-D Mapping
abstract
Traditional manual wheelchairs have a fixed seat with no movement or angle adjustment, which can seriously affect the user's comfort and greatly limit user experience. However, the electric wheelchair relies on strong intelligence and automatic features; it can not only realize the multidegree freedom adjustment of the human body and the seat but also has a rich and powerful man–machine control interface, which greatly facilitates and improves the user experience. This study upgraded a Permobil C400-powered wheelchair with multisensor data fusion technology to enrich its terrain recognition, tipping stability, and comfortability prediction. The tipping stability modeling of the wheelchair dummy system is carried out using multibody dynamics and vibration mechanics to obtain the tipping stability limit and the comfort evaluation of the wheelchair vibration acceleration on the human body during travel. Based on the elevation mapping method, the wheelchair can estimate the terrain from the local point of view at any point in time. At the same time, the RGB-D depth camera is connected to the robot operating system (ROS) system, and the open-source algorithm package RTAB-MAP is used to complete the MAP construction and collect the 3-D point-cloud terrain data. Then, the real 3-D terrain files are generated through the point-cloud stitching technology for stability simulation of the wheelchair–human system. The tipping stability and comfort indexes of the wheelchair–human system when passing over different physical terrains can be obtained. The experimental results show that the IMU data located on the human chest agree well with the simulation analysis data and are suitable for a variety of complex real-terrain conditions, verifying the accuracy of the wheelchair–human system dynamics model and the feasibility of the simulation analysis process. Thus, this modeling and simulation method can predict wheelchair stability and user comfortability well and ensure a high-performance experience.
Zongming Yang, Peng Yin 0001, Johnell O. Brooks, Bing Li 0008
IEEE Trans. Hum. Mach. Syst.3
2022 PSE-Match: A Viewpoint-Free Place Recognition Method With Parallel Semantic Embedding
abstract
Accurate localization on the autonomous driving cars is essential for autonomy and driving safety, especially for complex urban streets and search-and-rescue subterranean environments where high-accurate GPS is not available. However current odometry estimation may introduce the drifting problems in long-term navigation without robust global localization. The main challenges involve scene divergence under the interference of dynamic environments and effective perception of observation and object layout variance from different viewpoints. To tackle these challenges, we present PSE-Match, a viewpoint-free place recognition method based on parallel semantic analysis of isolated semantic attributes from 3D point-cloud models. Compared with the original point cloud, the observed variance of semantic attributes is smaller. PSE-Match incorporates a divergence place learning network to capture different semantic attributes parallelly through the spherical harmonics domain. Using both existing benchmark datasets and two in-field collected datasets, our experiments show that the proposed method achieves above 70% average recall with top one retrieval and above 95% average recall with top ten retrieval cases. And PSE-Match has also demonstrated an obvious generalization ability with limited training dataset.
Peng Yin 0001, Ziyue Feng, Anton Egorov, Bing Li 0008
IEEE Trans. Intell. Transp. Syst.1
2021 Improving Off-road Planning Techniques with Learned Costs from Physical Interactions
abstract
Autonomous ground vehicles have improved greatly over the past decades, but they still have their limitations when it comes to off-road environments. There is still a need for planning techniques that effectively handle physical interactions between a vehicle and its surroundings. We present a method of modifying a standard path planning algorithm to address these problems by incorporating a learned model to account for complexities that would be too hard to address manually. The model predicts how well a vehicle will be able to follow a potential plan in a given environment. These predictions are then used to assign costs to their associated paths, where the path predicted to be the most feasible will be output as the final path. This results in a planner that doesn't rely solely on engineered features to evaluate traversability of obstacles, and can also choose a better path based on an understanding of its own capability that it has learned from previous interactions. This modification was integrated into the Hybrid A* algorithm and experimental results demonstrated an improvement of 14.29% over the original version on a physical platform.
Matthew Sivaprakasam, Samuel Triest, Peng Yin 0001, Sebastian A. Scherer
ICRA4
2020 End-to-End 3D Point Cloud Learning for Registration Task Using Virtual Correspondences
abstract
3D Point cloud registration is still a very challenging topic due to the difficulty in finding the rigid transformation between two point clouds with partial correspondences, and it's even harder in the absence of any initial estimation information. In this paper, we present an end-to-end deep-learning based approach to resolve the point cloud registration problem. Firstly, the revised LPD-Net is introduced to extract features and aggregate them with the graph network. Secondly, the self-attention mechanism is utilized to enhance the structure information in the point cloud and the cross-attention mechanism is designed to enhance the corresponding information between the two input point clouds. Based on which, the virtual corresponding points can be generated by a soft pointer based method, and finally, the point cloud registration problem can be solved by implementing the SVD method. Comparison results in ModelNet40 dataset validate that the proposed approach reaches the state-of-the-art in point cloud registration tasks and experiment resutls in KITTI dataset validate the effectiveness of the proposed approach in real applications.
Huanshu Wei, Zhijian Qiao, Zhe Liu 0022, Chuanzhe Suo, Peng Yin 0001, Yueling Shen, Haoang Li, Hesheng Wang 0001
IROS5
2020 SeqSphereVLAD: Sequence Matching Enhanced Orientation-invariant Place Recognition
abstract
Human beings and animals are capable of recognizing places from a previous journey when viewing them under different environmental conditions (e.g., illuminations and weathers). This paper seeks to provide robots with a human-like place recognition ability using a new point cloud feature learning method. This is a challenging problem due to the difficulty of extracting invariant local descriptors from the same place under various orientation differences and dynamic obstacles. In this paper, we propose a novel lightweight 3D place recognition method, SeqSphereVLAD, which is capable of recognizing places from a previous trajectory regardless of the viewpoint and the temporary observation differences. The major contributions of our method lie in two modules: (1) the spherical convolution feature extraction module, which produces orientation-invariant local place descriptors, and (2) the coarse-to-fine sequence matching module, which ensures both accurate loop-closure detection and real-time performance. Despite the apparent simplicity, our proposed approach outperform the state-of-the-arts for place recognition under datasets that combine orientation and context differences. Compared with the arts, our method can achieve above 95% average recall for the best match with only 18% inference time of PointNet-based place recognition methods.
Peng Yin 0001, Fuying Wang, Anton Egorov, Jiafan Hou, Ji Zhang 0003, Howie Choset
IROS1
2019 LPD-Net: 3D Point Cloud Learning for Large-Scale Place Recognition and Environment Analysis
abstract
Point cloud based place recognition is still an open issue due to the difficulty in extracting local features from the raw 3D point cloud and generating the global descriptor, and it's even harder in the large-scale dynamic environments. In this paper, we develop a novel deep neural network, named LPD-Net (Large-scale Place Description Network), which can extract discriminative and generalizable global descriptors from the raw 3D point cloud. Two modules, the adaptive local feature extraction module and the graph-based neighborhood aggregation module, are proposed, which contribute to extract the local structures and reveal the spatial distribution of local features in the large-scale point cloud, with an end-to-end manner. We implement the proposed global descriptor in solving point cloud based retrieval tasks to achieve the large-scale place recognition. Comparison results show that our LPD-Net is much better than PointNetVLAD and reaches the state-of-the-art. We also compare our LPD-Net with the vision-based solutions to show the robustness of our approach to different weather and light conditions.
Zhe Liu 0022, Shunbo Zhou, Chuanzhe Suo, Peng Yin 0001, Wen Chen 0021, Hesheng Wang 0001, Haoang Li, Yun-Hui Liu 0001
ICCV4
2019 A Multi-modal Sensor Array for Safe Human-Robot Interaction and Mapping
abstract
In the future, human-robot interaction will include collaboration in close-quarters where the environment geometry is partially unknown. As a means for enabling such interaction, this paper presents a multi-modal sensor array capable of contact detection and localization, force sensing, proximity sensing, and mapping. The sensor array integrates Hall effect and time-of-flight (ToF) sensors in an I2C communication network. The design, fabrication, and characterization of the sensor array for a future in-situ collaborative continuum robot are presented. Possible perception benefits of the sensor array are demonstrated for accidental contact detection, mapping of the environment, selection of admissible zones for bracing, and constrained motion control of the end effector while maintaining a bracing constraint with an admissible rolling motion.
Colette Abah, Andrew L. Orekhov, Garrison L. H. Johnston, Peng Yin 0001, Howie Choset, Nabil Simaan
ICRA4
2019 MRS-VPR: a multi-resolution sampling based global visual place recognition method
abstract
Place recognition and loop closure detection are challenging for long-term visual navigation tasks. SeqSLAM is considered to be one of the most successful approaches to achieve long-term localization under varying environmental conditions and changing viewpoints. SeqSLAM uses a brute-force sequential matching method, which is computationally intensive. In this work, we introduce a multi-resolution sampling-based global visual place recognition method (MRS-VPR), which can significantly improve the matching efficiency and accuracy in sequential matching. The novelty of this method lies in the coarse-to-fine searching pipeline and a particle filter-based global sampling scheme, that can balance the matching efficiency and accuracy in the long-term navigation task. Moreover, our model works much better than SeqSLAM when the testing sequence is over a much smaller time scale than the reference sequence. Our experiments demonstrate that MRSVPR is efficient in locating short temporary trajectories within long-term reference ones without compromising on the accuracy compared to SeqSLAM.
Peng Yin 0001, Rangaprasad Arun Srivatsan, Xueqian Li, Hongda Zhang, Lu Li 0018, Zhenzhong Jia, Jianmin Ji
ICRA1
2019 A Multi-Domain Feature Learning Method for Visual Place Recognition
abstract
Visual Place Recognition (VPR) is an important component in both computer vision and robotics applications, thanks to its ability to determine whether a place has been visited and where specifically. A major challenge in VPR is to handle changes of environmental conditions including weather, season and illumination. Most VPR methods try to improve the place recognition performance by ignoring the environmental factors, leading to decreased accuracy decreases when environmental conditions change significantly, such as day versus night. To this end, we propose an end-to-end conditional visual place recognition method. Specifically, we introduce the multi-domain feature learning method (MDFL) to capture multiple attribute-descriptions for a given place, and then use a feature detaching module to separate the environmental condition-related features from those that are not. The only label required within this feature learning pipeline is the environmental condition. Evaluation of the proposed method is conducted on the multi-season NORDLAND dataset, and the multi-weather GTAV dataset. Experimental results show that our method improves the feature robustness against variant environmental conditions.
Peng Yin 0001, Xueqian Li, Yingli Li, Rangaprasad Arun Srivatsan, Lu Li 0018, Jianmin Ji
ICRA1
2018 Stabilize an Unsupervised Feature Learning for LiDAR-based Place Recognition
abstract
Place recognition is one of the major challenges for the LiDAR-based effective localization and mapping task. Traditional methods are usually relying on geometry matching to achieve place recognition, where a global geometry map need to be restored. In this paper, we accomplish the place recognition task based on an end-to-end feature learning framework with the LiDAR inputs. This method consists of two core modules, a dynamic octree mapping module that generates local 2D maps with the consideration of the robot's motion; and an unsupervised place feature learning module which is an improved adversarial feature learning network with additional assistance for the long-term place recognition requirement. More specially, in place feature learning, we present an additional Generative Adversarial Network with a designed Conditional Entropy Reduction module to stabilize the feature learning process in an unsupervised manner. We evaluate the proposed method on the Kitti dataset and North Campus Long-Term LiDAR dataset. Experimental results show that the proposed method outperforms state-of-the-art in place recognition tasks under long-term applications. What's more, the feature size and inference efficiency in the proposed method are applicable in real-time performance on practical robotic platforms.
Peng Yin 0001, Zhe Liu 0022, Lu Li 0018, Hadi Salman, Weiliang Xu 0001, Hesheng Wang 0001, Howie Choset
IROS1