Guohao Peng

dblp:254/2093 · DBLP profile ↗
← Back
25ranked-venue papers
5as first author
24since 2021 · last 2026
0000-0001-9967-4934ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 5 first-author · 19 since 2021Systems, architecture and hardware · 12 · 3 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021
YearPublicationVenuePosition
2026 Variational Bayesian Multi-Output Gaussian Process Regression for Metabolic Profiles Prediction With Microbiome Data
abstract
Understanding the pivotal role of the human microbiome in health necessitates accurate metabolite prediction, which is crucial for unraveling the intricate interplay between the gut microbiome and human health. This study introduces an innovative approach, Variational Bayesian Multi-Output Gaussian Process Regression (VBMOGPR), to address the challenges posed by the complex, high-dimensional nature of microbiome data. VBMOGPR predicts microbial metabolites, quantifies the model confidence, and incorporates uncertainty estimates. Employing a Bayesian framework with Automatic Relevance Determination (ARD) for feature selection enhances interpretability and performance. Comparative analysis across 14 datasets within a meta-database demonstrated the superiority of VBMOGPR, marking a significant advancement in metabolite prediction and its implications for microbiome impact on human health. In addition, we confirmed that VBMOGPR could tap the potential microbial metabolic association.
Qinghui Weng, Mingyi Hu, Guohao Peng, Wenwei Lu, Jinlin Zhu
IEEE Trans. Comput. Biol. Bioinform.3
2025 Overlapping Free: Anchorless UWB-Assisted Relative Pose Estimation for Multi-Robot Systems
abstract
Accurate Relative Pose Estimation (RPE) is critical for effective collaboration of multi-robot systems. Traditional methods using cameras or LiDARs heavily rely on overlapping Fields of View (FoV) between robots, which is highly demanding in practical applications and may hinder collaboration efficiency. To accommodate this issue, we propose Anchorless UWB-Assisted Relative Pose Estimation (AURPE), a novel approach that leverages ultra-wideband (UWB) technology in an anchorless setup to achieve multi-robot RPE without requiring overlapping FoVs or external infrastructure. AURPE first estimates the initial relative poses between robots using inter-robot UWB ranging combined with a Bayesian framework and constrained optimization. During robot operation, AURPE continuously refines the relative poses by integrating UWB measurements with LiDAR-inertial odometry (LIO) and employs a consensus voting mechanism to identify the most reliable pose estimates. Additionally, a pose graph-based backend optimization is incorporated to enhance the accuracy of both initial and real-time relative pose. Extensive simulations and real-world experiments demonstrate that AURPE achieves accurate RPE even in non-overlapping scenarios where traditional methods fail. Compared to state-of-the-art point cloud registration methods, AURPE shows superior performance in both accuracy and robustness, highlighting its potential to significantly enhance cooperative tasks in multi-robot systems operating in complex environments.
Yanpu Yun, Guohao Peng, Jun Zhang 0042, Yiyao Liu, Kaimin Mao, Danwei Wang
ICRA2
2025 LCSPose: Efficient, Accurate and Scalable Markerless 6-DoF Pose Estimation of a Quay Crane Spreader Based on LiDAR and Camera
abstract
Accurate Six Degrees of Freedom (6-DoF) pose estimation of Ship-To-Shore (STS) quay crane spreaders is crucial for ensuring safe and efficient container handling in port automation. However, existing pose estimation techniques face significant challenges, as camera-based systems either rely on markers, which are prone to damage, or struggle with depth estimation inaccuracies. Additionally, 3D sensor-based approaches, particularly point cloud registration (PCR), face challenges such as initial pose errors, high-latency inference, and difficulties in object identification based purely on geometric features. To address these limitations, we propose LCSPose, a LiDAR-camera fusion-based 6-DoF pose estimation method that is marker-free, accurate, efficient, and scalable. Our approach integrates three key modules: (1) a semantic-geometric segmentation module for spreader segmentation and outlier removal, (2) a spatial consistency template sampling module based on Spatial Consistency Score (SC-Score) for reliable template selection across varying distances, and (3) a multi-view coarse-to-fine pose refinement module which incorporates multi-view PCA alignment for robust initial posture prior estimation and iterative pose refinement strategy for long-range registration. Our method demonstrates a 60% improvement in registration recall over state-of-the-art (SOTA) PCR methods, achieving up to 6 cm in translation error and 0.19 degrees in rotation error, while maintaining real-time processing at 20Hz.
Jun Zhang 0042, Guohao Peng, Yanpu Yun, Yiyao Liu, Yuanzhe Wang, Danwei Wang
ICRA3
2025 DMoVGPE: predicting gut microbial associated metabolites profiles with deep mixture of variational Gaussian Process experts
abstract
BACKGROUND: Understanding the metabolic activities of the gut microbiome is vital for deciphering its impact on human health. While direct measurement of these metabolites through metabolomics is effective, it is often expensive and time-consuming. In contrast, microbial composition data obtained through sequencing is more accessible, making it a promising resource for predicting metabolite profiles. However, current computational models frequently face challenges related to limited prediction accuracy, generalizability, and interpretability. METHOD: Here, we present the Deep Mixture of Variational Gaussian Process Experts (DMoVGPE) model, designed to overcome these issues. DMoVGPE utilizes a dynamic gating mechanism, implemented through a neural network with fully connected layers and dropout for regularization, to select the most relevant Gaussian Process experts. During training, the gating network refines expert selection, dynamically adjusting their contribution based on the input features. The model also incorporates an Automatic Relevance Determination (ARD) mechanism, which assigns relevance scores to microbial features by evaluating their predictive power. Features linked to metabolite profiles are given smaller length scales to increase their influence, while irrelevant features are down-weighted through larger length scales, improving both prediction accuracy and interpretability. CONCLUSIONS: Through extensive evaluations on various datasets, DMoVGPE consistently achieves higher prediction performance than existing models. Furthermore, our model reveals significant associations between specific microbial taxa and metabolites, aligning well with findings from existing studies. These results highlight DMoVGPE's potential to provide accurate predictions and to uncover biologically meaningful relationships, paving the way for its application in disease research and personalized healthcare strategies.
Qinghui Weng, Mingyi Hu, Guohao Peng, Jinlin Zhu
BMC Bioinform.3
2025 From ensemble to knowledge distillation: Improving large-scale food recognition
Liming Nong, Guohao Peng, Jinlin Zhu
Eng. Appl. Artif. Intell.2
2025 Curb-Tracker: An Integrated Curb Following System for Autonomous Vehicles
Yuanzhe Wang, Guohao Peng, Zhenyu Wu 0001, Danwei Wang
IEEE Trans. Robotics3
2024 TransLoc4D: Transformer-Based 4D Radar Place Recognition
abstract
Place recognition is crucial for unmanned vehicles in terms of localization and mapping. Recent years have witnessed numerous explorations in the field, where 2D cameras and 3D LiDARs are mostly employed. Despite their admirable performance, they may encounter challenges in adverse weather such as rain and fog. Hopefully, 4D millimeter-wave radar emerges as a promising alternative, as its longer wavelength makes it virtually immune to interference from tiny particles of fog and rain. Therefore, in this work, we propose a novel 4D radar place recognition model, TransLoc4D, based on sparse convolutions and Transformer structures. Specifically, a MinkLoc4D back-bone is first proposed to leverage the multimodal information from 4D radar scans. Rather than merely capturing geometric structures of point clouds, MinkLoc4D additionally explores their intensity and velocity properties. After feature extraction, a Transformer layer is introduced to enhance local features before aggregation, where linear self-attention captures the long-range dependencies of the point cloud, alleviating its sparsity and noise. To validate TransLoc4D, we construct two datasets and set up benchmarks for 4D radar place recognition. Experiments vali-date the feasibility of TransLoc4D and demonstrate it can robustly deal with dynamic and adverse environments.
Guohao Peng, Heshan Li, Jun Zhang 0042, Zhenyu Wu 0001, Pengyu Zheng, Danwei Wang
CVPR1
2024 Cross-View Detection of Crowded Objects Based on Multi-Sensor Fusion
abstract
Traditional object detection methods are limited by single-sensor constraints, high computational requirements, and poor real-time performance. In addition, occlusion often occurs under the condition of restricted single-view. In this paper, we introduce a camera and LiDAR fusion-based object detection method, which achieves excellent detection performance under limited computational resources. We also explores a fusion detection method deployed with multi-view, which can effectively solve the occlusion issue encountered by single view. The proposed method is valuable for single view as well as multi-view in various application scenarios. Our fusion method significantly improves detection accuracy and reliability, and solves the problems of data discrepancy, interference between sensors, and occlusion due to restricted view. Simulations and extensive experiments show that our proposed object detection method exhibited high accuracy and relatively low computational time.
Zhipeng Gu, Guohao Peng, Yanpu Yun, Yiyao Liu, Zhenyu Wu 0001, Jun Zhang 0042, Xudong Suo, Danwei Wang
ICARCV2
2024 ACS-MM-Explore: Adaptive Circular Search Strategy for Multi-Modal Robot Exploration in Large-Scale Urban Environments
abstract
Autonomous exploration has become a crucial technology for mobile robots, and numerous broadly applicable algorithms have emerged. However, few exploration methods effectively utilize the features of specified types of areas to enhance the efficiency of autonomous exploration in a complex environment. In this paper, we propose ACS-MM-Explore, an adaptive-circular-search-based exploration framework for large-scale urban road environments, focusing on extracting and utilizing the boundaries of roads to enhance exploration efficiency. Our approach integrates a multi-modal traversabil-ity analysis module to distinguish between road and non-traversable areas on a 2D costmap. A novel mechanism for gen-erating exploration viewpoints is introduced, efficiently creating exploration viewpoints with a circular search process with an adaptive radius. An optimized viewpoint selection mechanism is included, taking into account the geographical and geomet-ric information of each viewpoint. The framework extends the move base and TEB local planner as a viewpoint-based navigation module. A comprehensive evaluation is concluded in a simulation environment, demonstrating the framework's effectiveness and robustness.
Kaimin Mao, Mingxing Wen, Jun Zhang 0042, Guohao Peng, Zhenyu Wu 0001, Danwei Wang
ICARCV5
2024 LB-R2R-Calib: Accurate and Robust Extrinsic Calibration of Multiple Long Baseline 4D Imaging Radars for V2X
abstract
As a new sensor, 4D radar (x, y, z, velocity) has great potential for V2X, due to its 3D point cloud, direct doppler velocity output, long distance ranging, low-cost, and more importantly, robust perception in all weathers. However, the extrinsic calibration of multiple long baseline 4D radars is rarely researched in V2X, which is the key to fuse multi-radars. The main reasons are three-folds: (1) New sensor. Thus, it is not surprising that little related work can be found. (2) Long baseline and large viewpoint-difference. Current works are mainly focused on unmanned vehicles, which is short baseline and small viewpoint-difference. (3) Sparse, noisy, and very cluttered 4D radar point cloud. Thus, it is challenging to rapidly and accurately locate the target and extract the feature. In this paper, LB-R2R-Calib (Long Baseline Radar to Radar extrinsic Calibration) is proposed to address these problems. The novelties are: (1) A new target is introduced: an eight-quadrant corner reflector enclosed by a foam sphere. The benefit is the target center is a viewpoint-invariant feature. Thus, it is ideal for large viewpoint-difference calibration. (2) A new feature extraction algorithm is proposed to rapidly locate the target and extract the target center from a very cluttered point cloud, as we observed some important characteristics of 4D radar. Experiments with two 4D radars in real environments with four configurations demonstrate our method is highly accurate and robust.
Jun Zhang 0042, Fangwei Zhang, Zhenyu Wu 0001, Guohao Peng, Yiyao Liu, Qiyang Lyu, Mingxing Wen, Danwei Wang
ICRA5
2023 CAHIR: Co-Attentive Hierarchical Image Representations for Visual Place Recognition
abstract
Robust visual place recognition (VPR) against significant appearance changes is crucial for the life-long operation of mobile robots. Focusing on this task, we propose a Co-Attentive Hierarchical Image Representations (CAHIR) framework for VPR, which unifies attention-sharing global and local descriptor generation into one encoding pipeline. The hierarchical descriptors are applied to a coarse-to-fine VPR system with global retrieval and local geometric verification. To explore high-quality local matches between task-relevant visual elements, a cross-attention mutual enhancement layer is introduced to strengthen the information interaction between the local descriptors. Through the proposed selective matching distillation, the mutual enhancement layer can learn from state-of-the-art local matchers in a distillation manner. After weighted cross-matching of the enhanced local descriptors, geometric verification is applied to evaluate the spatial consistency of the compared image pair. Experiments show CAHIR outperforms the existing global and local representations for VPR in terms of performance and efficiency. Quantitatively, it achieves state-of-the-art results on three city-scale benchmark datasets. Qualitatively, CAHIR proves to attach great importance to task-relevant visual elements and excels at finding local correspondences that are discriminative to the VPR task.
Guohao Peng, Heshan Li, Jun Zhang 0042, Mingxing Wen, Singh Rahul, Danwei Wang
ICRA1
2023 4DRadarSLAM: A 4D Imaging Radar SLAM System for Large-scale Environments based on Pose Graph Optimization
abstract
LiDAR-based SLAM may easily fail in adverse weathers (e.g., rain, snow, smoke, fog), while mmWave Radar remains unaffected. However, current researches are primarily focused on 2D$(x,y)$or 3D ($x, y$, doppler) Radar and 3D LiDAR, while limited work can be found for 4D Radar ($x, y, z$, doppler). As a new entrant to the market with unique characteristics, 4D Radar outputs 3D point cloud with added elevation information, rather than 2D point cloud; compared with 3D LiDAR, 4D Radar has noisier and sparser point cloud, making it more challenging to extract geometric features (edge and plane). In this paper, we propose a full system for 4D Radar SLAM consisting of three modules: 1) Front-end module performs scan-to-scan matching to calculate the odometry based on GICP, considering the probability distribution of each point; 2) Loop detection utilizes multiple rule-based loop pre-filtering steps, followed by an intensity scan context step to identify loop candidates, and odometry check to reject false loop; 3) Back-end builds a pose graph using front-end odometry, loop closure, and optional GPS data. Optimal pose is achieved through$\mathrm{g}2\mathrm{o}$. We conducted real experiments on two platforms and five datasets (ranging from 240m to 4.8km) and will make the code open-source to promote further research at: https://github.com/zhuge2333/4DRadarSLAM
Jun Zhang 0042, Huayang Zhuge, Zhenyu Wu 0001, Guohao Peng, Mingxing Wen, Yiyao Liu, Danwei Wang
ICRA4
2023 AdaptSeqVPR: An Adaptive Sequence-Based Visual Place Recognition Pipeline
abstract
Visual Place Recognition (VPR) is essential for autonomous robots and unmanned vehicles, as an accurate identification of visited places can trigger a loop closure to optimize the built map. The most prevalent methods tackle VPR as a single-frame retrieval task, which uses a CNN-based encoder to describe and compare each individual frame. These methods, however, overlook the temporal information between frames. Other methods improve this by searching the database with consecutive frames, which can greatly reduce false positives. Nevertheless, current sequence-based methods typically assume the consecutive image frames to be captured at an approximately constant speed, which is not always the case in practice. Therefore, we propose an adaptive sequence search strategy (AdaptSeq), which can dynamically alter the step size of adjacent frames in the retrieved sequence trajectory. Furthermore, to address false positive retrieval of input frames, we propose a CNN-based discriminator named DDsNet. It can determine whether the top retrieved candidates are true positives based on the learned statistics rather than an artificial threshold. Overall, we construct a novel sequence-based VPR pipeline named AdaptSeqVPR. It utilizes a CNN-based encoder for frame descriptions, and encompasses AdaptSeq and DDsNet for sequence matching. The experimental results indicate that our AdaptSeqVPR exhibits superior performance compared to the baseline SeqSLAM and SeqVLAD. Notably, our method can robustly handle the sequence-based VPR for vehicles traveling at non-uniform speeds in changing environments.
Heshan Li, Guohao Peng, Jun Zhang 0005, Sriram Vaikundam, Danwei Wang
IROS2
2023 LB-L2L-Calib 2.0: A Novel Online Extrinsic Calibration Method for Multiple Long Baseline 3D LiDARs Using Objects
abstract
In V2X (Vehicle-to-Everything), one important work is to extrinsically calibrate multiple 3D LiDARs, which are mounted with a long baseline and large viewpoint-difference at the road-side. Current solutions either require a specific target being set up (e.g., a sphere), or require specific features existing in the environment (e.g., mutually orthogonal planes). However, it is time-consuming, sometimes even inconvenient, to set up specific targets, e.g., at busy intersections and highways. Furthermore, specific features do not always exist in the traffic scenario. Thus, the current solutions are not feasible. To address this problem, a novel extrinsic calibration method is proposed in this paper, namely LB-L2L-Calib 2.0. It is the 2.0 version of our previous work. The novelties are: 1) We propose to use the easily accessible objects on the road as features for calibration (i.e., the vehicles). Thus, it is not necessary to set up any specific targets and we do not need to worry whether specific features exist or not. The key point is we observed that the 3D bounding box centers of the vehicles are viewpoint-invariant from different viewpoints, which makes them ideal features for long baseline and large viewpoint-difference calibration. 2) To establish correct correspondence between the bounding box centers detected from different LiDARs, we propose an exhaustive searching strategy. It can robustly output correct correspondence. Extensive experiments are performed in three scenarios (simulation: intersection, real: carpark and highway), with two types of LiDAR (Velodyne and Livox), demonstrating that LB-L2L-Calib 2.0 is robust, effective, and accurate.
Jun Zhang 0042, Qiao Yan, Mingxing Wen, Qiyang Lyu, Guohao Peng, Zhenyu Wu 0001, Danwei Wang
IROS5
2023 IMOVNN: incomplete multi-omics data integration variational neural networks for gut microbiome disease prediction and biomarker identification
abstract
The gut microbiome has been regarded as one of the fundamental determinants regulating human health, and multi-omics data profiling has been increasingly utilized to bolster the deep understanding of this complex system. However, stemming from cost or other constraints, the integration of multi-omics often suffers from incomplete views, which poses a great challenge for the comprehensive analysis. In this work, a novel deep model named Incomplete Multi-Omics Variational Neural Networks (IMOVNN) is proposed for incomplete data integration, disease prediction application and biomarker identification. Benefiting from the information bottleneck and the marginal-to-joint distribution integration mechanism, the IMOVNN can learn the marginal latent representation of each individual omics and the joint latent representation for better disease prediction. Moreover, owing to the feature-selective layer predicated upon the concrete distribution, the model is interpretable and can identify the most relevant features. Experiments on inflammatory bowel disease multi-omics datasets demonstrate that our method outperforms several state-of-the-art methods for disease prediction. In addition, IMOVNN has identified significant biomarkers from multi-omics data sources.
Mingyi Hu, Jinlin Zhu, Guohao Peng, Wenwei Lu, Zhenping Xie
Briefings Bioinform.3
2022 C-TM: Topo-metric Mapping and Localization based on Place Categorization and Place Recognition for a Delivery Robot on Footpath
abstract
In this work, C-TM is presented: a method to build a topo-metric map for delivery robot navigation in largescale city environments. This system automatically generates a compact map by only saving expensive LIDAR information at key locations. These locations form the nodes of a topological map. Nodes are identified using a Place-Categorization (PC) neural network which output the place category from RGB cameras. Inside nodes, we generate and save high quality LIDAR submaps. Global localization within the map is done with a Visual-Place-Recognition (VPR) neural network. The topo-metric map can be used for navigation on footpath. We deploy C-TM on a four-wheeled autonomous delivery robot and test the effectiveness in two environments, both day and night.
Timothy Chia, Jun Zhang 0042, Heshan Li, Guohao Peng, Mingxing Wen, Dawei Kee, P. G. C. N. Senarathne
ICARCV4
2022 Vision Based Sidewalk Navigation for Last-mile Delivery Robot
abstract
Navigating delivery robot along the sidewalk safely and robustly in a campus environment is extremely challenging due to the narrow motion space, appearance changes and unstable GPS localization signal under canopies of trees, etc. To that end, we have completed a systematic implementation for delivery robot sidewalk navigation, where a robust vision based navigation algorithm has been proposed. And it consists of three main modules: sidewalk segmentation, costmap generation and motion planning. More Specifically, the first module is to find the drivable area of the surrounding environment, where an image-based segmentation neural network has been developed to extract where the robot can traverse. Since it only takes as input immediate and local sensory data, thus releasing the high dependence on a prior map. Then, an inverse perspective mapping follows to generate a bird-eye-view of the drivable area and constructs the local occupancy grid map intuitively. Next, two different motion planners, control-based primitives (Dynamic Window Approach) and state-based primitives (state lattice planner), have been adopted to generate a trajectory candidate for navigating the robot along the sidewalk. Both simulation and real-world sidewalk navigation experiments have been conducted to test and evaluate their performance. The results show that our algorithm can precisely extract the sidewalk area for traversing, and the state-based primitive planner demonstrates superior performance in terms of trajectory length and time cost, achieving 14.3% and 18.7% improvement compared with control-based primitive planner.
Mingxing Wen, Jun Zhang 0042, Tairan Chen, Guohao Peng, Timothy Chia, Yingchong Ma
ICARCV4
2022 LB-L2L-Calib: Accurate and Robust Extrinsic Calibration for Multiple 3D LiDARs with Long Baseline and Large Viewpoint Difference
abstract
Multi-LiDAR system is an important part of V2X (Vehicle to Everything) to enhance the perception information for unmanned vehicles. To fuse the information from multiple 3D LiDARs, accurate extrinsic calibration between the LiDARs is essential. However, the existing multi-LiDAR calibration methods mainly focus on short baseline scenarios, where multiple LiDARs are closely mounted on a single platform (e.g., an unmanned vehicle). Besides, most methods typically use a planar target for calibration. Some of the methods require the motion of the multi-LiDAR system. The above conditions severely limit the application of these methods to V2X, where LiDARs are non-movable, the baseline and viewpoint difference between the LiDARs can be very large. In order to meet these challenges, we propose an accurate and robust extrinsic calibration method for long baseline multi-LiDAR systems, named LB-L2L-Calib (Large Baseline LiDAR to LiDAR extrinsic Calibration). (1) We use a sphere as the calibration target for multiple LiDARs with large viewpoint difference, leveraging the viewpoint-invariance of the sphere. (2) A improved sphere detection and sphere center estimation strategy is introduced to detect and extract the sphere center from a cluttered point cloud in large-scale outdoor scenario. (3) A extrinsic parameter regression scheme is introduced. Both simulation and real experiments demonstrate that LB-L2L-Calib is highly accurate and robust. Quantitative results show that the rotation and translation error is less than 0.01m and 0.01° (in simulation, Gauss noise 0.03m, the distance and viewpoint difference between two LiDARs is more than 30m and 90°).
Jun Zhang 0042, Qiyang Lyu, Guohao Peng, Zhenyu Wu 0001, Qiao Yan, Danwei Wang
ICRA3
2022 SectionKey: 3-D Semantic Point Cloud Descriptor for Place Recognition
abstract
Place recognition is seen as a crucial factor to correct cumulative errors in Simultaneous Localization and Mapping (SLAM) applications. Most existing studies focus on visual place recognition, which is inherently sensitive to environmental changes such as illumination, weather and seasons. Considering these facts, more recent attention has been attracted to use 3-D Light Detection and Ranging (LiDAR) scans for place recognition, which demonstrates more credibility by exerting accurate geometric information. Different from pure geometric-based studies, this paper proposes a novel global descriptor, named SectionKey, which leverages both semantic and geometric information to tackle the problem of place recognition in large-scale urban environments. The proposed descriptor is robust and invariant to viewpoint changes. Specifically, the encoded three-layers key serves as a pre-selection step and a ‘candidate center’ selection strategy is deployed before calculating the similarity score, thus improving the accuracy and efficiency significantly. Then, a two-step semantic iterative closest point (ICP) algorithm is applied to acquire the 3-D pose (x, y, θ) that is used to align the candidate point clouds with the query frame and calculate the similarity score. Extensive experiments have been conducted on public Semantic KITTI dataset to demonstrate the superior performance of our proposed system over state-of-the-art baselines.
Shutong Jin, Zhenyu Wu 0001, Jun Zhang 0042, Guohao Peng, Danwei Wang
IROS5
2022 LSDNet: A Lightweight Self-Attentional Distillation Network for Visual Place Recognition
abstract
Visual Place Recognition (VPR) has become an indispensable capacity for mobile robots to operate in large-scale environments. Existing methods in this field mostly focus on exploring high-performance encoding strategies, while few attempts are devoted to lightweight models that balance per-formance and computational cost. In this work, we propose a Lightweight Self-attentional Distillation Network (LSDNet) aiming to obtain advantages of both performance and efficiency. (1) From a performance perspective, an attentional encoding strategy is proposed to integrate crucial information in the scene. It extends the NetVlad architecture with a self-attention module to facilitate non-local information interaction between local features. Through further visual word vector rescaling, the final image representation can benefit from both non-local spatial integration and cluster-wise weighting. (2) From an efficiency perspective, LSDNet is built upon a lightweight back-bone. To maintain comparable performance to large backbone models, a dual distillation strategy is introduced. It prompts LSDNet to learn both encoding patterns in the hidden space and feature distributions in the encoding space from the teacher model. Through distillation-augmented training, LSDNet is able to rival the teacher model and outperform SOTA global representations with the same lightweight backbone.
Guohao Peng, Heshan Li, Zhenyu Wu 0001, Danwei Wang
IROS1
2021 Attentional Pyramid Pooling of Salient Visual Residuals for Place Recognition
abstract
The core of visual place recognition (VPR) lies in how to identify task-relevant visual cues and embed them into dis- criminative representations. Focusing on these two points, we propose a novel encoding strategy named Attentional Pyramid Pooling of Salient Visual Residuals (APPSVR). It incorporates three types of attention modules to model the saliency of local features in individual, spatial and cluster dimensions respectively. (1) To inhibit task-irrelevant local features, a semantic-reinforced local weighting scheme is employed for local feature refinement; (2) To leverage the spatial context, an attentional pyramid structure is constructed to adaptively encode regional features according to their relative spatial saliency; (3) To distinguish the different importance of visual clusters to the task, a parametric normalization is proposed to adjust their contribution to image descriptor generation. Experiments demonstrate APPSVR outperforms the existing techniques and achieves a new state-of-the-art performance on VPR benchmark datasets. The visualization shows the saliency map learned in a weakly supervised manner is largely consistent with human cognition.
Guohao Peng, Jun Zhang 0042, Heshan Li, Danwei Wang
ICCV1
2021 Semantic Reinforced Attention Learning for Visual Place Recognition
abstract
Large-scale visual place recognition (VPR) is inherently challenging because not all visual cues in the image are beneficial to the task. In order to highlight the task-relevant visual cues in the feature embedding, the existing attention mechanisms are either based on artificial rules or trained in a thorough data-driven manner. To fill the gap between the two types, we propose a novel Semantic Reinforced Attention Learning Network (SRALNet), in which the inferred attention can benefit from both semantic priors and data-driven fine-tuning. The contribution lies in two-folds. (1) To suppress misleading local features, an interpretable local weighting scheme is proposed based on hierarchical feature distribution. (2) By exploiting the interpretability of the local weighting scheme, a semantic constrained initialization is proposed so that the local attention can be reinforced by semantic priors. Experiments demonstrate that our method outperforms state-of-the-art techniques on city-scale VPR benchmark datasets.
Guohao Peng, Yufeng Yue, Jun Zhang 0042, Zhenyu Wu 0001, Danwei Wang
ICRA1
2021 MSTSL: Multi-Sensor Based Two-Step Localization in Geometrically Symmetric Environments
abstract
Symmetric environment is one of the most intractable and challenging scenarios for mobile robots to accomplish global localization tasks, due to the highly similar geometrical structures and insufficient distinctive features. Existing localization solutions in such scenarios either depend on pre-deployed infrastructures which are expensive, inflexible, and hard to maintain; or rely on single sensor-based methods whose initialization module is incapable to provide enough unique information. Thus, this paper proposes a novel Multi-Sensor based Two-Step Localization framework named MSTSL, which addresses the problem of mobile robot global localization in geometrically symmetric environments by utilizing the measured magnetic field, 2-D LiDAR, and wheel odometry information. The proposed system mainly consists of two steps: 1) Magnetic Field-based Initialization, and 2) LiDAR-based Localization. Based on the pre-built magnetic field database, multiple initial hypotheses poses can firstly be determined by the proposed two-stage initialization algorithm. Then, utilizing the obtained multiple initial hypotheses, the robot can be localized more accurately by LiDAR-based localization. Extensive experiments demonstrate the practical utility and accuracy of the proposed system over the alternative approaches in real-world scenarios.
Zhenyu Wu 0001, Yufeng Yue, Mingxing Wen, Jun Zhang 0042, Guohao Peng, Danwei Wang
ICRA5
2021 Dual-Domain-Based Adversarial Defense With Conditional VAE and Bayesian Network
abstract
Adversarial examples can be imperceptible to human eyes but can easily fool deep models. Such intrigue property has raised security issues for real-world industrial deep learning systems. To combat those malicious attacks, a novel defense strategy has been proposed based on the conditional variational autoencoder (CVAE) and Bayesian network (BN). The main contribution lies in the provided systematic dual-domain-based defense framework, which covers three modules named detection, diagnosis, and recovery. Specifically, the CVAE is first introduced for latent- and residual-domain generation. Subsequently, a composite and hierarchical BN detector is proposed to conduct the adversary detection through feature validation and output justification. Afterwards, a diagnosis strategy has been constructed for residual domain and different attacks can be evaluated in the unified framework. Finally, a two-step recovery mechanism is established on the CVAE that can effectively restore the feature representations and the network predictions from various adversaries. The feasibility of the entire defense diagram has been extensively demonstrated on three real-world recognition problems.
Jinlin Zhu, Guohao Peng, Danwei Wang
IEEE Trans. Ind. Informatics2
2020 Conditional Gaussian Distribution Learning for Open Set Recognition
abstract
Deep neural networks have achieved state-of-the-art performance in a wide range of recognition/classification tasks. However, when applying deep learning to real-world applications, there are still multiple challenges. A typical challenge is that unknown samples may be fed into the system during the testing phase and traditional deep neural networks will wrongly recognize the unknown sample as one of the known classes. Open set recognition is a potential solution to overcome this problem, where the open set classifier should have the ability to reject unknown samples as well as maintain high classification accuracy on known classes. The variational auto-encoder (VAE) is a popular model to detect unknowns, but it cannot provide discriminative representations for known classification. In this paper, we propose a novel method, Conditional Gaussian Distribution Learning (CGDL), for open set recognition. In addition to detecting unknown samples, this method can also classify known samples by forcing different latent features to approximate different Gaussian models. Meanwhile, to avoid information hidden in the input vanishing in the middle layers, we also adopt the probabilistic ladder architecture to extract high-level abstract features. Experiments on several standard image datasets reveal that the proposed method significantly outperforms the baseline method and achieves new state-of-the-art results.
Xin Sun 0015, Zhenning Yang, Chi Zhang 0007, Keck Voon Ling, Guohao Peng
CVPR5