EDBT 2026 Demo / reviewers in the wild / expert
Heshan Li
dblp:25/560
· DBLP profile ↗
6ranked-venue papers
1as first author
6since 2021 · last 2024
0009-0007-3291-0112ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | TransLoc4D: Transformer-Based 4D Radar Place RecognitionabstractPlace recognition is crucial for unmanned vehicles in terms of localization and mapping. Recent years have witnessed numerous explorations in the field, where 2D cameras and 3D LiDARs are mostly employed. Despite their admirable performance, they may encounter challenges in adverse weather such as rain and fog. Hopefully, 4D millimeter-wave radar emerges as a promising alternative, as its longer wavelength makes it virtually immune to interference from tiny particles of fog and rain. Therefore, in this work, we propose a novel 4D radar place recognition model, TransLoc4D, based on sparse convolutions and Transformer structures. Specifically, a MinkLoc4D back-bone is first proposed to leverage the multimodal information from 4D radar scans. Rather than merely capturing geometric structures of point clouds, MinkLoc4D additionally explores their intensity and velocity properties. After feature extraction, a Transformer layer is introduced to enhance local features before aggregation, where linear self-attention captures the long-range dependencies of the point cloud, alleviating its sparsity and noise. To validate TransLoc4D, we construct two datasets and set up benchmarks for 4D radar place recognition. Experiments vali-date the feasibility of TransLoc4D and demonstrate it can robustly deal with dynamic and adverse environments. Guohao Peng, Heshan Li, Jun Zhang 0042, Zhenyu Wu 0001, Pengyu Zheng, Danwei Wang |
CVPR | 2 |
| 2023 | CAHIR: Co-Attentive Hierarchical Image Representations for Visual Place RecognitionabstractRobust visual place recognition (VPR) against significant appearance changes is crucial for the life-long operation of mobile robots. Focusing on this task, we propose a Co-Attentive Hierarchical Image Representations (CAHIR) framework for VPR, which unifies attention-sharing global and local descriptor generation into one encoding pipeline. The hierarchical descriptors are applied to a coarse-to-fine VPR system with global retrieval and local geometric verification. To explore high-quality local matches between task-relevant visual elements, a cross-attention mutual enhancement layer is introduced to strengthen the information interaction between the local descriptors. Through the proposed selective matching distillation, the mutual enhancement layer can learn from state-of-the-art local matchers in a distillation manner. After weighted cross-matching of the enhanced local descriptors, geometric verification is applied to evaluate the spatial consistency of the compared image pair. Experiments show CAHIR outperforms the existing global and local representations for VPR in terms of performance and efficiency. Quantitatively, it achieves state-of-the-art results on three city-scale benchmark datasets. Qualitatively, CAHIR proves to attach great importance to task-relevant visual elements and excels at finding local correspondences that are discriminative to the VPR task. Guohao Peng, Heshan Li, Jun Zhang 0042, Mingxing Wen, Singh Rahul, Danwei Wang |
ICRA | 2 |
| 2023 | AdaptSeqVPR: An Adaptive Sequence-Based Visual Place Recognition PipelineabstractVisual Place Recognition (VPR) is essential for autonomous robots and unmanned vehicles, as an accurate identification of visited places can trigger a loop closure to optimize the built map. The most prevalent methods tackle VPR as a single-frame retrieval task, which uses a CNN-based encoder to describe and compare each individual frame. These methods, however, overlook the temporal information between frames. Other methods improve this by searching the database with consecutive frames, which can greatly reduce false positives. Nevertheless, current sequence-based methods typically assume the consecutive image frames to be captured at an approximately constant speed, which is not always the case in practice. Therefore, we propose an adaptive sequence search strategy (AdaptSeq), which can dynamically alter the step size of adjacent frames in the retrieved sequence trajectory. Furthermore, to address false positive retrieval of input frames, we propose a CNN-based discriminator named DDsNet. It can determine whether the top retrieved candidates are true positives based on the learned statistics rather than an artificial threshold. Overall, we construct a novel sequence-based VPR pipeline named AdaptSeqVPR. It utilizes a CNN-based encoder for frame descriptions, and encompasses AdaptSeq and DDsNet for sequence matching. The experimental results indicate that our AdaptSeqVPR exhibits superior performance compared to the baseline SeqSLAM and SeqVLAD. Notably, our method can robustly handle the sequence-based VPR for vehicles traveling at non-uniform speeds in changing environments. Heshan Li, Guohao Peng, Jun Zhang 0005, Sriram Vaikundam, Danwei Wang |
IROS | 1 |
| 2022 | C-TM: Topo-metric Mapping and Localization based on Place Categorization and Place Recognition for a Delivery Robot on FootpathabstractIn this work, C-TM is presented: a method to build a topo-metric map for delivery robot navigation in largescale city environments. This system automatically generates a compact map by only saving expensive LIDAR information at key locations. These locations form the nodes of a topological map. Nodes are identified using a Place-Categorization (PC) neural network which output the place category from RGB cameras. Inside nodes, we generate and save high quality LIDAR submaps. Global localization within the map is done with a Visual-Place-Recognition (VPR) neural network. The topo-metric map can be used for navigation on footpath. We deploy C-TM on a four-wheeled autonomous delivery robot and test the effectiveness in two environments, both day and night. Timothy Chia, Jun Zhang 0042, Heshan Li, Guohao Peng, Mingxing Wen, Dawei Kee, P. G. C. N. Senarathne |
ICARCV | 3 |
| 2022 | LSDNet: A Lightweight Self-Attentional Distillation Network for Visual Place RecognitionabstractVisual Place Recognition (VPR) has become an indispensable capacity for mobile robots to operate in large-scale environments. Existing methods in this field mostly focus on exploring high-performance encoding strategies, while few attempts are devoted to lightweight models that balance per-formance and computational cost. In this work, we propose a Lightweight Self-attentional Distillation Network (LSDNet) aiming to obtain advantages of both performance and efficiency. (1) From a performance perspective, an attentional encoding strategy is proposed to integrate crucial information in the scene. It extends the NetVlad architecture with a self-attention module to facilitate non-local information interaction between local features. Through further visual word vector rescaling, the final image representation can benefit from both non-local spatial integration and cluster-wise weighting. (2) From an efficiency perspective, LSDNet is built upon a lightweight back-bone. To maintain comparable performance to large backbone models, a dual distillation strategy is introduced. It prompts LSDNet to learn both encoding patterns in the hidden space and feature distributions in the encoding space from the teacher model. Through distillation-augmented training, LSDNet is able to rival the teacher model and outperform SOTA global representations with the same lightweight backbone. Guohao Peng, Heshan Li, Zhenyu Wu 0001, Danwei Wang |
IROS | 3 |
| 2021 | Attentional Pyramid Pooling of Salient Visual Residuals for Place RecognitionabstractThe core of visual place recognition (VPR) lies in how to identify task-relevant visual cues and embed them into dis- criminative representations. Focusing on these two points, we propose a novel encoding strategy named Attentional Pyramid Pooling of Salient Visual Residuals (APPSVR). It incorporates three types of attention modules to model the saliency of local features in individual, spatial and cluster dimensions respectively. (1) To inhibit task-irrelevant local features, a semantic-reinforced local weighting scheme is employed for local feature refinement; (2) To leverage the spatial context, an attentional pyramid structure is constructed to adaptively encode regional features according to their relative spatial saliency; (3) To distinguish the different importance of visual clusters to the task, a parametric normalization is proposed to adjust their contribution to image descriptor generation. Experiments demonstrate APPSVR outperforms the existing techniques and achieves a new state-of-the-art performance on VPR benchmark datasets. The visualization shows the saliency map learned in a weakly supervised manner is largely consistent with human cognition. Guohao Peng, Jun Zhang 0042, Heshan Li, Danwei Wang |
ICCV | 3 |