EDBT 2026 Demo / reviewers in the wild / expert
Zhenyu Li 0010
dblp:58/5750-10
· DBLP profile ↗
12ranked-venue papers
8as first author
12since 2021 · last 2026
0009-0005-2344-870XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 5 first-author · 6 since 2021Systems, architecture and hardware · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Quadruplet-attention transformer for scale-invariant robot place recognition
Zhenyu Li 0010, Pengjie Xu |
Expert Syst. Appl. | 1 |
| 2026 | Hybrid State Space Modeling for Sequence-Based Robot Localization Under Challenging Environments
Zhenyu Li 0010, Tianyi Shang |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2026 | Seeing Through the Rain: Multistage Attention With Depth-Guided Restoration for Robust Localization
Zhenyu Li 0010, Tianyi Shang, Jinwei Qiao, Pengbo Liu 0002, Zhaojun Deng |
IEEE Trans. Ind. Informatics | 1 |
| 2026 | Vehicle-Scene Interaction: A Text-Driven 3-D Lidar Place Recognition Method for Autonomous DrivingabstractEnvironment description-based vehicle-scene interaction for vehicle localization in large-scale point cloud maps constructed through multi-sensor systems is critically significant for the advancement of large-scale intelligent transportation systems, such as delivery vehicles operating in the ‘last mile’. However, current vehicle-scene interaction encounter challenges due to the inability of point cloud encoders to effectively capture local details and long-range spatial relationships and a significant modality gap between text and point cloud representations. To address these challenges, we present Des4Pos, a novel two-stage text-driven 3D Lidar place recognition framework. In the coarse stage, the point-cloud encoder utilizes the Multi-scale Fusion Attention Mechanism (MFAM) to enhance local geometric features, followed by a bidirectional Long Short-Term Memory (LSTM) module to strengthen global spatial relationships. Concurrently, the Stepped Text Encoder (STE) integrates cross-modal prior knowledge from CLIP and aligns text and point-cloud features using this prior knowledge, effectively bridging modality discrepancies. In the fine stage, we introduce a Cascaded Residual Attention (CRA) module to fuse cross-modal features and predict relative localization offsets, thereby achieving greater localization precision. Experiments on the KITTI360Pose test set demonstrate that Des4Pos achieves state-of-the-art performance in text-to-point-cloud place recognition at the expense of a slight increase in the runtime. Specifically, it attains a top-1 accuracy of 40% and a top-10 accuracy of 77% under a 5-meter radius threshold, surpassing the best previous sota method Text2Loc by 8% and 7%, respectively. Our code and datasets are publicly available athttps://github.com/nuozimiaowu/Des4Pos Tianyi Shang, Zhenyu Li 0010, Pengjie Xu, Zhaojun Deng |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2025 | MambaPlace: Text-to-Point-Cloud Cross-Modal Place Recognition with Attention Mamba MechanismsabstractVision-Language Place Recognition (VLPR) enhances robot localization performance by incorporating natural language descriptions from images. By utilizing language information, VLPR directs robot place matching, overcoming the constraint of solely depending on vision. However, general multimodal information integration methods are not well equipped to capture the dynamics of cross-modal interactions, especially in the presence of complex intra-modal and inter-modal correlations. To this end, this paper proposes a novel coarse-to-fine and end-to-end connected cross-modal place recognition framework, called MambaPlace. In the coarse-localization stage, the text description and 3D point cloud are encoded by the pre-trained T5 and instance encoder, respectively. They are then processed using Text-Attention Mamba (TAM) and Point Cloud Multi-Strategy Scanning Mamba (MSSM), with the latter mimicking the eye’s focusing mechanism, for data enhancement and alignment. In the subsequent fine-localization stage, the features of the text description and 3D point cloud are cross-modally fused and further enhanced through Cascaded Cross-Attention Mamba (CCAM). Finally, we predict the positional offset from the fused text-point cloud features, achieving the most accurate localization. Extensive experiments show that MambaPlace achieves improved localization accuracy on the KITTI360Pose dataset compared to the state-of-the-art methods. Specifically, as shown in Fig. 1, when ϵ<5, MambaPlace achieves 5% higher test accuracy compared to the existing state-of-the-art. Tianyi Shang, Zhenyu Li 0010, Pengjie Xu, Jinwei Qiao |
IROS | 2 |
| 2025 | Bridging Text and Vision: A Multi-View Text-Vision Registration Approach for Cross-Modal Place RecognitionabstractMobile robots necessitate advanced natural language understanding capabilities to accurately identify locations and perform tasks such as package delivery. However, traditional visual place recognition (VPR) methods rely solely on single-view visual information and cannot interpret human language descriptions. To overcome this challenge, we bridge text and vision by proposing a multiview (360° views of the surroundings) text-vision registration approach called Text4VPR for place recognition task, which is the first method that exclusively utilizes textual descriptions to match a database of images. Text4VPR employs the frozen T5 language model to extract global textual embeddings. Additionally, it utilizes the Sinkhorn algorithm with temperature coefficient to assign local tokens to their respective clusters, thereby aggregating visual descriptors from images. During the training stage, Text4VPR emphasizes the alignment between individual text-image pairs for precise textual description. In the inference stage, Text4VPR uses the Cascaded Cross-Attention Cosine Alignment (CCCA) to address the internal mismatch between text and image groups. Subsequently, Text4VPR performs precisely place match based on the descriptions of text-image groups. On Street360Loc, the first text to image VPR dataset we created, Text4VPR builds a robust baseline, achieving a leading top-1 accuracy of 56% and a leading top-10 accuracy of 91% within a 5-meter radius on the test set, which indicates that localization from textual descriptions to images is not only feasible but also holds significant potential for further advancement, as shown in Figure 1. Our code is available at https://github.com/nuozimiaowu/Text4VPR. Tianyi Shang, Zhenyu Li 0010, Pengjie Xu, Jinwei Qiao, Zihan Ruan, Weijun Hu |
IROS | 2 |
| 2025 | Toward Robust Visual Place Recognition for Mobile Robots With an End-to-End Dark-Enhanced NetabstractRecent years have witnessed a fast evolution and promising performance of the vision transformer (ViT)-based place recognizer, which aims at building a general system. State-of-the-arts (SOTAs) can hardly carry on their superiority at low light so far, thereby considerably blocking the broadening of visual place recognition-related mobile robot applications. To perform robust visual place recognition in low-light scenes, this article proposes an end-to-end trainable dark-enhanced Net, which tries to alleviate the impact of poor illumination and environmental noise. Specifically, a lightweight dark enhancement module, i.e.,$\sf ResEM$, is firstly trained to efficiently improve image illumination quality by residual-based adversarial learning. A dual-level sampling pyramid transformer, i.e.,$\sf DSPFormer$, is then constructed to extract discriminative features through aggregating reconstructed descriptors. Moreover, to improve the performance and reliability of place recognition, a reranking method based on cross-entropy loss is used for final place matching. To provide a comprehensive evaluation, we also build two challenging place benchmarks, namely,$\sf SimPlace$and$\sf DarkPlace$. Evaluations of both the public benchmarks and the newly built benchmarks show that the task-inspired design enables the recognizer to achieve significant performance improvements in the nighttime for robot place recognition compared to other top-ranked place recognizers. Zhenyu Li 0010, Tianyi Shang, Pengjie Xu, Zhaojun Deng |
IEEE Trans. Ind. Informatics | 1 |
| 2025 | Multi-Modal Attention Perception for Intelligent Vehicle Navigation Using Deep Reinforcement LearningabstractIn this paper, we propose a new framework for collision-free intelligent vehicle navigation, aiming to successfully avoid obstacles using deep reinforcement learning. The navigation system separates perception and control and utilizes multimodal perception to achieve reliable online interaction with the surroundings. This allows for direct policy learning to generate flexible actions and avoid collisions. Our navigation system establishes a connection between the virtual environment and the real world, allowing learning policies in the virtual environment to be implemented in real-world environments through transfer learning. Our approach aims to integrate camera, Lidar, and Inertial Measurement Unit (IMU) data to construct a multimodal perception-based environment model, which is a state input for reinforcement learning. In this process, we utilize a series of cross-domain self-attention layers to enhance visual and Lidar perception, promoting significant improvements in perception. We also introduce the recurrent deduction to facilitate global decision-making based on local perception. Additionally, we introduce the Self-Assessment Gradient model into the reinforcement learning process to further optimize the learning policy. The experimental results demonstrate the proposed approach reduces the disparity between the virtual environment and the real world, highlighting its superiority over other state-of-the-art methods. Our project page is publicly available athttps://github.com/CV4RA/MMAP-DRL-Nav. Zhenyu Li 0010, Tianyi Shang, Pengjie Xu |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2025 | Feature-Level Knowledge Distillation for Place Recognition Based on Soft-Hard Labels Teaching ParadigmabstractThe motivation of visual place recognition (VPR) is to enable robots to identify and localize specific places within an environment using visual cues, facilitating navigation, mapping, and context-aware applications. On the other hand, deeper networks impose an extra computing strain on robots and greatly impede the development of real-time robot applications. Most groundbreaking studies either focus on learning from only one teacher in their distilled approach to learning, ignoring the possibility of students learning from multiple teachers, or failing to reveal teachers place varying importance on specific examples. To cope with the above issues, we propose a novel adaptive soft-hard label teaching feature-level knowledge distillation learning framework, namely ASHT-KD, for all-day mobile robot VPR tasks. This framework learns a compact and quick all-day place recognizer through knowledge transfer from several teachers to a limited number of students. Specifically, depending on the complexity of the environments, teachers can impart knowledge to two types of students in two teaching modes: soft-label teaching and hard-label teaching, which corresponds to one type of student being required to learn a new and uncomplicated environment (query image), while the other type of students are forced to learn a more complex environment (database images). To balance computational memory and performance, the teacher network is designed to be a two-level sampling ViT pipeline, while the Siamese student network is constructed to be a lightweight pipeline consisting only of one-level down-sampling ViT for place matching. In addition, a cross-entropy loss network is introduced to further improve the VPR performance by strengthening the correlation of feature representations from the Siamese network. Extensive experiments demonstrate the effectiveness and superiority of ASHT-KD. The practicability of ASHT-KD is also verified through outdoor testing. Zhenyu Li 0010, Pengjie Xu, Zhenbiao Dong, Zhaojun Deng |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2024 | Pyramid transformer-based triplet hashing for robust visual place recognition
Zhenyu Li 0010, Pengjie Xu |
Comput. Vis. Image Underst. | 1 |
| 2024 | Reinforcement learning-based distributed impedance control of robots for compliant operation in tight interaction tasksabstractIt is challenging to achieve compliant operation in tight interaction tasks, where closed-loop constraints are formed between the robots or between a robot and the environment. Complex dynamic interactions and uncertain parameters degrade the performance of model-based controllers. In this paper, a reinforcement learning-based distributed impedance control approach is proposed for these tight interaction tasks. Two aspects are considered to ensure compliant operation for the robots. First, a distributed impedance model is established through the design of reasonable independence states and nodes networks. Second, the reinforcement learning agent is designed to make decision for adjusting impedance parameters . The trained parameters are integrated into the designed impedance model, and a model-based controller is then employed for compliant control for the robots. The effectiveness is validated under two different tight interaction scenarios via co-simulation. Compliant operation can be achieved whether between robots or between a robot and the environment. Pengjie Xu, Zhenyu Li 0010, Tianrui Zhao, Lin Zhang 0024, Yanzheng Zhao |
Eng. Appl. Artif. Intell. | 2 |
| 2024 | CSPFormer: A cross-spatial pyramid transformer for visual place recognition
Zhenyu Li 0010, Pengjie Xu |
Neurocomputing | 1 |