Yanqing Shen

dblp:147/4248 · DBLP profile ↗
← Back
17ranked-venue papers
3as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 3 first-author · 10 since 2021Systems, architecture and hardware · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Computer networks · 2Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 FlowCalib: Targetless Infrastructure LiDAR-Camera Extrinsic Calibration Based on Optical Flow and Scene Flow
abstract
Recently, multi-sensor fusion-based vehicle infrastructure cooperative perception has aroused extensive attention due to the demands for the safety of autonomous driving and traffic monitoring. An accurate calibration between different sensors is a critical foundation for most sensor fusion systems. For LiDAR-camera calibration, high accuracy can be achieved with the help of artificial calibration targets, such as a checkerboard. However, unlike autonomous vehicles, roadside sensors monitor traffic scenes with continuous traffic flow from a fixed viewpoint, posing challenges for conventional calibration methods. There, a calibration method suitable for roadside scenes is required for infrastructure sensors. In this paper, we propose FlowCalib, a novel targetless infrastructure LiDAR-camera spatial calibration method through alignment of scene flow and optical flow. The main idea is to leverage the inherent consistency of moving objects in traffic flow across two types of sensor data. Firstly, the moving objects are extracted by optical flow and scene flow. Then, the extrinsic parameters are obtained in two steps: rough calibration and calibration refinement. In rough calibration, the center and motion flow of each moving instance are calculated by clustering methods separately in the point cloud and image. Based on this, the possible initial value set of extrinsic parameters is estimated by two-step parameter sampling. The initial parameters are obtained by distance of center and motion flow in point cloud and image based scoring. Subsequently, the extrinsic parameters are refined by optimization of instance alignment loss and flow alignment loss of moving objects. In the end, quantitative and qualitative experiments are conducted to validate the effectiveness of the algorithm across both simulated datasets and real-world datasets.
Renwei Hai, Yanqing Shen, Shi-tao Chen, Jingmin Xin, Nanning Zheng 0001
IEEE Trans. Intell. Transp. Syst.2
2025 ForestLPR: LiDAR Place Recognition in Forests Attentioning Multiple BEV Density Images
abstract
Place recognition is essential to maintain global consistency in large-scale localization systems. While research in urban environments has progressed significantly using LiDARs or cameras, applications in natural forest-like environments remain largely under-explored. Furthermore, forests present particular challenges due to high self-similarity and substantial variations in vegetation growth over time. In this work, we propose a robust LiDAR-based place recognition method for natural forests, ForestLPR. We hypothesize that a set of cross-sectional images of the forest’s geometry at different heights contains the information needed to recognize revisiting a place. The cross-sectional images are represented by bird’s-eye view (BEV) density images of horizontal slices of the point cloud at different heights. Our approach utilizes a visual transformer as the shared backbone to produce sets of local descriptors and introduces a multi-BEV interaction module to attend to information at different heights adaptively. It is followed by an aggregation layer that produces a rotation-invariant place descriptor. We evaluated the efficacy of our method extensively on real-world data from public benchmarks as well as robotic datasets and compared it against the state-of-the-art (SOTA) methods. The results indicate that ForestLPR has consistently good performance on all evaluations and achieves an average increase of 7.38% and 9.11% on Recall@1 over the closest competitor on intra-sequence loop closure detection and inter-sequence re-localization, respectively, validating our hypothesis1.
Yanqing Shen, Turcan Tuna, Marco Hutter 0001, Cesar Dario Cadena Lerma, Nanning Zheng 0001
CVPR1
2025 StructVPR++: Distill Structural and Semantic Knowledge With Weighting Samples for Visual Place Recognition
abstract
Visual place recognition is a challenging task for autonomous driving and robotics, which is usually considered as an image retrieval problem. A commonly used two-stage strategy involves global retrieval followed by re-ranking using patch-level descriptors. Most deep learning-based methods in an end-to-end manner cannot extract global features with sufficient semantic information from RGB images. In contrast, re-ranking can utilize more explicit structural and semantic information in one-to-one matching process, but it is time-consuming. To bridge the gap between global retrieval and re-ranking and achieve a good trade-off between accuracy and efficiency, we propose StructVPR++, a framework that embeds structural and semantic knowledge into RGB global representations via segmentation-guided distillation. Our key innovation lies in decoupling label-specific features from global descriptors, enabling explicit semantic alignment between image pairs without requiring segmentation during deployment. Furthermore, we introduce a sample-wise weighted distillation strategy that prioritizes reliable training pairs while suppressing noisy ones. Experiments on four benchmarks demonstrate that StructVPR++ surpasses state-of-the-art global methods by 5-23% in Recall@1 and even outperforms many two-stage approaches, achieving real-time efficiency with a single RGB input.
Yanqing Shen, Sanping Zhou, Jingwen Fu, Ruotong Wang 0005, Shi-tao Chen, Nanning Zheng 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2025 PRGS: Patch-to-Region Graph Search for Visual Place Recognition
Weiliang Zuo, Liguo Liu, Yanqing Shen, Fuhua Xiang, Jingmin Xin, Nanning Zheng 0001
Pattern Recognit.4
2025 RankTuning: Cross-Image Partial Tuning Strategies for Rank Optimization in Visual Place Recognition
abstract
Aiming to estimate the location, a common strategy of Visual Place Recognition (VPR) involves utilizing global retrieval to get top-k candidates first and performing local feature matching in candidates for reranking. Although local reranking methods bring performance gains, they need a lot of computational overhead. To narrow the performance gap between global retrieval and local reranking methods with little cost, one method is to rerank candidates with global features. However, previous works only utilized the information from positive samples in candidates, ignoring the fact that negative samples can also provide useful information. To this end, we propose RankTuning, a method that aggregates all the information from candidates using global features for reranking. Specifically, we design a cross-image interaction module that allows all candidates to interact with others to enhance the discriminative power of features. Furthermore, to drive the training of this module, we propose Generalized Recall loss to handle hard samples with a better gradient strategy. Experimental results demonstrate that our method can be easily inserted into existing architectures and achieve state-of-the-art performance. Meanwhile, our method does not require additional storage overhead, and the matching latency is only 6.3% of that of the current fastest local reranking method. The code is released athttps://github.com/LKELN/RankTuning.git
Liguo Liu, Weiliang Zuo, Jingwen Fu, Yanqing Shen, Jingmin Xin, Nanning Zheng 0001
IEEE Trans. Intell. Transp. Syst.4
2024 Complementing Onboard Sensors with Satellite Maps: A New Perspective for HD Map Construction
abstract
High-definition (HD) maps play a crucial role in autonomous driving systems. Recent methods have attempted to construct HD maps in real-time using vehicle onboard sensors. Due to the inherent limitations of onboard sensors, which include sensitivity to detection range and susceptibility to occlusion by nearby vehicles, the performance of these methods significantly declines in complex scenarios and long-range detection tasks. In this paper, we explore a new perspective that boosts HD map construction through the use of satellite maps to complement onboard sensors. We initially generate the satellite map tiles for each sample in nuScenes and release a complementary dataset for further research. To enable better integration of satellite maps with existing methods, we propose a hierarchical fusion module, which includes feature-level fusion and BEV-level fusion. The feature-level fusion, composed of a mask generator and a masked cross-attention mechanism, is used to refine the features from onboard sensors. The BEV-level fusion mitigates the coordinate differences between features obtained from onboard sensors and satellite maps through an alignment module. The experimental results on the augmented nuScenes showcase the seamless integration of our module into three existing HD map construction methods. The satellite maps and our proposed module notably enhance their performance in both HD map semantic segmentation and instance detection tasks. Our code will be available at https://github.com/xjtu-csgao/SatforHDMap.
Wenjie Gao 0001, Jiawei Fu 0001, Yanqing Shen, Haodong Jing, Shi-tao Chen, Nanning Zheng 0001
ICRA3
2023 StructVPR: Distill Structural Knowledge with Weighting Samples for Visual Place Recognition
abstract
Visual place recognition (VPR) is usually considered as a specific image retrieval problem. Limited by existing training frameworks, most deep learning-based works cannot extract sufficiently stable global features from RGB images and rely on a time-consuming re-ranking step to exploit spatial structural information for better performance. In this paper, we propose StructVPR, a novel training architecture for VPR, to enhance structural knowledge in RGB global features and thus improve feature stability in a constantly changing environment. Specifically, StructVPR uses segmentation images as a more definitive source of structural knowledge input into a CNN network and applies knowledge distillation to avoid online segmentation and inference of seg-branch in testing. Considering that not all samples contain high-quality and helpful knowledge, and some even hurt the performance of distillation, we partition samples and weigh each sample's distillation loss to enhance the expected knowledge precisely. Finally, StructVPR achieves impressive performance on several benchmarks using only global retrieval and even outperforms many two-stage approaches by a large margin. After adding additional re-ranking, ours achieves state-of-the-art performance while maintaining a low computational cost.
Yanqing Shen, Sanping Zhou, Jingwen Fu, Ruotong Wang 0005, Shi-tao Chen, Nanning Zheng 0001
CVPR1
2023 MLF-DET: Multi-Level Fusion for Cross-Modal 3D Object Detection
Zewei Lin, Yanqing Shen, Sanping Zhou, Shi-tao Chen, Nanning Zheng 0001
ICANN (7)2
2023 InteractionNet: Joint Planning and Prediction for Autonomous Driving with Transformers
abstract
Planning and prediction are two important modules of autonomous driving and have experienced tremendous advancement recently. Nevertheless, most existing methods regard planning and prediction as independent and ignore the correlation between them, leading to the lack of consideration for interaction and dynamic changes of traffic scenarios. To address this challenge, we propose InteractionNet, which leverages transformer to share global contextual reasoning among all traffic participants to capture interaction and interconnect planning and prediction to achieve joint. Besides, InteractionNet deploys another transformer to help the model pay extra attention to the perceived region containing critical or unseen vehicles. InteractionNet outperforms other baselines in several benchmarks, especially in terms of safety, which benefits from the joint consideration of planning and forecasting. The code will be available at https://github.com/fujiawei0724/InteractionNet.
Jiawei Fu 0001, Yanqing Shen, Zhiqiang Jian, Shi-tao Chen, Jingmin Xin, Nanning Zheng 0001
IROS2
2023 Onboard Sensors-Based Self-Localization for Autonomous Vehicle With Hierarchical Map
abstract
Localization is a fundamental and crucial module for autonomous vehicles. Most of the existing localization methodologies, such as signal-dependent methods (RTK-GPS and Bluetooth), simultaneous localization and mapping (SLAM), and map-based methods, have been utilized in outdoor autonomous driving vehicles and indoor robot positioning. However, they suffer from severe limitations, such as signal-blocked scenes of GPS, computing resource occupation explosion in large-scale scenarios, intolerable time delay, and registration divergence of SLAM/map-based methods. In this article, a self-localization framework, without relying on GPS or any other wireless signals, is proposed. We demonstrate that the proposed homogeneous normal distribution transform algorithm and two-way information interaction mechanism could achieve centimeter-level localization accuracy, which reaches the requirement of autonomous vehicle localization for instantaneity and robustness. In addition, benefitting from hardware and software co-design, the proposed localization approach is extremely light-weighted enough to be operated on an embedded computing system, which is different from other LiDAR localization methods relying on high-performance CPU/GPU. Experiments on a public dataset (Baidu Apollo SouthBay dataset) and real-world verified the effectiveness and advantages of our approach compared with other similar algorithms.
Yanqing Shen, Yuedong Yang, Xiaodong Deng, Shi-tao Chen, Jingmin Xin, Nanning Zheng 0001
IEEE Trans. Cybern.2
2022 TransVPR: Transformer-Based Place Recognition with Multi-Level Attention Aggregation
abstract
Visual place recognition is a challenging task for applications such as autonomous driving navigation and mobile robot localization. Distracting elements presenting in complex scenes often lead to deviations in the perception of visual place. To address this problem, it is crucial to integrate information from only task-relevant regions into image representations. In this paper, we introduce a novel holistic place recognition model, TransVPR, based on vision Transformers. It benefits from the desirable property of the self-attention operation in Transformers which can naturally aggregate task-relevant features. Attentions from multiple levels of the Transformer, which focus on different regions of interest, are further combined to generate a global image representation. In addition, the output tokens from Transformer layers filtered by the fused attention mask are considered as key-patch descriptors, which are used to perform spatial matching to re-rank the candidates retrieved by the global image features. The whole model allows end-to-end training with a single objective and image-level supervision. TransVPR achieves state-of-the-art performance on several real-world benchmarks while maintaining low computational time and storage requirements.
Ruotong Wang 0005, Yanqing Shen, Weiliang Zuo, Sanping Zhou, Nanning Zheng 0001
CVPR2
2022 A New Attention-Based LSTM for Image Captioning
Fen Xiao, Wenfeng Xue, Yanqing Shen, Xieping Gao 0001
Neural Process. Lett.3
2020 HeLPS: Heterogeneous LiDAR-based Positioning System for Autonomous Vehicle
abstract
LiDAR-based positioning systems are widely used in unmanned systems. However, affected by the high computational complexity of high-precision positioning algorithms, the current positioning system is supported by hardware with low power efficiency and thus hard to integrate into many platforms. In this paper, we analyze features of the positioning system in autonomous driving application, design, and apply Heterogeneous LiDAR-based Positioning System, HeLPS, with software-hardware co-design methodology to achieve better efficiency. Our contributions can be concluded in three aspects. Firstly, we design the CPU-FPGA heterogeneous positioning system accelerating Iterative Closest Point (ICP) algorithm and achieves improvements on both speed and power-efficiency. Secondly, we exploit the spatial locality in the point cloud and design a new compressed data structure for fast neighbor accessing. The experiment reports a significant speedup comparing with other data structures. Lastly, we explore the data access pattern in positioning application and develop a specific cache system and out-of-order execution system reducing memory burden. Our system is deployed on a small and cheap Xilinx Zynq7000 ARM+FPGA platform, which achieves 983.3x speedup compared with Cortex-A9 CPU, and 31.8x speedup compared with i7-7820 CPU, with only 2.37W power consumption.
Yuedong Yang, Xiaodong Deng, Yanqing Shen, Shi-tao Chen, Nanning Zheng 0001
IECON4
2019 Robust Extrinsic Parameter Calibration of 3D LIDAR Using Lie Algebras
abstract
In the field of autonomous driving, multi-beam light detection and ranging (3D LIDAR) system and global navigation satellite system/integrated inertial navigation system (GNSS/INS) are widely used in high-definition map construction, localization and obstacle detection. As 3D LIDAR system and INS have their own coordinate systems, the calibration of the two mentioned systems is required. In this paper, a novel algorithm for calibrating the coordinate system of 3D LIDAR and INS is proposed, which consists of three parts. The first procedure is to project two point clouds to the world coordinate system based on the initial transform matrix between 3D LIDAR and INS with the real-time data from INS. Then optimal point-to-point correspondences can be found between two frames of point cloud data through registration method. Finally, the loss function is constructed with the sum of the Euclidean distances of the corresponding points and optimized by using perturbation model of Lie algebras, so as to obtain the optimal transform matrix. With different given initial calibration parameters, test results of both simulation and real experiments validate the proposed algorithm and quantify its accuracy and robustness.
Yanqing Shen, Tangyike Zhang, Songyi Zhang, Yongbo Huo, Shi-tao Chen, Nanning Zheng 0001
IV2
2019 DAA: Dual LSTMs with adaptive attention for image captioning
Fen Xiao, Yanqing Shen, Xieping Gao 0001
Neurocomputing4
2018 Egocentric network focused community aware multicast routing for DTNs
Guoxing Jiang, Yanqing Shen, Yan Dong 0001
Wirel. Networks2
2014 Delivery ratio- and buffered time-constrained: Multicasting for Delay Tolerant Networks
Guoxing Jiang, Yanqing Shen
J. Netw. Comput. Appl.3