VLDB 2026 Research / reviewers in the wild / expert
Lei Pu
dblp:221/2571
· DBLP profile ↗
24ranked-venue papers
4as first author
21since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Robust visual tracking via implicit memory-guided re-detection
Chuangye Xu, Sugang Ma, Xiaobao Yang 0001, Lei Pu |
Eng. Appl. Artif. Intell. | 5 |
| 2026 | Implicit Motion State Modeling for Efficient and Effective Video-Level Object Tracking
Xianxin Jia, Sugang Ma, Lei Pu |
Expert Syst. Appl. | 4 |
| 2026 | Unified Spatio-Temporal Tracking via Adaptive Embedding and Temporal Context Modeling
Xianxin Jia, Sugang Ma, Lei Pu |
Knowl. Based Syst. | 4 |
| 2026 | USGA: unified intra- and cross-scale features with global-local aggregation for long-term tracking
Xianxin Jia, Shuai Hu, Sugang Ma, Xiaobao Yang 0001, Lei Pu |
Multim. Syst. | 7 |
| 2026 | Discovering Multi-Frequency Embedding for Visible-Infrared Person Re-IdentificationabstractVisible-Infrared Person Re-identification (VI-ReID) is critical for round-the-clock surveillance systems yet is hindered by significant modality discrepancies. Existing methods often fail to fully exploit frequency domain information, focusing predominantly on spatial domain feature learning or limited frequency decompositions. To address this, we propose the Multi-Frequency Embedding Network (MFENet), a feature-level method that operates in the frequency domain through multi-frequency decomposition to learn discriminative and modality-invariant features. Specifically, the HiLo-Frequency Modulation (HiLo-FM) module efficiently extracts low-frequency features via frequency-domain filtering and high-frequency details through lightweight multiscale convolutions, followed by attention-based fusion. The Frequency-Aware Diversity Enhancer (FADE) module further enriches feature discriminability by weighting multi-frequency components and learning diverse features through multi-branch architectures. To further enhance the performance of our method, we introduce two innovative loss functions. The Cross-Modality Soft Retrieval (CMSR) loss prioritizes cross-modality consistency over intra-modality similarity, while the Cross-Modality Ranking Regularization (CMRR) loss enhances feature diversity through differentiable rank correlation optimization. Extensive experiments demonstrate the state-of-the-art performance of our method, achieving 61.06% Rank-1 and 67.75% mAP in the challenging IR to VIS mode on the largest VI-ReID benchmark LLCM, surpassing existing methods by significant margins without resorting to reranking or additional labeled data. Code is available at https://github.com/GuHY777/MFENet-VIReID. Hongyang Gu, Ruitao Lu, Lei Pu, Siming Han |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | ReIDMamba: Learning Discriminative Features With Visual State Space Model for Person Re-IdentificationabstractExtracting robust discriminative features is a critical challenge in person re-identification (ReID). While Transformer-based methods have successfully addressed some limitations of convolutional neural networks (CNNs), such as their local processing nature and information loss resulting from convolution and downsampling operations, they still face the scalability issue due to the quadratic increase in memory and computational requirements with the length of the input sequence. To overcome this, we propose a pure Mamba-based person ReID framework named ReIDMamba. Specifically, we have designed a Mamba-based strong baseline that effectively leverages fine-grained, discriminative global features by introducing multiple class tokens. To further enhance robust features learning within Mamba, we have carefully designed two novel techniques. First, the multi-granularity feature extractor (MGFE) module, designed with a multi-branch architecture and class token fusion, effectively forms multi-granularity features, enhancing both discrimination ability and fine-grained coverage. Second, the ranking-aware triplet regularization (RATR) is introduced to reduce redundancy in features from multiple branches, enhancing the diversity of multi-granularity features by incorporating both intra-class and inter-class diversity constraints, thus ensuring the robustness of person features. To our knowledge, this is the pioneering work that integrates a purely Mamba-driven approach into ReID research. Our proposed ReIDMamba model boasts only one-third the parameters of TransReID, along with lower GPU memory usage and faster inference throughput. Experimental results demonstrate ReIDMamba's superior and promising performance, achieving state-of-the-art performance on five person ReID benchmarks. Code is available athttps://github.com/GuHY777/ReIDMamba. Hongyang Gu, Qisong Yang, Lei Pu, Siming Han, Yao Ding 0010 |
IEEE Trans. Multim. | 3 |
| 2025 | Progressive Multi-Scale Vision Transformer for Hierarchical Myocardial Segmentation in Cardiac MRIabstractMyocardial infarction remains a global health challenge. Accurate myocardial segmentation in late gadolinium enhancement cardiac magnetic resonance imaging (LGE-CMRI) is critical for diagnosis and treatment planning. Although deep learning architectures have demonstrated excellent segmentation performance in conventional CMRI, their accuracy significantly declines in LGE-CMRI due to low tissue contrast and complex background interference. To address these challenges, we propose a Progressive Multi-scale Vision Transformer (PMVT) for myocardial segmentation in LGE-CMRI, which enhances spatial representation capabilities and improves adaptability in complex scenarios through multi-scale feature fusion and interaction. Specifically, PMVT includes a Multi-scale Progressive Attention Decoder (MPSD) for modeling both long and short-term dependencies, and a Multi-layer Hybrid Context Purification (MHCP) that combines different combinations of four prediction heads for prediction and loss calculation, effectively suppressing background interference. Experiments demonstrate that the proposed PMVT outperforms state-of-the-art (SOTA) models, achieving a Dice score of 89.61% for myocardial segmentation (a 1.56% improvement over the current SOTA). This result highlights its considerable potential for clinical applications in automated LGE-CMRI analysis. Lei Pu, Yangjie Li, Yuanwei Xu, Jingfeng Jiang, Jinshan Tang, Nan Mu |
SMC | 1 |
| 2025 | Frequency-aware fusion for improved video object segmentation
Chenxu Wang 0012, Sugang Ma, Xiaobao Yang 0001, Lei Pu |
Neurocomputing | 6 |
| 2025 | Instance-aware global re-detection for precise and efficient long-term visual tracking
Xianxin Jia, Sugang Ma, Xiaobao Yang 0001, Lei Pu |
Neurocomputing | 6 |
| 2025 | Integrating multi-scale appearance and motion cues for visual tracking via spatio-temporal prompt
Xianxin Jia, Shuai Hu, Sugang Ma, Xiaobao Yang 0001, Lei Pu |
Knowl. Based Syst. | 8 |
| 2025 | Tgcpn: two-level grid context propagation network for 3D small object detection
Lei Pu, Xuemiao Xu, Chang'an Yi, Yuexia Zhou, Yewen Xu |
Pattern Anal. Appl. | 2 |
| 2024 | DS-MSFF-Net: Dual-path self-attention multi-scale feature fusion network for CT image segmentation
Lei Pu, Liming Wan |
Appl. Intell. | 2 |
| 2024 | Multi-object tracking algorithm based on interactive attention network and adaptive trajectory reconnection
Sugang Ma, Shuaipeng Duan, Wangsheng Yu, Lei Pu, Xiangmo Zhao |
Expert Syst. Appl. | 5 |
| 2024 | SOCF: A correlation filter for real-time UAV tracking based on spatial disturbance suppression and object saliency-aware
Sugang Ma, Bo Zhao 0035, Wangsheng Yu, Lei Pu, Xiaobao Yang 0001 |
Expert Syst. Appl. | 5 |
| 2023 | Ms-AMPool: Down-Sampling Method for Dense Prediction Tasks
Shukai Yang, Yufeng Chen 0006, Lei Pu |
ICANN (2) | 4 |
| 2023 | Robust Visual Object Tracking Based on Feature Channel Weighting and Game TheoryabstractAlthough the discriminative correlation filter‐ (DCF)‐based tracker improves tracking performance, some object representation issues can still be further optimized. On the one hand, the DCF tracker’s deep convolutional features contain many noisy channels, and assigning the same weights to multiple channels cannot distinguish the importance of different channels. On the other hand, a simple weighted fusion approach cannot fully utilize the benefits of different feature types. We propose a visual object tracking algorithm based on adaptive channel weighting and feature game fusion to solve these problems. In this study, an adaptive channel weighting strategy is designed to assign suitable weights to each channel based on the average energy ratio of the target and background regions in the feature channels and prune the channels with low weights to improve feature robustness and reduce computational complexity. Simultaneously, the game theory concept is introduced in the multifeature fusion. The handcrafted features are combined with shallow and deep convolutional features according to feature complementarity. Then, the two combined features are seen as two sides of the game, continuously gamed during the tracking process to generate a feature model with a higher representation capacity. Extensive experiments are conducted on four mainstream visual tracking benchmark datasets, including OTB2015, VOT2018, LaSOT, and UAV123. The experimental results show that the proposed algorithm performs outstandingly compared to the state‐of‐the‐art trackers. Sugang Ma, Bo Zhao 0035, Wangsheng Yu, Lei Pu, Lei Zhang 0166 |
Int. J. Intell. Syst. | 5 |
| 2023 | UcUNet: A lightweight and precise medical image segmentation network based on efficient large kernel U-shaped convolutional module design
Shukai Yang, Yufeng Chen 0006, Youtao Jiang, Quan Feng, Lei Pu |
Knowl. Based Syst. | 6 |
| 2023 | DASGC-Unet: An Attention Network for Accurate Segmentation of Liver CT Images
Yufeng Chen 0006, Lei Pu, Youdong He, Huaijiang Sun |
Neural Process. Lett. | 3 |
| 2022 | Semantic-aware spatial regularization correlation filter for visual trackingabstractAbstract Correlation filters with convolutional neural network (CNN) features have been successfully applied to visual tracking owing to their impressive combined capability for object representation. Unfortunately, further performance improvement is limited due to unwanted boundary effects of the circular structure. In this work, through an in‐depth study of the features’ characteristics, the authors propose a novel tracking strategy to achieve simultaneous filter matching and regularization with CNN features when tracking is on the fly. With a feature decomposed transform matrix, a spatial semantic regularization is generated to reduce the boundary effect effectively during filter optimization. Before each output, the regularized filter is then back performed to match with the extracted features of a search region to find the optimum candidate. Specifically, the most important advantage of the proposed spatial semantic map is to initialize only in the first frame as all the other tracking strategies. Besides, the authors design a novel updating strategy to tackle the cases where the object is occluded or disappeared in the scene. At this time, the maximum of the map is small, even negative. A substantial experiment has been carried out on the popular benchmark tracking datasets; the reliable results have demonstrated that the authors’ method is able to outperform most of the state‐of‐the‐art tracking works in both accuracy and robustness. Yufei Zha, Peng Zhang 0005, Lei Pu, Lichao Zhang 0001 |
IET Comput. Vis. | 3 |
| 2022 | Robust visual tracking via adaptive feature channel selectionabstractDiscriminative correlation filters (DCFs) have shown promising tracking performance in recent years thanks to the powerful representation ability of deep features. However, a large number of target-irrelevant channels in deep features limits the tracking performance and increases the computational cost. To eliminate the negative impact of noisy channels and improve the utilization efficiency of deep features in DCF-based trackers, we present an adaptive feature channel selection method for robust visual tracking. Our method adaptively chooses the most discriminative channels to learn a more robust target appearance model, which is achieved by evaluating the energy relationship between background and foreground in each feature channel. Moreover, according to the feedback of channel selection, an adaptive model update strategy is proposed to alleviate the model degradation problem caused by incorrect model updating. Extensive experimental results obtained on five popular tracking benchmarks demonstrate the effectiveness of the proposed algorithm and its superiority over the state-of-the-art trackers. Sugang Ma, Lei Zhang 0166, Xiaobao Yang 0001, Lei Pu, Xiangmo Zhao |
Int. J. Intell. Syst. | 5 |
| 2021 | SiamDA: Dual attention Siamese network for real-time visual tracking
Lei Pu, Xinxi Feng, Wangsheng Yu, Yufei Zha |
Signal Process. Image Commun. | 1 |
| 2020 | MHASiam: Mixed High-Order Attention Siamese Network for Real-Time Visual Tracking
Lei Pu, Xinxi Feng, Wangsheng Yu, Yufei Zha, Zhiqiang Jiao |
PRCV (2) | 1 |
| 2019 | Deep Correlation Filter based Real-Time Tracker
Lei Pu, Xinxi Feng, Wangsheng Yu, Yufei Zha, Sugang Ma |
FUSION | 1 |
| 2018 | Robust occlusion-aware part-based visual tracking with object scale adaptation
Xin Wang 0026, Wangsheng Yu, Lei Pu, Zefenfen Jin, Xianxiang Qin |
Pattern Recognit. | 4 |