VLDB 2026 Research / reviewers in the wild / expert
Zefeng Lu
dblp:230/0601
· DBLP profile ↗
12ranked-venue papers
10as first author
12since 2021 · last 2026
0000-0002-3139-4546ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 5 since 2021Security and privacy · 4 · 4 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | STAR: Spatial-Temporal Attention Reasoning model for dynamic logistics network routing in Cyber-Physical Internet
Zefeng Lu, Zhiheng Zhao, George Q. Huang |
Adv. Eng. Informatics | 1 |
| 2026 | Weakly Supervised Learning Meets VIS-NIR Re-Identification: Progressive Cross-Modality Alignment for Representation Learning
Zefeng Lu, Yunfei He, Meijun Chen, Hui Lu 0005, Yuyu He 0001, Zhihong Tian 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2025 | Prompt-Guided Transformer and MLLM Interactive Learning for Text-Based Pedestrian SearchabstractAiming to retrieve pedestrian images based on a textual description query, Text-Based Pedestrian Search (TBPS) gains increasingly attention due to its applications in security surveillance. As a fine-grained classification task, TBPS requires identifying images of individuals with different semantic contexts yet the same identity, as well as distinguishing images of individuals who share similar appearances but distinct identities. Consequently, TBPS is challenged by semantic variations in positive pairs and appearance similarity between negative pairs. To tackle these challenges, we propose the Prompt-guided Transformer and MLLM Interactive learning (PTMI) model to learn identity-discriminative representations across different modalities. PTMI consists of three components: the Prompt-guided Transformer (Promformer), MLLM Interactive Learning (MIL) and Dual-branch Cross-modal Learning (DCL). Firstly, the Promformer is designed to handle semantic variations in positive pairs by introducing learnable prompts, composing of three types: instance-shared, instance-specific and layer-specific. Optimized by Cross-modal Intra-class Consistency (CIC) loss, these prompts minimize intra-class variations and retrieve positive images with various semantics. Secondly, the MIL component is introduced to address appearance similarity between negative pairs by focusing on key image patches and description words filtering by the local discriminator. Powered by Multimodal Large Language Model (MLLM), the local discriminator adopts soft attention to highlight important image regions and descriptive words, which preserves semantic information while emphasize discriminative details. Lastly, the DCL integrates global and local branches to bridge modality discrepancies. The global branch employs SDM loss for heterogeneous distribution alignment, while the local branch applies Anchor-Based Contrastive (ABC) loss for instance-level contrastive learning. Unlike conventional contrastive loss, ABC loss leverages MLLM features as anchors to decouple modality and semantic differences, enhancing alignment efficiency. Extensive experiments on three TBPS datasets have validated the effectiveness of PTMI. Zefeng Lu, Ronghao Lin, Yap-Peng Tan, Haifeng Hu 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2025 | Disentangling Modality and Posture Factors: Memory-Attention and Orthogonal Decomposition for Visible-Infrared Person Re-IdentificationabstractStriving to match the person identities between visible (VIS) and near-infrared (NIR) images, VIS-NIR reidentification (Re-ID) has attracted increasing attention due to its wide applications in low-light scenes. However, owing to the modality and pose discrepancies exhibited in heterogeneous images, the extracted representations inevitably comprise various modality and posture factors, impacting the matching of cross-modality person identity. To solve the problem, we propose a disentangling modality and posture factors (DMPFs) model to disentangle modality and posture factors by fusing the information of features memory and pedestrian skeleton. Specifically, the DMPF comprises three modules: three-stream features extraction network (TFENet), modality factor disentanglement (MFD), and posture factor disentanglement (PFD). First, aiming to provide memory and skeleton information for modality and posture factors disentanglement, the TFENet is designed as a three-stream network to extract VIS-NIR image features and skeleton features. Second, to eliminate modality discrepancy across different batches, we maintain memory queues of previous batch features through the momentum updating mechanism and propose MFD to integrate features in the whole training set by memory-attention layers. These layers explore intramodality and intermodality relationships between features from the current batch and memory queues under the optimization of the optimal transport (OT) method, which encourages the heterogeneous features with the same identity to present higher similarity. Third, to decouple the posture factors from representations, we introduce the PFD module to learn posture-unrelated features with the assistance of the skeleton features. Besides, we perform subspace orthogonal decomposition on both image and skeleton features to separate the posture-related and identity-related information. The posture-related features are adopted to disentangle the posture factors from representations by a designed posture-features consistency (PfC) loss, while the identity-related features are concatenated to obtain more discriminative identity representations. The effectiveness of DMPF is validated through comprehensive experiments on two VIS-NIR pedestrian Re-ID datasets. Zefeng Lu, Ronghao Lin, Haifeng Hu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Mind the Inconsistent Semantics in Positive Pairs: Semantic Aligning and Multimodal Contrastive Learning for Text-Based Pedestrian SearchabstractAiming at retrieving pedestrian images based on a provided textual description query, Text-Based Pedestrian Search (TBPS) has gained attention due to its implications in public security tasks such as suspect tracking. Nevertheless, the modality discrepancies between textual descriptions and visual images pose a challenge in aligning semantic information between these two modalities. Moreover, the text description annotated on a particular pedestrian image may not align with the content of other images sharing the same identity, due to variations in viewpoint. These text-image pairs exhibiting inconsistent semantics, termed weak positive pairs, have a discernible impact on the model’s performance. To address these challenges, we propose a Semantic Aligning and Multimodal Contrastive learning (SAMC) model to capture cross-modality identity-invariant features, including three modules: Multi-modality Features Fusion (MFF), Semantic-aligning Optimal Transport (SOT), and Multi-modality Contrastive Learning (MCL). Firstly, the MFF is designed to fuse textual and visual information and extract identity-discriminative multimodal features using self- and cross-attention mechanisms. The multimodal features act as anchors, bridging the gap between the two modalities and enhancing the identity-invariance of unimodal features. Secondly, the SOT is designed to address the semantic misalignment issue between textual descriptions and visual images. Utilizing the Optimal Transport (OT) theory, SOT encourages high features similarity between positive samples from different modalities, thereby exploring semantic relationships between image and text data without requiring extra supervised labels. Lastly, the MCL is introduced to narrow the modality gap, compelling two unimodal features towards the identity-discriminative multimodal features through contrastive learning. Different temperature coefficients are employed for strong and weak positive pairs to mitigate the inconsistency in text-image pair correlation. The effectiveness of SAMC is validated by extensive comprehensive experiments on three TBPS datasets. Zefeng Lu, Ronghao Lin, Haifeng Hu 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2024 | Tri-Level Modality-Information Disentanglement for Visible-Infrared Person Re-IdentificationabstractAiming to match the person identity between daytime VISible (VIS) and nighttime Near-InfraRed (NIR) images, VIS-NIR re-identification (Re-ID) has attracted increasing attention due to its wide applications in low-light scenes. However, dramatic modality discrepancies between VIS and NIR images lead to a considerable intra-class gap in the feature space, which impacts identity matching. To bridge the modality gap, we propose a Tri-level Modality-information Disentanglement (TMD) to disentangle modality information at the levels of raw image, features distribution and instance features. Our model consists of three key modules, including Style-Aligned Converter (SAC), Two-Steps Wasserstein Loss (TSWL) and Self-supervised Orthogonal Disentanglement (SOD) to handle the modality information at the three levels. Firstly, aiming at reducing modality discrepancy at image-level, the SAC is introduced to generate style-aligned images by the designed style converter and$\mathcal {A}$-distance learning approach. The SAC can effectively alleviate the style discrepancy between VIS and NIR images with a negligible increase in model complexity. Secondly, considering the heterogeneity of VIS and NIR feature distribution caused by the structure- and style-misaligned raw images, we propose the TSWL to decrease the VIS-NIR gap at distribution-level by two distribution alignment steps. Specifically, after generating style-consistent images, we eliminate modality-related discrepancy by aligning the distribution between structure-aligned original and generated VIS/NIR images and bridge the modality-unrelated gap by aligning the style-consistent generated VIS-NIR images. Thirdly, focusing on further reducing the modality discrepancy at instance-level, the SOD is presented to construct orthogonal constraints between the extracted modality- and identity-related features. Since the modality-related factors are disentangled from the instance features, the proposed TMD efficiently learns the modality-unrelated and identity-discriminative representations, which are productive to conduct person Re-ID task on the VIS-NIR images. Comprehensive experiments are carried out on two cross-modality pedestrian Re-ID datasets to demonstrate the effectiveness of TMD. Zefeng Lu, Ronghao Lin, Haifeng Hu 0001 |
IEEE Trans. Multim. | 1 |
| 2023 | IMperm: a fast and comprehensive IMmune Paired-End Reads Merger for sequencing dataabstractThe adaptive immune receptor repertoire (AIRR), consisting of T- and B-cell receptors, is the core component of the immune system. The AIRR sequencing is commonly used in cancer immunotherapy and minimal residual disease (MRD) detection of leukemia and lymphoma. The AIRR is captured by primers and sequenced to yield paired-end (PE) reads. The PE reads could be merged into one sequence by the overlapped region between them. However, the wide range of AIRR data raises the difficulty, so a special tool is required. We developed a software package for IMmune PE reads merger of sequencing data, named IMperm. We used the k-mer-and-vote strategy to pin down the overlapped region rapidly. IMperm could handle all types of PE reads, eliminate adapter contamination and successfully merge low-quality and minor/non-overlapping reads. Compared with existing tools, IMperm performed better in both simulated and sequencing data. Notably, IMperm was well suited to processing the data of MRD detection in leukemia and lymphoma and detected 19 novel MRD clones in 14 patients with leukemia from previously published data. Additionally, IMperm can handle PE reads from other sources, and we demonstrated its effectiveness on two genomic and one cell-free deoxyribonucleic acid datasets. IMperm is implemented in the C programming language and consumes little runtime and memory. It is freely available at https://github.com/zhangwei2015/IMperm. Wei Zhang 0178, Jia Ju, Teng Xiong, Chaohui Li, Shixin Lu, Zefeng Lu, Liya Lin, Shuaicheng Li 0001 |
Briefings Bioinform. | 8 |
| 2023 | Modality and Camera Factors Bi-Disentanglement for NIR-VIS Object Re-IdentificationabstractAiming to match object identities across different modalities of images, the challenging task namely cross-modality object re-identification (NIR-VIS object Re-ID), has attracted increasing attention due to its wide application in low-light scenes. However, dramatic modality-dependent and camera-related discrepancies between Near-InfraRed-spectrum (NIR) and VISible-spectrum (VIS) images lead to a considerable intra-class gap in feature space. To address the problem, we propose a novel Modality and Camera factors Bi-Disentanglement (MCBD) model to learn modality-independent and camera-unrelated features for NIR-VIS object Re-ID. Our model consists of three key modules, including Confused Modality Generation (CMG), Modality-independent Information Distillation (MID), and Cameras Factor Disentanglement (CFD). Firstly, aiming at aligning image style between NIR and VIS data, the CMG utilizes a designed channel-interactive generator to generate confused modality images which preserves the structure information of original images. Besides, CMG is trained with confused adversarial learning which bridges the modality gap at the image level. Nevertheless, training the model with the confused modality images discards identity-related information such as color and contrast, which is not conducive to the extraction of distinctive features. To solve this problem, the MID is presented to distill out modality-independent information by feeding the original images to the model and reconstructing corresponding NIR and VIS modalities features to the confused modality images. Finally, due to the complexity of camera-related information in images, identity representation inevitably contains camera-related elements such as background and perspective information, which may interfere the matching process. To address this issue, the CFD is introduced to disentangle camera-related factors by the designed three-streams network and two factor decoupling losses, i.e., Camera-Camera Factor loss (CCF) and Identity-Camera Factor (ICF) loss. Comprehensive experiments are carried out on two cross-modality pedestrian Re-ID datasets and a cross-modality vehicle Re-ID dataset to demonstrate that the MCBD is effective in cross-modality object Re-ID task. Zefeng Lu, Ronghao Lin, Haifeng Hu 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2023 | Graph-Based Progressive Fusion Network for Multi-Modality Vehicle Re-IdentificationabstractVehicle re-identification (Re-ID) is a critical task in intelligent transportation, aiming to match vehicle images of the same identity captured by non-overlapping cameras. However, it is difficult to achieve satisfactory results based on RGB images alone in darkness. Therefore, it is of great importance to consider multi-modality vehicle re-identification. Currently, the proposed works deal with different modality features through direct summation and fusion based on heat map, which however ignores the relationship between them. Meanwhile, there is a huge gap between the different modalities, which needs to be reduced. In this paper, to solve the above two problems, we propose a Graph-based Progressive Fusion Network (GPFNet) using a graph convolutional network to adaptively fuse multi-modality features in an end-to-end learning framework. GPFNet consists of a CNN feature extraction module (FEM), a GCN feature fusion module (FFM), and a loss function module (LFM). Firstly, in FEM, we employ a multi-stream network architecture to extract single-modality features and common-modality features and employ a random modality substitution module to extract mixed-modality features. Secondly, in FFM, we design an efficient graph structure to associate the features of different modalities and adopt a progressive two-stage strategy to fuse them. Finally, in LFM, we use GCN-aware multi-modality loss to constrain the features. For reducing modality differences and contributing better initial mixed-modality features to FFM, we propose random modality substitution as a data enhancement method for multi-modality datasets. Extensive experiments on multi-modality vehicle Re-ID datasets RGBN300 and RGBNT100 show that our model achieves state-of-the-art performance. Qiaolin He, Zefeng Lu, Haifeng Hu 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2023 | MART: Mask-Aware Reasoning Transformer for Vehicle Re-IdentificationabstractAs a significant topic in Intelligent Transportation Systems (ITS), vehicle Re-Identification (Re-ID) has attracted increasing research attention. However, the variation of shooting scenes and the similar appearance among the vehicles with the same type and color lead to large intra-class variances and small inter-class variances, respectively. To address the problems, we propose a novel Mask-Aware Reasoning Transformer (MART) to extract the background-unrelated global features and perspective-invariant local features. The MART contains three effective modules including Foreground Global Features Extraction (FGFE), Mask-guided Local Features Extraction (MLFE) and Cross-images Local Features Reasoning (CLFR). Firstly, due to the complexity of background information in images, identity representation inevitably contains background elements, which may impact identity matching. To address this issue, we propose the FGFE to extract background-independent global features by introducing the mask semantic information to the inputs of Vision Transformer (ViT). Secondly, to fill the gap that the previous local features extraction methods cannot be directly applied to ViT, the MLFE is presented to extract distinctive local features by recombining token features according to vehicle mask. Thirdly, when local components are invisible in the image due to the occlusion problem, the corresponding local information is absent from the image, leading to unreliable local features. To solve this problem, the CLFR is proposed to reason the occluded local features by exploiting the correlation between cross-image local features. We carry out comprehensive experiments to illustrate the effectiveness of the MART on two challenging datasets. Zefeng Lu, Ronghao Lin, Haifeng Hu 0001 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2023 | Mask-Aware Pseudo Label Denoising for Unsupervised Vehicle Re-IdentificationabstractAs a significant part of Intelligent Transportation System (ITS), vehicle Re-Identification (Re-ID) aims to retrieve all target vehicle images captured from non-overlapping cameras. Though the Re-ID methods based on supervised learning have achieved rapid progress, they are still difficult to be applied in real scenarios due to the domain bias between the training set and real scenarios. Recently, methods based on unsupervised learning have been proposed to address the problem of domain bias by exploring techniques of pseudo-label generation. However, these methods suffer from pseudo-label noise. To solve this problem, we propose the Mask-Aware Pseudo Label Denoising framework (MAPLD) consisting of three key components, i.e., Mask-Aware Feature Extraction (MAFE), Adaptive Threshold Neighborhood Consistency (ATNC), and Compact Loss (CL). Firstly, the MAFE is proposed to improve the distinguishability of feature representation and widen the gap in feature space among vehicles with different IDs. Next, the ATNC is introduced to filter out pseudo-label noise of hard negative samples by comparing the image ID of the samples in their neighborhood set i.e., neighborhood consistency. Moreover, the threshold of neighborhood consistency is adaptively adjusted according to feature similarity ranking, which is robust to hyper-parameter variation. Finally, consisting of regression term and compact term, the CL is designed to drive the cluster more compact and alleviate the impact of outliers of hard positive samples. Extensive experiments on VeRi-776 and VeRi-Wild datasets demonstrate that MAPLD can generate reliable pseudo-labels and achieve superior performance in unsupervised target-only and unsupervised domain adaptation tasks. Zefeng Lu, Ronghao Lin, Qiaolin He, Haifeng Hu 0001 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2022 | Identity-Unrelated Information Decoupling Model for Vehicle Re-IdentificationabstractAs an indispensable part of intelligent transportation system (ITS), vehicle re-identification (Re-ID) aims to retrieve all target vehicle images captured from non-overlapping cameras. However, this task remains very challenging due to the variation of camera perspective and the similar appearance among the vehicles with the same type and color, i.e., large intra-class variances and small inter-class variances. Previous methods have made a great progress on vehicle Re-ID by leveraging local details and aligning local features, this issue is still far from being solved. In this work, we propose to decouple identity-unrelated information from vehicle representation, tackling the problems of camera perspective variation and vehicle appearance similarity. The keypoint of this method is to learn a distinguishable feature embedding that is independent of identity-unrelated information. Specially, the novel Identity-Unrelated Information Decoupling (IUID) paradigm is designed to learn invariant features of the vehicle with the same ID in different scenes. In our approach, identity-unrelated information can be divided into two kinds of information, i.e., camera perspective information and background information. For the former, through a feature-level camera generative adversarial module, we can decouple camera perspective information from the feature embedding after extracting invariant features across different cameras perspective. For the latter, we propose a vehicle-mask transformer to enhance the attention of the model on local details while reducing the impact of background information. Extensive experiments on two public datasets demonstrate the superiority of IUID over the current state-of-the-arts methods. Zefeng Lu, Ronghao Lin, Xulei Lou, Lifeng Zheng, Haifeng Hu 0001 |
IEEE Trans. Intell. Transp. Syst. | 1 |