VLDB 2026 Research / reviewers in the wild / expert
Zi Wang 0013
dblp:78/8711-13
· DBLP profile ↗
12ranked-venue papers
4as first author
12since 2021 · last 2026
0000-0002-8001-0318ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 2 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Security and privacy · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Progressive Multi-modal Knowledge Distillation for Multi-spectral Object Re-identificationabstractIn the field of multi-spectral object re-identification (ReID), multi-modal knowledge and modal-specific knowledge exhibit complementary advantages when handling hard samples, but existing methods rarely integrate this collaborative information. Knowledge distillation is a direct approach for transferring information, however, heterogeneity in model architectures and variations in sample hardness can undermine the stability and controllability of knowledge transfer. To alleviate these limitations, we propose the novel Progressive Multi-modal Knowledge Distillation (PMKD) framework that enables multi-stage knowledge transfer guided by hard sample awareness. In the multi-modal knowledge transfer stage, the source model (pre-trained on multi-modal data) disseminates its learned multi-modal collaborative knowledge to multiple independently modal-specific target models, guiding their adaptation to hard samples within training batches. In the modal-specific knowledge retention stage, the independent models enriched with multi-modal knowledge guide the training phase. The architectural consistency between source-target models ensures more lossless knowledge transfer, effectively mitigating the risk of capability drift, and preserving inherent competence. Moreover, the entire progressive multi-modal knowledge distillation is regulated by the proposed hardness-aware distillation loss, which automatically adapts distillation intensity through hard sample mining, thereby ensuring stable transfer of hard sample handling capabilities. Extensive experiments on benchmark multi-spectral ReID datasets validate the effectiveness and superior performance of the proposed method. Aihua Zheng, Zi Wang 0013, Jin Tang 0001 |
AAAI | 3 |
| 2026 | Semantic-Driven Visual Progressive Refinement for Aerial-Ground Person ReID: A Challenging Large-Scale BenchmarkabstractAerial-Ground Person Re-IDentification (AGPReID) aims to extract identity-discriminative representations from heterogeneous perspectives across different platforms in complex real-world environments. However, existing methods primarily focus on visual appearance modeling and make insufficient use of semantic attribute priors, which limits their ability to bridge the aerial-ground view gap. To address this limitation, we propose a Semantic-driven Visual Progressive Refinement framework for AGPReID (SVPR-ReID), which effectively leverages textual attribute priors to guide the extraction of fine-grained visual cues. Specifically, we design a View-Decoupled Feature Extractor that incorporates view-aware textual prompts to decouple view-invariant identity features. Then, to alleviate inter-class ambiguity, we propose an Attribute-Scattered Mixture-of-Experts module that integrates attribute semantics into the visual space, thereby improving discrimination among visually similar pedestrians. Finally, we design a Context-Vision Progressive Refinement module for progressive refinement of attribute and view-invariant features, obtaining robust cross-view identity representations. In particular, we contribute a comprehensive benchmark for AGPReID, named CP2108, which contains 142,817 images of 2,108 identities annotated with 22 attributes. Notably, it includes 191 identities captured across different times, enabling both short- and long-term ReID evaluation, addressing the limitation of existing datasets that focus only on short-term scenarios. Extensive experimental results validate the effectiveness of our SVPR-ReID on four AGPReID datasets. Aihua Zheng, Xixi Wan, Zi Wang 0013, Jin Tang 0001, Bin Luo 0001 |
AAAI | 4 |
| 2026 | Multi-level alignment network for unsupervised domain adaptive multi-modality object re-identification
Yusong Sheng, Yuhe Ding, Aihua Zheng, Zi Wang 0013, Jin Tang 0001 |
Knowl. Based Syst. | 5 |
| 2026 | Reliable Multi-Modal Object Re-Identification via Modality-Aware Graph ReasoningabstractMulti-modal data provides abundant and diverse object information, crucial for effective modal interactions in Re-Identification (ReID) task. However, existing approaches often overlook the quality variations in local features and fail to fully leverage the complementary information across modalities, particularly in cases where features are of low quality. In this paper, we propose to address this issue by leveraging a novel graph reasoning model, termed the Modality-aware Graph Reasoning Network (MGRNet). Specifically, we first construct modality-aware graphs to enhance the extraction of fine-grained local details by effectively capturing and modeling the relationships between patches. Subsequently, the selective graph nodes swap operation is employed to alleviate the adverse effects of low-quality local features by considering both local and global information, enhancing the representation of discriminative information. Finally, the swapped modality-aware graphs are fed into the local-aware graph reasoning module, which propagates multi-modal information to yield a reliable feature representation. Another advantage of the proposed graph reasoning approach is its ability to reconstruct missing modal information by exploiting inherent structural relationships, thereby minimizing disparities between different modalities. Experimental results on four benchmarks (RGBNT201, Market1501-MM, RGBNT100, MSVR310) indicate that the proposed method achieves state-of-the-art performance in multi-modal object ReID. The code for our method will be available upon acceptance. Xixi Wan, Aihua Zheng, Zi Wang 0013, Bo Jiang 0002, Jin Tang 0001, Jixin Ma 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2026 | REMIND: Retrieval-Augmented Reconstruction With Dual Memories for Modality-Missing Object Re-IdentificationabstractTo address the modality-missing object Re-Identification (Re-ID) task, a common strategy is to compensate for absent information by exploiting available modalities. However, existing reconstruction-based approaches suffer from two major limitations: 1) they often overlook modality-specific cues inherent in the missing modality; 2) they typically adopt a single-path reconstruction strategy. These issues result in incomplete representations and constrain the capacity to model complex semantic mappings across heterogeneous modalities. To address these challenges, we propose REMIND, a novel framework for modality-missing object Re-Identification, namely REtrieval-AugMented ReconstructIoN With Dual Memories. Specifically, we design a Dual Memory Construction module that, guided by information-theoretic insights, extracts modality-specific and modality-common features through two complementary branches and stores them in dedicated memory banks. These memory banks serve as structured prior knowledge to guide the reconstruction process, ensuring that the features of missing modalities are preserved even under modality-missing conditions. In addition, we have developed a retrieval-augmented missing reconstruction module that enhances the expressiveness and robustness of the reconstruction through multi-path reconstruction and perturbation mechanisms. Adaptive fusion techniques are employed for integration, simultaneously improving the expressiveness and robustness of the reconstructed features. Through the synergy of information-theoretically motivated regularization and retrieval-enhanced reconstruction, REMIND achieves robust feature recovery and delivers highly discriminative representations for reliable modality-missing Re-ID. Extensive experiments on several multi-modal object Re-ID benchmarks demonstrate the effectiveness and superiority of REMIND under various missing modality scenarios. The code is publicly available at: https://github.com/skye-1201/REMIND. Zhendong Xu, Zi Wang 0013, Aihua Zheng, Chenglong Li 0002, Jin Tang 0001 |
IEEE Trans. Image Process. | 2 |
| 2025 | Keypoint-guided feature enhancement and alignment for cross-resolution vehicle re-identification
Aihua Zheng, Zi Wang 0013, Chenglong Li 0002, Xiaofei Sheng |
Eng. Appl. Artif. Intell. | 4 |
| 2025 | Prototype-Based Diversity and Integrity Learning for All-Day Multi-Modal Person Re-IdentificationabstractRecent multi-modal person re-identification methods have improved model performance by leveraging complementary information from multiple spectra. However, existing methods cannot ensure feature stability under varying illumination and rely on inflexible paired data, remaining inadequate against real-world cross-time retrieval and modality-missing challenges. To solve these, we first propose diversity representation that augments illumination-sensitive images to simulate diverse lighting conditions via illumination augmentation and enriches instance features using modality-specific prototypes via multiple interaction modules. Secondly, we propose integrity reconstruction that leverages prototypes and available instance features to recover information, the reconstruction module effectively utilizes identity and modality cues to address unpredictable missing problems. In addition, we build a more comprehensive dataset (AllDay843) to alleviate the inadequate dataset diversity, which comprises 91,371 images of 843 identities captured by multi-modal cameras across various periods throughout the day, while incorporating numerous real-world challenges. By integrating diversity representation and integrity reconstruction, the proposed Prototype-Based Diversity and Integrity learning network (PDINet) establishes excellence on the AllDay843 dataset, surpassing existing state-of-the-art approaches. The data and codes are available in https://github.com/ziwang1121/PDINet. Zi Wang 0013, Chenglong Li 0002, Aihua Zheng, Jin Tang 0001, Bin Luo 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2024 | Heterogeneous Test-Time Training for Multi-Modal Person Re-identificationabstractMulti-modal person re-identification (ReID) seeks to mitigate challenging lighting conditions by incorporating diverse modalities. Most existing multi-modal ReID methods concentrate on leveraging complementary multi-modal information via fusion or interaction. However, the relationships among heterogeneous modalities and the domain traits of unlabeled test data are rarely explored. In this paper, we propose a Heterogeneous Test-time Training (HTT) framework for multi-modal person ReID. We first propose a Cross-identity Inter-modal Margin (CIM) loss to amplify the differentiation among distinct identity samples. Moreover, we design a Multi-modal Test-time Training (MTT) strategy to enhance the generalization of the model by leveraging the relationships in the heterogeneous modalities and the information existing in the test data. Specifically, in the training stage, we utilize the CIM loss to further enlarge the distance between anchor and negative by forcing the inter-modal distance to maintain the margin, resulting in an enhancement of the discriminative capacity of the ultimate descriptor. Subsequently, since the test data contains characteristics of the target domain, we adapt the MTT strategy to optimize the network before the inference by using self-supervised tasks designed based on relationships among modalities. Experimental results on benchmark multi-modal ReID datasets RGBNT201, Market1501-MM, RGBN300, and RGBNT100 validate the effectiveness of the proposed method. The codes can be found at https://github.com/ziwang1121/HTT. Zi Wang 0013, Huaibo Huang, Aihua Zheng, Ran He 0001 |
AAAI | 1 |
| 2024 | Parallel Augmentation and Dual Enhancement for Occluded Person Re-IdentificationabstractOccluded person re-identification (Re-ID), the task of searching for the same person’s images in occluded environments, has attracted lots of attention in the past decades. Recent approaches concentrate on improving performance on occluded data by data/feature augmentation or using extra models to predict occlusions. However, they ignore the imbalance problem in this task and can not fully utilize the information from the training data. To alleviate these two issues, we propose a simple yet effective method with Parallel Augmentation and Dual Enhancement (PADE), which is robust on both occluded and non-occluded data and does not require any auxiliary clues. First, we design a parallel augmentation mechanism (PAM) to generate more suitable occluded data to mitigate the negative effects of unbalanced data. Second, we propose the global and local dual enhancement strategy (DES) to promote the context information and details. Experimental results on three widely used occluded datasets and two non-occluded datasets validate the effectiveness of our method. The code is available at PADE (GitHub). Zi Wang 0013, Huaibo Huang, Aihua Zheng, Chenglong Li 0002, Ran He 0001 |
ICASSP | 1 |
| 2023 | Iterative embedding distillation for open world vehicle recognition
Junxian Duan, Xiang Wu 0001, Yibo Hu 0001, Chaoyou Fu, Zi Wang 0013, Ran He 0001 |
Pattern Recognit. | 5 |
| 2022 | Interact, Embed, and EnlargE: Boosting Modality-Specific Representations for Multi-Modal Person Re-identificationabstractMulti-modal person Re-ID introduces more complementary information to assist the traditional Re-ID task. Existing multi-modal methods ignore the importance of modality-specific information in the feature fusion stage. To this end, we propose a novel method to boost modality-specific representations for multi-modal person Re-ID: Interact, Embed, and EnlargE (IEEE). First, we propose a cross-modal interacting module to exchange useful information between different modalities in the feature extraction phase. Second, we propose a relation-based embedding module to enhance the richness of feature descriptors by embedding the global feature into the fine-grained local information. Finally, we propose multi-modal margin loss to force the network to learn modality-specific information for each modality by enlarging the intra-class discrepancy. Superior performance on multi-modal Re-ID dataset RGBNT201 and three constructed Re-ID datasets validate the effectiveness of the proposed method compared with the state-of-the-art approaches. Zi Wang 0013, Chenglong Li 0002, Aihua Zheng, Ran He 0001, Jin Tang 0001 |
AAAI | 1 |
| 2021 | Robust Multi-Modality Person Re-identificationabstractTo avoid the illumination limitation in visible person re-identification (Re-ID) and the heterogeneous issue in cross-modality Re-ID, we propose to utilize complementary advantages of multiple modalities including visible (RGB), near infrared (NI) and thermal infrared (TI) ones for robust person Re-ID. A novel progressive fusion network is designed to learn effective multi-modal features from single to multiple modalities and from local to global views. Our method works well in diversely challenging scenarios even in the presence of missing modalities. Moreover, we contribute a comprehensive benchmark dataset, RGBNT201, including 201 identities captured from various challenging conditions, to facilitate the research of RGB-NI-TI multi-modality person Re-ID. Comprehensive experiments on RGBNT201 dataset comparing to the state-of-the-art methods demonstrate the contribution of multi-modality person Re-ID and the effectiveness of the proposed approach, which launch a new benchmark and a new baseline for multi-modality person Re-ID. Aihua Zheng, Zi Wang 0013, Zi-Han Chen, Chenglong Li 0002, Jin Tang 0001 |
AAAI | 2 |