Zaiyu Pan

dblp:205/7734 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
12since 2021 · last 2026
0000-0003-2946-6296ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Security and privacy · 5 · 2 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021
YearPublicationVenuePosition
2026 Progressive collaborative adversarial learning with missing modality for comprehensive multi-modal hand biometrics dataset
Hai Yuan, Zhengwen Shen, Zaiyu Pan
Expert Syst. Appl.6
2025 A Mutual Distillation Learning Framework for Multimodal Biometric Recognition with Uncertain Missing Modality
abstract
Currently, multimodal biometric recognition technology is receiving increasing attention. Most existing multimodal biometric recognition techniques require complete multimodal data during both the training and testing phases. However, due to various limitations, it is challenging to obtain complete and high-quality multimodal biometric data. To address this problem, we proposed a Mutual Distillation Framework (MDF) for palmprint and palmvein based multimodal biometric recognition with uncertain missing modality. Specifically, we firstly design an intra-inter modal data missing augmentation mechanism to generate heterogeneous missing samples. Moreover, a mutual knowledge distillation module is introduced to establish correlation of beneficial semantic information across different missing scenarios. In particular, we incorporate a Sample Semantic Discrepancy Guided mechanism within the mutual distillation framework, which calculates the semantic discrepancy between samples under corresponding heterogeneous missing patterns to encourage the model to focus on samples containing more complementary information. Experimental results demonstrate that the proposed model outperforms state-of-the-art incomplete multimodal learning models across three multimodal biometric benchmark datasets.
Shuangtian Jiang, Hai Yuan, Jun Wang 0071, Zaiyu Pan
IJCB5
2025 LMBR-Net: A Lightweight Multimodal Biometric Recognition Network via Joint Progressive Dynamic Sparsity
abstract
Multimodal biometric recognition technology has shown great potential in enhancing both recognition accuracy and robustness. However, challenges persist in controlling network parameter size and preserving fine-grained features. While existing lightweight methods excel in unimodal tasks, they often struggle to effectively balance parameter scale with fusion performance in multimodal scenarios. In this paper, we propose a lightweight multimodal recognition network based on joint progressive dynamic sparsity (LMBR-Net). Our method introduces the Joint Progressive Dynamic Sparsity Module (JPDSM), which dynamically reduces parameter redundancy while preserving complementary information across modalities. Additionally, the Lightweight Collaborative Sparse Feature Fusion Module (LCSFM) effectively captures global channel information, improving inter-modality fusion and ensuring consistency in the feature space while retaining fine-grained details across modalities. Experimental results demonstrate that the proposed network achieves superior recognition accuracy on multiple public datasets, maintaining a parameter count of approximately 3.64M, thus achieving an effective balance between network complexity and recognition performance.
Jun Wang 0071, Penghao Jin, Zaiyu Pan
IJCB5
2025 Rethinking Early-Fusion Strategy for Palmprint and Palm Vein Fusion Recognition
abstract
Multimodal biometric recognition has shown great potential in identity authentication tasks. Currently, most existing multimodal biometric recognition methods mainly employ multiple separate feature extraction networks for multimodal biometrics, leading to nearly multiple times the inference time compared to a single-stream feature extraction network. This increased inference time has hindered the widespread employment of multimodal biometric recognition in embedded devices for autonomous systems. To this end, we rethink the early-fusion strategy for palmprint and palm vein fusion recognition. First, an Image-Level Adaptive Fusion module (ILAF) is proposed to achieve the dynamic fusion of palmprint and palm vein images in pixel space, and then the fused images are inputted into the single-stream feature extraction network for identity recognition. Second, Modality-Aware Knowledge Distillation (MAKD) is presented, which includes PalmVein-Aware Knowledge Distillation (PVKD) and PalmPrint-Aware Knowledge Distillation (PPKD). This approach enables the single-stream feature extraction network to learn more modality-specific information of different modalities, thereby enhancing the discriminative ability of multimodal fusion representation. Experimental results on four multimodal biometric benchmark datasets demonstrate that our proposed model outperforms state-of-the-art multimodal biometric recognition models.
Penghao Jin, Jun Wang 0071, Zaiyu Pan
IJCB4
2025 Dynamic interaction and router selection network for multi-modality biometric recognition
Hai Yuan, Zaiyu Pan, Zhengwen Shen, Jun Wang 0071
Knowl. Based Syst.4
2025 Hierarchical Cross-Modal Image Generation for Multimodal Biometric Recognition With Missing Modality
abstract
Multimodal biometric recognition has shown great potential in identity authentication tasks and has attracted increasing interest recently. Currently, most existing multimodal biometric recognition algorithms require test samples with complete multimodal data. However, it often encounters the problem of missing modality data and thus suffers severe performance degradation in practical scenarios. To this end, we proposed a hierarchical cross-modal image generation for palmprint and palmvein based multimodal biometric recognition with missing modality. First, a hierarchical cross-modal image generation model is designed to achieve the pixel alignment of different modalities and reconstruct the image information of missing modality. Specifically, a cross-modal texture transfer network is utilized to implement the texture style transformation between different modalities, and then a cross-modal structure generation network is proposed to establish the correlation mapping of structural information between different modalities. Second, multimodal dynamic sparse feature fusion model is presented to obtain more discriminative and reliable representations, which can also enhance the robustness of our proposed model to dynamic changes in image quality of different modalities. The proposed model is evaluated on three multimodal biometric benchmark datasets, and experimental results demonstrate that our proposed model outperforms recent mainstream incomplete multimodal learning models.
Zaiyu Pan, Shuangtian Jiang, Hai Yuan, Jun Wang 0071
IEEE Trans. Inf. Forensics Secur.1
2024 Synergizing Global and Local Knowledge via Dynamic Focus Mechanism for Low-Light Image Enhancement
Shuyu Han, Zhengwen Shen, Yulian Li, Zaiyu Pan, Jun Wang 0071
PRCV (9)4
2024 HEFANet: hierarchical efficient fusion and aggregation segmentation network for enhanced rgb-thermal urban scene parsing
Zhengwen Shen, Zaiyu Pan, Yuchen Weng, Yulian Li, Jiangyu Wang, Jun Wang 0071
Appl. Intell.2
2023 Cross-domain collaborative learning for single image deraining
Zaiyu Pan, Jun Wang 0071, Zhengwen Shen, Shuyu Han, Jihong Zhu 0001
Expert Syst. Appl.1
2023 Disentangled Representation and Enhancement Network for Vein Recognition
abstract
Recently, vein recognition has been paid more considerable attention in biometric recognition fields. In the process of vein image acquisition, due to the influence of external factors such as illumination change, the texture information of vein images with the same identity information may change, which enhances the difference of intra-class and extremely degrades the performance of vein recognition systems. To address this problem, we proposed a Disentangled Representation and Enhancement Network for vein recognition called DRE-Net. First, robust vein shape masks are obtained by the designed vein segmentation algorithms, which are utilized as label information in DRE-Net to acquire the shape features of vein images. Second, a disentangled representation network which contains two encoders and two decoders is designed to disentangle the texture features and shape features of vein images. Besides, a Multi-Scale Attention Residual Block (MSARB) is presented to better mine the vein network information and enhance the representation ability of DRE-Net for vein images. Finally, a Weight-Guided Feature Enhancement Module (WGFEM) is presented to obtain more discriminative representations for vein recognition by reducing the importance of texture features and increasing the importance of shape features, which decreases the difference of intra-class caused by illumination change. Extensive experiments have been carried out on three benchmark vein databases, and the experimental results demonstrate that our proposed model outperforms the state-of-the-art methods.
Zaiyu Pan, Jun Wang 0071, Zhengwen Shen, Shuyu Han
IEEE Trans. Circuits Syst. Video Technol.1
2022 CTFusion: Convolutions Integrate with Transformers for Multi-modal Image Fusion
Zhengwen Shen, Jun Wang 0071, Zaiyu Pan, Jiangyu Wang, Yulian Li
PRCV (1)3
2021 Multi-Scale Deep Representation Aggregation for Vein Recognition
abstract
The recent success of Deep Convolutional Neural Network (DCNN) for various computer vision tasks such as image recognition has already demonstrated its robust feature representation ability. However, the limitation of training database on small scale vein recognition tasks restricts its performance because the recognition result of DCNN depends heavily on the number of trainsets. This motivates the design of a Multi-Scale Deep Representation Aggregation (MSDRA) model based on a pre-trained DCNN for vein recognition. First, the multi-scale feature maps are extracted by a pre-trained DCNN model. Second, a local mean threshold approach is designed to preliminarily remove the noisy information of multi-scale feature maps and generate the selected feature maps. Third, we propose an Unsupervised Vein Information Mining (UVIM) method to localize vein information of selected feature maps for generating a binary vein information mask, and then the vein information mask is utilized to keep useful deep representation and discard the background information. Finally, the discriminative multi-scale deep representations, which are generated by using the vein information mask to aggregate multi-scale feature maps, are concatenated into the final compact feature vectors, and then a Support Vector Machine (SVM) is introduced for final recognition. Our proposed model outperforms the state-of-the-art methods on two benchmark vein databases. Moreover, an additional experiment using the subset of PolyU Palmprint database illustrates the system's generalization ability and robustness.
Zaiyu Pan, Jun Wang 0071, Guoqing Wang 0001, Jihong Zhu 0001
IEEE Trans. Inf. Forensics Secur.1