VLDB 2026 Research / reviewers in the wild / expert
Weisong Zhao
dblp:135/6303
· DBLP profile ↗
6ranked-venue papers
5as first author
6since 2021 · last 2025
0000-0003-3957-8590ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Security and privacy · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | GenFIQA: Generative Face Image Quality Assessment via Identity-conditioned Diffusion ModelabstractFace recognition (FR) systems are widely deployed but often struggle due to unconstrained image-capturing conditions. Face image quality assessment (FIQA), applied before recognition, mitigates these challenges by filtering out unreliable samples. Current leading FIQA methods evaluate image quality based on the characteristics observed within the FR model pipeline. However, they leave out the inherent differences in identity embeddings between high-and low-quality face images. To this end, we propose Gen-FIQA, which utilizes a generative model to probe and amplify this difference. Specifically, we extract the identity embedding from an input image using a pre-trained FR model, and then use it as a conditioning signal to generate several face images of the same identity. This generation process leverages the inherent prior in the generative model to translate the difference in identity embedding space back to pixel space. To quantify these differences, the quality score is computed as the average cosine similarity between embeddings from the original and generated images. To improve computational efficiency, we further distill GenFIQA into a lightweight regression-based variant, GenFIQA(R). Extensive experiments across five benchmark datasets and four FR models demonstrate the superiority of our methods over thirteen state-of-the-art FIQA methods. Zheyu Yan, Weisong Zhao, Kai Pang, Xiangyu Zhu 0001, Xiaoyu Zhang 0002, Zhen Lei 0001 |
IJCB | 2 |
| 2025 | Global Cross-Entropy Loss for Deep Face RecognitionabstractContemporary deep face recognition techniques predominantly utilize the Softmax loss function, designed based on the similarities between sample features and class prototypes. These similarities can be categorized into four types: in-sample target similarity, in-sample non-target similarity, out-sample target similarity, and out-sample non-target similarity. When a sample feature from a specific class is designated as the anchor, the similarity between this sample and any class prototype is referred to as in-sample similarity. In contrast, the similarity between samples from other classes and any class prototype is known as out-sample similarity. The terms target and non-target indicate whether the sample and the class prototype used for similarity calculation belong to the same identity or not. The conventional Softmax loss function promotes higher in-sample target similarity than in-sample non-target similarity. However, it overlooks the relation between in-sample and out-sample similarity. In this paper, we propose Global Cross-Entropy loss (GCE), which promotes 1) greater in-sample target similarity over both the in-sample and out-sample non-target similarity, and 2) smaller in-sample non-target similarity to both in-sample and out-sample target similarity. In addition, we propose to establish a bilateral margin penalty for both in-sample target and non-target similarity, so that the discrimination and generalization of the deep face model are improved. To bridge the gap between training and testing of face recognition, we adapt the GCE loss into a pairwise framework by randomly replacing some class prototypes with sample features. We designate the model trained with the proposed Global Cross-Entropy loss as GFace. Extensive experiments on several public face benchmarks, including LFW, CALFW, CPLFW, CFP-FP, AgeDB, IJB-C, IJB-B, MFR-Ongoing, and MegaFace, demonstrate the superiority of GFace over other methods. Additionally, GFace exhibits robust performance in general visual recognition task. Weisong Zhao, Xiangyu Zhu 0001, Haichao Shi, Xiaoyu Zhang 0002, Guoying Zhao 0001, Zhen Lei 0001 |
IEEE Trans. Image Process. | 1 |
| 2024 | Masked Face TransformerabstractThe COVID-19 pandemic makes wearing masks mandatory. Existing CNN-based face recognition (FR) systems suffer from severe performance degradation as masks occlude the vital facial regions. Recently, Vision Transformers have shown promising performance in various vision tasks with quadratic computation costs. Swin Transformer first proposes a successive window attention mechanism allowing the cross-window connection and more computational efficiency. Despite its potential, the deployment of Swin Transformer in masked face recognition encounters two challenges: 1) the attention range is insufficient to capture locally compatible face regions. 2) Masked face recognition can be defined as an occlusion-robust classification task with a known occlusion position, i.e., the position of the mask is minor-varying, which is overlooked but efficient in improving the model’s recognition accuracy. To alleviate the above problem, we propose a Masked Face Transformer (MFT) with Masked Face-compatible Attention (MFA). The proposed MFA 1) introduces two additional window partition configurations, e.g., row shift and column shift, to enlarge the attention range in Swin with invariant computation costs, and 2) suppresses the interaction between the masked and non-masked regions to retain their discrepancies. Additionally, as mask occlusion leads to a separation between the masked and non-masked samples of the same identity, we propose to explore the relationship between them by a ClassFormer module to enhance intra-class aggregation. Extensive experiments show that MFT outperforms state-of-the-art masked face recognition methods in both simulated and real masked face testing datasets. Weisong Zhao, Xiangyu Zhu 0001, Haichao Shi, Xiaoyu Zhang 0002, Zhen Lei 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2023 | Grouped Knowledge Distillation for Deep Face RecognitionabstractCompared with the feature-based distillation methods, logits distillation can liberalize the requirements of consistent feature dimension between teacher and student networks, while the performance is deemed inferior in face recognition. One major challenge is that the light-weight student network has difficulty fitting the target logits due to its low model capacity, which is attributed to the significant number of identities in face recognition. Therefore, we seek to probe the target logits to extract the primary knowledge related to face identity, and discard the others, to make the distillation more achievable for the student network. Specifically, there is a tail group with near-zero values in the prediction, containing minor knowledge for distillation. To provide a clear perspective of its impact, we first partition the logits into two groups, i.e., Primary Group and Secondary Group, according to the cumulative probability of the softened prediction. Then, we reorganize the Knowledge Distillation (KD) loss of grouped logits into three parts, i.e., Primary-KD, Secondary-KD, and Binary-KD. Primary-KD refers to distilling the primary knowledge from the teacher, Secondary-KD aims to refine minor knowledge but increases the difficulty of distillation, and Binary-KD ensures the consistency of knowledge distribution between teacher and student. We experimentally found that (1) Primary-KD and Binary-KD are indispensable for KD, and (2) Secondary-KD is the culprit restricting KD at the bottleneck. Therefore, we propose a Grouped Knowledge Distillation (GKD) that retains the Primary-KD and Binary-KD but omits Secondary-KD in the ultimate KD loss calculation. Extensive experimental results on popular face recognition benchmarks demonstrate the superiority of proposed GKD over state-of-the-art methods. Weisong Zhao, Xiangyu Zhu 0001, Xiaoyu Zhang 0002, Zhen Lei 0001 |
AAAI | 1 |
| 2023 | Cross-Architecture Distillation for Face RecognitionabstractTransformers have emerged as the superior choice for face recognition tasks, but their insufficient platform acceleration hinders their application on mobile devices. In contrast, Convolutional Neural Networks (CNNs) capitalize on hardware-compatible acceleration libraries. Consequently, it has become indispensable to preserve the distillation efficacy when transferring knowledge from a Transformer-based teacher model to a CNN-based student model, known as Cross-Architecture Knowledge Distillation (CAKD). Despite its potential, the deployment of CAKD in face recognition encounters two challenges: 1) the teacher and student share disparate spatial information for each pixel, obstructing the alignment of feature space, and 2) the teacher network is not trained in the role of a teacher, lacking proficiency in handling distillation-specific knowledge. To surmount these two constraints, 1) we first introduce a Unified Receptive Fields Mapping module (URFM) that maps pixel features of the teacher and student into local features with unified receptive fields, thereby synchronizing the pixel-wise spatial information of teacher and student. Subsequently, 2) we develop an Adaptable Prompting Teacher network (APT) that integrates prompts into the teacher, enabling it to manage distillation-specific knowledge while preserving the model's discriminative capacity. Extensive experiments on popular face benchmarks and two large-scale verification sets demonstrate the superiority of our method. Weisong Zhao, Xiangyu Zhu 0001, Zhixiang He, Xiaoyu Zhang 0002, Zhen Lei 0001 |
ACM Multimedia | 1 |
| 2022 | Consistent Sub-Decision Network for Low-Quality Masked Face RecognitionabstractThe COVID-19 pandemic makes wearing masks mandatory in supermarkets, pharmacies, public transport, etc. Existing facial recognition systems encounter severe performance degradation as the masks occlude key facial regions. Recently, simulation-based methods are proposed to generate masked faces from unmasked faces. However, among simulated faces, there are low-quality samples with negative occlusion, which leads to ambiguous or absent facial features. In this paper, we propose a consistent sub-decision network to obtain sub-decisions that correspond to different facial regions and constrain sub-decisions by weighted bidirectional KL divergence to make the network concentrate on the upper faces without occlusion. In addition, we perform knowledge distillation to drive the masked face embeddings towards an approximation of the original data distribution to mitigate the information loss. Experiments show that the proposed method performs better than the baseline on public masked face recognition datasets, i.e., RMFD, MFR2, and MLFW. Weisong Zhao, Xiangyu Zhu 0001, Haichao Shi, Xiaoyu Zhang 0002, Zhen Lei 0001 |
IEEE Signal Process. Lett. | 1 |