Yu Zhou 0049

dblp:36/2728-49 · DBLP profile ↗
← Back
12ranked-venue papers
6as first author
11since 2021 · last 2026
0000-0003-1937-5331ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CaMST: Certainty-aware matching self-training for semi-supervised few-shot learning
Kunlei Jing, Hebo Ma, Yu Zhou 0049
Neurocomputing4
2026 Parameter-free discrete clustering via adaptive hypergraph fusion
Yu Zhou 0049, Ben Yang, Xuetao Zhang 0001, Badong Chen
Inf. Sci.1
2026 Multi-modal deep facial expression recognition framework combining knowledge distillation and retrieval-augmented generation
abstract
In recent years, significant progress has been made in facial expression recognition (FER) methods based on deep learning. However, existing models still face challenges in terms of computational efficiency and generalization performance when dealing with diverse emotional expressions and complex environmental variations. Recently, large-scale vision-language pre-training models such as CLIP have achieved remarkable success in multi-modal learning. Their rich visual and textual representations offer valuable insights for downstream tasks. Consequently, transferring the knowledge to develop efficient and accurate facial expression recognition (FER) systems has emerged as a key research direction. To the end, this paper proposes a novel model, termed Knowledge Distillation and Retrieval-Augmented Generation (KDRAG), which combines Distillation and Retrieval-Augmented Generation (RAG) techniques to improve the efficiency and accuracy of FER. Through knowledge distillation, the teacher model (ViT-L/14) transfers its rich knowledge to the smaller student model (ViT-B/32). An additional linear projection layer is added to map the teacher model’s output features to the student model’s feature dimensions for feature alignment. Moreover, the RAG mechanism is developed to enhance the emotional understanding of students by retrieving text descriptions related to the input image. Additionally, this framework combines soft loss (from the teacher model’s knowledge) and hard loss (from the true targets of the labels) to enhance the model’s generalization ability. Extensive experimental results on multiple datasets demonstrate that the KDRAG framework can achieve significant improvements in accuracy and computational efficiency, providing new insights for real-time FER systems.
Beibei Jiang, Yu Zhou 0049
J. Vis. Commun. Image Represent.2
2026 Correntropy-induced multi-hypergraph discrete clustering
Yu Zhou 0049, Haoxin Wu, Jiajing Hu
Signal Process.1
2026 Efficient Structure-Aware Discrete Clustering via Multi-Order Anchor Graphs
Ben Yang, Xuetao Zhang 0001, Yu Zhou 0049, Haoxin Wu, Feiping Nie 0001, Badong Chen
IEEE Trans. Knowl. Data Eng.3
2025 Real-time facial expression recognition via quaternion Gabor convolutional neural network
Yu Zhou 0049, Liyuan Guo, Beibei Jiang, Kunlei Jing
J. Vis. Commun. Image Represent.1
2025 Delving Into Quaternion Wavelet Transformer for Facial Expression Recognition in the Wild
abstract
The Facial Expression Recognition (FER) technique has increasingly matured over time. However, recognizing facial expressions in wild environments poses great challenges in achieving promising performance. The main obstacles arise from various factors, such as illumination changes, head pose variations, and occlusions. To overcome interferences from external environments and improve recognition accuracy, we propose a novel Quaternion Wavelet TRansformer (QWTR) model for FER in the wild. Specifically, we present a Quaternion Value Transformer (QVT) network that combines quaternion multi-head attention with quaternion CNN to capture emotional cues from global and local perception. To preserve the color structure while enhancing image contrast and brightness, we introduce a Quaternion Histogram Equalization (QHE) representation to transform color images into quaternion matrices representation. After that, to alleviate the impact of head pose and occlusion together with feature redundancy, a Quaternion Wavelet Feature Selection (QWFS) scheme is designed to decompose quaternion features and select the most correlated signals. Extensive experiments have been conducted on four in-the-wild FER datasets and several specific FER benchmarks under various conditions. The qualitative and quantitative results demonstrate thatQWTRoutperforms other state-of-the-art methods in FER benchmarks, e.g., 68.37% vs. 66.31% accuracy on the AffectNet dataset.
Yu Zhou 0049, Jialun Pei, Weixin Si, Harry Qin, Pheng-Ann Heng
IEEE Trans. Multim.1
2024 Cross-Domain Facial Expression Recognition by Combining Transfer Learning and Face-Cycle Generative Adversarial Network
Yu Zhou 0049, Ben Yang, Zhenni Liu
Multim. Tools Appl.1
2024 Robust spectral embedded bilateral orthogonal concept factorization for clustering
Ben Yang, Jinghan Wu, Yu Zhou 0049, Xuetao Zhang 0001, Zhiping Lin 0001, Feiping Nie 0001, Badong Chen
Pattern Recognit.3
2024 Quaternion Deformable Local Binary Pattern and Pose-Correction Facial Decomposition for Color Facial Expression Recognition in the Wild
abstract
Facial expression recognition (FER) in the wild is a more challenging topic than that under laboratory-controlled conditions. The major obstacles of FER in the wild are head pose variations, illumination changes, and different skin colors. To address these problems, we propose a framework named quaternion deformable local binary pattern (QDLBP)-Net for color FER in the wild. First, to eliminate the interferences of head pose variations, a pose-correction facial decomposition (PCFD) strategy is proposed to correct the head pose and decompose the facial image into five emotion-related regions. Then, to handle the problems of illumination changes and different skin colors, an effective feature descriptor named “QDLBP” is developed. QDLBP extracts color quaternion features from each emotional region, which not only computes the strength of emotional features, but also maintains the spectral correlation between color channels. Finally, a quaternion classification network (QC-Net) is proposed to classify the quaternion features from five emotional regions into seven basic expressions. The experimental results on three in-the-wild FER datasets and two nonfrontal pose variation datasets exhibit the effectiveness and superiority of QDLBP-Net by showing clear performance improvements over other state-of-the-art (SOTA) FER methods.
Yu Zhou 0049, Guangzhi Ma, Enmin Song
IEEE Trans. Comput. Soc. Syst.2
2023 Quaternion Orthogonal Transformer for Facial Expression Recognition in the Wild
abstract
Facial expression recognition (FER) is a challenging topic in artificial intelligence. Recently, many researchers have attempted to introduce Vision Transformer (ViT) to the FER task. However, ViT cannot fully utilize emotional features extracted from raw images and requires a lot of computing resources. To overcome these problems, we propose a quaternion orthogonal transformer (QOT) for FER. Firstly, to reduce redundancy among features extracted from pre-trained ResNet-50, we use the orthogonal loss to decompose and compact these features into three sets of orthogonal sub-features. Secondly, three orthogonal sub-features are integrated into a quaternion matrix, which maintains the correlations between different orthogonal components. Finally, we develop a quaternion vision transformer (Q-ViT) for feature classification. The Q-ViT adopts quaternion operations instead of the original operations in ViT, which improves the final accuracies with fewer parameters. Experimental results on three in-the-wild FER datasets show that the proposed QOT outperforms several state-of-the-art models and reduces the computations.Codes are available at https://github.com/Gabrella/QOT.
Yu Zhou 0049, Liyuan Guo
ICASSP1
2020 Deformable Quaternion Gabor Convolutional Neural Network For Color Facial Expression Recognition
abstract
In facial expression recognition (FER), convolutional neural networks (CNNs) have been shown great capability of learning features. In this paper, we propose a new CNN framework for FER in color images, which incorporates deformable Gabor filters into a quaternion CNN. Deformable Gabor filters reinforce the network's ability of extracting different orientations of facial wrinkles information. Quaternion CNNs have greater advantages over the regular CNNs in handling the coupling between color channels. The proposed deformable quaternion Gabor convolutional neural network (DQG-CNN) not only learns FER feature representation excellently, but also processes spectral correlation between color channels naturally. Moreover, it can also effectively reduce training complexity compared to other reference models. Experimental results on three benchmark color datasets Oulu-CASIA, MMI, and SFEW, demonstrate that the proposed DGQ-CNN outperforms other state-of-the-art methods clearly.
Yu Zhou 0049, Hong Liu 0005, Enmin Song
ICIP2