Xinglong Mao

dblp:304/1504 · DBLP profile ↗
← Back
9ranked-venue papers
1as first author
9since 2021 · last 2025
0000-0003-0019-2295ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
YearPublicationVenuePosition
2025 DGRGaze: A Difference-Guided Gaze Estimation Framework Based on 6D Rotation Matrix Representation
abstract
Gaze estimation aims to infer a person’s gaze direction from images, which has wide applications in human-computer interaction. Recently, full-face-based gaze estimation methods have become increasingly prominent, owing to their efficiency and adaptability. However, irrelevant information in facial images limits estimation accuracy. To address this challenge, we propose DGRGaze, a difference-guided gaze estimation framework based on 6D rotation matrix representation. By incorporating an auxiliary task, namely predicting the gaze differences between facial image pairs, DGRGaze becomes more sensitive to gaze-related features. Additionally, a novel 6D rotation matrix representation is introduced to enhance the efficiency of network learning and resolve ambiguities in large-angle scenarios. Furthermore, we design a multi-task loss function based on the geodesic distance of the rotation matrix, enabling more precise quantification of gaze differences. Extensive experiments demonstrate our method achieves state-of-the-art performance on benchmark datasets, proving its effectiveness.
Xiaohao Wang, Sirui Zhao, Xinglong Mao, Tong Xu 0001, Enhong Chen
ICIP3
2025 MER-CLIP: AU-Guided Vision-Language Alignment for Micro-Expression Recognition
abstract
As a critical psychological stress response, micro-expressions (MEs) are fleeting and subtle facial movements revealing genuine emotions. Automatic ME recognition (MER) holds valuable applications in fields such as criminal investigation and psychological diagnosis. The Facial Action Coding System (FACS) encodes expressions by identifying activations of specific facial action units (AUs), serving as a key reference for ME analysis. However, current MER methods typically limit AU utilization to defining regions of interest (ROIs) or relying on specific prior knowledge, often resulting in limited performance and poor generalization. To address this, we integrate the CLIP model's powerful cross-modal semantic alignment capability into MER and propose a novel approach namely MER-CLIP. Specifically, we convert AU labels into detailed textual descriptions of facial muscle movements, guiding fine-grained spatiotemporal ME learning by aligning visual dynamics and textual AU-based representations. Additionally, we introduce an Emotion Inference Module to capture the nuanced relationships between ME patterns and emotions with higher-level semantic understanding. To mitigate overfitting caused by the scarcity of ME data, we put forward LocalStaticFaceMix, an effective data augmentation strategy blending facial images to enhance facial diversity while preserving critical ME features. Finally, comprehensive experiments on four benchmark ME datasets confirm the superiority of MER-CLIP. Notably, UF1 scores on CAS(ME)$^{3}$reach 0.7832, 0.6544, and 0.4997 for 3-, 4-, and 7-class classification tasks, significantly outperforming previous methods.
Xinglong Mao, Sirui Zhao, Peiming Li, Tong Xu 0001, Enhong Chen
IEEE Trans. Affect. Comput.2
2024 TGMAE: Self-supervised Micro-Expression Recognition with Temporal Gaussian Masked Autoencoder
abstract
Micro-expressions (MEs) are fleeting, subtle, and involuntary facial expressions that can reveal genuine emotions of human beings. Although many advanced supervised deep learning efforts have been devoted to ME recognition (MER), they are severely limited by the lack of sufficient well-labeled ME data when learning discriminative ME features. To address this problem, we propose a novel self-supervised ME representation learning method based on Temporal Gaussian Masked Autoencoder, termed TGMAE. Specifically, a Temporal Gaussian Masking strategy is customized to construct a challenging spatiotemporal ME movement reconstruction task, which can effectively assist the model in perceiving ME features from abundant unlabeled ME data. Additionally, to bridge the semantic gap between encoded features for reconstruction and emotion features for recognition, a bridging classifier is introduced for downstream MER. Comprehensive experiments demonstrate the remarkable performance of TGMAE, significantly surpassing the second-best method with a maximum improvement of 4.96% in UF1 and 6.92% in UAR.
Xinglong Mao, Sirui Zhao, Chaoyou Fu, Tong Xu 0001, Enhong Chen
ICME2
2024 A Multi-scale Feature Learning Network with Optical Flow Correction for Micro- and Macro-expression Spotting
abstract
Recently, automatic micro-expression (ME) analysis has attracted increasing attention, since ME is a spontaneous facial expression that can truly reflect the emotional state an individual tries to conceal. As a crucial step in ME analysis, Micro- and Macro-expression (MaE) spotting aims to sequentially identify the occurrence intervals of MEs and MaEs within a long video sequence. However, the subtle spatiotemporal movements of MEs and the scarcity of well-labeled data pose great challenges for accurately spotting them. To this end, this paper proposes a novel spotting framework based on Multi-scale Feature Learning Network with Optical Flow Correction. Specifically, we first integrate the pre-trained VideoMAE and customized convolutional layers as a visual feature extraction module to learn the facial motion features in long video sequences. Then, to comprehensively locate and identify the existing ME and MaE segments, we introduce a multi-scale candidate segment generation method based on the ActionFormer. In particular, a multi-start points optical flow filtering method is proposed to improve the precision of expression spotting. Finally, we conduct comprehensive experiments on the MEGC2024 spotting task, and the experimental results demonstrate the effectiveness of our method, which ranks second in this task. The implemented code is also publicly available at https://github.com/zzy188zzy/megc_spotting_code.
Zhengye Zhang, Sirui Zhao, Xinglong Mao, Hao Wang 0076, Tong Xu 0001, Enhong Chen
ACM Multimedia3
2024 H2LMER: A Cross Frame-Rate Representation Alignment Framework for Micro-expression Recognition
Xinglong Mao, Sirui Zhao, Hao Wang 0076, Tong Xu 0001, Enhong Chen
PRCV (11)1
2024 DFME: A New Benchmark for Dynamic Facial Micro-Expression Recognition
abstract
One of the most important subconscious reactions, micro-expression (ME), is a spontaneous, subtle, and transient facial expression that reveals human beings' genuine emotion. Therefore, automatically recognizing ME (MER) is becoming increasingly crucial in the field of affective computing, providing essential technical support for lie detection, clinical psychological diagnosis, and public safety. However, the ME data scarcity has severely hindered the development of advanced data-driven MER models. Despite the recent efforts by several spontaneous ME databases to alleviate this problem, there is still a lack of sufficient data. Hence, in this paper, we overcome the ME data scarcity problem by collecting and annotating a dynamic spontaneous ME database with the largest current ME data scale called DFME (Dynamic Facial Micro-expressions). Specifically, the DFME database contains 7,526 well-labeled ME videos spanning multiple high frame rates, elicited by 671 participants and annotated by more than 20 professional annotators over three years. Furthermore, we comprehensively verify the created DFME, including using influential spatiotemporal video feature learning models and MER models as baselines, and conduct emotion classification and ME action unit classification experiments. The experimental results demonstrate that the DFME database can facilitate research in automatic MER, and provide a new benchmark for this field. DFME will be published via https://mea-lab-421.github.io.
Sirui Zhao, Huaying Tang, Xinglong Mao, Hao Wang 0076, Tong Xu 0001, Enhong Chen
IEEE Trans. Affect. Comput.3
2023 Adaptive Graph Attention Network with Temporal Fusion for Micro-Expressions Recognition
abstract
Automatic micro-expression recognition (MER) has essential applications in the psychological field. Graph-based models, due to their advantages in analyzing regionalized faces, have become a powerful method for MER. However, how to construct a graph from ME videos remains to be studied. To solve this problem, we design an adaptive graph attention network with temporal fusion to model the dynamic relationships between facial regions of interest (ROIs). Specifically, we first propose adaptive graph attention to establish learnable spatial graphs from ME videos. Then, we adopt an optical-flow-based feature as the suitable input for the graph network. In addition, an implicit semantic data augmentation algorithm is employed and improved as a data-driven weighted loss for better performance. Extensive experiments on SMIC-HS, CASME II and SAMM datasets have demonstrated the effectiveness of the proposed method, and it achieves to be the first graph-based model where UF1 and UAR both exceed 0.90 for 3-classes MER on CASME II. Code will be available at https://github.com/MEA-LAB-421/ICME2023-Recognition.
Hao Wang 0076, Yifan Xu 0011, Xinglong Mao, Tong Xu 0001, Sirui Zhao, Enhong Chen
ICME4
2022 ABPN: Apex and Boundary Perception Network for Micro- and Macro-Expression Spotting
abstract
Recently, Micro expression~(ME) has achieved remarkable progress in a wide range of applications, since it's an involuntary facial expression that reflects personal psychological state truly. In the procedure of ME analysis, spotting ME is an essential step, and is non trivial to be detected from a long interval video because of the short duration and low intensity issues. To alleviate this problem, in this paper, we propose a novel Micro- and Macro-Expression~(MaE) Spotting framework based on Apex and Boundary Perception Network~(ABPN), which mainly consists of three parts, i.e., video encoding module ~(VEM), probability evaluation module~(PEM), and expression proposal generation module~(EPGM). Firstly, we adopt Main Directional Mean Optical Flow (MDMO) algorithm and calculate optical flow differences to extract facial motion features in VEM, which can alleviate the impact of head movement and other areas of the face on ME spotting. Then, we extract temporal features with one-dimension convolutional layers and introduce PEM to infer the auxiliary probability that each frame belongs to an apex or boundary frame. With these frame-level auxiliary probabilities, the EPGM further combines the frames from different categories to generate expression proposals for the accurate localization. Besides, we conduct comprehensive experiments on MEGC2022 spotting task, and demonstrate that our proposed method achieves significant improvement with the comparison of state-of-the-art baselines on rm CAS(ME)2 and SAMM-LV datasets. The implemented code is also publicly available at https://github.com/wenhaocold/USTC_ME_Spotting.
Wenhao Leng, Sirui Zhao, Xinglong Mao, Hao Wang 0076, Tong Xu 0001, Enhong Chen
ACM Multimedia5
2021 FAMGAN: Fine-grained AUs Modulation based Generative Adversarial Network for Micro-Expression Generation
abstract
Micro-expressions (MEs) are significant and effective clues to reveal the true feelings and emotions of human beings, and thus MEs analysis is widely used in different fields such as medical diagnosis, interrogation and security. However, it is extremely difficult to elicit and label MEs, resulting in a lack of sufficient MEs data for MEs analysis. To address this challenge and inspired by the current face generation technology, in this paper we introduce Generative Adversarial Network based on fine-grained Action Units (AUs) modulation to generate MEs sequence (FAMGAN). Specifically, after comprehensively analyzing the factors that lead to inaccurate AU values detection, we performed fine-grained AUs modulation, which includes carefully eliminating the various noises and dealing with the asymmetry of AUs intensity. Additionally, we incorporate super-resolution into our model to enhance the quality of the generated images. Through experiments, we show that the system achieves very competitive results on the Micro-Expression Grand Challenge (MEGC2021).
Yifan Xu 0011, Sirui Zhao, Huaying Tang, Xinglong Mao, Tong Xu 0001, Enhong Chen
ACM Multimedia4