Sirui Zhao

dblp:288/9976 · DBLP profile ↗
← Back
30ranked-venue papers
5as first author
30since 2021 · last 2026
0000-0001-8103-0321ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 17 · 2 first-author · 17 since 2021Artificial intelligence and machine learning · 9 · 3 first-author · 9 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Computer networks · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Unsupervised lightweight 3D convolutional network for enhanced infrared imaging in wearable devices
Biao Zhu, Sirui Zhao, Zhengye Zhang, Enhong Chen
Frontiers Comput. Sci.3
2026 Emotionally Controllable Audio-driven Talking Face Generation
abstract
Talking face generation is a technique that synthesizes realistic, speech-synchronized facial animations from static images driven by audio or textual input. In recent years, notable advancements have been achieved in the generation of realistic lip-synchronized talking face videos. Although these methods have demonstrated highly realistic results in lip-sync generation, they rarely focus on the expression of emotions, which is essential for achieving emotionally rich and lifelike talking face videos. In this article, we propose a novel emotionally controllable audio-driven talking face generation framework, termed ECATFG. Specifically, the proposed ECATFG mainly consists of three modules: template video generator, landmark generator, and rendering module. Firstly, we integrate a customized motion transfer module into StyleGAN to generate template videos that encapsulate the target emotions and head movements. Subsequently, a Mamba-based landmark generator is proposed to accurately model the correlation between input audio and facial landmarks. Afterwards, in the rendering module, we employ a framework that combines both AdaIN and SPADE layers to transform audio-predicted landmarks into highly realistic images, guided by reference images and facial contour landmarks. Finally, comprehensive experiments validate the effectiveness of ECATFG, demonstrating its capability to generate high-quality talking face videos that not only achieve exceptional lip-synchronization accuracy but also effectively convey a wide range of emotions.
Yifan Xu 0011, Sirui Zhao, Tong Xu 0001, Enhong Chen
ACM Trans. Multim. Comput. Commun. Appl.2
2025 Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
abstract
In the quest for artificial general intelligence, Multi-modal Large Language Models (MLLMs) have emerged as a focal point in recent advancements. However, the predominant focus remains on developing their capabilities in static image understanding. The potential of MLLMs to process sequential visual data is still insufficiently explored, highlighting the lack of a comprehensive, high-quality assessment of their performance. In this paper, we introduce Video-MME, the first-ever full-spectrum, Multi-Modal Evaluation benchmark of MLLMs in Video analysis. Our work distinguishes from existing benchmarks through four key features: 1) Diversity in video types, spanning 6 primary visual domains with 30 subfields to ensure broad scenario generalizability; 2) Duration in temporal dimension, encompassing both short-, medium-, and long-term videos, ranging from 11 seconds to 1 hour, for robust contextual dynamics; 3) Breadth in data modalities, integrating multi-modal inputs besides video frames, including subtitles and audios, to unveil the all-round capabilities of MLLMs; 4) Quality in annotations, utilizing rigorous manual labeling by expert annotators to facilitate precise and reliable model assessment. With Video-MME, we extensively evaluate various state-of-the-art MLLMs, and reveal that Gemini 1.5 Pro is the best-performing commercial model, significantly outperforming the open-source models with an average accuracy of 75%, compared to 71.9% for GPT-4o. The results also demonstrate that Video-MME is a universal benchmark that applies to both image and video MLLMs. Further analysis indicates that subtitle and audio information could significantly enhance video understanding. Besides, a decline in MLLM performance is observed as video duration increases for all models. Our dataset along with these findings underscores the need for further improvements in handling longer sequences and multi-modal data, shedding light on future MLLM development. Project page: https://video-mme.github.io.
Chaoyou Fu, Yuhan Dai, Yongdong Luo, Shuhuai Ren, Renrui Zhang, Yunhang Shen, Mengdan Zhang, Peixian Chen, Shaohui Lin, Sirui Zhao, Ke Li 0015, Tong Xu 0001, Xiawu Zheng, Enhong Chen, Caifeng Shan, Ran He 0001, Xing Sun 0001
CVPR14
2025 JoyLive: Efficient Audio-Driven Portrait Animation by 3D Implict Keypoints
Siyuan Jin, Sirui Zhao, Yifan Xu 0011, Mengduo Wu, Tong Xu 0001
ICIC (15)2
2025 DGRGaze: A Difference-Guided Gaze Estimation Framework Based on 6D Rotation Matrix Representation
abstract
Gaze estimation aims to infer a person’s gaze direction from images, which has wide applications in human-computer interaction. Recently, full-face-based gaze estimation methods have become increasingly prominent, owing to their efficiency and adaptability. However, irrelevant information in facial images limits estimation accuracy. To address this challenge, we propose DGRGaze, a difference-guided gaze estimation framework based on 6D rotation matrix representation. By incorporating an auxiliary task, namely predicting the gaze differences between facial image pairs, DGRGaze becomes more sensitive to gaze-related features. Additionally, a novel 6D rotation matrix representation is introduced to enhance the efficiency of network learning and resolve ambiguities in large-angle scenarios. Furthermore, we design a multi-task loss function based on the geodesic distance of the rotation matrix, enabling more precise quantification of gaze differences. Extensive experiments demonstrate our method achieves state-of-the-art performance on benchmark datasets, proving its effectiveness.
Xiaohao Wang, Sirui Zhao, Xinglong Mao, Tong Xu 0001, Enhong Chen
ICIP2
2025 PhysFFTFormer: A Frequency Domain-based Vision Transformer for Efficient Remote Physiological Measurement
abstract
Remote Photoplethysmography (rPPG) is a non-contact technique for extracting physiological signals from facial videos. Recently, Transformer-based architectures have exhibited remarkable performance in rPPG estimation, owing to their excellent long-range spatial-temporal modeling capacities. However, challenges persist in applying Transformer-based rPPG methods, where quadratic computational costs and inadequate feature modeling diversity remain formidable. To address these challenges, we leverage the power of the Frequency Domain-based Vision Transformer and propose an end-to-end model PhysFFTFormer. Specifically, by integrating customized Frequency-Domain Spatiotemporal Self-Attention and Frequency-Domain Discriminative Feed-Forward modules, PhysFFTFormer efficiently captures spatial-temporal dependencies with reduced computational complexity. To extract rich and diverse feature representations, we design a dual-pathway architecture to utilize both raw and differential video frames. Furthermore, the Frequency-Domain Spatiotemporal Cross-Attention module is introduced to enhance information exchange and enable feature complementation between the two pathways. Extensive experiments on multiple benchmark datasets demonstrate PhysFFTFormer’s state-of-the-art performance, robustness, and potential for real-world non-contact health monitoring.
Sirui Zhao, Tong Xu 0001, Yu Sun 0021, Hao Wang 0076, Suojuan Zhang, Enhong Chen
ICME2
2025 DepFormer: A Unified Framework with Bimodal Collaborative Transformer for Depression Detection
abstract
Depression is an increasingly prevalent mental health issue worldwide, especially among the elderly, where effective early detection is crucial for timely intervention. While recent multimodal approaches demonstrate promise in leveraging visual and auditory cues for automatic depression recognition, existing methods often fail to extract fine-grained, depression-specific patterns from pre-extracted features and underutilize cross-modal interactions. To address these challenges, we propose DepFormer, a unified framework that incorporates a Bimodal Collaborative Transformer(BCT) for cross-modal representation learning and a personalized fusion module to enhance individual-specific modeling. The architecture comprises: (1) unimodal feature extraction, (2) bimodal collaborative representation learning via the BCT, (3) personalized feature fusion, and (4) final depression classification. Critically, the BCT employs symmetric bidirectional branches comprising an Audio-to-Video Transformer and a Video-to-Audio Transformer to enable mutual enhancement and complementary learning across modalities. Extensive experiments validate DepFormer's effectiveness, securing first place in the MPDD Challenge 2025 (Elderly Track), underscoring its strong practical potential.
Sirui Zhao, Tong Xu 0001, Enhong Chen
ACM Multimedia2
2025 MER-CLIP: AU-Guided Vision-Language Alignment for Micro-Expression Recognition
abstract
As a critical psychological stress response, micro-expressions (MEs) are fleeting and subtle facial movements revealing genuine emotions. Automatic ME recognition (MER) holds valuable applications in fields such as criminal investigation and psychological diagnosis. The Facial Action Coding System (FACS) encodes expressions by identifying activations of specific facial action units (AUs), serving as a key reference for ME analysis. However, current MER methods typically limit AU utilization to defining regions of interest (ROIs) or relying on specific prior knowledge, often resulting in limited performance and poor generalization. To address this, we integrate the CLIP model's powerful cross-modal semantic alignment capability into MER and propose a novel approach namely MER-CLIP. Specifically, we convert AU labels into detailed textual descriptions of facial muscle movements, guiding fine-grained spatiotemporal ME learning by aligning visual dynamics and textual AU-based representations. Additionally, we introduce an Emotion Inference Module to capture the nuanced relationships between ME patterns and emotions with higher-level semantic understanding. To mitigate overfitting caused by the scarcity of ME data, we put forward LocalStaticFaceMix, an effective data augmentation strategy blending facial images to enhance facial diversity while preserving critical ME features. Finally, comprehensive experiments on four benchmark ME datasets confirm the superiority of MER-CLIP. Notably, UF1 scores on CAS(ME)$^{3}$reach 0.7832, 0.6544, and 0.4997 for 3-, 4-, and 7-class classification tasks, significantly outperforming previous methods.
Xinglong Mao, Sirui Zhao, Peiming Li, Tong Xu 0001, Enhong Chen
IEEE Trans. Affect. Comput.3
2025 MF-GSLAE: A Multi-Factor User Representation Pre-Training Framework for Dual-Target Cross-Domain Recommendation
abstract
Recently, the dual-target cross-domain recommendation has been an emerging research problem, which aims to improve the performances of both source and target domains by transferring the preferences of overlapping users. Most of the existing work adopted a coarse-grained manner to detach general users’ preferences and associate them with domain-specific information for enhancing user representation learning, which fails to depict the differences in users’ diverse preferences and aggregate relevant preferences with improper propagation. To this end, in this article, we propose a multi-factor user representation pre-training framework, dubbed MF-GSLAE, with a focus on fine-grained preference learning and transferring. Specifically, we first propose a fine-grained factor representation pre-training paradigm. It projects the behavior records of both domains into several subspaces and introduces a compactness regularization to generate multiple fine-grained preference factors. Furthermore, we propose a multi-factor graph structure learning method within linear complexity to efficiently construct preference connections on different scales of users, which could aggregate the intrinsic relationship of user preferences in immediate embedding spaces to capture high-order information. Following the pre-training, we subsequently design a factor selection module with the bootstrapping mechanism to adaptively choose the corresponding domain-related preferences and transfer domain-shared information through partial overlapping factors for addressing the negative transfer problem. Finally, the optimization objectives of both domains are formalized in a multi-task learning framework and derive the learned user representation in an end-to-end training manner. Extensive experimental results on several publicly available datasets have not only demonstrated the effectiveness of the learned user representations with the comparison of state-of-the-art baselines but also indicated the interpretability and robustness. The code of our work is publicly available at https://github.com/USTC-StarTeam/MF-GSLAE .
Hao Wang 0076, Mingjia Yin, Luankang Zhang, Sirui Zhao, Enhong Chen
ACM Trans. Inf. Syst.4
2024 TGMAE: Self-supervised Micro-Expression Recognition with Temporal Gaussian Masked Autoencoder
abstract
Micro-expressions (MEs) are fleeting, subtle, and involuntary facial expressions that can reveal genuine emotions of human beings. Although many advanced supervised deep learning efforts have been devoted to ME recognition (MER), they are severely limited by the lack of sufficient well-labeled ME data when learning discriminative ME features. To address this problem, we propose a novel self-supervised ME representation learning method based on Temporal Gaussian Masked Autoencoder, termed TGMAE. Specifically, a Temporal Gaussian Masking strategy is customized to construct a challenging spatiotemporal ME movement reconstruction task, which can effectively assist the model in perceiving ME features from abundant unlabeled ME data. Additionally, to bridge the semantic gap between encoded features for reconstruction and emotion features for recognition, a bridging classifier is introduced for downstream MER. Comprehensive experiments demonstrate the remarkable performance of TGMAE, significantly surpassing the second-best method with a maximum improvement of 4.96% in UF1 and 6.92% in UAR.
Xinglong Mao, Sirui Zhao, Chaoyou Fu, Tong Xu 0001, Enhong Chen
ICME3
2024 Dataset Regeneration for Sequential Recommendation
abstract
The sequential recommender (SR) system is a crucial component of modern recommender systems, as it aims to capture the evolving preferences of users. Significant efforts have been made to enhance the capabilities of SR systems. These methods typically follow the model-centric paradigm, which involves developing effective models based on fixed datasets. However, this approach often overlooks potential quality issues and flaws inherent in the data. Driven by the potential of data-centric AI, we propose a novel data-centric paradigm for developing an ideal training dataset using a model-agnostic dataset regeneration framework called DR4SR. This framework enables the regeneration of a dataset with exceptional cross-architecture generalizability. Additionally, we introduce the DR4SR+ framework, which incorporates a model-aware dataset personalizer to tailor the regenerated dataset specifically for a target model. To demonstrate the effectiveness of the data-centric paradigm, we integrate our framework with various model-centric methods and observe significant performance improvements across four widely adopted datasets. Furthermore, we conduct in-depth analyses to explore the potential of the data-centric paradigm and provide valuable insights. The code can be found at https://github.com/USTC-StarTeam/DR4SR.
Mingjia Yin, Hao Wang 0076, Wei Guo 0006, Yong Liu 0020, Suojuan Zhang, Sirui Zhao, Defu Lian, Enhong Chen
KDD6
2024 Speak From Heart: An Emotion-Guided LLM-Based Multimodal Method for Emotional Dialogue Generation
abstract
Recent advancements in Large Language Models~(LLMs) have greatly enhanced the generation capabilities of dialogue systems. However, progress on emotional expression during dialogues might be still limited, especially when capturing and processing the multimodal cues for emotional expression. Therefore, it is urgent to fully adapt the multimodal understanding ability and transferability of LLMs to enhance the emotional-oriented multimodal processing capabilities. To that end, in this paper, we propose a novel Emotion-Guided Multimodal Dialogue model based on LLM, termed ELMD. Specifically, to enhance the emotional expression ability of LLMs, our ELMD customizes an emotional retrieval module, which mainly provides appropriate response demonstration for LLM in understanding emotional context. Subsequently, a two-stage training strategy is proposed, founded on previous demonstration support, to support uncovering nuanced emotions behind multimodal information and constructing natural responses. Comprehensive experiments demonstrate the effectiveness and superiority of ELMD.
Chenxiao Liu, Zheyong Xie, Sirui Zhao, Tong Xu 0001, Minglei Li 0001, Enhong Chen
ICMR3
2024 A Multi-scale Feature Learning Network with Optical Flow Correction for Micro- and Macro-expression Spotting
abstract
Recently, automatic micro-expression (ME) analysis has attracted increasing attention, since ME is a spontaneous facial expression that can truly reflect the emotional state an individual tries to conceal. As a crucial step in ME analysis, Micro- and Macro-expression (MaE) spotting aims to sequentially identify the occurrence intervals of MEs and MaEs within a long video sequence. However, the subtle spatiotemporal movements of MEs and the scarcity of well-labeled data pose great challenges for accurately spotting them. To this end, this paper proposes a novel spotting framework based on Multi-scale Feature Learning Network with Optical Flow Correction. Specifically, we first integrate the pre-trained VideoMAE and customized convolutional layers as a visual feature extraction module to learn the facial motion features in long video sequences. Then, to comprehensively locate and identify the existing ME and MaE segments, we introduce a multi-scale candidate segment generation method based on the ActionFormer. In particular, a multi-start points optical flow filtering method is proposed to improve the precision of expression spotting. Finally, we conduct comprehensive experiments on the MEGC2024 spotting task, and the experimental results demonstrate the effectiveness of our method, which ranks second in this task. The implemented code is also publicly available at https://github.com/zzy188zzy/megc_spotting_code.
Zhengye Zhang, Sirui Zhao, Xinglong Mao, Hao Wang 0076, Tong Xu 0001, Enhong Chen
ACM Multimedia2
2024 H2LMER: A Cross Frame-Rate Representation Alignment Framework for Micro-expression Recognition
Xinglong Mao, Sirui Zhao, Hao Wang 0076, Tong Xu 0001, Enhong Chen
PRCV (11)3
2024 Woodpecker: hallucination correction for multimodal large language models
Shukang Yin, Chaoyou Fu, Sirui Zhao, Tong Xu 0001, Hao Wang 0076, Dianbo Sui, Yunhang Shen, Ke Li 0015, Xing Sun 0001, Enhong Chen
Sci. China Inf. Sci.3
2024 DFME: A New Benchmark for Dynamic Facial Micro-Expression Recognition
abstract
One of the most important subconscious reactions, micro-expression (ME), is a spontaneous, subtle, and transient facial expression that reveals human beings' genuine emotion. Therefore, automatically recognizing ME (MER) is becoming increasingly crucial in the field of affective computing, providing essential technical support for lie detection, clinical psychological diagnosis, and public safety. However, the ME data scarcity has severely hindered the development of advanced data-driven MER models. Despite the recent efforts by several spontaneous ME databases to alleviate this problem, there is still a lack of sufficient data. Hence, in this paper, we overcome the ME data scarcity problem by collecting and annotating a dynamic spontaneous ME database with the largest current ME data scale called DFME (Dynamic Facial Micro-expressions). Specifically, the DFME database contains 7,526 well-labeled ME videos spanning multiple high frame rates, elicited by 671 participants and annotated by more than 20 professional annotators over three years. Furthermore, we comprehensively verify the created DFME, including using influential spatiotemporal video feature learning models and MER models as baselines, and conduct emotion classification and ME action unit classification experiments. The experimental results demonstrate that the DFME database can facilitate research in automatic MER, and provide a new benchmark for this field. DFME will be published via https://mea-lab-421.github.io.
Sirui Zhao, Huaying Tang, Xinglong Mao, Hao Wang 0076, Tong Xu 0001, Enhong Chen
IEEE Trans. Affect. Comput.1
2024 Exploiting Instance-level Relationships in Weakly Supervised Text-to-Video Retrieval
abstract
Text-to-Video Retrieval is a typical cross-modal retrieval task that has been studied extensively under a conventional supervised setting. Recently, some works have sought to extend the problem to a weakly supervised formulation, which can be more consistent with real-life scenarios and more efficient in annotation cost. In this context, a new task called Partially Relevant Video Retrieval (PRVR) is proposed, which aims to retrieve videos that are partially relevant to a given textual query, i.e., the videos containing at least one semantically relevant moment. Formulating the task as a Multiple Instance Learning (MIL) ranking problem, prior arts rely on heuristics algorithms such as a simple greedy search strategy and deal with each query independently. Although these early explorations have achieved decent performance, they may not fully utilize the bag-level label and only consider the local optimum, which could result in suboptimal solutions and inferior final retrieval performance. To address this problem, in this paper, we propose to exploit the relationships between instances to boost retrieval performance. Based on this idea, we creatively put forward: (1) a new matching scheme for pairing queries and their related moments in the video; and (2) a new loss function to facilitate cross-modal alignment between two views of an instance. Extensive validations on three publicly available datasets have demonstrated the effectiveness of our solution and verified our hypothesis that modeling instance-level relationships is beneficial in the MIL ranking setting. Our code will be publicly available at https://github.com/xjtupanda/BGM-Net .
Shukang Yin, Sirui Zhao, Hao Wang 0076, Tong Xu 0001, Enhong Chen
ACM Trans. Multim. Comput. Commun. Appl.2
2023 APGL4SR: A Generic Framework with Adaptive and Personalized Global Collaborative Information in Sequential Recommendation
abstract
The sequential recommendation system has been widely studied for its promising effectiveness in capturing dynamic preferences buried in users' sequential behaviors. Despite the considerable achievements, existing methods usually focus on intra-sequence modeling while overlooking exploiting global collaborative information by inter-sequence modeling, resulting in inferior recommendation performance. Therefore, previous works attempt to tackle this problem with a global collaborative item graph constructed by pre-defined rules. However, these methods neglect two crucial properties when capturing global collaborative information, i.e., adaptiveness and personalization, yielding sub-optimal user representations. To this end, we propose a graph-driven framework, named Adaptive and Personalized Graph Learning for Sequential Recommendation (APGL4SR), that incorporates adaptive and personalized global collaborative information into sequential recommendation systems. Specifically, we first learn an adaptive global graph among all items and capture global collaborative information with it in a self-supervised fashion, whose computational burden can be further alleviated by the proposed SVD-based accelerator. Furthermore, based on the graph, we propose to extract and utilize personalized item correlations in the form of relative positional encoding, which is a highly compatible manner of personalizing the utilization of global collaborative information. Finally, the entire framework is optimized in a multi-task learning paradigm, thus each part of APGL4SR can be mutually reinforced. As a generic framework, APGL4SR can not only outperform other baselines with significant margins, but also exhibit promising versatility, the ability to learn a meaningful global collaborative graph, and the ability to alleviate the dimensional collapse issue of item embeddings.
Mingjia Yin, Hao Wang 0076, Likang Wu, Sirui Zhao, Wei Guo 0006, Yong Liu 0020, Ruiming Tang, Defu Lian, Enhong Chen
CIKM5
2023 AU-aware graph convolutional network for Macroand Micro-expression spotting
abstract
Automatic Micro-Expression (ME) spotting in long videos is a crucial step in ME analysis but also a challenging task due to the short duration and low intensity of MEs. When solving this problem, previous works generally lack in considering the structures of human faces and the correspondence between expressions and relevant facial muscles. To address this issue for better performance of ME spotting, this paper seeks to extract finer spatial features by modeling the relationships between facial Regions of Interest (ROIs). Specifically, we propose a graph convolutional-based network, called Action-Unit-aWare Graph Convolutional Network (AUW-GCN). In addition, to incorporate prior knowledge and address the issue of small datasets, AU-related statistics are encoded into the network. Comprehensive experiments show that our results outperform baseline methods consistently and achieve new SOTA performance in two benchmark datasets, CAS(ME)2and SAMM-LV. Our code is available at https://github.com/xjtupanda/AUW-GCN.
Shukang Yin, Tong Xu 0001, Sirui Zhao, Enhong Chen
ICME5
2023 Adaptive Graph Attention Network with Temporal Fusion for Micro-Expressions Recognition
abstract
Automatic micro-expression recognition (MER) has essential applications in the psychological field. Graph-based models, due to their advantages in analyzing regionalized faces, have become a powerful method for MER. However, how to construct a graph from ME videos remains to be studied. To solve this problem, we design an adaptive graph attention network with temporal fusion to model the dynamic relationships between facial regions of interest (ROIs). Specifically, we first propose adaptive graph attention to establish learnable spatial graphs from ME videos. Then, we adopt an optical-flow-based feature as the suitable input for the graph network. In addition, an implicit semantic data augmentation algorithm is employed and improved as a data-driven weighted loss for better performance. Extensive experiments on SMIC-HS, CASME II and SAMM datasets have demonstrated the effectiveness of the proposed method, and it achieves to be the first graph-based model where UF1 and UAR both exceed 0.90 for 3-classes MER on CASME II. Code will be available at https://github.com/MEA-LAB-421/ICME2023-Recognition.
Hao Wang 0076, Yifan Xu 0011, Xinglong Mao, Tong Xu 0001, Sirui Zhao, Enhong Chen
ICME6
2023 Vehicle color recognition based on smooth modulation neural network with multi-scale feature fusion
Mingdi Hu, Long Bai 0007, JiuLun Fan 0001, Sirui Zhao, Enhong Chen
Frontiers Comput. Sci.4
2023 PEDM: A Multi-task Learning Model for Persona-aware Emoji-embedded Dialogue Generation
abstract
As a vivid and linguistic symbol, Emojis have become a prevailing medium interspersed in text-based communication (e.g., social media and chit-chat) to express emotions, attitudes, and situations. Generally speaking, a social-oriented chatbot that can generate appropriate Emoji-embedded responses would be much more competitive, making communications more fun, engaging, and human-like. However, the current Emoji-related research is still in its infancy, leading to an awkward situation of data deficiency. How to develop an Emoji-embedded dialogue system while addressing the lack of data will be interesting and meaningful for the application of future AI. To bridge this gap, we propose a multi-task learning method for persona-aware Emoji-embedded dialogue generation in this article. Specifically, as the benchmark of model training and evaluation, which includes 1.2 million Emoji-embedded tweets and 1.1 million post-response pairs, we first construct a dataset named EmojiTweet to handle the data deficiency problem. Then, a Seq2Seq-based model with multi-task learning is designed to simultaneously learn response generation and Emoji embedding from the constructed non-Emoji dialogue and Emoji-embedded monologue data. Afterward, we incorporate persona factors into our model by adopting persona fusion and personalized bias methods to deliver personalized dialogues with more accurately selected Emojis. Finally, we conduct extensive experiments, where the experimental results and evaluations demonstrate that our model has three key benefits: improved dialogue quality, higher user engagement, and not relying on large-scale Emoji-embedded dialogue data representing specific personas. EmojiTweet will be published publicly via https://mea-lab-421.github.io/EmojiTweet/ .
Sirui Zhao, Hongyu Jiang, Hanqing Tao, Rui Zha, Kun Zhang 0015, Tong Xu 0001, Enhong Chen
ACM Trans. Multim. Comput. Commun. Appl.1
2022 AVT: Au-Assisted Visual Transformer for Facial Expression Recognition
abstract
Facial expression recognition (FER) has made significant progress over the past few years. But how to overcome the problem of high inter-class similarity and large intra-class difference in FER is still challenging. To address this problem, we propose a novel FER framework called AU-assisted Visual Transformer (AVT) by incorporating facial action units (AU) information into Visual Transformer, which mainly consists of three modules: Local Feature Extraction (LFE) module, Global Relationship Modeling (GRM) module and AU Fusion Module (AFM). Specifically, the LFE module aims to extract local facial expression features by using a deep convolutional neural network, the GRM module is a multi-layer Transformer encoder that captures the relation between local facial regions and obtains a global understanding of the face, and in particular, the AFM introduces fine-grained AU feature and fuses it with expression feature for final classification. Extensive experiments are conducted on RAF-DB and FERPlus datasets, and our AVT achieves competitive results compared to previous state-of-the-art methods, demonstrating the effectiveness of our approach.
Rijin Jin, Sirui Zhao, Zhongkai Hao, Yifan Xu 0011, Tong Xu 0001, Enhong Chen
ICIP2
2022 ABPN: Apex and Boundary Perception Network for Micro- and Macro-Expression Spotting
abstract
Recently, Micro expression~(ME) has achieved remarkable progress in a wide range of applications, since it's an involuntary facial expression that reflects personal psychological state truly. In the procedure of ME analysis, spotting ME is an essential step, and is non trivial to be detected from a long interval video because of the short duration and low intensity issues. To alleviate this problem, in this paper, we propose a novel Micro- and Macro-Expression~(MaE) Spotting framework based on Apex and Boundary Perception Network~(ABPN), which mainly consists of three parts, i.e., video encoding module ~(VEM), probability evaluation module~(PEM), and expression proposal generation module~(EPGM). Firstly, we adopt Main Directional Mean Optical Flow (MDMO) algorithm and calculate optical flow differences to extract facial motion features in VEM, which can alleviate the impact of head movement and other areas of the face on ME spotting. Then, we extract temporal features with one-dimension convolutional layers and introduce PEM to infer the auxiliary probability that each frame belongs to an apex or boundary frame. With these frame-level auxiliary probabilities, the EPGM further combines the frames from different categories to generate expression proposals for the accurate localization. Besides, we conduct comprehensive experiments on MEGC2022 spotting task, and demonstrate that our proposed method achieves significant improvement with the comparison of state-of-the-art baselines on rm CAS(ME)2 and SAMM-LV datasets. The implemented code is also publicly available at https://github.com/wenhaocold/USTC_ME_Spotting.
Wenhao Leng, Sirui Zhao, Xinglong Mao, Hao Wang 0076, Tong Xu 0001, Enhong Chen
ACM Multimedia2
2022 Fine-grained Micro-Expression Generation based on Thin-Plate Spline and Relative AU Constraint
abstract
As a typical psychological stress reaction, micro-expression (ME) is usually quickly leaked on a human face and can reveal the true feeling and emotional cognition. Therefore,automatic ME analysis (MEA) has essential applications in safety, clinical and other fields. However, the lack of adequate ME data has severely hindered MEA research. To overcome this dilemma and encouraged by current image generation techniques, this paper proposes a fine-grained ME generation method to enhance ME data in terms of data volume and diversity. Specifically, we first estimate non-linear ME motion using thin-plate spline transformation with a dense motion network. Then, the estimated ME motion transformations, including optical flow and occlusion masks, are sent to the generation network to synthesize the target facial micro-expression. In particular, we obtain the relative action units (AUs) of the source ME to the target face as a constraint to encourage the network to ignore expression-irrelevant movements, thereby generating fine-grained MEs. Through comparative experiments on CASME II, SMIC and SAMM datasets, we demonstrate the effectiveness and superiority of our method. Source code is provided in https://github.com/MEA-LAB-421/MEGC2022-Generation.
Sirui Zhao, Shukang Yin, Huaying Tang, Rijin Jin, Yifan Xu 0011, Tong Xu 0001, Enhong Chen
ACM Multimedia1
2022 ME-PLAN: A deep prototypical learning with local attention network for dynamic micro-expression recognition
Sirui Zhao, Huaying Tang, Yangsong Zhang 0001, Hao Wang 0076, Tong Xu 0001, Enhong Chen, Cuntai Guan
Neural Networks1
2022 Causal Narrative Comprehension: A New Perspective for Emotion Cause Extraction
abstract
Emotion Cause Extraction (ECE) aims to reveal the cause clauses behind a given emotion expressed in a text, which has become an emerging topic in broad research communities, such as affective computing and natural language processing. Despite the fact that current methods about the ECE task have made great progress in text semantic understanding from lexicon- and sentence-level, they always ignore the certain causal narratives of emotion text. Significantly, these causal narratives are presented in the form of semantic structure and highly helpful for structure-level emotion cause understanding. Nevertheless, causal narrative is just an abstract narratological concept and its involving semantics is quite different from the common sequential information. Thus, how to properly model and utilize such particular narrative information to boost the ECE performance still remains an unresolved challenge. To this end, in this paper, we propose a novel Causal Narrative Comprehension Model (CNCM) for emotion cause extraction, which learns and leverages causal narrative information smartly to address the above problem. Specifically, we develop a Narrative-aware Causal Association (NCA) unit, which mines the narrative cue about emotional results and uses the semantic correlation between causes and results to model causal narratives of documents. Besides, we design a Result-aware Emotion Attention (REA) unit to make full use of the known result of causal narrative for multiple understanding about emotional causal associations. Through the ingenious combination and collaborative utilization of these two units, we could better identify the emotion cause in the text with causal narrative comprehension. Extensive experiments on the public English and Chinese benchmark datasets of ECE task have validated the effectiveness of CNCM with significant margin by comparing with the state-of-the-art baselines, which demonstrates the potential of narrative information in long text understanding.
Kun Zhang 0015, Shulan Ruan, Hanqing Tao, Sirui Zhao, Hao Wang 0076, Qi Liu 0003, Enhong Chen
IEEE Trans. Affect. Comput.5
2021 FAMGAN: Fine-grained AUs Modulation based Generative Adversarial Network for Micro-Expression Generation
abstract
Micro-expressions (MEs) are significant and effective clues to reveal the true feelings and emotions of human beings, and thus MEs analysis is widely used in different fields such as medical diagnosis, interrogation and security. However, it is extremely difficult to elicit and label MEs, resulting in a lack of sufficient MEs data for MEs analysis. To address this challenge and inspired by the current face generation technology, in this paper we introduce Generative Adversarial Network based on fine-grained Action Units (AUs) modulation to generate MEs sequence (FAMGAN). Specifically, after comprehensively analyzing the factors that lead to inaccurate AU values detection, we performed fine-grained AUs modulation, which includes carefully eliminating the various noises and dealing with the asymmetry of AUs intensity. Additionally, we incorporate super-resolution into our model to enhance the quality of the generated images. Through experiments, we show that the system achieves very competitive results on the Micro-Expression Grand Challenge (MEGC2021).
Yifan Xu 0011, Sirui Zhao, Huaying Tang, Xinglong Mao, Tong Xu 0001, Enhong Chen
ACM Multimedia2
2021 A two-stage 3D CNN based learning method for spontaneous micro-expression recognition
Sirui Zhao, Hanqing Tao, Yangsong Zhang 0001, Tong Xu 0001, Kun Zhang 0015, Zhongkai Hao, Enhong Chen
Neurocomputing1
2021 An end-to-end 3D convolutional neural network for decoding attentive mental state
Yangsong Zhang 0001, Huan Cai, Li Nie, Peng Xu 0001, Sirui Zhao, Cuntai Guan
Neural Networks5