VLDB 2026 Research / reviewers in the wild / expert
Shangfei Wang
dblp:15/2254
· DBLP profile ↗
138ranked-venue papers
48as first author
43since 2021 · last 2026
0000-0003-1164-9895ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 84 · 11 first-author · 25 since 2021Artificial intelligence and machine learning · 76 · 30 first-author · 30 since 2021Human-computer interaction and ubiquitous computing · 8 · 2 first-authorDatabases, data management, data science and information retrieval · 6 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 6 · 4 first-author · 2 since 2021Computer networks · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ConsistTalk: Intensity Controllable Temporally Consistent Talking Head Generation with Diffusion Noise SearchabstractRecent advancements in video diffusion models have significantly enhanced audio-driven portrait animation. However, current methods still suffer from flickering, identity drift, and poor audio-visual synchronization. These issues primarily stem from entangled appearance-motion representations and unstable inference strategies. In this paper, we introduce ConsistTalk, a novel intensity-controllable and temporally consistent talking head generation framework with diffusion noise search inference. First, we propose an optical flow-guided temporal module (OFT) that decouples motion features from static appearance by leveraging facial optical flow, thereby reducing visual flicker and improving temporal consistency. Second, we present an Audio-to-Intensity (A2I) model obtained through multimodal teacher-student knowledge distillation. By transforming audio and facial velocity features into a frame-wise intensity sequence, the A2I model enables joint modeling of audio and visual motion, resulting in more natural dynamics. This further enables fine-grained, frame-wise control of motion dynamics while maintaining tight audio-visual synchronization. Third, we introduce a diffusion noise initialization strategy (IC-Init). By enforcing explicit constraints on background coherence and motion continuity during inference-time noise search, we achieve better identity preservation and refine motion dynamics compared to the current autoregressive strategy. Extensive experiments demonstrate that ConsistTalk significantly outperforms prior methods in reducing flicker, preserving identity, and delivering temporally stable, high-fidelity talking head videos. Zhenjie Liu, Jianzhang Lu, Cong Liang 0002, Shangfei Wang |
AAAI | 5 |
| 2026 | FINE: Factorized Multimodal Sentiment Analysis via Mutual INformation EstimationabstractMultimodal sentiment analysis remains a challenging task due to the inherent heterogeneity across modalities. Such heterogeneity often manifests as asynchronous signals, imbalanced information between modalities, and interference from task-irrelevant noise, hindering the learning of robust and accurate sentiment representations. To address these issues, we propose a factorized multimodal fusion framework that first disentangles each modality into shared and unique representations, and then suppresses task-irrelevant noise within both to retain only sentiment-critical representations. This fine-grained decomposition improves representation quality by reducing redundancy, prompting cross-modal complementarity, and isolating task-relevant sentiment cues. Rather than manipulating the feature space directly, we adopt a mutual information–based optimization strategy to guide the factorization process in a more stable and principled manner. To further support feature extraction and long-term temporal modeling, we introduce two auxiliary modules: a Mixture of Q-Formers, placed before factorization, which precedes the factorization and uses learnable queries to extract fine-grained affective features from multiple modalities, and a Dynamic Contrastive Queue, placed after factorization, which stores latest high-level representations for contrastive learning, enabling the model to capture long-range discriminative patterns and improve class-level separability. Extensive experiments on multiple public datasets demonstrate that our method consistently outperforms existing approaches, validating the effecti veness and robustness of the proposed framework. Yadong Liu 0003, Shangfei Wang |
AAAI | 2 |
| 2026 | Learning Knowledge from Textual Descriptions for 3D Human Pose EstimationabstractMainstream 3D human pose estimation methods directly predict 3D coordinates of joints from 2D keypoints, suffering from severe depth ambiguity. Pose textual descriptions contain abundant semantic information, which facilitates the model to learn the spatial relationship among different body parts, partially alleviating this issue. Leveraging this insight, we propose a 3D human pose estimation method assisted by textual descriptions. Specifically, we utilize an automatic captioning pipeline to generate textual descriptions of 3D poses based on spatial relations among joints. These descriptions include details regarding angles, distances, relative positions, pitch\&roll and ground-contacts. Subsequently, text features are extracted from these descriptions using a language model, while a 3D human pose estimation model extracts pose features. Aligning the pose features with the text features allows for a more targeted optimization of the estimation model. Therefore, we systematically introduce three alignment approaches to effectively align features extracted by two models operating in entirely different domains. Our method incorporates prior knowledge derived from the textual descriptions into the estimation model and can be seamlessly applied to various existing framework. Experimental results on the Human3.6M and MPI-INF-3DHP datasets demonstrate that our method surpasses state-of-the-art methods. Yi Wu 0019, Jingtian Li, Shangfei Wang, Meng Mao, Linxiang Tan |
AAAI | 3 |
| 2026 | Knowledge-Guided Open-Set Facial Expression Recognition via Action Unit Reasoning
Caichao Zhang, Shangfei Wang |
ICIC (7) | 2 |
| 2025 | From Traits to Empathy: Personality-Aware Multimodal Empathetic Response GenerationabstractEmpathetic dialogue systems improve user experience across various domains. Existing approaches mainly focus on acquiring affective and cognitive knowledge from text, but neglect the unique personality traits of individuals and the inherently multimodal nature of human face-to-face conversation. To this end, we enhance the dialogue system with the ability to generate empathetic responses from a multimodal perspective, and consider the diverse personality traits of users. We incorporate multimodal data, such as images and texts, to understand the user’s emotional state and situation. Concretely, we first identify the user’s personality trait. Then, the dialogue system comprehends the user’s emotions and situation by the analysis of multimodal inputs. Finally, the response generator models the correlations among the personality, emotion, and multimodal data, to generate empathetic responses. Experiments on the MELD dataset and the MEDIC dataset validate the effectiveness of the proposed approach. Jiaqiang Wu, Xuandong Huang, Zhouan Zhu, Shangfei Wang |
COLING | 4 |
| 2025 | Integrating Visual Modalities with Large Language Models for Mental Health SupportabstractCurrent work of mental health support primarily utilizes unimodal textual data and often fails to understand and respond to users’ emotional states comprehensively. In this study, we introduce a novel framework that enhances Large Language Model (LLM) performance in mental health dialogue systems by integrating multimodal inputs. Our framework uses visual language models to analyze facial expressions and body movements, then combines these visual elements with dialogue context and counseling strategies. This approach allows LLMs to generate more nuanced and supportive responses. The framework comprises four components: in-context learning via computation of semantic similarity; extraction of facial expression descriptions through visual modality data; integration of external knowledge from a knowledge base; and delivery of strategic guidance through a strategy selection module. Both automatic and human evaluations confirm that our approach outperforms existing models, delivering more empathetic, coherent, and contextually relevant mental health support responses. Zhouan Zhu, Shangfei Wang, Jiaqiang Wu |
COLING | 2 |
| 2025 | Context-Assisted Low-Light Face Detection through Global and Local Image EnhancementabstractCurrent low-light face detection usually enhances images first and then detects faces. Image enhancement focuses on the global enhancement of the whole image and is suitable for the human perspective. However, face detection requires high quality of local face regions and has different optimization objectives from image enhancement task. Such mismatches between the two tasks result in inferior performance of low-light face detection. To solve these obstacles, we propose end-to-end low-light face detection through global and local enhancement. Specifically, the proposed method consists of two components, i.e., image enhancement and face detection. For image enhancement, both global and local enhancement are explored. Global enhancement utilizes illumination constraints to enhance the overall image illumination, while local enhancement employs a generative adversarial network to improve the illumination of face regions during training. Through a combination of global and local enhancement, we reduce the gap between image enhancement and face detection and optimize the two tasks jointly. Furthermore, regions of the human body may help face detection, especially when face regions are small and blurred. Therefore, we regard body region as context to assist face detection. We use the context assistance module and optimize the annotation of the body regions with the aid of human structure prior knowledge. Experiments on the Dark Face dataset demonstrate the effectiveness of the proposed method. Xiangyu Miao, Caichao Zhang, Yanan Chang, Shangfei Wang |
FG | 4 |
| 2025 | BTPose: 3D Pose Estimation from Bone to Pose with Efficient Multi-hypothesis Aggregation
Jingtian Li, Yi Wu 0019, Shangfei Wang, Meng Mao, Linxiang Tan |
ICIC (5) | 3 |
| 2025 | EmoDETective: Detecting, Exploring, and Thinking Emotional Cause in Videos
Xuandong Huang, Yuzhe Zhou 0001, Jiashu Li, Shiqian Lu, Shangfei Wang |
ACM Multimedia | 5 |
| 2025 | EmIT: Emotional Interaction control in Text-to-image diffusion modelsabstractAlthough current work of text-to-Image generation can preliminarily generate images from the descriptions of human-object interactions, it fails to consider the emotions involved in human-object interactions. While people often experience emotions when using objects or interacting with them. Therefore, in this paper, we propose Emotional Interaction Generation task, a novel image generation task, which generates emotionally expressive human-object interaction images from given prompts, human-object interaction (HOI) region, and emotions. First, we construct a new emotional interaction dataset, called EmotionHOI, which including 47,776 images with content prompt, emotions and human-object interaction bounding box. Second, we propose an emotion-aware text-to-image diffusion model, named EmIT, for emotional interaction generation. Specifically, EmIT consists of three components: (1) an emotion interaction tokenizer that encodes subject, object, action, and emotion into structured tokens; (2) an Emo-Interaction Self-Attention that preliminarily guides the latent space to conduct hybrid learning with emotional interaction tokens; and (3) a Hierarchical Emotion-Visual Cross-Attention that further focus on grounding affect-such as pose, gaze, or interaction intensity-into specific spatial regions and capture subtle emotional variations. These components jointly model interaction semantics and emotional context, enabling EmIT to generate images that are both behaviorally coherent and emotionally expressive. Experimental results on the EmotionHOI dataset demonstrate the superiority of the proposed model. Haofan Zhang, Shangfei Wang |
ACM Multimedia | 2 |
| 2025 | Facial Action Unit Recognition Enhanced by Text Descriptions of FACSabstractAlthough the descriptions of facial action units (AUs) provide crucial semantic knowledge for representation learning from facial images, they have not been fully explored for facial action unit recognition. In this paper, we propose a method that effectively explores the knowledge existing in AU descriptions to enhance AU recognition. Specifically, the proposed method consists of three components, i.e., AU recognition network, global representation alignment, and AU representation alignment. The AU recognition network extracts global features and AU-specific features for AU prediction from images. To leverage AU textual descriptions fully, we design two-level representation alignment for AU recognition. The global representation alignment component closes the distance between the global facial features and its corresponding positive global embedding extracted from textual descriptions. Then, the AU-specific features are aligned with the positive AU textual embedding by the AU representation alignment component. Negative textual embedding generation strategies are also designed to further boost the two-level representation alignment. Through the two-level alignment, AU textual descriptions guide image representation learning of the AU recognition network. Experiments on two benchmark datasets and one in-the-wild dataset demonstrate the efficacy of the description-enhanced AU recognition method, compared with the state-of-the-art works. Yanan Chang, Caichao Zhang, Yi Wu 0019, Shangfei Wang |
IEEE Trans. Affect. Comput. | 4 |
| 2025 | Empathetic Response Generation Through Multi-ModalityabstractDespite remarkable advancements in empathetic response generation (ERG) area, existing research has centered on achieving affective and cognitive empathy by perceiving users' emotions and deducing contextual information from knowledge databases. Human communication combines textual, visual, and audio cues to interpret other's intentions. However, previous ERG works have focused on text-based methods and neglected contextual information within audiovisual data. To bridge the gap, we propose fostering empathy with users by integrating audiovisual and text modalities. First, the proposed method uses a cross-modal attention mechanism to perceive users' emotions from the multi-modal conversation. It integrates multi-modal data with the perceived emotions during response generation process, so that the generated responses resonate with users at the affective level by mirroring their emotions. Second, we import guidance text that focuses on visual context or user experiences and provides contextual information, thus enhancing cognitive empathy. The proposed method aligns multi-modal dialogue history and guidance text through the multi-source attention mechanism. Finally, the proposed method produces empathetic responses by understanding users' backgrounds and emotions. Experiments on three multi-modal datasets, e.g., MELD, IEMOCAP, and MEDIC, demonstrate that the proposed method outperforms state-of-the-art works. Jiaqiang Wu, Shangfei Wang, Yanan Chang, Zhouan Zhu |
IEEE Trans. Affect. Comput. | 2 |
| 2024 | Progressive Target Refinement by Self-distillation for Human Pose Estimation
Jingtian Li, Yi Wu 0019, Shangfei Wang |
ACCV (8) | 4 |
| 2024 | One-to-Many Appropriate Reaction Mapping Modeling with Discrete Latent VariableabstractIn dyadic interaction, listener reaction generation can be treated as a one-to-many mapping problem since multiple listener reactions can correspond to a given speaker action. The existing methods have not modeled the diversity of contextual factors well and fail to generate diverse appropriate listener reactions. In response, we introduce discrete latent variables to tackle this one-to-many mapping problem. We conducted experiments on the datasets provided by the REACT2024 Challenge, and the results demonstrated that our approach is capable of generating appropriate listening reactions with higher diversity. Our method achieved first place in the offline track and second in the online track. Zhenjie Liu, Cong Liang 0002, Haofan Zhang, Yadong Liu 0003, Caichao Zhang, Jialin Gui, Shangfei Wang |
FG | 8 |
| 2024 | 1DFormer: A Transformer Architecture Learning 1D Landmark Representations for Facial Landmark Tracking
Shijie Huan, Shangfei Wang, Jinshui Hu, Cong Liu 0006 |
IJCAI | 3 |
| 2024 | Temporal Enhancement for Video Affective Content AnalysisabstractWith the popularity and advancement of the Internet and video-sharing platforms, video affective content analysis has greatly developed. Temporal information is crucial for this task. Nevertheless, existing methods often overlook the fact that there is substantial irrelevant information in videos and that the importance of modalities is uneven for emotional tasks. This could result in noise from both temporal fragments and modalities, reducing the model's ability to identify crucial temporal fragments and recognize emotions. To tackle the above issues, we propose a Temporal Enhancement (TE) method in this paper. Specifically, we utilize three encoders for extracting features at various levels and employ temporal sampling to enhance the temporal data, thereby enriching video representation and improving the model's robustness to noise. Subsequently, we design a cross-modal temporal enhancement module to enhance temporal information for every modal feature. This module interacts with multiple modalities simultaneously to emphasize critical temporal fragments while suppressing irrelevant ones. The experimental results on four benchmark datasets show that the proposed temporal enhancement method achieves state-of-the-art video affective content analysis performance. Moreover, the effectiveness of each module is confirmed through ablation experiments. Xin Li 0123, Shangfei Wang, Xuandong Huang |
ACM Multimedia | 2 |
| 2024 | FreqMark: Invisible Image Watermarking via Frequency Based Optimization in Latent SpaceabstractInvisible watermarking is essential for safeguarding digital content, enabling copyright protection and content authentication.
However, existing watermarking methods fall short in robustness against regeneration attacks.
In this paper, we propose a novel method called FreqMark that involves unconstrained optimization of the image latent frequency space obtained after VAE encoding. Specifically, FreqMark embeds the watermark by optimizing the latent frequency space of the images and then extracts the watermark through a pre-trained image encoder. This optimization allows a flexible trade-off between image quality with watermark robustness and effectively resists regeneration attacks.
Experimental results demonstrate that FreqMark offers significant advantages in image quality and robustness, permits flexible selection of the encoding bit number, and achieves a bit accuracy exceeding 90\% when encoding a 48-bit hidden message under various attack scenarios. Yiyang Guo, Mude Hui, Hanzhong Guo, Chuangjian Cai, Le Wan, Shangfei Wang |
NeurIPS | 8 |
| 2024 | Pose-robust personalized facial expression recognition through unsupervised multi-source domain adaptation
Shangfei Wang, Yanan Chang, Meng Mao |
Pattern Recognit. | 1 |
| 2024 | A Multi-Stage Visual Perception Approach for Image Emotion AnalysisabstractMost current methods for image emotion analysis suffer from the affective gap, in which features directly extracted from images are supervised by a single emotional label, which may not align with users' perceived emotions. To effectively address this limitation, this paper introduces a novel multi-stage perception approach inspired by the human staged emotion perception process. The proposed approach comprises three perception modules: entity perception, attribute perception, and emotion perception. The entity perception module identifies entities in images, while the attribute perception module captures the attribute content associated with each entity. Finally, the emotion perception module combines entity and attribute information to extract emotion features. Pseudo-labels of entities and attributes are generated through image segmentation and vision-language models to provide auxiliary guidance for network learning. A progressive understanding of entities and attributes allows the network to hierarchically extract semantic-level features for emotion analysis. Comprehensive experiments on image emotion classification, regression, and distribution learning demonstrate the superior performance of our multi-stage perception network. Jicai Pan, Jinqiao Lu, Shangfei Wang |
IEEE Trans. Affect. Comput. | 3 |
| 2024 | VAD: A Video Affective Dataset With DanmuabstractAlthough video affective content analysis has great potential in many applications, it has not been thoroughly studied due to limited datasets. In this paper, we construct a large-scale video affective dataset with danmu (VAD). It consists of 19,267 elaborately segmented video clips from user-generated videos. The VAD dataset is annotated by the crowdsourcing platform with discrete valence, arousal, and primary emotions, as well as the comparison of valence and arousal between two consecutive video clips. Unlike previous datasets, including only video clips, our proposed dataset also provides danmu, which is the real-time comment from users as they watch a video. Danmu provides extra information for video affective content analysis. As a preliminary assessment of the usability of our dataset, an analysis of inter-annotator consistency for each label is conducted using weighted Fleiss' Kappa, regular Fleiss' Kappa, intraclass correlation coefficient, and percent consensus. Besides, we also perform a statistical analysis of labels and danmu. Finally, video affective content analysis is conducted on our dataset and three typical methods (i.e., TFN, MulT, and MISA) are leveraged to provide benchmarks. We also demonstrate that danmu can significantly improve the performance of the video affective content analysis task on some labels. Our dataset is available for research purposes. Shangfei Wang, Xin Li 0123, Feiyi Zheng, Jicai Pan, Yanan Chang, Zhouan Zhu, Yufei Xiao |
IEEE Trans. Affect. Comput. | 1 |
| 2024 | Pose-Aware Facial Expression Recognition Assisted by Expression DescriptionsabstractAlthough expression descriptions provide additional information about facial behaviors despite of different poses, and pose features are beneficial to adapt to pose variety, neither has been fully leveraged in facial expression recognition. This paper proposes a pose-aware text-assisted facial expression recognition method using cross-modality attention. Specifically, the method contains three components. The pose feature extractor extracts pose-related features from facial images, and then cooperates with a fully-connected layer for pose classification. When poses can be clearly discriminated and classified, features obtained from the extractor can represent the corresponding poses. To eliminate bias due to appearance and illumination, cluster centers are taken as the final pose features. The text feature extractor obtains embeddings from expression descriptions. These descriptions are first passed through Intra-Exp attention to obtain preliminary embeddings. To leverage the correlations among expressions, all expression embeddings are then concatenated and passed through Inter-Exp attention. The cross-modality module attempts to learn attention maps that distinguish the importance of facial regions by using prior knowledge about poses and expression descriptions. The image features weighted by the attention maps are utilized to recognize pose and expression jointly. Experiments on three benchmark datasets demonstrate the superiority of the proposed method. Shangfei Wang, Yi Wu 0019, Yanan Chang, Meng Mao |
IEEE Trans. Affect. Comput. | 1 |
| 2023 | Human Pose Estimation with Shape Aware LossabstractAlthough the mean square error (mse) of heatmap is an intuitive loss for heatmap-based human pose estimation, the joints localization accuracy may not be improved when heatmap mse reduces. In this paper, we show that a great cause for such misalignment is the unnecessary requirement from heatmap mse on the irrelevant Gaussian parameter, i.e. maximum. The coordinate prediction is precise as long as the probability distribution held by the predicted heatmap is a well-shaped Gaussian distribution and has the same center as the ground truth. However, heatmap mse unnecessarily requires the Gaussian distribution to hold the same maximum as the ground truth. Correspondingly, we introduce mse on the image gradients of the target and predicted heatmap (referred to as gradmap mse) to focus on the shape of the heatmap. Combining heatmap and gradmap mse, we propose a simple yet effective Shape Aware Loss (SAL) method. Being model-agnostic, our method can benefit various existing models. We apply SAL to the three latest network architectures and obtain performance improvements for all of them. Comparisons of the visualized predicted heatmaps further prove the effectiveness of the proposed method. Shangfei Wang |
FG | 2 |
| 2023 | Low-Resolution Face Recognition Enhanced by High-Resolution Facial ImagesabstractDespite recent advances in high-resolution (HR) face recognition, recognizing identities from low-resolution (LR) facial images remains challenging due to the absence of facial shape and detail. Current research focuses solely on reducing the distribution discrepancy between the HR and LR embeddings from the output layer, rather than thoroughly investigating the superiority of HR facial images for improved performance. In this paper, we propose a novel low-resolution face recognition method enhanced by the guidance of high-resolution facial images in both feature map space and embedding space. Specifically, in feature map space, the similarity constraint across the multilayer feature maps is adopted to align the intermediate features of facial images. Then we introduce multiple generators to recover HR images from extracted feature maps and utilize the reconstructed loss to supplement the missing facial details in LR images. In embedding space, we propose a supervised auxiliary contrastive loss to encourage the paired HR and LR embedding from the same class to be pulled together, whereas those from different classes are pushed apart. The one-to-many matching strategy and the adaptive weight adjustment strategy are applied to make the network adapt to the inputs of different resolutions. Experiments on four benchmark datasets with both synthesized and realistic LR facial images demonstrate the superiority of the proposed method to state-of-the-art. Haihan Wang, Shangfei Wang |
FG | 2 |
| 2023 | Privacy-Protected Facial Expression Recognition Augmented by High-Resolution Facial ImagesabstractCloud-based expression recognition from high-resolution facial images may put the subjects’ privacy at risk. We identify two kinds of privacy leakage, the appearance leakage in which the visual appearances of subjects are disclosed and the identity-pattern leakage in which the identity information of subjects is dug out. To address both leakages, we propose privacy-protected facial expression recognition from low-resolution facial images with the help of high-resolution facial images. Specifically, to prevent appearance leakage, we propose to extract identity-invariant representations from downsampled images, from which the visually distinguishable appearances cannot be recovered. To prevent identity-pattern leakage, we propose to eliminate the identity information from the extracted representations by leveraging the disentangled representations of high-resolution images as privileged information. After training, our method can fully capture identity-invariant representations from downsampled images for expression recognition without the requirement of high-resolution samples. These privacy-protected representations can be safely transmitted through the Internet. Experimental results in different scenarios demonstrate that the proposed method protects privacy without significantly inhibiting facial expression recognition. Cong Liang 0002, Shangfei Wang |
ICME | 2 |
| 2023 | UniFaRN: Unified Transformer for Facial Reaction GenerationabstractWe propose the Unified Transformer for Facial Reaction GeneratioN (UniFaRN) framework for facial reaction prediction in dyadic interactions. Given the video and audio of one side, the task is to generate facial reactions of the other side. The challenge of the task lies in the fusion of multi-modal inputs and balancing appropriateness and diversity. We adopt the Transformer architecture to tackle the challenge by leveraging its flexibility of handling multi-modal data and ability to control the generation process. By successfully capturing the correlations between multi-modal inputs and outputs with unified layers and balancing the performance with sampling methods, we have won first place in the REACT2023 challenge. Cong Liang 0002, Haofan Zhang, Bing Tang, Junshan Huang, Shangfei Wang |
ACM Multimedia | 6 |
| 2023 | Progressive Visual Content Understanding Network for Image Emotion ClassificationabstractMost existing methods for image emotion classification extract features directly from images supervised by a single emotional label. However, this approach has a limitation known as the affective gap which restricts the capability of these features as they do not always align with the emotions perceived by users. To effectively bridge the affective gap, this paper proposes a visual content understanding network inspired by the human staged emotion perception process. The proposed network is comprised of three perception modules designed to extract multi-level information. Firstly, an entity perception module extracts entities from images. Secondly, an attribute perception module extracts the attribute content of each entity. Thirdly, an emotion perception module extracts emotion features based on both the entity and attribute information. We generate pseudo-labels of entities and attributes through image segmentation and vision-language models to provide auxiliary guidance for network learning. The progressive entity and attribute understanding enable the network to hierarchically extract semantic-level features for emotion analysis. Extensive experiments demonstrate that our progressive learning network achieves superior performance on various benchmark datasets for image emotion classification. Jicai Pan, Shangfei Wang |
ACM Multimedia | 2 |
| 2023 | Patch-Aware Representation Learning for Facial Expression RecognitionabstractExisting methods for facial expression recognition (FER) lack the utilization of prior facial knowledge, primarily focusing on expression-related regions while disregarding explicitly processing expression-independent information. This paper proposes a patch-aware FER method that incorporates facial keypoints to guide the model and learns precise representations through two collaborative streams, addressing these issues. First, facial keypoints are detected using a facial landmark detection algorithm, and the facial image is divided into equal-sized patches using the Patch Embedding Module. Then, a correlation is established between the keypoints and patches using a simplified conversion relationship. Two collaborative streams are introduced, each corresponding to a specific mask strategy. The first stream masks patches corresponding to the keypoints, excluding those along the facial contour, with a certain probability. The resulting image embedding is input into the Encoder to obtain expression-related features. The features are passed through the Decoder and Classifier to reconstruct the masked patches and recognize the expression, respectively. The second stream masks patches corresponding to all the above keypoints. The resulting image embedding is input into the Encoder and Classifier successively, with the resulting logit approximating a uniform distribution. Through the first stream, the Encoder learns features in the regions related to expression, while the second stream enables the Encoder to better ignore expression-independent information, such as the background, facial contours, and hair. Experiments on two benchmark datasets demonstrate that the proposed method outperforms state-of-the-art methods. Yi Wu 0019, Shangfei Wang, Yanan Chang |
ACM Multimedia | 2 |
| 2023 | MEDIC: A Multimodal Empathy Dataset in CounselingabstractAlthough empathic interaction between counselor and client is fundamental to success in the psychotherapeutic process, there are currently few datasets to aid a computational approach to empathy understanding. In this paper, we construct a multimodal empathy dataset collected from face-to-face psychological counseling sessions. The dataset consists of 771 video clips. We also propose three labels (i.e., expression of experience, emotional reaction, and cognitive reaction) to describe the degree of empathy between counselors and their clients. Expression of experience describes whether the client has expressed experiences that can trigger empathy, and emotional and cognitive reactions indicate the counselor's empathic reactions. As an elementary assessment of the usability of the constructed multimodal empathy dataset, an interrater reliability analysis of annotators' subjective evaluations for video clips is conducted using the intraclass correlation coefficient and Fleiss' Kappa. Results prove that our data annotation is reliable. Furthermore, we conduct empathy prediction using three typical methods, including the tensor fusion network, the sentimental words aware fusion network, and a simple concatenation model. The experimental results show that empathy can be well predicted on our dataset. Our dataset is available for research purposes. Zhouan Zhu, Jicai Pan, Xin Li 0123, Yufei Xiao, Yanan Chang, Feiyi Zheng, Shangfei Wang |
ACM Multimedia | 8 |
| 2023 | Dual Learning for Joint Facial Landmark Detection and Action Unit RecognitionabstractFacial landmark detection and action unit (AU) recognition are two essential tasks in facial analysis. Previous works rarely consider the relationship between these complementary tasks. In this article, we introduce a novel multi-task dual learning framework to exploit the relationship between facial landmark detection and AU recognition while simultaneously addressing both tasks. When both tasks share middle-level features, common patterns can be exploited and middle- and high-level features can be used to perform facial landmark detection and AU recognition, respectively. In addition, a dual learning mechanism is designed to convert the predicted landmarks and AUs of the label space to the corresponding facial image of the image space, further exploring the strong correlations between the tasks. By jointly training the proposed method at both the feature and label levels, each task improves the other. Experiments on two benchmark databases demonstrate that the proposed method can leverage dependencies to boost the generalization of both tasks. Shangfei Wang, Yanan Chang, Can Wang 0007 |
IEEE Trans. Affect. Comput. | 1 |
| 2023 | Emotional Attention Detection and Correlation Exploration for Image Emotion Distribution LearningabstractCurrent works on image emotion distribution learning typically extract visual representations from the holistic image or explore emotion-related regions in the image from a global-wise perspective. However, different regions of an image contribute differently to the arousal of each emotion. Existing works do not deeply explore corresponding emotion-aware regions of each emotion in the image, nor do they fully capture the relationship between each emotion-aware region and the emotion labels. In this article, we propose a novel attention based emotion distribution learning method, which can explore the emotion-related regions of images from the perspective of each emotion category, and can conduct region relationship learning. Specifically, we introduce a semantic guided attention detection network to generate class-wise attention maps for each emotion and a global-wise attention map for the holistic image. Meanwhile, an emotional graph-based network is adopted to capture the correlation between each region and the emotion distribution. Experiments on several benchmark datasets demonstrate the superiority of the proposed method compared to related works. Shangfei Wang |
IEEE Trans. Affect. Comput. | 2 |
| 2022 | Occluded Facial Expression Recognition Using Self-supervised Learning
Heyan Ding, Shangfei Wang |
ACCV (4) | 3 |
| 2022 | Knowledge-Driven Self-Supervised Representation Learning for Facial Action Unit RecognitionabstractFacial action unit (AU) recognition is formulated as a supervised learning problem by recent works. However, the complex labeling process makes it challenging to provide AU annotations for large amounts of facial images. To remedy this, we utilize AU labeling rules defined by the Facial Action Coding System (FACS) to design a novel knowledge-driven self-supervised representation learning framework for AU recognition. The representation encoder is trained using large amounts of facial images without AU annotations. AU labeling rules are summarized from FACS to design facial partition manners and determine correlations between facial regions. The method utilizes a backbone network to extract local facial area representations and a project head to map the representations into a low-dimensional latent space. In the latent space, a contrastive learning component leverages the inter-area difference to learn AU-related local representations while maintaining intra-area instance discrimination. Correlations between facial regions summarized from AU labeling rules are also explored to further learn representations using a predicting learning component. Evaluation on two benchmark databases demonstrates that the learned representation is powerful and data-efficient for AU recognition. Yanan Chang, Shangfei Wang |
CVPR | 2 |
| 2022 | Adversarial Stacking Ensemble for Facial Landmark TrackingabstractCurrent approaches for facial landmark tracking predict facial landmarks through either a single tracker or an ensemble of trackers. However, the conventional ensemble is not designed for facial landmark tracking and can not capture spatial and temporal patterns of facial landmarks efficiently. In this paper, we propose to extend the conventional stacking with an adversarial training strategy to better suit the facial landmark tracking task. Specifically, the meta learner attempts to distinguish the predictions from the base learners with the ground truths, while the base learners attempt to confuse the meta learner by predicting landmarks close to the ground truths. The adversary between the two levels of learners forces them to fully capture the inherent spatial and temporal patterns of facial landmarks. Moreover, to promote the diversity of different base trackers, we design two classification tasks at both the feature level and prediction level. Experimental results on the 300VW dataset and the TF dataset demonstrate the effectiveness of our method, and we achieve state-of-the-art performances on both datasets. Shangfei Wang, Yanan Chang |
ICPR | 3 |
| 2022 | Knowledge Guided Representation Disentanglement for Face Recognition from Low Illumination ImagesabstractLow illumination face recognition is challenging as details are lacking due to lighting conditions. Retinex theory points out that images can be divided into reflectance with color constancy and ambient illumination. Inspired by this, we propose a knowledge-guided representation disentanglement method to disentangle facial images into face-related and illumination-related features, and then leverage the disentangled face-related features for face recognition. Specifically, the proposed method consists of two components: feature disentanglement and face classifier. Following Retinex, high-dimensional face-related features and ambient illumination-related features are extracted from facial images. Reconstruction and crossreconstruction methods are used to make sure the integrity and accuracy of the disentangled features. Furthermore, we find that the influence of illumination changes on illumination-related features should be invariant for faces of different identities, so we design an illumination offset loss to satisfy the prior invariance for better disentanglement. Finally high-dimensional face-related features are mapped to low-dimensional features through the face classifier for use in face recognition task. Experimental results on low illumination and NIR-VIS datasets demonstrate the superiority and effectiveness of our proposed method. Xiangyu Miao, Shangfei Wang |
ACM Multimedia | 2 |
| 2022 | Representation Learning through Multimodal Attention and Time-Sync Comments for Affective Video Content AnalysisabstractAlthough temporal patterns inherent in visual and audio signals are crucial for affective video content analysis, they have not been thoroughly explored yet. In this paper, we propose a novel Temporal-Aware Multimodal (TAM) method to fully capture the temporal information. Specifically, we design a cross-temporal multimodal fusion module that applies attention-based fusion to different modalities within and across video segments. As a result, it fully captures the temporal relations between different modalities. Furthermore, a single emotion label lacks supervision for learning representation of each segment, making temporal pattern mining difficult. We leverage time-synchronized comments (TSCs) as auxiliary supervision, since these comments are easily accessible and contain rich emotional cues. Two TSC-based self-supervised tasks are designed: the first aims to predict the emotional words in a TSC from video representation and TSC contextual semantics, and the second predicts the segment in which the TSC appears by calculating the correlation between video representation and TSC embedding. These self-supervised tasks are used to pre-train the cross-temporal multimodal fusion module on a large-scale video-TSC dataset, which is crawled from the web without labeling costs. These self-supervised pre-training tasks prompt the fusion module to perform representation learning on segments including TSC, thus capturing more temporal affective patterns. Experimental results on three benchmark datasets show that the proposed fusion module achieves state-of-the-art results in affective video content analysis. Ablation studies verify that after TSC-based pre-training, the fusion module learns more segments' affective patterns and achieves better performance. Jicai Pan, Shangfei Wang |
ACM Multimedia | 2 |
| 2022 | Two-Stage Multi-Scale Resolution-Adaptive Network for Low-Resolution Face RecognitionabstractLow-resolution face recognition is challenging due to uncertain input resolutions and the lack of distinguishing details in low-resolution (LR) facial images. Resolution-invariant representations must be learned for optimal performance. Existing methods for this task mainly minimize the distance between the representations of the low-resolution (LR) and corresponding high-resolution (HR) image pairs in a common subspace. However, these works only focus on introducing various distance metrics at the final layer and between HR-LR image pairs. They do not fully utilize the intermediate layers or multi-resolution supervision, yielding only modest performance. In this paper, we propose a novel two-stage multi-scale resolution-adaptive network to learn more robust resolution-invariant representations. In the first stage, the structural patterns and the semantic patterns are distilled from HR images to provide sufficient supervision for LR images. A curriculum learning strategy facilitates the training of HR and LR image matching, smoothly decreasing the resolution of LR images. In the second stage, a multi-resolution contrastive loss is introduced on LR images to enforce intra-class clustering and inter-class separation of the LR representations. By introducing multi-scale supervision and multi-resolution LR representation clustering, our network can produce robust representations despite uncertain input sizes. Experimental results on eight benchmark datasets demonstrate the effectiveness of the proposed method. Code will be released at https://github.com/hhwang98/TMR. Haihan Wang, Shangfei Wang |
ACM Multimedia | 2 |
| 2022 | Dual Learning for Facial Action Unit Detection Under Nonfull AnnotationabstractMost methods for facial action unit (AU) recognition typically require training images that are fully AU labeled. Manual AU annotation is time intensive. To alleviate this, we propose a novel dual learning framework and apply it to AU detection under two scenarios, that is, semisupervised AU detection with partially AU-labeled and fully expression-labeled samples, and weakly supervised AU detection with fully expression-labeled samples alone. We leverage two forms of auxiliary information. The first is the probabilistic duality between the AU detection task and its dual task, in this case, the face synthesis task given AU labels. We also take advantage of the dependencies among multiple AUs, the dependencies between expression and AUs, and the dependencies between facial features and AUs. Specifically, the proposed method consists of a classifier, an image generator, and a discriminator. The classifier and generator yield face-AU-expression tuples, which are forced to coverage of the ground-truth distribution. This joint distribution also includes three kinds of inherent dependencies: 1) the dependencies among multiple AUs; 2) the dependencies between expression and AUs; and 3) the dependencies between facial features and AUs. We reconstruct the inputted face and AU labels and introduce two reconstruction losses. In a semisupervised scenario, the supervised loss is also incorporated into the full objective for AU-labeled samples. In a weakly supervised scenario, we generate pseudo paired data according to the domain knowledge about expression and AUs. Semisupervised and weakly supervised experiments on three widely used datasets demonstrate the superiority of the proposed method for AU detection and facial synthesis tasks over current works. Shangfei Wang, Heyan Ding, Guozhu Peng |
IEEE Trans. Cybern. | 1 |
| 2021 | Pose-Invariant Facial Expression RecognitionabstractPose-invariant facial expression recognition is quite challenging due to variations in facial appearance and self-occlusion caused by head rotations. In this paper, we propose an adversarial multi-view subspace learning method for pose-robust facial expression recognition. Specifically, a deep neural network is trained from face images of a certain pose to learn facial representations. Then, an adversarial strategy is adopted to force statistical similarity among the learned representations from facial images with different poses. Simultaneously, an expression classifier is trained on the learned pose-robust facial representations. Through adversarial learning, the proposed method leverages inherent dependencies among multiple pose facial images to construct pose-robust image representations and a classifier during training. The ensemble method is adopted to combine the predictions of multiple deep neural networks and the common expression classifier, so pose estimation is not required. Experimental results on four benchmark databases demonstrate the superiority of the proposed method to state-of-the-art works. Guang Liang, Shangfei Wang |
FG | 2 |
| 2021 | Micro-Expression Recognition Enhanced by Macro-Expression from Spatial-Temporal DomainabstractFacial micro-expression recognition has attracted much attention due to its objectiveness to reveal the true emotion of a person. However, the limited micro-expression datasets have posed a great challenge to train a high performance micro-expression classifier. Since micro-expression and macro-expression share some similarities in both spatial and temporal facial behavior patterns, we propose a macro-to-micro transformation framework for micro-expression recognition. Specifically, we first pretrain two-stream baseline model from micro-expression data and macro-expression data respectively, named MiNet and MaNet. Then, we introduce two auxiliary tasks to align the spatial and temporal features learned from micro-expression data and macro-expression data. In spatial domain, we introduce a domain discriminator to align the features of MiNet and MaNet. In temporal domain, we introduce relation classifier to predict the correct relation for temporal features from MaNet and MiNet. Finally, we propose contrastive loss to encourage the MiNet to give closely aligned features to all entries from the same class in each instance. Experiments on three benchmark databases demonstrate the superiority of the proposed method. Bin Xia 0012, Shangfei Wang |
IJCAI | 2 |
| 2021 | Multi-task face analyses through adversarial learning
Shangfei Wang, Longfei Hao, Guang Liang |
Pattern Recognit. | 1 |
| 2021 | Deep Facial Action Unit Recognition and Intensity Estimation from Partially Labelled DataabstractResearch on facial action unit (AU) analysis typically require facial images that are labelled with those action units. While unlabelled facial images abound, labelling those images with action units or intensity is costly and time-consuming. Our approach makes it possible to analyze facial AUs when only some of the images have been labelled. We use many facial images to learn a deep framework that is able to take advantage of the facial representations. A restricted Boltzmann machine uses the available AU annotations to learn the AU label or intensity distribution. We train a support vector machine for AU recognition and a support vector regression for AU intensity estimation by maximizing the log likelihood of the AU mapping functions, taking into account the learned multiple AU distribution for all training data, while simultaneously diminishing errors between the predicted action units and ground-truth action unit occurrence or intensities for all labelled data. We perform experiments on two databases. The results demonstrate the superiority of a deep neural network for learning facial features, as well as the benefit of action unit label or intensity constraints for action unit occurrence recognition or intensity estimation in fully or semi-supervised scenarios. Shangfei Wang, Bowen Pan |
IEEE Trans. Affect. Comput. | 1 |
| 2021 | Capturing Emotion Distribution for Multimedia Emotion TaggingabstractMultimedia collections usually induce multiple emotions in audiences. The data distribution of multiple emotions can be leveraged to facilitate the learning process of emotion tagging, yet has not been thoroughly explored. To address this, we propose adversarial learning to fully capture emotion distributions for emotion tagging of multimedia data. The proposed multimedia emotion tagging approach includes an emotion classifier and a discriminator. The emotion classifier predicts emotion labels of multimedia data from their content. The discriminator distinguishes the predicted emotion labels from the ground truth labels. The emotion classifier and the discriminator are trained simultaneously in competition with each other. By jointly minimizing the traditional supervised loss and maximizing the distribution similarity between the predicted emotion labels and the ground truth emotion labels, the proposed multimedia emotion tagging approach successfully captures both the mapping function between multimedia content and emotion labels as well as prior distribution in emotion labels, and thus achieves state-of-the-art performance for multiple emotion tagging, as demonstrated by the experimental results on four benchmark databases. Shangfei Wang, Guozhu Peng, Zhuangqiang Zheng |
IEEE Trans. Affect. Comput. | 1 |
| 2021 | Video Affective Content Analysis by Exploring Domain KnowledgeabstractFilm grammar is often used to invoke certain emotional experiences from audiences through changing visual, speech, and musical elements of videos. Such film grammar, referred to as domain knowledge, is of great importance for video affective content analysis but has not been thoroughly examined in research. In this paper, we propose an improved method for emotion recognition and regression from videos through exploring domain knowledge. We first investigate the domain knowledge of visual, speech, and musical elements, and infer probabilistic dependencies between elements and emotions from the summarized film grammar. Then, we transfer the summarized dependencies between elements and emotions as constraints, and formulate video affective content analysis, including both emotion recognition and emotion regression from video content, as a constrained optimization problem. Experiments on the LIRIS-ACCEDE database, the FilmStim database, and the DEAP database demonstrate that the proposed video affective content analysis method can successfully leverage well-established film grammar to improve emotion recognition and regression from video content. Shangfei Wang, Can Wang 0007, Tanfang Chen, Yangyang Shu |
IEEE Trans. Affect. Comput. | 1 |
| 2020 | Unpaired Multimodal Facial Expression Recognition
Bin Xia 0012, Shangfei Wang |
ACCV (5) | 2 |
| 2020 | Occluded Facial Expression Recognition with Step-Wise Assistance from Unpaired Non-Occluded ImagesabstractAlthough facial expression recognition has improved in recent years, it is still very challenging to recognize expressions from occluded facial images in the wild. Due to the lack of large-scale facial expression datasets with diversity of the type and position of occlusions, it is very difficult to learn robust occluded expression classifier directly from limited occluded images. Considering facial images without occlusions usually provide more information for facial expression recognition compared to occluded facial images, we propose a step-wise learning strategy for occluded facial expression recognition that utilizes unpaired non-occluded images as guidance in the feature and label space. Specifically, we first measure the complexity of non-occluded data using distribution density in a feature space and split data into three subsets. In this way, the occluded expression classifier can be guided by basic samples first, and subsequently leverage more meaningful and discriminative samples. Complementary adversarial learning techniques are applied in the global-level and local-level feature space throughout, forcing the distribution of the occluded features to be close to the distribution of the non-occluded features. We also take the variability of the different images' transferability into account via adaptive classification loss. Loss inequality regularization is imposed in the label space to calibrate the output values of the occluded network. Experimental results show that our method improves performance on both synthesized occluded databases and realistic occluded databases. Bin Xia 0012, Shangfei Wang |
ACM Multimedia | 2 |
| 2020 | Learning from Macro-expression: a Micro-expression Recognition FrameworkabstractAs one of the most important forms of psychological behaviors, micro-expression can reveal the real emotion. However, the existing labeled micro-expression samples are limited to train a high performance micro-expression classifier. Since micro-expression and macro-expression share some similarities in facial muscle movements and texture changes, in this paper we propose a micro-expression recognition framework that leverages macro-expression samples as guidance. Specifically, we first introduce two Expression-Identity Disentangle Network, named MicroNet and MacroNet, as the feature extractor to disentangle expression-related features for micro and macro expression samples. Then MacroNet is fixed and used to guide the fine-tuning of MicroNet from both label and feature space. Adversarial learning strategy and triplet loss are added upon feature level between the MicroNet and MacroNet, so the MicroNet can efficiently capture the shared features of micro-expression and macro-expression samples. Loss inequality regularization is imposed to the label space to make the output of MicroNet converge to that of MicroNet. Comprehensive experiments on three public spontaneous micro-expression databases, i.e., SMIC, CASME2 and SAMM demonstrate the superiority of the proposed method. Bin Xia 0012, Shangfei Wang, Enhong Chen |
ACM Multimedia | 3 |
| 2020 | Exploiting Multi-Emotion Relations at Feature and Label Levels for Emotion TaggingabstractThe dependence among emotions is crucial to boost emotion tagging. In this paper, we propose a novel emotion tagging method, that thoroughly explores emotion relations from both the feature and label levels. Specifically, a graph convolutional network is introduced to inject local dependence among emotions into the model at the feature level, while an adversarial learning strategy is applied to constrain the joint distribution of multiple emotions at the label level. In addition, a new balanced loss function that mitigates the adverse effects of intra-class and inter-class imbalance is introduced to deal with the imbalance of emotion labels. Experimental results on several benchmark databases demonstrate the superiority of the proposed method compared to state-of-the-art works. Shangfei Wang |
ACM Multimedia | 2 |
| 2020 | Exploiting Self-Supervised and Semi-Supervised Learning for Facial Landmark Tracking with Unlabeled DataabstractCurrent work of facial landmark tracking usually requires large amounts of fully annotated facial videos to train a landmark tracker. To relieve the burden of manual annotations, we propose a novel facial landmark tracking method that makes full use of unlabeled facial videos by exploiting both self-supervised and semi-supervised learning mechanisms. First, self-supervised learning is adopted for representation learning from unlabeled facial videos. Specifically, a facial video and its shuffled version are fed into a feature encoder and a classifier. The feature encoder is used to learn visual representations, and the classifier distinguishes the input videos as the original or the shuffled ones. The feature encoder and the classifier are trained jointly. Through self-supervised learning, the spatial and temporal patterns of a facial video are captured at representation level. After that, the facial landmark tracker, consisting of the pre-trained feature encoder and a regressor, is trained semi-supervisedly. The consistencies among the tracking results of the original, the inverse and the disturbed facial sequences are exploited as the constraints on the unlabeled facial videos, and the supervised loss is adopted for the labeled videos. Through semi-supervised end-to-end training, the tracker captures sequential patterns inherent in facial videos despite small amount of manual annotations. Experiments on two benchmark datasets show that the proposed framework outperforms state-of-the-art semi-supervised facial landmark tracking methods, and also achieves advanced performance compared to fully supervised facial landmark tracking methods. Shangfei Wang, Enhong Chen |
ACM Multimedia | 2 |
| 2020 | Attentive One-Dimensional Heatmap Regression for Facial Landmark Detection and TrackingabstractAlthough heatmap regression is considered a state-of-the-art method to locate facial landmarks, it suffers from huge spatial complexity and is prone to quantization error. To address this, we propose a novel attentive one-dimensional heatmap regression method for facial landmark localization. First, we predict two groups of 1D heatmaps to represent the marginal distributions of the x and y coordinates. These 1D heatmaps reduce spatial complexity significantly compared to current heatmap regression methods, which use 2D heatmaps to represent the joint distributions of x and y coordinates. With much lower spatial complexity, the proposed method can output high-resolution 1D heatmaps despite limited GPU memory, significantly alleviating the quantization error. Second, a co-attention mechanism is adopted to model the inherent spatial patterns existing in x and y coordinates, and therefore the joint distributions on the x and y axes are also captured. Third, based on the 1D heatmap structures, we propose a facial landmark detector capturing spatial patterns for landmark detection on an image; and a tracker further capturing temporal patterns with a temporal refinement mechanism for landmark tracking. Experimental results on four benchmark databases demonstrate the superiority of our method. Shangfei Wang, Enhong Chen, Cong Liang 0002 |
ACM Multimedia | 2 |
| 2020 | A Novel Dynamic Model Capturing Spatial and Temporal Patterns for Facial Expression AnalysisabstractFacial expression analysis could be greatly improved by incorporating spatial and temporal patterns present in facial behavior, but the patterns have not yet been utilized to their full advantage. We remedy this via a novel dynamic model-an interval temporal restricted Boltzmann machine (IT-RBM) - that is able to capture both universal spatial patterns and complicated temporal patterns in facial behavior for facial expression analysis. We regard a facial expression as a multifarious activity composed of sequential or overlapping primitive facial events. Allen's interval algebra is implemented to portray these complicated temporal patterns via a two-layer Bayesian network. The nodes in the upper-most layer are representative of the primitive facial events, and the nodes in the lower layer depict the temporal relationships between those events. Our model also captures inherent universal spatial patterns via a multi-value restricted Boltzmann machine in which the visible nodes are facial events, and the connections between hidden and visible nodes model intrinsic spatial patterns. Efficient learning and inference algorithms are proposed. Experiments on posed and spontaneous expression distinction and expression recognition demonstrate that our proposed IT-RBM achieves superior performance compared to state-of-the art research due to its ability to incorporate these facial behavior patterns. Shangfei Wang, Zhuangqiang Zheng, Jiajia Yang 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2020 | Exploring Domain Knowledge for Facial Expression-Assisted Action Unit Activation RecognitionabstractCurrent works on facial action unit (AU) activation recognition typically include supervised training using AU-annotated training images. Compared to facial expression labeling, AU annotation is a time-consuming, expensive, and error-prone process. Domain knowledge refers to the strong probabilistic dependencies between facial expressions and AUs, as well as dependencies among AUs. To take advantage of this, we avoid the time-consuming process of AU annotation and introduce a new AU activation recognition method that learns AU classifiers from domain knowledge, and requires only expression-annotated facial images. Specifically, we first generate pseudo AU labels according to the probabilistic dependencies between expressions and AUs as well as correlations among AUs summarized from domain knowledge. Then, we propose to use a Restricted Boltzmann Machine to model AU label prior distribution from the generated pseudo AU data. After that, we train AU classifiers from expression-annotated facial images and the learned prior model by maximizing the log likelihood of AU classifiers with regard to the learned AU label prior. The proposed AU activation recognition can also be extended to semi-supervised learning scenarios with partially AU-annotated facial images. Experimental results on four benchmark databases demonstrate the effectiveness of the proposed approach in learning AU classifiers from domain knowledge. Shangfei Wang, Guozhu Peng |
IEEE Trans. Affect. Comput. | 1 |
| 2020 | Capturing Joint Label Distribution for Multi-Label Classification Through Adversarial LearningabstractLabel correlations are important for multi-label learning. Although current multi-label learning approaches can exploit first-order, second-order, and high-order label dependencies, they fail to exploit complete label correlations, which are included in the joint label distribution of the ground truth labels. However, directly modeling the complex and unknown joint label distribution is very challenging, if not impossible. In this paper, we propose an adversarial learning framework to enforce similarity between joint distribution of the ground truth multi-labels and the predicted multiple labels. Specifically, the proposed multi-label learning method includes a multi-label classifier and a label discriminator. The classifier minimizes error between predicted labels and corresponding ground truth labels and gives the discriminator room for error. The object of the discriminator is to distinguish the predicted labels from the ground truth labels. The classifier and discriminator are trained simultaneously through an alternate process. By adversarial learning, the joint label distribution of the predicted multi-labels converges to the joint distribution inherent in the ground truth multi-labels, and thus boosts the performance of multi-label learning as demonstrated in the experiments on 11 benchmark databases. Shangfei Wang, Guozhu Peng, Zhuangqiang Zheng |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2020 | Knowledge-Augmented Multimodal Deep Regression Bayesian Networks for Emotion Video TaggingabstractThe immanent dependencies between audio and visual modalities extracted from video content and the well-established film grammar (i.e., domain knowledge) are important for emotion video recognition and regression. However, these tools have yet to be exploited successfully. Therefore, we propose a multimodal deep regression Bayesian network (MMDRBN) to capture the relationship between audio and visual modalities for emotion video tagging. We then modify the structure of the MMDRBN to incorporate domain knowledge. A regression Bayesian network (RBN) is formed from one latent layer, one visible layer and directed links from the latent layer to the visible layer. RBN is able to fully represent the data, since it captures the dependencies not only among the visible variables but also among the latent variables given visible variables. For the MMDRBN, first, we learn several layers of RBNs using audio and visual modalities, and then stack these RBNs to form two deep networks. A joint representation is obtained from the top layers of the two deep networks, capturing the deep dependencies between audio and visual modalities. We also summarize the main audio and visual elements used by filmmakers to convey emotions and formulate them as semantical meaningful middle-level representation, i.e., attributes. Through these attributes, we construct the knowledge-augmented MMDRBN, which learns a hybrid middle-level video representation using video data and the summarized attributes. Experimental results of both emotion recognition and regression from videos on the LIRIS-ACCEDE database demonstrate that the proposed model can successfully capture the intrinsic connections between audio and visual modalities, and integrate the middle-level representation learning from video data and semantical attributes summarized from film grammar. Thus, it achieves superior performance on emotion video tagging compared to state-of-the-art methods. Shangfei Wang, Longfei Hao |
IEEE Trans. Multim. | 1 |
| 2020 | Posed and Spontaneous Expression Distinction Using Latent Regression Bayesian NetworksabstractFacial spatial patterns can help distinguish between posed and spontaneous expressions, but this information has not been thoroughly leveraged by current studies. We present several latent regression Bayesian networks (LRBNs) to capture the patterns existing in facial landmark points and to use those points to differentiate posed from spontaneous expressions. The visible nodes of the LRBN represent facial landmark points. Through learning, the LRBN captures the probabilistic dependencies among landmark points as well as latent variables given observations, successfully modeling the spatial patterns inherent in expressions. Current methods tend to ignore gender and expression categories, although these factors can influence spatial patterns. Therefore, we propose to incorporate this as a kind of privileged information. We construct several LRBNs to capture spatial patterns from spontaneous and posed facial expressions given expression-related factors. Facial landmark points are used during testing to classify samples as either posed or spontaneous, depending on which LRBN has the largest likelihood. We conduct experiments to showcase the superiority of the proposed approach in both modeling spatial patterns and classifying expressions as either posed or spontaneous. Shangfei Wang, Longfei Hao |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2019 | Image Aesthetic Assessment Assisted by Attributes through Adversarial LearningabstractThe inherent connections among aesthetic attributes and aesthetics are crucial for image aesthetic assessment, but have not been thoroughly explored yet. In this paper, we propose a novel image aesthetic assessment assisted by attributes through both representation-level and label-level. The attributes are used as privileged information, which is only required during training. Specifically, we first propose a multitask deep convolutional rating network to learn the aesthetic score and attributes simultaneously. The attributes are explored to construct better feature representations for aesthetic assessment through multi-task learning. After that, we introduce a discriminator to distinguish the predicted attributes and aesthetics of the multi-task deep network from the ground truth label distribution embedded in the training data. The multi-task deep network wants to output aesthetic score and attributes as close to the ground truth labels as possible. Thus the deep network and the discriminator compete with each other. Through adversarial learning, the attributes are explored to enforce the distribution of the predicted attributes and aesthetics to converge to the ground truth label distribution. Experimental results on two benchmark databases demonstrate the superiority of the proposed method to state of the art work. Bowen Pan, Shangfei Wang, Qisheng Jiang 0003 |
AAAI | 2 |
| 2019 | Dual Semi-Supervised Learning for Facial Action Unit RecognitionabstractCurrent works on facial action unit (AU) recognition typically require fully AU-labeled training samples. To reduce the reliance on time-consuming manual AU annotations, we propose a novel semi-supervised AU recognition method leveraging two kinds of readily available auxiliary information. The method leverages the dependencies between AUs and expressions as well as the dependencies among AUs, which are caused by facial anatomy and therefore embedded in all facial images, independent on their AU annotation status. The other auxiliary information is facial image synthesis given AUs, the dual task of AU recognition from facial images, and therefore has intrinsic probabilistic connections with AU recognition, regardless of AU annotations. Specifically, we propose a dual semi-supervised generative adversarial network for AU recognition from partially AU-labeled and fully expressionlabeled facial images. The proposed network consists of an AU classifier C, an image generator G, and a discriminator D. In addition to minimize the supervised losses of the AU classifier and the face generator for labeled training data, we explore the probabilistic duality between the tasks using adversary learning to force the convergence of the face-AU-expression tuples generated from the AU classifier and the face generator, and the ground-truth distribution in labeled data for all training data. This joint distribution also includes the inherent AU dependencies. Furthermore, we reconstruct the facial image using the output of the AU classifier as the input of the face generator, and create AU labels by feeding the output of the face generator to the AU classifier. We minimize reconstruction losses for all training data, thus exploiting the informative feedback provided by the dual tasks. Within-database and cross-database experiments on three benchmark databases demonstrate the superiority of our method in both AU recognition and face synthesis compared to state-of-the-art works. Guozhu Peng, Shangfei Wang |
AAAI | 2 |
| 2019 | Integrating Facial Images, Speeches and Time for Empathy PredictionabstractWe propose a multi-modal method for the One-Minute Empathy Prediction competition. First, we use bottleneck residual and fully-connected network to encode facial images and speeches of the listener. Second, we propose to use the current time stage as a temporal feature and encoded it into the proposed multi-modal network. Third, we select a subset training data based on its performance of empathy prediction on the validation data. Experimental results on the testing set show that the proposed method outperforms the baseline methods significantly according to the CCC metric (0.14 vs 0.06). Yonggan Fu, Runlong Wu, Heyan Ding, Shangfei Wang |
FG | 6 |
| 2019 | Capturing Spatial and Temporal Patterns for Facial Landmark Tracking through Adversarial LearningabstractThe spatial and temporal patterns inherent in facial feature points are crucial for facial landmark tracking, but have not been thoroughly explored yet. In this paper, we propose a novel deep adversarial framework to explore the shape and temporal dependencies from both appearance level and target label level. The proposed deep adversarial framework consists of a deep landmark tracker and a discriminator. The deep landmark tracker is composed of a stacked Hourglass network as well as a convolutional neural network and a long short-term memory network, and thus implicitly capture spatial and temporal patterns from facial appearance for facial landmark tracking. The discriminator is adopted to distinguish the tracked facial landmarks from ground truth ones. It explicitly models shape and temporal dependencies existing in ground truth facial landmarks through another convolutional neural network and another long short-term memory network. The deep landmark tracker and the discriminator compete with each other. Through adversarial learning, the proposed deep adversarial landmark tracking approach leverages inherent spatial and temporal patterns to facilitate facial landmark tracking from both appearance level and target label level. Experimental results on two benchmark databases demonstrate the superiority of the proposed approach to state-of-the-art work. Shangfei Wang, Guozhu Peng, Bowen Pan |
IJCAI | 2 |
| 2019 | KDSL: a Knowledge-Driven Supervised Learning Framework for Word Sense DisambiguationabstractWe propose KDSL, a new word sense disambiguation (WSD) framework that utilizes knowledge to automatically generate sense-labeled data for supervised learning. First, from WordNet, we automatically construct a semantic knowledge base called DisDict, which provides refined feature words that highlight the differences among word senses, i.e., synsets. Second, we automatically generate new sense-labeled data by DisDict from unlabeled corpora. Third, these generated data, together with manually labeled data and unlabeled data, are fed to a neural framework conducting supervised and unsupervised learning jointly to model the semantic relations among synsets, feature words and their contexts. The experimental results show that KDSL outperforms several representative state-of-the-art methods on various major benchmarks. Interestingly, it performs relatively well even when manually labeled data is unavailable, thus provides a potential solution for similar tasks in a lack of manual annotations. Shangfei Wang, Jianmin Ji, Ruili Wang 0001 |
IJCNN | 4 |
| 2019 | Occluded Facial Expression Recognition Enhanced through Privileged InformationabstractIn this paper, we propose a novel approach of occluded facial expression recognition under the help of non-occluded facial images. The non-occluded facial images are used as privileged information, which is only required during training, but not required during testing. Specifically, two deep neural networks are first trained from occluded and non-occluded facial images respectively. Then the non-occluded network is fixed and is used to guide the fine-tuning of the occluded network from both label space and feature space. Similarity constraint and loss inequality regularization are imposed to the label space to make the output of occluded network converge to that of the non-occluded network. Adversarial leaning is adopted to force the distribution of the learned features from occluded facial images to be close to that from non-occluded facial images. Furthermore, a decoder network is employed to reconstruct the non-occluded facial images from occluded features. Under the guidance of non-occluded facial images, the occluded network is expected to learn better features and classifier during training. Experiments on the benchmark databases with both synthesized and realistic occluded facial images demonstrate the superiority of the proposed method to state-of-the-art. Bowen Pan, Shangfei Wang, Bin Xia 0012 |
ACM Multimedia | 2 |
| 2019 | Identity- and Pose-Robust Facial Expression Recognition through Adversarial Feature LearningabstractExisting facial expression recognition methods either focus on pose variations or identity bias, but not both simultaneously. This paper proposes an adversarial feature learning method to address both of these issues. Specifically, the proposed method consists of five components: an encoder, an expression classifier, a pose discriminator, a subject discriminator, and a generator. An encoder extracts feature representations, and an expression classifier tries to perform facial expression recognition using the extracted feature representations. The encoder and the expression classifier are trained collaboratively, so that the extracted feature representations are discriminative for expression recognition. A pose discriminator and a subject discriminator classify the pose and the subject from the extracted feature representations respectively. They are trained adversarially with the encoder. Thus, the extracted feature representations are robust to poses and subjects. A generator reconstructs facial images to further favor the feature representations. Experiments on five benchmark databases demonstrate the superiority of the proposed method to state-of-the-art work. Shangfei Wang, Guang Liang |
ACM Multimedia | 2 |
| 2019 | Content-Based Video Emotion Tagging Augmented by Users' Multiple Physiological ResponsesabstractThe intrinsic interactions among a video's emotion tag, its content, and a user's spontaneous responses while consuming the video can be leveraged to improve video emotion tagging, but such interactions have not been thoroughly exploited yet. In this paper, we propose a novel content-based video emotion tagging approach augmented by users' multiple physiological responses, which are only required during training. Specifically, a better emotion tagging model is constructed by introducing similarity constraints on the classifiers from video content and multiple physiological signals available during training. Maximum margin classifiers are adopted and efficient learning algorithms of the proposed model are also developed. Furthermore, the proposed video emotion tagging approach is extended to utilize incomplete physiological signals, since these signals are often corrupted by artifacts. Experiments on four benchmark databases demonstrate the effectiveness of the proposed method for implicitly integrating multiple physiological responses, and its superior performance to existing methods using both complete and incomplete multiple physiological signals. Shangfei Wang |
IEEE Trans. Affect. Comput. | 1 |
| 2019 | Capturing Feature and Label Relations Simultaneously for Multiple Facial Action Unit RecognitionabstractAlthough both feature dependencies and label dependencies are crucial for facial action unit (AU) recognition, little work addresses them simultaneously till now. In this paper, we propose a 4-layer Restricted Boltzmann Machine (RBM) to simultaneously capture feature level and label level dependencies to recognize multiple AUs. The middle hidden layer of the 4-layer RBM model captures dependencies among image features for multiple AUs, while the top latent units capture the high-order semantic dependencies among AU labels. Furthermore, we extend the proposed 4-layer RBM for facial expression-augmented AU recognition, since AU relations are influenced by expressions. By introducing facial expression nodes in the middle visible layer, facial expressions, which are only required during training, facilitate the estimation of both feature dependencies and label dependencies among AUs. Efficient learning and inference algorithms for the extended model are also developed. Experimental results on three benchmark databases, i.e., the CK+ database, the DISFA database and the SEMAINE database, demonstrate that the proposed approaches can successfully capture complex AU relationships from features and labels jointly, and the expression labels available only during training are benefit for AU recognition during testing for both posed and spontaneous facial expressions. Shangfei Wang, Guozhu Peng |
IEEE Trans. Affect. Comput. | 1 |
| 2019 | Facial Action Unit Recognition and Intensity Estimation Enhanced Through Label DependenciesabstractThe inherent dependencies among facial action units (AU) caused by the underlying anatomic mechanism are essential for the proper recognition of AUs and estimation of intensity levels, but they have not been exploited to their full potential. We are proposing novel methods to recognize AUs and estimate intensity via hybrid Bayesian networks. The upper two layers are latent regression Bayesian networks (LRBNs), and the lower layers are Bayesian networks (BNs). The visible nodes of the LRBN layers are representations of ground-truth AU occurrences or AU intensities. Through the directed connections from latent layer and visible layer, an LRBN can successfully represent relationships between multiple AUs or AU intensities. The lower layers include Bayesian networks with two nodes for AU recognition, and Bayesian networks with three nodes for AU intensity estimation. The bottom layers incorporate measurements from facial images with AU dependencies for intensity estimation and AU recognition. Efficient learning algorithms of the hybrid Bayesian networks are proposed for AU recognition as well as intensity estimation. Furthermore, the proposed hybrid Bayesian network models are extended for facial expression-assisted AU recognition and intensity estimation, as AU relationships are closely related to facial expressions. We test our methods on three benchmark databases for AU recognition and two benchmark databases for intensity estimation. The results demonstrate that the proposed approaches faithfully model the complex and global inherent AU dependencies, and the expression labels available only during training can boost the estimation of AU dependencies for both AU recognition and intensity estimation. Shangfei Wang, Longfei Hao |
IEEE Trans. Image Process. | 1 |
| 2019 | Weakly Supervised Dual Learning for Facial Action Unit RecognitionabstractCurrent research on facial action unit (AU) recognition typically requires fully AU-annotated facial images. Compared to facial expression labeling, AU annotation is a time-consuming, expensive, and error-prone process. Inspired by dual learning, we propose a novel weakly supervised dual learning mechanism to train facial action unit classifiers from expression-annotated images. Specifically, we consider AU recognition from facial images as the main task, and face synthesis given AUs as the auxiliary task. For AU recognition, we force the recognized AUs to satisfy the expression-dependent and expression-independent AU dependencies, i.e., the domain knowledge about expressions and AUs. For face synthesis given AUs, we minimize the difference between the synthetic face and the ground truth face, which has identical recognized and given AUs. By optimizing the dual tasks simultaneously, we successfully leverage their intrinsic connections as well as domain knowledge about expressions and AUs to facilitate the learning of AU classifiers from expression-annotated image. Furthermore, we extend the proposed weakly supervised dual learning mechanism to semi-supervised dual learning scenarios with partially AU-annotated images. Experimental results on three benchmark databases demonstrate the effectiveness of the proposed approach for both tasks. Shangfei Wang, Guozhu Peng |
IEEE Trans. Multim. | 1 |
| 2018 | Weakly Supervised Facial Action Unit Recognition Through Adversarial TrainingabstractCurrent works on facial action unit (AU) recognition typically require fully AU-annotated facial images for supervised AU classifier training. AU annotation is a time-consuming, expensive, and error-prone process. While AUs are hard to annotate, facial expression is relatively easy to label. Furthermore, there exist strong probabilistic dependencies between expressions and AUs as well as dependencies among AUs. Such dependencies are referred to as domain knowledge. In this paper, we propose a novel AU recognition method that learns AU classifiers from domain knowledge and expression-annotated facial images through adversarial training. Specifically, we first generate pseudo AU labels according to the probabilistic dependencies between expressions and AUs as well as correlations among AUs summarized from domain knowledge. Then we propose a weakly supervised AU recognition method via an adversarial process, in which we simultaneously train two models: a recognition model R, which learns AU classifiers, and a discrimination model D, which estimates the probability that AU labels generated from domain knowledge rather than the recognized AU labels from R. The training procedure for R maximizes the probability of D making a mistake. By leveraging the adversarial mechanism, the distribution of recognized AUs is closed to AU prior distribution from domain knowledge. Furthermore, the proposed weakly supervised AU recognition can be extended to semi-supervised learning scenarios with partially AU-annotated images. Experimental results on three benchmark databases demonstrate that the proposed method successfully leverages the summarized domain knowledge to weakly supervised AU classifier learning through an adversarial process, and thus achieves state-of-the-art performance. Guozhu Peng, Shangfei Wang |
CVPR | 2 |
| 2018 | Facial Action Unit Recognition Augmented by Their DependenciesabstractDue to the underlying anatomic mechanism that govern facial muscular interactions, there exist inherent dependencies between facial action units (AU). Such dependencies carry crucial information for AU recognition, yet have not been thoroughly exploited. Therefore, in this paper, we propose a novel AU recognition method with a three-layer hybrid Bayesian network, whose top two layers consist of a latent regression Bayesian network (LRBN), and the bottom two layers are Bayesian networks. The LRBN is a directed graphical model consisting of one latent layer and one visible layer. Specifically, the visible nodes of LRBN represent the ground-truth AU labels. Due to the "explaining away" effect in Bayesian networks, LRBN is able to capture both the dependencies among the latent variables given the observation and the dependencies among visible variables. Such dependencies successfully and faithfully represent relations among multiple AUs. The bottom two layers are two node Bayesian networks, connecting the ground truth AU labels and their measurements. Efficient learning and inference algorithms are also proposed. Furthermore, we extend the proposed hybrid Bayesian network model for facial expression-assisted AU recognition, since AU relations are influenced by expressions. By introducing facial expression nodes in the middle visible layer, facial expressions, which are only required during training, facilitate the estimation of label dependencies among AUs. Experimental results on three benchmark databases, i.e. the CK+ database, the SEMAINE database, and the BP4D database, demonstrate that the proposed approaches can successfully capture complex AU relationships, and the expression labels available only during training are benefit for AU recognition during testing. Longfei Hao, Shangfei Wang, Guozhu Peng |
FG | 2 |
| 2018 | Facial Expression Recognition Enhanced by Thermal Images through Adversarial LearningabstractCurrently, fusing visible and thermal images for facial expression recognition requires two modalities during both training and testing. Visible cameras are commonly used in real-life applications, and thermal cameras are typically only available in lab situations due to their high price. Thermal imaging for facial expression recognition is not frequently used in real-world situations. To address this, we propose a novel thermally enhanced facial expression recognition method which uses thermal images as privileged information to construct better visible feature representation and improved classifiers by incorporating adversarial learning and similarity constraints during training. Specifically, we train two deep neural networks from visible images and thermal images. We impose adversarial loss to enforce statistical similarity between the learned representations of two modalities, and a similarity constraint to regulate the mapping functions from visible and thermal representation to expressions. Thus, thermal images are leveraged to simultaneously improve visible feature representation and classification during training. To mimic real-world scenarios, only visible images are available during testing. We further extend the proposed expression recognition method for partially unpaired data to explore thermal images' supplementary role in visible facial expression recognition when visible images and thermal images are not synchronously recorded. Experimental results on the MAHNOB Laughter database demonstrate that our proposed method can effectively regularize visible representation and expression classifiers with the help of thermal images, achieving state-of-the-art recognition performance. Bowen Pan, Shangfei Wang |
ACM Multimedia | 2 |
| 2018 | Personalized Multiple Facial Action Unit Recognition through Generative Adversarial Recognition NetworkabstractPersonalized facial action unit (AU) recognition is challenging due to subject-dependent facial behavior. This paper proposes a method to recognize personalized multiple facial AUs through a novel generative adversarial network, which adapts the distribution of source domain facial images to that of target domain facial images and detects multiple AUs by leveraging AU dependencies. Specifically, we use a generative adversarial network to generate synthetic images from source domain; the synthetic images have a similar appearance to the target subject and retain the AU patterns of the source images. We simultaneously leverage AU dependencies to train a multiple AU classifier. Experimental results on three benchmark databases demonstrate that the proposed method can successfully realize unsupervised domain adaptation for individual AU detection, and thus outperforms state-of-the-art AU detection methods. Shangfei Wang |
ACM Multimedia | 2 |
| 2018 | Learning with privileged information for multi-Label classification
Shangfei Wang, Tanfang Chen, Xiaoxiao Shi |
Pattern Recognit. | 1 |
| 2018 | Thermal Augmented Expression RecognitionabstractVisible facial images provide geometric and appearance patterns of facial expressions and are sensitive to illumination changes. Thermal facial images record facial temperature distribution and are robust to light conditions. Therefore, expression recognition is enhanced by visible and thermal image fusion. In most cases, only visible images are available due to the widespread popularity of visible cameras and the high cost of thermal cameras. Thus, we propose a novel visible expression recognition method by using thermal infrared (IR) data as privileged information, which is only available during training. Specifically, we first learn a deep model for visible images and thermal images. Then we use the learned feature representations to train support vector machine (SVM) classifiers for expression classification. We jointly refine the deep models as well as the SVM classifiers for both thermal images and visible images by imposing the constraint that the outputs of the SVM classifiers from two views are similar. Thermal IR images during training are then exploited to construct better facial representations and expression classifiers from visible images. We extend the proposed thermal augmented expression recognition method for partially unpaired data, acknowledging that visible images and thermal images maybe not be recorded synchronously. Experimental resulton the MAHNOB laughter database demonstrate that the proposed thermal augmented expression recognition method can effectively exploit thermal IR images' supplementary role for visible facial expression recognition during training to obtain better facial representations and a better visible expression classifier. The proposed thermal augmented expression recognition method achieves state-of-the-art expression recognition performance for both paired and unpaired facial images. Shangfei Wang, Bowen Pan, Huaping Chen 0001 |
IEEE Trans. Cybern. | 1 |
| 2018 | Weakly Supervised Facial Action Unit Recognition With Domain KnowledgeabstractCurrent facial action unit (AU) recognition typically includes supervised training, where the fully AU annotated training images are required. Due to the nuances of facial appearance and individual differences, AU annotation is a time-consuming, expensive, and error-prone process. Facial expression is relatively simple to label, since facial expressions describe facial behavior globally and the number of expressions appearing on a face is much less than that of AUs. Furthermore, there exist strong dependencies between AUs and expressions, referred to as domain knowledge. Such domain knowledge is inherent in facial anatomy and facial behavior. Therefore, in this paper, we propose a novel weakly supervised AU recognition method to jointly learn multiple AU classifiers with expression annotations but without any AU annotations by leveraging domain knowledge. Specifically, we first summarize the expression-dependent AU ranking from the domain knowledge of conditional probabilities of AUs given expressions. Then, we formulate the weakly supervised AU recognition as a multilabel ranking problem and propose an efficient learning algorithm to solve it. Furthermore, we extend the proposed weakly supervised AU recognition method to a semi-supervised learning scenario when partial AU labeled samples are available. Experimental results on three benchmark databases demonstrate that the proposed method can successfully exploit domain knowledge for multiple AU recognition and, thus, outperforms both state-of-the-art weakly supervised AU recognition method and the semi-supervised AU recognition method. Shangfei Wang, Guozhu Peng |
IEEE Trans. Cybern. | 1 |
| 2017 | Differentiating Between Posed and Spontaneous Expressions with Latent Regression Bayesian NetworkabstractSpatial patterns embedded in human faces are crucial for differentiating posed expressions from spontaneous ones, yet they have not been thoroughly exploited in the literature. To tackle this problem, we present a generative model, i.e., Latent Regression Bayesian Network (LRBN), to effectively capture the spatial patterns embedded in facial landmark points to differentiate between posed and spontaneous facial expressions. The LRBN is a directed graphical model consisting of one latent layer and one visible layer. Due to the “explaining away“ effect in Bayesian networks, LRBN is able to capture both the dependencies among the latent variables given the observation and the dependencies among visible variables. We believe that such dependencies are crucial for faithful data representation. Specifically, during training, we construct two LRBNs to capture spatial patterns inherent in displacements of landmark points from spontaneous facial expressions and posed facial expressions respectively. During testing, the samples are classified into posed or spontaneous expressions according to their likelihoods on two models. Efficient learning and inference algorithms are proposed. Experimental results on two benchmark databases demonstrate the advantages of the proposed approach in modeling spatial patterns as well as its superior performance to the existing methods in differentiating between posed and spontaneous expressions. Siqi Nie, Shangfei Wang |
AAAI | 3 |
| 2017 | Capturing Dependencies among Labels and Features for Multiple Emotion Tagging of Multimedia DataabstractIn this paper, we tackle the problem of emotion tagging of multimedia data by modeling the dependencies among multiple emotions in both the feature and label spaces. These dependencies, which carry crucial top-down and bottom-up evidence for improving multimedia affective content analysis, have not been thoroughly exploited yet. To this end, we propose two hierarchical models that independently and dependently learn the shared features and global semantic relationships among emotion labels to jointly tag multiple emotion labels of multimedia data. Efficient learning and inference algorithms of the proposed models are also developed. Experiments on three benchmark emotion databases demonstrate the superior performance of our methods to existing methods. Shangfei Wang |
AAAI | 2 |
| 2017 | Emotion recognition through integrating EEG and peripheral signalsabstractThe inherent dependencies among multiple physiological signals are crucial for multimodal emotion recognition, but have not been thoroughly exploited yet. This paper propose to use restricted Boltzmann machine (RBM) to model such dependencies.Specifically, the visible nodes of RBM represent EEG and peripheral physiological signals, and thus the connections between visible nodes and hidden nodes capture the intrinsic relations among multiple physiological signals. The RBM generates new representation from multiple physiological signals. Then, a support vector machine is adopted to recognize users' emotion states from the generated features. Furthermore, we extend the proposed fusion method for incomplete datas, since physiological signals are often corrupted due to artifacts. Specifically, we pre-train the RBM using all the complete data, then we update missing values and RBM parameters to minimize free energy of visible vectors using both complete and incomplete data. Experiments on two benchmark databases demonstrate the effectiveness of the proposed methods. Yangyang Shu, Shangfei Wang |
ICASSP | 2 |
| 2017 | Personalized video emotion tagging through a topic modelabstractThe inherent dependencies among video content, personal characteristics, and perceptual emotion are crucial for personalized video emotion tagging, but have not been thoroughly exploited. To address this, we propose a novel topic model to capture such inherent dependencies. We assume that there are several potential human factors, or “topics,” that affect the personal characteristics and the personalized emotion responses to videos. During training, the proposed topic model exploits the latent space to model the relationships among personal characteristics, video content and video tagging behaviors. After learning, the proposed model can generate meaningful latent topics, which help personalized video emotion tagging. Efficient learning and inference algorithms of the model are proposed. Experimental results on the CP-QAE-I database demonstrate the effectiveness of the proposed approach in modeling complex relationships among video content, personal characteristics, and perceptual emotion, as well as its good performance in personalized video emotion. Shangfei Wang, Zhen Gao 0003 |
ICASSP | 2 |
| 2017 | A Multimodal Deep Regression Bayesian Network for Affective Video Content AnalysesabstractThe inherent dependencies between visual elements and aural elements are crucial for affective video content analyses, yet have not been successfully exploited. Therefore, we propose a multimodal deep regression Bayesian network (MMDRBN) to capture the dependencies between visual elements and aural elements for affective video content analyses. The regression Bayesian network (RBN) is a directed graphical model consisting of one latent layer and one visible layer. Due to the explaining away effect in Bayesian networks (BN), RBN is able to capture both the dependencies among the latent variables given the observation and the dependencies among visible variables. We propose a fast learning algorithm to learn the RBN. For the MMDRB-N, first, we learn several RBNs layer-wisely from visual modality and audio modality respectively. Then we stack these RBNs and obtain two deep networks. After that, a joint representation is extracted from the top layers of the two deep networks, and thus captures the high order dependencies between visual modality and audio modality. In order to predict the valence or arousal score of video contents, we initialize a feed-forward inference network from the MMDRBN whose inference is intractable by minimizing the KullbackCLeibler (KL)divergence between the two networks. The back propagation algorithm is adopted for finetuning the inference network. Experimental results on the LIRIS-ACCEDE database demonstrate that the proposed MMDRBN successfully captures the dependencies between visual and audio elements, and thus achieves better performance compared with state-of-the-art work. Shangfei Wang, Longfei Hao |
ICCV | 2 |
| 2017 | Deep Facial Action Unit Recognition from Partially Labeled DataabstractCurrent work on facial action unit (AU) recognition requires AU-labeled facial images. Although large amounts of facial images are readily available, AU annotation is expensive and time consuming. To address this, we propose a deep facial action unit recognition approach learning from partially AU-labeled data. The proposed approach makes full use of both partly available ground-truth AU labels and the readily available large scale facial images without annotation. Specifically, we propose to learn label distribution from the ground-truth AU labels, and then train the AU classifiers from the large-scale facial images by maximizing the log likelihood of the mapping functions of AUs with regard to the learnt label distribution for all training data and minimizing the error between predicted AUs and ground-truth AUs for labeled data simultaneously. A restricted Boltzmann machine is adopted to model AU label distribution, a deep neural network is used to learn facial representation from facial images, and the support vector machine is employed as the classifier. Experiments on two benchmark databases demonstrate the effectiveness of the proposed approach. Shangfei Wang, Bowen Pan |
ICCV | 2 |
| 2017 | Deep multimodal network for multi-label classificationabstractCurrent multimodal deep learning approaches rarely explicitly exploit the dependencies inherent in multiple labels, which are crucial for multimodal multi-label classification. In this paper, we propose a multimodal deep learning approach for multi-label classification. Specifically, we introduce deep networks for feature representation learning and construct classifiers with the objective function which is constrained with dependencies among both labels and modals. We further propose effective training algorithm to learn deep networks and classifiers jointly. Thus, we explicitly leverage the relations among labels and modals to facilitate multimodal multi-label classification. Experiments of multi-label classification and cross-modal retrieval on the Pascal VOC dataset and the La-belMe dataset demonstrate the effectiveness of the proposed approach. Tanfang Chen, Shangfei Wang |
ICME | 2 |
| 2017 | Exploring Domain Knowledge for Affective Video Content AnalysesabstractThe well-established film grammar is often used to change visual and audio elements of videos to invoke audiences' emotional experience. Such film grammar, referred to as domain knowledge, is crucial for affective video content analyses, but has not been thoroughly explored yet. In this paper, we propose a novel method to analyze video affective content through exploring domain knowledge. Specifically, take visual elements as an example, we first infer probabilistic dependencies between visual elements and emotions from the summarized film grammar. Then, we transfer the domain knowledge as constraints, and formulate affective video content analyses as a constrained optimization problem. Experiments on the LIRIS-ACCEDE database and the DEAP database demonstrate that the proposed affective content analyses method can successfully leverage well-established film grammar for better emotion classification from video content. Tanfang Chen, Shangfei Wang |
ACM Multimedia | 3 |
| 2017 | Capturing Spatial and Temporal Patterns for Distinguishing between Posed and Spontaneous ExpressionsabstractSpatial and temporal patterns inherent in facial behavior carry crucial information for posed and spontaneous expressions distinction, but have not been thoroughly exploited yet. To address this issue, we propose a novel dynamic model, termed as interval temporal restricted Boltzmann machine (IT-RBM), to jointly capture global spatial patterns and complex temporal patterns embedded in posed and spontaneous expressions respectively for distinguishing between posed and spontaneous expressions. Specifically, we consider a facial expression as a complex activity that consists of temporally overlapping or sequential primitive facial events, which are defined as the motion of feature points. We propose using the Allen's Interval Algebra to represent the complex temporal patterns existing in facial events through a two-layer Bayesian network. Furthermore, we propose employing multi-value restricted Boltzmann machine to capture intrinsic global spatial patterns among facial events. Experimental results on three benchmark databases, the UvA-NEMO smile database, the DISFA+ database, and theSPOS database, demonstrate the proposed interval temporal restricted Boltzmann machine can successfully capture the intrinsic spatial-temporal patterns in facial behavior, and thus outperform state-of-the art work of posed and spontaneous expressions distinction. Shangfei Wang |
ACM Multimedia | 2 |
| 2017 | Expression-assisted facial action unit recognition under incomplete AU annotation
Shangfei Wang |
Pattern Recognit. | 1 |
| 2017 | Feature and label relation modeling for multiple-facial action unit classification and intensity estimation
Shangfei Wang, Zhen Gao 0003 |
Pattern Recognit. | 1 |
| 2016 | Facial Expression Intensity Estimation Using Ordinal InformationabstractPrevious studies on facial expression analysis have been focused on recognizing basic expression categories. There is limited amount of work on the continuous expression intensity estimation, which is important for detecting and tracking emotion change. Part of the reason is the lack of labeled data with annotated expression intensity since expression intensity annotation requires expertise and is time consuming. In this work, we treat the expression intensity estimation as a regression problem. By taking advantage of the natural onset-apex-offset evolution pattern of facial expression, the proposed method can handle different amounts of annotations to perform frame-level expression intensity estimation. In fully supervised case, all the frames are provided with intensity annotations. In weakly supervised case, only the annotations of selected key frames are used. While in unsupervised case, expression intensity can be estimated without any annotations. An efficient optimization algorithm based on Alternating Direction Method of Multipliers (ADMM) is developed for solving the optimization problem associated with parameter learning. We demonstrate the effectiveness of proposed method by comparing it against both fully supervised and unsupervised approaches on benchmark facial expression datasets. Rui Zhao 0015, Shangfei Wang |
CVPR | 3 |
| 2016 | Emotion recognition from peripheral physiological signals enhanced by EEGabstractCurrent multi-modal emotion recognition from physiological signals requires electroencephalogram(EEG) signals and peripheral physiological signals during both training and test. Compared with the peripheral physiological signals, it is more difficult to obtain EEG signals in our daily life. Therefore, we propose a novel approach to recognize emotions from peripheral signals by using EEG features as privileged information, which is only available during training. During training, first, peripheral physiological features and EEG features are extracted. Then, we construct a new peripheral physiological feature space using canonical correlation analysis with the help of EEG features. Finally we train a support vector machine(SVM) to map the new peripheral physiological features to the emotion labels. During test, only peripheral physiological features are used to recognize emotions from the constructed peripheral physiological feature space with the trained SVM model. The experimental results on two benchmark databases show that our proposed approach using EEG features as privileged information outperforms the method which recognizes emotions merely from the peripheral physiological signals. Zhen Gao 0003, Shangfei Wang |
ICASSP | 3 |
| 2016 | Implicit hybrid video emotion tagging by integrating video content and users' multiple physiological responsesabstractThe intrinsic interactions among a video's emotion tag, its content, and a user's spontaneous response while consuming the video can be leveraged to improve video emotion tagging, but this capability has not been thoroughly exploited yet. In this paper, we propose an implicit hybrid video emotion tagging approach by integrating video content and users' multiple physiological responses, which are only required during training. Specifically, multiple physiological signals during training construct a better emotion tagging model from video content. We add similarity constraints on the classifier mapping functions during training to capture the relationships among different kinds of features. We modify the traditional support vector machine with these constraints to improve video emotion tagging. Efficient learning algorithms of the proposed model are also developed. Experiments on three benchmark databases demonstrate the effectiveness and superior performance of our proposed method for implicitly integrating multiple physiological responses to improve video emotion tagging. Shangfei Wang, Chongliang Wu, Zhen Gao 0003, Xiaoxiao Shi |
ICPR | 2 |
| 2016 | Multiple Facial Action Unit recognition by learning joint features and label relationsabstractAlthough both feature dependencies and label dependencies are crucial for facial action unit (AU) recognition, little work addresses them simultaneously till now. To address this limitation, we propose a 4-layer Restrict Boltzmann Machine (RBM) to simultaneously capture feature level and AU level dependencies to recognize multiple AUs. Specifically, the bottom two layers of the RBM model capture dependencies among image features, while the top two layers capture the high order dependencies among AU labels. An efficient learning algorithm is introduced to jointly learn all layers to leverage the interactions among different layers. Experiments on two benchmark databases demonstrate the effectiveness of the proposed approach in modelling complex AU relationships from both features and labels jointly, and its improved performance over the existing methods. Shangfei Wang |
ICPR | 2 |
| 2016 | Employing subjects' information as privileged information for emotion recognition from EEG signalsabstractCurrent research of emotion recognition from electroencephalogram (EEG) signals rarely considers common patterns embodied in multiple subjects and individual patterns for each subject simultaneously. Therefore, in this paper, we propose a novel emotion recognition approach using subjects or subject groups as privileged information, which is only available during training. First, five frequency features are extracted from each channel of the EEG signals, and features are selected by statistical tests. Then, we propose two three-node Bayesian networks to capture the joint probability distribution function of emotion labels, EEG features, and subjects or subject groups during training. Through the learned joint probability distribution, the Bayesian networks model both common and individual emotion patterns simultaneously. During testing, emotion labels can be estimated from EEG features only by marginalized over the privileged information, i.e. subjects or subject groups. Experimental results on three benchmark databases, i.e. the MAHNOB-HCI database, the DEAP database and the USTC-ERVS database, demonstrate that our approach incorporating subjects and clusters achieves better emotion recognition performance than training a classifier for each subject, as well as training a classifier without subject information on the whole dataset. Shangfei Wang, Yachen Zhu, Zhen Gao 0003, Lihua Yue |
ICPR | 2 |
| 2016 | Multiple facial action unit recognition enhanced by facial expressionsabstractFacial expressions and facial action units (AU) respectively describe facial behavior globally and locally. Therefore, the dependencies between expressions and AUs carry crucial information for facial action unit recognition, yet have not been thoroughly exploited. In this paper, we propose a novel facial action unit recognition method enhanced by facial expressions, which are only required during training. Specifically, we propose a three-layer restricted Boltzmann machine (RBM) to capture the probabilistic dependencies among expressions and AUs. The parameters of the RBM model are learned by maximizing the log conditional likelihood with gradient ascent. After that, the learned RBM model combines AU measurements with the AU-expression relations it captures to perform multiple AU recognition through probabilistic inference. Experimental results on three benchmark databases, i.e. the CK+ database, the ISL database and the BP4D database, demonstrate the effectiveness of our method on capturing the joint relations among AUs and expression to improve AU recognition. Shangfei Wang |
ICPR | 3 |
| 2016 | Emotion Recognition from EEG Signals Enhanced by User's ProfileabstractThe main stream of current research of emotion recognition from electroencephalogram (EEG) signals infers users' emotion directly from EEG signals, without considering users' profile, which carries crucial information for emotion recognition. To address this, we propose a novel approach of emotion recognition from EEG signals enhanced by users' profile, which is only required during training. Specifically, we propose to use discriminative Restricted Boltzmann Machine (DRBM) to capture the inherent relationships among users' profile, EEG signals and emotion labels. During training, the proposed RBM classifier learns the mapping function from EEG signals and users' profile to emotion labels through discriminative learning algorithm. During testing, users' emotion can be recognized from EEG signals through marginalizing over users' profile, since RBM is a generative model. Experimental results on three benchmark databases, i.e. the DEAP database, the MAHNOB-HCI database and the USTC-ERVS database, demonstrate the effectiveness of the proposed approach in leveraging users' profile for enhancing emotion recognition from EEG signals and the superior emotion recognition performance of the proposed approach over existing approaches. Tanfang Chen, Shangfei Wang, Zhen Gao 0003, Chongliang Wu |
ICMR | 2 |
| 2016 | Facial Expression Recognition with Deep two-view Support Vector MachineabstractThis paper proposes a novel deep two-view approach to learn features from both visible and thermal images and leverage the commonality among visible and thermal images for facial expression recognition from visible images. The thermal images are used as privileged information, which is required only during training to help visible images learn better features and classifier. Specifically, we first learn a deep model for visible images and thermal images respectively, and use the learned feature representations to train SVM classifiers for expression classification. We then jointly refine the deep models as well as the SVM classifiers for both thermal images and visible images by imposing the constraint that the outputs of the SVM classifiers from two views are similar. Therefore, the resulting representations and classifiers capture the inherent connections among visible facial image, infrared facial image and target expression labels, and hence improve the recognition performance for facial expression recognition from visible images during testing. Experimental results on the benchmark expression database demonstrate the effectiveness of our proposed method. Chongliang Wu, Shangfei Wang, Bowen Pan, Huaping Chen 0001 |
ACM Multimedia | 2 |
| 2016 | Posed and Spontaneous Expression Recognition Through Restricted Boltzmann Machine
Chongliang Wu, Shangfei Wang |
MMM (1) | 2 |
| 2016 | Capturing global spatial patterns for distinguishing posed and spontaneous expressions
Shangfei Wang, Chongliang Wu |
Comput. Vis. Image Underst. | 1 |
| 2016 | Gender recognition from visible and thermal infrared facial images
Shangfei Wang, Zhen Gao 0003, Menghua He |
Multim. Tools Appl. | 1 |
| 2016 | Facial expression recognition through modeling age-related spatial patterns
Shangfei Wang, Zhen Gao 0003 |
Multim. Tools Appl. | 1 |
| 2015 | Posed and spontaneous facial expression differentiation using deep Boltzmann machinesabstractCurrent works on differentiating between posed and spontaneous facial expressions usually use features that are handcrafted for expression category recognition. Till now, no features have been specifically designed for differentiating between posed and spontaneous facial expressions. Recently, deep learning models have been proven to be efficient for many challenging computer vision tasks, and therefore in this paper we propose using the deep Boltzmann machine to learn representations of facial images and to differentiate between posed and spontaneous facial expressions. First, faces are located from images. Then, a two-layer deep Boltzmann machine is trained to distinguish posed and spon-tanous expressions. Experimental results on two benchmark datasets, i.e. the SPOS and USTC-NVIE datasets, demonstrate that the deep Boltzmann machine performs well on posed and spontaneous expression differentiation tasks. Comparison results on both datasets show that our method has an advantage over the other methods. Chongliang Wu, Shangfei Wang |
ACII | 3 |
| 2015 | Emotion Recognition from EEG Signals using Hierarchical Bayesian Network with Privileged InformationabstractCurrent work of emotion recognition from electroencephalogram (EEG) signals mainly focuses on the generality among users, ignoring users' specificity. However, users' emotion is a subjective phenomenon with both common and specific characteristics. Therefore, we propose a novel emotion recognition method using hierarchical Bayesian network to handle generality and specificity of emotions simultaneously. Specifically, by modeling the prior distributions of parameters for each subject, the classifier for a subject is learned together with those of others, with a shared representation. In addition, by marginalizing over the node of subjects, the subject information is used as privileged information, which is only required during training to build a better classifier. Experimental results on the MAHNOB-HCI and DEAP databases demonstrate that our model with the subject id as privileged information can improve the emotion recognition performance. Zhen Gao 0003, Shangfei Wang |
ICMR | 2 |
| 2015 | Multiple Aesthetic Attribute Assessment by Exploiting Relations Among Aesthetic AttributesabstractCurrent research of aesthetic assessment for images assumes one aesthetic score or one aesthetic label for an image, ignoring the relations of multiple aesthetic-related attributes. However, most images can be described by multiple aesthetic attributes simultaneously. Therefore, in this paper, we propose multiple aesthetic attribute prediction and classification by modeling relations among aesthetic attributes through Bayesian Networks (BN). In order to realize continuous aesthetic attribute prediction, each aesthetic attribute is represented by a three-node BN, including the discretized aesthetic attribute label, the predicted aesthetic attribute score, and the measurement of the aesthetic attribute score. In addition, the relations among multiple aesthetic attributes are modeled by another discrete BN, whose structure and conditional probabilities are learned from the training data. The attribute measurements are obtained by an existing image-driven regression method. With the learned BN, we infer the true discrete label and continuous score for each attribute by combining the relations among attributes with the previously obtained measurements. Experiments on the Memorability dataset show the superiority of our proposed approach to current image-driven methods for both multiple continuous aesthetic attribute score prediction and multiple discrete aesthetic attribute label classification, indicating the effectiveness of the captured relations for aesthetic quality assessment. Zhen Gao 0003, Shangfei Wang |
ICMR | 2 |
| 2015 | Expression Recognition from Visible Images with the Help of Thermal ImagesabstractMost facial expression recognition research focused on visible spectrum, which is sensitive to illumination changes. While thermal images, recording facial temperature distribution, are robust to light conditions. Therefore, expression recognition by visible and thermal image fusion is promising. However, in most cases, only visible images are available, since thermal cameras are much more expensive than visible cameras, which are popular in our daily life. Thus, in this paper, we propose a novel visible expression recognition approach by using thermal infrared data as privileged information, which is only available during training. First, active appearance model parameters and three defined head motion features are extracted from visible spectrum images, and several thermal statistical features are extracted from thermal infrared images. Second, feature selection is performed using the F-test statistic. Third, a new visible feature space is constructed using canonical correlation analysis under the help of thermal infrared images. After that, a support vector machine is adopted as the classifier on the constructed visible feature space. Experiments on the NVIE and Equinox database show the effectiveness of the proposed methods, and demonstrate that thermal infrared images' supplementary role for visible facial expression recognition. Xiaoxiao Shi, Shangfei Wang, Yachen Zhu |
ICMR | 2 |
| 2015 | Facial Action Unit Classification with Hidden Knowledge under Incomplete AnnotationabstractFacial action unit (AU) recognition is an important task for facial expression analysis. Traditional AU recognition methods typically include a supervised training, where the AU annotated training images are needed. AU annotation is a time consuming, expensive, and error prone process. While AU is hard to annotate, facial expression is relatively easy to label. To take advantage of this, we introduce a new learning method that trains an AU classifier using images with incomplete AU annotation but with complete expression labels. The goal is to use expression labels as hidden knowledge to complement the missing AU labels. Towards this goal, we propose to construct a Bayesian Network (BN) to capture the relationships between facial expression and AUs. Structural Expectation Maximum is used to learn the structure and parameters of the BN when the AU labels are missing. Given the learned BNs and measurements of AUs and expression, we can then perform AU recognition within the BN through a probabilistic inference. Experimental results on the CK+ and ISL databases demonstrate the effectiveness of our method. Jun Wang 0130, Shangfei Wang |
ICMR | 2 |
| 2015 | Learning with privileged information using Bayesian networks
Shangfei Wang, Menghua He, Yachen Zhu |
Frontiers Comput. Sci. | 1 |
| 2015 | Implicit video emotion tagging from audiences' facial expression
Shangfei Wang, Zhilei Liu, Yachen Zhu, Menghua He |
Multim. Tools Appl. | 1 |
| 2015 | Multiple emotional tagging of multimedia data by exploiting dependencies among emotions
Shangfei Wang, Zhaoyu Wang 0003 |
Multim. Tools Appl. | 1 |
| 2015 | Posed and spontaneous expression recognition through modeling their spatial patterns
Shangfei Wang, Chongliang Wu, Menghua He, Jun Wang 0130 |
Mach. Vis. Appl. | 1 |
| 2015 | Video Affective Content Analysis: A Survey of State-of-the-Art MethodsabstractVideo affective content analysis has been an active research area in recent decades, since emotion is an important component in the classification and retrieval of videos. Video affective content analysis can be divided into two approaches: direct and implicit. Direct approaches infer the affective content of videos directly from related audiovisual features. Implicit approaches, on the other hand, detect affective content from videos based on an automatic analysis of a user's spontaneous response while consuming the videos. This paper first proposes a general framework for video affective content analysis, which includes video content, emotional descriptors, and users' spontaneous nonverbal responses, as well as the relationships between the three. Then, we survey current research in both direct and implicit video affective content analysis, with a focus on direct video affective content analysis. Lastly, we identify several challenges in this field and put forward recommendations for future research. Shangfei Wang |
IEEE Trans. Affect. Comput. | 1 |
| 2015 | Multiple Emotion Tagging for Multimedia Data by Exploiting High-Order Dependencies Among EmotionsabstractIn this paper, a novel approach of multiple emotional multimedia tagging is proposed, which explicitly models the higher-order relations among emotions. First, multimedia features are extracted from the multimedia data. Second, a traditional multi-label classifier is used to obtain the measurements of the multi-emotion labels. Then, we propose a three-layer restricted Boltzmann machine (TRBM) model to capture the higher-order relations among emotion labels, as well as the relations between labels and measurements . Finally , the TRBM model is used to infer the samples' multi- emotion labels by combining the emotion measurements with the dependencies among multi- emotions . Experimental results on four databases demonstrate that our method is more effective than both feature -driven methods and current model-based methods, which capture the pairwise relations among labels by the Bayesian network (BN). Furthermore , the comparison of BN models and the proposed TRBM model verifies that the patterns captured by the latent units of TRBM contain not only all the dependencies captured by the BN but also many other dependencies that the BN cannot capture. Shangfei Wang, Jun Wang 0130, Ziheng Wang 0001 |
IEEE Trans. Multim. | 1 |
| 2014 | Emotion recognition from users' EEG signals with the help of stimulus VIDEOSabstractIn this paper, we propose a novel approach to recognize users' emotions from electroencephalogram (EEG) signals by using stimulus videos as privileged information, which is only available during training. Firstly, five frequency features are extracted from each channel of the EEG signals, and several audio/visual features are extracted from video stimulus. Secondly, features are selected by statistical analyses. Then, a new EEG feature space is constructed using Canonical Correlation Analysis under the help of video content. Finally, a support vector machine is adopted as the classifier on the constructed EEG feature space. Experimental results on two benchmark databases demonstrate that video content, as the context, can improve the emotion recognition performance when employed as privileged information. Yachen Zhu, Shangfei Wang |
ICME | 2 |
| 2014 | Early Facial Expression Recognition Using Hidden Markov ModelsabstractAlthough it is often necessary to recognize users' expressions as soon as possible after it starts and before it ends in many applications, few methods have been proposed explicitly for early facial expression recognition. In this paper, we propose an early facial expression recognition method by using Hidden Markov Model. The relative displacement of the feature points between the current frame and the neutral frame are extracted as the facial features. During training, an iterative algorithm is introduced to find a classification entropy threshold and model parameters of early HMM. During testing, an image sequence is assigned an expression category when the entropy of the expression likelihood obtained from early HMMs is below the threshold by gradually increasing sequence length. Experimental results on CK+ and MMI databases show the effectiveness of our approach. Jun Wang 0130, Shangfei Wang |
ICPR | 2 |
| 2014 | Multi-label Learning with Missing LabelsabstractIn multi-label learning, each sample can be assigned to multiple class labels simultaneously. In this work, we focus on the problem of multi-label learning with missing labels (MLML), where instead of assuming a complete label assignment is provided for each sample, only partial labels are assigned with values, while the rest are missing or not provided. The positive (presence), negative (absence) and missing labels are explicitly distinguished in MLML. We formulate MLML as a transductive learning problem, where the goal is to recover the full label assignment for each sample by enforcing consistency with available label assignments and smoothness of label assignments. Along with an exact solution, we also provide an effective and efficient approximated solution. Our method shows much better performance than several state-of-the-art methods on several benchmark data sets. Baoyuan Wu, Zhilei Liu, Shangfei Wang, Bao-Gang Hu |
ICPR | 3 |
| 2014 | Multiple-Facial Action Unit Recognition by Shared Feature Learning and Semantic Relation ModelingabstractIn this paper, we propose multiple facial action unit recognition by modeling their relations from both features and target labels. First, a multi-task feature learning method is adopted to divide action unit recognition tasks into several groups, and then learn the shared features for each group. Second, a Bayesian network is used to model the co-existent and mutual-exclusive semantic relations among action units from the target labels of facial images. After that, the learned Bayesian network employs the recognition results of the multi-task learning, and realizes multiple facial action recognition by probabilistic inference. Experiments on the extended Cohn-Kanade database and the Denver Intensity of Spontaneous Facial Actions database demonstrate the effectiveness of our approach. Yachen Zhu, Shangfei Wang, Lihua Yue |
ICPR | 2 |
| 2014 | Emotion recognition from thermal infrared images using deep Boltzmann machine
Shangfei Wang, Menghua He, Zhen Gao 0003 |
Frontiers Comput. Sci. | 1 |
| 2014 | Fusion of visible and thermal images for facial expression recognition
Shangfei Wang, Yue Wu 0002, Menghua He |
Frontiers Comput. Sci. | 1 |
| 2014 | Exploiting multi-expression dependences for implicit multi-emotion video tagging
Shangfei Wang, Zhilei Liu, Jun Wang 0130, Zhaoyu Wang 0003 |
Image Vis. Comput. | 1 |
| 2014 | Hybrid video emotional tagging using users' EEG and video content
Shangfei Wang, Yachen Zhu, Guobing Wu |
Multim. Tools Appl. | 1 |
| 2014 | Enhancing multi-label classification by modeling dependencies among labels
Shangfei Wang, Jun Wang 0130, Zhaoyu Wang 0003 |
Pattern Recognit. | 1 |
| 2013 | Active Labeling of Facial Feature PointsabstractAlthough considerable progress has been made in the field of facial feature point detection and tracking, accurate feature point tracking is still very challenging. Manually feature point labeling and correction are time consuming and labor intensive. To alleviate this problem, an active feature point labeling method is proposed in this paper. First, the spatial relations among feature points are modeled by a Bayesian Network. Second, the mutual information between a feature point and the remaining feature points is calculated in two steps: in the first step, to identify the most informative facial region, the mutual information between one facial sub-region and the other sub-regions is calculated, in the second step, the mutual information between one feature point and the other feature points in the most informative facial sub-region is established to rank the facial feature points. Users provide labels of the feature points according to their mutual information in descending order. After that, the human corrections and the image measurements are integrated by the Bayesian Network to produce the refined annotations. Simulative experiments on the extended Cohn-Kanade (CK+) database demonstrate the effectiveness of our approach. Menghua He, Shangfei Wang |
ACII | 2 |
| 2013 | Analyses of the Differences between Posed and Spontaneous Facial ExpressionsabstractThis paper presents comprehensive analyses of the differences between posed and spontaneous expressions from visible images. First, geometric and appearance features are extracted from the difference images between apex and onset facial images. Secondly, the differences between the posed and spontaneous facial expressions are analyzed through hypothetical testing methods from three aspects: on overall samples, on samples with different genders, and on samples with different expressions. Thirdly, Bayesian networks (BNs) are used to classify posed versus spontaneous expressions from the same three aspects. Statistical analyses on the NVIE database demonstrate the importance of the geometric and appearance features for discriminating posed and spontaneous expressions. Gender effect exists on the differences between posed and spontaneous expressions. It is easier to distinguish posed happiness from spontaneous happiness than other expressions. Recognition experimental results confirm the observations of statistical analyses in most cases. Menghua He, Shangfei Wang, Zhilei Liu |
ACII | 2 |
| 2013 | Facial Expression Recognition Using Deep Boltzmann Machine from Thermal Infrared ImagesabstractFacial expression recognition from thermal infrared images has attracted more and more attentions in recent years. However, the features adopted in current work are either temperature statistical parameters extracted from the facial regions of interest or several hand-crafted features that are commonly used in visible spectrum. Till now there is no image features specially defined for thermal infrared images. In this paper, we are the first to propose using the Deep Boltzmann Machine to learn thermal features for expression recognition from thermal long wavelength infrared images. First, the face are located and normalized from the thermal infrared images. Then, a Deep Boltzmann Machine model composed of two layers is proposed. The parameters of the Deep Boltzmann Machine model are further fine-tuned for facial expression recognition after pre-training of feature learning. Comparison experimental results on the NVIE database demonstrate that our approach outperforms other approaches using temperature statistic features or hand-crafted features borrowed from visible domain. The learned features from the forehead, mouth, and cheek are more reliable for discriminating disgust, fear, and happiness compared with other facial areas. Shangfei Wang, Wuwei Lan, Huan Fu |
ACII | 2 |
| 2013 | Emotional Influence on SSVEP Based BCIabstractThe objective of the paper is to investigate the effect of subject's emotional states on Brain Computer Interface (BCI) performance. Two psycho-physiological experiments are designed and implemented. The first one induces subjects' emotion using video clips first, then involves subjects' in SSVEP task. The second one induces subjects' emotions and SSVEP simultaneously by flickering IAPS pictures in four directions. used to recognize the performed BCI tasks. Based on the performances of learned classifiers, we analyzed the influence of emotion using two statistical tests. The McNamara's test serves to assess if emotion has any influences on mental task performing while Wilcox on signed-rank test analyses if emotion has a positive or detrimental effect on ability to achieve a SSVEP task. Obtained results suggest influence of emotional states: the positive and neutral emotions influence BCI performance similarly, while the negative emotion tends to deteriorate classification accuracy. Yachen Zhu, Xilan Tian, Guobing Wu, Gilles Gasso, Shangfei Wang, Stéphane Canu |
ACII | 5 |
| 2013 | Capturing Complex Spatio-temporal Relations among Facial Muscles for Facial Expression RecognitionabstractSpatial-temporal relations among facial muscles carry crucial information about facial expressions yet have not been thoroughly exploited. One contributing factor for this is the limited ability of the current dynamic models in capturing complex spatial and temporal relations. Existing dynamic models can only capture simple local temporal relations among sequential events, or lack the ability for incorporating uncertainties. To overcome these limitations and take full advantage of the spatio-temporal information, we propose to model the facial expression as a complex activity that consists of temporally overlapping or sequential primitive facial events. We further propose the Interval Temporal Bayesian Network to capture these complex temporal relations among primitive facial events for facial expression modeling and recognition. Experimental results on benchmark databases demonstrate the feasibility of the proposed approach in recognizing facial expressions based purely on spatio-temporal relations among facial muscles, as well as its advantage over the existing methods. Ziheng Wang 0001, Shangfei Wang |
CVPR | 2 |
| 2013 | Capturing Global Semantic Relationships for Facial Action Unit RecognitionabstractIn this paper we tackle the problem of facial action unit (AU) recognition by exploiting the complex semantic relationships among AUs, which carry crucial top-down information yet have not been thoroughly exploited. Towards this goal, we build a hierarchical model that combines the bottom-level image features and the top-level AU relationships to jointly recognize AUs in a principled manner. The proposed model has two major advantages over existing methods. 1) Unlike methods that can only capture local pair-wise AU dependencies, our model is developed upon the restricted Boltzmann machine and therefore can exploit the global relationships among AUs. 2) Although AU relationships are influenced by many related factors such as facial expressions, these factors are generally ignored by the current methods. Our model, however, can successfully capture them to more accurately characterize the AU relationships. Efficient learning and inference algorithms of the proposed model are also developed. Experimental results on benchmark databases demonstrate the effectiveness of the proposed approach in modelling complex AU relationships as well as its superior AU recognition performance over existing approaches. Ziheng Wang 0001, Shangfei Wang |
ICCV | 3 |
| 2013 | Eye localization from thermal infrared images
Shangfei Wang, Zhilei Liu, Peijia Shen |
Pattern Recognit. | 1 |
| 2013 | Analyses of a Multimodal Spontaneous Facial Expression DatabaseabstractCreating a large and natural facial expression database is a prerequisite for facial expression analysis and classification. It is, however, not only time consuming but also difficult to capture an adequately large number of spontaneous facial expression images and their meanings because no standard, uniform, and exact measurements are available for database collection and annotation. Thus, comprehensive first-hand data analyses of a spontaneous expression database may provide insight for future research on database construction, expression recognition, and emotion inference. This paper presents our analyses of a multimodal spontaneous facial expression database of natural visible and infrared facial expressions (NVIE). First, the effectiveness of emotion-eliciting videos in the database collection is analyzed with the mean and variance of the subjects' self-reported data. Second, an interrater reliability analysis of raters' subjective evaluations for apex expression images and sequences is conducted using Kappa and Kendall's coefficients. Third, we propose a matching rate matrix to explore the agreements between displayed spontaneous expressions and felt affective states. Lastly, the thermal differences between the posed and spontaneous facial expressions are analyzed using a paired-samples t-test. The results of these analyses demonstrate the effectiveness of our emotion-inducing experimental design, the gender difference in emotional responses, and the coexistence of multiple emotions/expressions. Facial image sequences are more informative than apex images for both expression and emotion recognition. Labeling an expression image or sequence with multiple categories together with their intensities could be a better approach than labeling the expression image or sequence with one dominant category. The results also demonstrate both the importance of facial expressions as a means of communication to convey affective states and the diversity of the displayed manifestations of felt emotions. There are indeed some significant differences between the temperature difference data of most posed and spontaneous facial expressions, many of which are found in the forehead and cheek regions. Shangfei Wang, Zhilei Liu, Zhaoyu Wang 0003, Guobing Wu, Peijia Shen, Xufa Wang |
IEEE Trans. Affect. Comput. | 1 |
| 2013 | Simultaneous Facial Feature Tracking and Facial Expression RecognitionabstractThe tracking and recognition of facial activities from images or videos have attracted great attention in computer vision field. Facial activities are characterized by three levels. First, in the bottom level, facial feature points around each facial component, i.e., eyebrow, mouth, etc., capture the detailed face shape information. Second, in the middle level, facial action units, defined in the facial action coding system, represent the contraction of a specific set of facial muscles, i.e., lid tightener, eyebrow raiser, etc. Finally, in the top level, six prototypical facial expressions represent the global facial muscle movement and are commonly used to describe the human emotion states. In contrast to the mainstream approaches, which usually only focus on one or two levels of facial activities, and track (or recognize) them separately, this paper introduces a unified probabilistic framework based on the dynamic Bayesian network to simultaneously and coherently represent the facial evolvement in different levels, their interactions and their observations. Advanced machine learning methods are introduced to learn the model based on both training data and subjective prior knowledge. Given the model and the measurements of facial motions, all three levels of facial activities are simultaneously recognized through a probabilistic inference. Extensive experiments are performed to illustrate the feasibility and effectiveness of the proposed model on all three level facial activities. Shangfei Wang, Yongping Zhao |
IEEE Trans. Image Process. | 2 |
| 2012 | Posed and spontaneous expression distinguishment from infrared thermal images
Zhilei Liu, Shangfei Wang |
ICPR | 2 |
| 2012 | Bias analyses of spontaneous facial expression database
Zhaoyu Wang 0003, Shangfei Wang, Yachen Zhu |
ICPR | 2 |
| 2012 | Similarity Measurement and Feature Selection Using Genetic Algorithm
Shangfei Wang |
ISNN (2) | 1 |
| 2012 | A qualitative and quantitative study of color emotion using valence-arousal
Shangfei Wang |
Frontiers Comput. Sci. | 1 |
| 2011 | Emotion Recognition Using Hidden Markov Models from Facial Temperature Sequence
Zhilei Liu, Shangfei Wang |
ACII (2) | 2 |
| 2011 | Spontaneous Facial Expression Recognition Based on Feature Point TrackingabstractIn recent years, facial expression recognition has attracted a lot of attention because of its importance in human-computer interaction. However, most previous work has focus on posed expression. In this paper, we propose a spontaneous facial expression recognition method based on feature point tracking. First, all expression sequences are normalized according to their pupils' coordinates. Second, 23 points are labeled manually in the onset and apex frames. Then Kalman filter is used for tracking. Two kinds of features, point displacement features and points distance variation features, are extracted. Finally, Hidden Markov Model is employed as classifier. Experiments have been conducted on the spontaneous facial expression database of USTC-NVIE. The results indicate that Kalman filter point tracking method could detect the right place of points, and the points distance variation features are more suitable than the point displacement features for the facial expression recognition. Shangfei Wang, Yanpeng Lv |
ICIG | 2 |
| 2011 | Affective Classification in Video Based on Semi-supervised Learning
Shangfei Wang, Yongjie Hu |
ISNN (3) | 1 |
| 2010 | Musical perceptual similarity estimation using interactive genetic algorithmabstractThis paper proposes a new approach to estimate the emotional perceptual similarity of music using interactive genetic algorithm. Different combinations of measure function and feature weights construct the searching space, and users' subjective similarity evaluations of musical pieces are used as the fitness. The approach tries to search for the optimal combination of measure function and feature weights to better reflect human's perception. A comparative emotion detection experiment is designed to explore the effectiveness of our approach in our MIDI database. The one is the simple K-Nearest Neighbor (KNN) and the other is the modified KNN using the optimal combination obtained. Experimental results show that our method outperforms the simple KNN classifier, which confirms its usefulness of better reflecting human's perception. Shangfei Wang |
IEEE Congress on Evolutionary Computation | 1 |
| 2010 | An emotional harmony generation systemabstractThis paper proposes an emotional harmony composition system using a series of predefined rules and interactive genetic algorithm. Two kinds of rules are defined: One is harmony composition rules used for composing harmony and the other is harmony emotion rules which may indicate the mapping between harmonic elements and a specific emotion of either happiness or sadness. Interactive genetic algorithm is applied to incorporate users' subjective perception in the generated harmony. Both subjective and objective experiments are conducted to prove the effectiveness of our approach. Shangfei Wang |
IEEE Congress on Evolutionary Computation | 2 |
| 2010 | Infrared Face Recognition Based on Histogram and K-Nearest Neighbor Classification
Shangfei Wang, Zhilei Liu |
ISNN (2) | 1 |
| 2010 | A Natural Visible and Infrared Facial Expression Database for Expression Recognition and Emotion InferenceabstractTo date, most facial expression analysis has been based on visible and posed expression databases. Visible images, however, are easily affected by illumination variations, while posed expressions differ in appearance and timing from natural ones. In this paper, we propose and establish a natural visible and infrared facial expression database, which contains both spontaneous and posed expressions of more than 100 subjects, recorded simultaneously by a visible and an infrared thermal camera, with illumination provided from three different directions. The posed database includes the apex expressional images with and without glasses. As an elementary assessment of the usability of our spontaneous database for expression recognition and emotion inference, we conduct visible facial expression recognition using four typical methods, including the eigenface approach [principle component analysis (PCA)], the fisherface approach [PCA + linear discriminant analysis (LDA)], the Active Appearance Model (AAM), and the AAM-based + LDA. We also use PCA and PCA+LDA to recognize expressions from infrared thermal images. In addition, we analyze the relationship between facial temperature and emotion through statistical analysis. Our database is available for research purposes. Shangfei Wang, Zhilei Liu, Siliang Lv, Yanpeng Lv, Guobing Wu, Peng Peng 0003, Fei Chen 0013, Xufa Wang |
IEEE Trans. Multim. | 1 |
| 2006 | User Fatigue Reduction by an Absolute Rating Data-trained Predictor in IECabstractPredicting IEC users’ evaluation characteristics is one way of reducing users’ fatigue. However, users’ relative evaluation appears as noise to the algorithm which learns and predicts the users’ evaluation characteristics. This paper introduces the idea of absolute scale to improve the performance of predicting users’ subjective evaluation characteristics in IEC, and thus it will accelerate EC convergence and reduce users’ fatigue. We first evaluate the effectiveness of the proposed method using seven benchmark functions instead of a human user. The experimental results show that the convergence speed of an IEC using the proposed absolute rating data-trained predictor is much faster than that of an IEC using a conventional predictor training with relative rating data. Next, the proposed algorithm is used in an individual emotion fashion image retrieval system. Experimental results of sign tests demonstrate that the proposed algorithm can alleviate user fatigue and has a good performance in individual emotional image retrieval. Shangfei Wang, Xufa Wang, Hideyuki Takagi |
IEEE Congress on Evolutionary Computation | 1 |
| 2005 | Emotion Semantics Image Retrieval: An Brief Overview
Shangfei Wang, Xufa Wang |
ACII | 1 |
| 2005 | Case-Based Facial Action Units Recognition Using Interactive Genetic Algorithm
Shangfei Wang, Jia Xue |
ACII | 1 |