Huiyuan Yang

dblp:150/4087 · DBLP profile ↗
← Back
27ranked-venue papers
13as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 7 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 7 first-author · 5 since 2021Computer networks · 8 · 4 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 You Only Need One Stage: Novel-View Synthesis from a Single Blind Face Image
abstract
We propose a novel one-stage method, NVB-Face, for generating consistent Novel-View images directly from a single Blind Face image. Existing approaches to novel-view synthesis for objects or faces typically require a high-resolution RGB image as input. When dealing with degraded images, the conventional pipeline follows a two-stage process: first restoring the image to high resolution, then synthesizing novel views from the restored result. However, this approach is highly dependent on the quality of the restored image, often leading to inaccuracies and inconsistencies in the final output. To address this limitation, we extract single-view features directly from the blind face image and introduce a feature manipulator that transforms these features into 3D-aware, multi-view latent representations. Leveraging the powerful generative capacity of a diffusion model, our framework synthesizes high-quality, consistent novel-view face images. Experimental results show that our method significantly outperforms traditional two-stage approaches in both consistency and fidelity.
Taoyue Wang, Xiang Zhang 0030, Huiyuan Yang, Lijun Yin 0001
AAAI4
2025 Indirect Lossy Source Coding With Observed Source Reconstruction: Nonasymptotic Bounds and Second-Order Asymptotics
abstract
This paper considers the joint compression of a pair of correlated sources, where the encoder is allowed to access only one of the sources. The objective is to recover both sources under separate distortion constraints for each source while minimizing the rate. This problem generalizes the indirect lossy source coding problem by also requiring the recovery of the observed source. In this paper, we aim to study the nonasymptotic and second-order asymptotic properties of this problem. Specifically, we begin by deriving nonasymptotic achievability and converse bounds valid for general sources and distortion measures. The source dispersion (Gaussian approximation) is then determined through asymptotic analysis of the nonasymptotic bounds. We further examine the case of erased fair coin flips (EFCF) and provide its specific nonasymptotic achievability and converse bounds. Numerical results under the EFCF case demonstrate that our second-order asymptotic approximation closely approximates the optimum rate at appropriately large blocklengths.
Huiyuan Yang, Yuxuan Shi 0001, Shuo Shao 0001, Xiaojun Yuan 0002
IEEE Trans. Commun.1
2025 UAV-Enabled Asynchronous Federated Learning
abstract
To exploit unprecedented data generation in mobile edge networks, federated learning (FL) has emerged as a promising alternative to the conventional centralized machine learning (ML). By collectively training a unified learning model on edge devices, FL bypasses the need of direct data transmission, thereby addressing problems such as latency issues and privacy concerns inherent in centralized ML. However, in practical deployment FL suffers from low learning efficiency due to the involved straggler issue and huge uplink overhead. In this paper, we develop a UAV-enabled over-the-air asynchronous FL (UAV-AFL) framework to address this problem. This framework significantly enhance the learning efficiency by supporting the UAV as the parameter server (UAV-PS) in collecting data over-the-air and updating model continuously. We conduct a convergence analysis to quantitatively capture the impact of model asynchrony, device selection and communication errors on the UAV-AFL learning efficiency. Based on this analysis, a unified communication-learning problem is formulated to maximize asymptotical learning accuracy by optimizing the UAV-PS trajectory, device selection and over-the-air transceiver design. Simulation results reveal valuable insights for the system design and demonstrate that the proposed UAV-AFL scheme achieves substantially improvement in learning efficiency compared with the state-of-the-art approaches.
Zhiyuan Zhai, Xiaojun Yuan 0002, Xin Wang 0003, Huiyuan Yang
IEEE Trans. Wirel. Commun.4
2024 Multimodal Channel-Mixing: Channel and Spatial Masked AutoEncoder on Facial Action Unit Detection
abstract
Recent studies have focused on utilizing multi-modal data to develop robust models for facial Action Unit (AU) detection. However, the heterogeneity of multi-modal data poses challenges in learning effective representations. One such challenge is extracting relevant features from multiple modalities using a single feature extractor. Moreover, previous studies have not fully explored the potential of multi-modal fusion strategies. In contrast to the extensive work on late fusion, there are limited investigations on early fusion for channel information exploration. This paper presents a novel multi-modal reconstruction network, named Multimodal Channel-Mixing (MCM), as a pre-trained model to learn robust representation for facilitating multi-modal fusion. The approach follows an early fusion setup, integrating a Channel-Mixing module, where two out of five channels are randomly dropped. The dropped channels then are reconstructed from the remaining channels using masked autoencoder. This module not only reduces channel redundancy, but also facilitates multi-modal learning and reconstruction capabilities, resulting in robust feature learning. The encoder is fine-tuned on a downstream task of automatic facial action unit detection. Pre-training experiments were conducted on BP4D+, followed by fine-tuning on BP4D and DISFA to assess the effectiveness and robustness of the proposed framework. The results demonstrate that our method meets and surpasses the performance of state-of-the-art baseline methods.
Xiang Zhang 0030, Huiyuan Yang, Taoyue Wang, Lijun Yin 0001
WACV2
2024 Disagreement Matters: Exploring Internal Diversification for Redundant Attention in Generic Facial Action Analysis
abstract
This paper demonstrates the effectiveness of a diversification mechanism for building a more robust multi-attention system in generic facial action analysis. While previous multi-attention (e.g., visual attention and self-attention) research on facial expression recognition (FER) and Action Unit (AU) detection have been thoroughly studied to focus on ”external attention diversification”, where attention branches localize different facial areas, we delve into the realm of ”internal attention diversification” and explore the impact of diverse attention patterns within the same Region of Interest (RoI). Our experiments reveal that variability in attention patterns significantly impacts model performance, indicating that unconstrained multi-attention plagued by redundancy and over-parameterization, leading to sub-optimal results. To tackle this issue, we propose a compact module that guides the model to achieve self-diversified multi-attention. Our method is applied to both CNN-based and Transformer-based models, benchmarked on popular databases such as BP4D and DISFA for AU detection, as well as CK+, MMI, BU-3DFE, and BP4D+ for facial expression recognition. We also evaluate the mechanism on Self-attention and Channel-wise attention designs for improving their adaptive capabilities in multi-modal feature fusion tasks. The multi-modal evaluation is conducted on BP4D, BP4D+, and our newly developed large-scale comprehensive emotion database BP4D++, which contains well-synchronized and aligned sensor modalities, addressing the scarcity of annotations and identities in human affective computing. We plan to release the new database to the research community, fostering further advancements in this field.
Zheng Zhang 0023, Xiang Zhang 0030, Taoyue Wang, Huiyuan Yang, Umur A. Ciftci, Jeffrey F. Cohn, Lijun Yin 0001
IEEE Trans. Affect. Comput.6
2023 Analysis and Optimization of Model Aggregation Performance in Federated Learning
abstract
Recently, federated learning (FL), which replaces data sharing with model sharing, has emerged as an efficient and privacy-friendly machine learning paradigm. One of the main challenges in FL is the huge communication cost for model aggregation. Many compression/quantization schemes have been proposed to reduce the communication cost of model aggregation, with remarkable results. However, the following two questions remain unanswered: What are the performance limits of model aggregation? How much can FL convergence performance be improved by modifying existing model aggregation schemes? In this paper, we manage to answer these two questions. Specifically, for the first question, we put forth a general model aggregation performance analysis framework based on the rate-distortion theory. We then derive an inner bound of the rate-distortion region of model aggregation. For the second question, we develop an algorithm to search for the minimum achievable aggregation distortion, which, combined with the achievability scheme of the derived inner bound, induces a method to approximate the limits of FL convergence performance numerically. Numerical results demonstrate that the baseline model aggregation schemes still have great potential for further improvement.
Huiyuan Yang, Tian Ding, Xiaojun Yuan 0002
ICC1
2023 Weakly-Supervised Text-driven Contrastive Learning for Facial Behavior Understanding
abstract
Contrastive learning has shown promising potential for learning robust representations by utilizing unlabeled data. However, constructing effective positive-negative pairs for contrastive learning on facial behavior datasets remains challenging. This is because such pairs inevitably encode the subject-ID information, and the randomly constructed pairs may push similar facial images away due to the limited number of subjects in facial behavior datasets. To address this issue, we propose to utilize activity descriptions, coarse-grained information provided in some datasets, which can provide high-level semantic information about the image sequences but is often neglected in previous studies. More specifically, we introduce a two-stage Contrastive Learning with Text-Embeded framework for Facial behavior understanding (CLEF). The first stage is a weakly-supervised contrastive learning method that learns representations from positive-negative pairs constructed using coarse-grained activity information. The second stage aims to train the recognition of facial expressions or facial action units by maximizing the similarity between the image and the corresponding text label names. The proposed CLEF achieves state-of-the-art performance on three in-the-lab datasets for AU recognition and three in-the-wild datasets for facial expression recognition.
Xiang Zhang 0030, Taoyue Wang, Huiyuan Yang, Lijun Yin 0001
ICCV4
2023 Federated Learning With Lossy Distributed Source Coding: Analysis and Optimization
abstract
Recently, federated learning (FL), which replaces data sharing with model sharing, has emerged as an efficient and privacy-friendly machine learning (ML) paradigm. One of the main challenges in FL is the huge communication cost for model aggregation. Many compression/quantization schemes have been proposed to reduce the communication cost for model aggregation. However, the following question remains unanswered: What is the fundamental trade-off between the communication cost and the FL convergence performance? In this paper, we manage to answer this question. Specifically, we first put forth a general framework for model aggregation performance analysis based on the rate-distortion theory. Under the proposed analysis framework, we derive an inner bound of the rate-distortion region of model aggregation. We then conduct an FL convergence analysis to connect the aggregation distortion and the FL convergence performance. We formulate an aggregation distortion minimization problem to improve the FL convergence performance. Two algorithms are developed to solve the above problem. Numerical results on aggregation distortion, convergence performance, and communication cost demonstrate that the baseline model aggregation schemes still have great potential for further improvement.
Huiyuan Yang, Tian Ding, Xiaojun Yuan 0002
IEEE Trans. Commun.1
2023 Over-the-Air Federated Multi-Task Learning Over MIMO Multiple Access Channels
abstract
With the explosive growth of data and wireless devices, federated learning (FL) over wireless medium has emerged as a promising technology for large-scale distributed intelligent systems. Yet, the urgent demand for ubiquitous intelligence will generate a large number of concurrent FL tasks, which may seriously aggravate the scarcity of communication resources. By exploiting the analog superposition of electromagnetic waves, over-the-air computation (AirComp) is an appealing solution to alleviate the burden of communication required by FL. However, sharing frequency-time resources in over-the-air computation inevitably brings about the problem of inter-task interference, which poses a new challenge that needs to be appropriately addressed. In this paper, we study over-the-air federated multi-task learning (OA-FMTL) over the multiple-input multiple-output (MIMO) multiple access (MAC) channel. We propose a novel model aggregation method for the alignment of local gradients of different devices, which alleviates the straggler problem in over-the-air computation due to the channel heterogeneity. We establish a communication-learning analysis framework for the proposed OA-FMTL scheme by considering the spatial correlation between devices, and formulate an optimization problem for the design of transceiver beamforming and device selection. To solve this problem, we develop an algorithm by using alternating optimization (AO) and fractional programming (FP), which effectively mitigates the impact of inter-task interference on the FL learning performance. We show that due to the use of the new model aggregation method, device selection is no longer essential, thereby avoiding the heavy computational burden involved in selecting active devices. Numerical results demonstrate the validity of the analysis and the superb performance of the proposed scheme.
Chenxi Zhong, Huiyuan Yang, Xiaojun Yuan 0002
IEEE Trans. Wirel. Commun.2
2022 UAV-Assisted Hierarchical Aggregation for Over-the-Air Federated Learning
abstract
With huge amounts of data explosively increasing on the mobile edge, over-the-air federated learning (OA-FL) emerges as a promising technique to reduce communication costs and privacy leak risks. However, when devices in a relatively large area cooperatively train a machine learning model, the attendant straggler issue will significantly reduce the learning performance. In this paper, we propose an unmanned aerial vehicle (UAV) assisted OA-FL system, where the UAV acts as a parameter server (PS) to aggregate the local gradients hierarchically for global model updating. Under this UAV-assisted hierarchical aggregation scheme, we carry out a gradient-correlation-aware FL performance analysis. We then formulate a mean squared error (MSE) minimization problem to tune the UAV trajectory and the global aggregation coefficients based on the analysis results. An algorithm based on alternating optimization (AO) and successive convex approximation (SCA) is developed to solve the formulated problem. Simulation results demonstrate the great potential of our UAV-assisted hierarchical aggregation scheme.
Xiangyu Zhong, Xiaojun Yuan 0002, Huiyuan Yang, Chenxi Zhong
GLOBECOM3
2022 Multi-Task Federated Learning with Over-the-Air Computation for MIMO Interference Channels
abstract
Although Federated learning (FL) over wireless medium is a promising technology, a large number of concurrent FL tasks, generated by the urgent demand for ubiquitous intelligence, may seriously aggravate the scarcity of communication resources. By exploiting the analog superposition of electromagnetic waves, over-the-air computation (AirComp) is an appealing solution to alleviate the burden of communication required by FL. However, sharing frequency-time resources in AirComp inevitably brings about the problem of inter-task interference, which poses a new challenge. In this paper, we study over-the-air multi-task FL (OA-MTFL) over the multiple-input multiple-output (MIMO) interference channel. We establish a communication-learning analytical framework for the proposed OA-MTFL scheme by considering the spatial correlation between devices, formulate an optimization problem of designing transceiver beamforming and device selection, and develop an efficient algorithm to solve it. The numerical results demonstrate the outstanding performance of the proposed scheme.
Chenxi Zhong, Huiyuan Yang, Xiaojun Yuan 0002
ISNCC2
2022 Multiscale feature fusion for surveillance video diagnosis
Fanglin Chen 0001, Weihang Wang 0005, Huiyuan Yang, Wenjie Pei, Guangming Lu 0002
Knowl. Based Syst.3
2021 Exploiting Semantic Embedding and Visual Feature for Facial Action Unit Detection
abstract
Recent study on detecting facial action units (AU) has utilized auxiliary information (i.e., facial landmarks, relationship among AUs and expressions, web facial images, etc.), in order to improve the AU detection performance. As of now, no semantic information of AUs has yet been explored for such a task. As a matter of fact, AU semantic descriptions provide much more information than the binary AU labels alone, thus we propose to exploit the Semantic Embedding and Visual feature (SEV-Net) for AU detection. More specifically, AU semantic embeddings are obtained through both Intra-AU and Inter-AU attention modules, where the Intra-AU attention module captures the relation among words within each sentence that describes individual AU, and the Inter-AU attention module focuses on the relation among those sentences. The learned AU semantic embeddings are then used as guidance for the generation of attention maps through a cross-modality attention network. The generated cross-modality attention maps are further used as weights for the aggregated feature. Our proposed method is unique in that the semantic features are exploited as the first of this kind. The approach has been evaluated on three public AU-coded facial expression databases, and has achieved a superior performance than the state-of-the-art peer methods.
Huiyuan Yang, Lijun Yin 0001, Jiuxiang Gu
CVPR1
2021 Your "Attention" Deserves Attention: A Self-Diversified Multi-Channel Attention for Facial Action Analysis
abstract
Visual attention has been extensively studied for learning fine-grained features in both facial expression recognition (FER) and Action Unit (AU) detection. A broad range of previous research has explored how to use attention modules to localize detailed facial parts (e,g. facial action units), learn discriminative features, and learn inter-class correlation. However, few related works pay attention to the robustness of the attention module itself. Through experiments, we found neural attention maps initialized with different feature maps yield diverse representations when learning to attend the identical Region of Interest (ROI). In other words, similar to general feature learning, the representational quality of attention maps also greatly affects the performance of a model, which means unconstrained attention learning has lots of randomnesses. This uncertainty lets conventional attention learning fall into sub-optimal. In this paper, we propose a compact model to enhance the representational and focusing power of neural attention maps and learn the “inter-attention” correlation for refined attention maps, which we term the “Self-Diversified Multi-Channel Attention Network (SMA-Net)”. The proposed method is evaluated on two benchmark databases (BP4D and DISFA) for AU detection and four databases (CK+, MMI, BU-3DFE, and BP4D+) for facial expression recognition. It achieves superior performance compared to the state-of-the-art methods.
Huiyuan Yang, Geran Zhao, Lijun Yin 0001
FG3
2020 RE-Net: A Relation Embedded Deep Model for AU Occurrence and Intensity Estimation
Huiyuan Yang, Lijun Yin 0001
ACCV (5)1
2020 An EEG-Based Multi-Modal Emotion Database with Both Posed and Authentic Facial Actions for Emotion Analysis
abstract
Emotion is an experience associated with a particular pattern of physiological activity along with different physiological, behavioral and cognitive changes. One behavioral change is facial expression, which has been studied extensively over the past few decades. Facial behavior varies with a person's emotion according to differences in terms of culture, personality, age, context, and environment. In recent years, physiological activities have been used to study emotional responses. A typical signal is the electroencephalogram (EEG), which measures brain activity. Most of existing EEG-based emotion analysis has overlooked the role of facial expression changes. There exits little research on the relationship between facial behavior and brain signals due to the lack of dataset measuring both EEG and facial action signals simultaneously. To address this problem, we propose to develop a new database by collecting facial expressions, action units, and EEGs simultaneously. We recorded the EEGs and face videos of both posed facial actions and spontaneous expressions from 29 participants with different ages, genders, ethnic backgrounds. Differing from existing approaches, we designed a protocol to capture the EEG signals by evoking participants' individual action units explicitly. We also investigated the relation between the EEG signals and facial action units. As a baseline, the database has been evaluated through the experiments on both posed and spontaneous emotion recognition with images alone, EEG alone, and EEG fused with images, respectively. The database will be released to the research community to advance the state of the art for automatic emotion recognition.
Xiang Zhang 0030, Huiyuan Yang, Wenna Duan, Weiying Dai, Lijun Yin 0001
FG3
2020 Set Operation Aided Network for Action Units Detection
abstract
As a large number of parameters exist in deepmodel based methods, training such models usually requires many fully AU-annotated facial images. This is true with regard to the number of frames in two widely used datasets: BP4D[31] and DISFA [18], while those frames were captured from a small number of subjects (41, 27 respectively). This is problematic, as subjects produce highly consistent facial muscle movements, adding more frames per subject would only adds more close points in the feature space, and thus the classifier does not benefit from those extra frames. Data augmentation methods can be applied to alleviate the problem to a certain degree, but they fail to augment new subjects. We propose a novel Set Operation Aided Network (SO-Net) for action units detection. Specifically, new features and the corresponding labels are generated by adding set operations to both the feature and label spaces. The generated new features can be treated as a representation of a hypothetical image. As a result, we can implicitly obtain training examples beyond what was originally observed in the dataset. Therefore, the deep model is forced to learn subject-independent features, and is generalizable to unseen subjects. SO-Net is end-to-end trainable, and can be flexibly plugged in any CNN model during training. We evaluate the proposed method on two public datasets, BP4D and DISFA. The experiment shows a state-of-the-art performance, demonstrating the effectiveness of the proposed method.
Huiyuan Yang, Taoyue Wang, Lijun Yin 0001
FG1
2020 Two-Timescale Optimization for Intelligent Reflecting Surface Aided D2D Underlay Communication
abstract
The performance of a device-to-device (D2D) underlay communication system is limited by the co-channel interference between cellular users (CUs) and D2D devices. To address this challenge, an intelligent reflecting surface (IRS) aided D2D underlay system is studied in this paper. A two-timescale optimization scheme is proposed to reduce the required channel training and feedback overhead, where transmit beamforming at the base station (BS) and power control at the D2D transmitter are adapted to instantaneous effective channel state information (CSI); and the IRS phase shifts are adapted to slow-varying channel mean. Based on the two-timescale optimization scheme, we aim to maximize the D2D ergodic rate subject to a given outage probability constrained signal-to-interference-plus-noise ratio (SINR) target for the CU. The two-timescale problem is decoupled into two sub-problems, and the two sub-problems are solved iteratively with closed-form expressions. Numerical results verify that the two-timescale based optimization performs better than several baselines, and also demonstrate a favorable trade-off between system performance and CSI overhead.
Huiyuan Yang, Xiaojun Yuan 0002, Ying-Chang Liang
GLOBECOM2
2020 Reconfigurable Intelligent Surface Aided Constant-Envelope Wireless Power Transfer
abstract
By reconfiguring the propagation environment of electromagnetic waves artificially, reconfigurable intelligent surfaces (RISs) have been regarded as a promising and revolutionary hardware technology to improve the energy and spectrum efficiency of wireless networks. In this paper, we study a RIS aided multiuser multiple-input single-output (MISO) wireless power transfer (WPT) system, where the transmitter is equipped with a constant-envelope analog beamformer. We formulate a novel problem to maximize the total received power of all the users by jointly optimizing the beamformer at transmitter and the phase shifts at the RISs, subject to the individual minimum received power constraints of users. We further solve the problem iteratively with a closed-form expression for each step. Numerical results show the performance gain of deploying RIS and the effectiveness of the proposed algorithm.
Huiyuan Yang, Xiaojun Yuan 0002, Jun Fang 0001, Ying-Chang Liang
GLOBECOM1
2020 Adaptive Multimodal Fusion for Facial Action Units Recognition
abstract
Multimodal facial action units (AU) recognition aims to build models that are capable of processing, correlating, and integrating information from multiple modalities (i.e., 2D images from a visual sensor, 3D geometry from 3D imaging, and thermal images from an infrared sensor). Although the multimodel data can provide rich information, there are two challenges that have to be addressed when learning from multimodal data: 1) the model must capture the complex cross-modal interactions in order to utilize the additional and mutual information effectively; 2) the model must be robust enough in the circumstance of unexpected data corruptions during testing, in case of a certain modality missing or being noisy. In this paper, we propose a novel A daptive M ultimodal F usion method (AMF ) for AU detection, which learns to select the most relevant feature representations from different modalities by a re-sampling procedure conditioned on a feature scoring module. The feature scoring module is designed to allow for evaluating the quality of features learned from multiple modalities. As a result, AMF is able to adaptively select more discriminative features, thus increasing the robustness to missing or corrupted modalities. In addition, to alleviate the over-fitting problem and make the model generalize better on the testing data, a cut-switch multimodal data augmentation method is designed, by which a random block is cut and switched across multiple modalities. We have conducted a thorough investigation on two public multimodal AU datasets, BP4D and BP4D+, and the results demonstrate the effectiveness of the proposed method. Ablation studies on various circumstances also show that our method remains robust to missing or noisy modalities during tests.
Huiyuan Yang, Taoyue Wang, Lijun Yin 0001
ACM Multimedia1
2019 Learning Temporal Information From A Single Image For AU Detection
abstract
Automatic Facial Action Units (AUs) detection is the recognition of the facial appearance changes caused by the contraction or relaxation of one or more related facial muscles. Compared to the sequence-based methods, a decreased performance is observed for the static image-based AU detection, due to the loss of temporal information. To solve this problem, we propose a novel method that implicitly learns temporal information from a single image for AU detection by adding a hidden optical-flow layer to concatenate two Convolutional Neural Networks (CNNs) models: optical-flow net (OF-Net) and AU detection net (AU-Net). The OF-Net is designed to estimate the facial appearance changes (optical flow) from a single input image through unsupervised learning. The AU-Net accepts the estimated optical-flow as input and predicts the AU occurrence. By training both OF-Net and AU-Net jointly, our model achieves better performance than training them separately, as the AU-Net provides semantic constraints for the optical-flow learning and helps generate more meaningful optical-flow. In return, the estimated optical-flow, which reflects facial appearance changes, benefits the AU-Net. Our proposed method has been evaluated on two benchmarks: BP4D and DISFA, and the experiments show significant performance improvement as compared to the state-of-the-art methods.
Huiyuan Yang, Lijun Yin 0001
FG1
2019 Multi-Modality Empowered Network for Facial Action Unit Detection
abstract
This paper presents a new thermal empowered multi-task network (TEMT-Net) to improve facial action unit detection. Our primary goal is to leverage the situation that the training set has multi-modality data while the application scenario only has one modality. Thermal images are robust to illumination and face color. In the proposed multi-task framework, we utilize both modality data. Action unit detection and facial landmark detection are correlated tasks. To utilize the advantage and the correlation of different modalities and different tasks, we propose a novel thermal empowered multi-task deep neural network learning approach for action unit detection, facial landmark detection and thermal image reconstruction simultaneously. The thermal image generator and facial landmark detection provide regularization on the learned features with shared factors as the input color images. Extensive experiments are conducted on the BP4D and MMSE databases, with the comparison to the state of the art methods. The experiments show that the multi-modality framework improves the AU detection significantly.
Peng Liu 0039, Zheng Zhang 0023, Huiyuan Yang, Lijun Yin 0001
WACV3
2018 Facial Expression Recognition by De-Expression Residue Learning
abstract
A facial expression is a combination of an expressive component and a neutral component of a person. In this paper, we propose to recognize facial expressions by extracting information of the expressive component through a de-expression learning procedure, called De-expression Residue Learning (DeRL). First, a generative model is trained by cGAN. This model generates the corresponding neutral face image for any input face image. We call this procedure de-expression because the expressive information is filtered out by the generative model; however, the expressive information is still recorded in the intermediate layers. Given the neutral face image, unlike previous works using pixel-level or feature-level difference for facial expression classification, our new method learns the deposition (or residue) that remains in the intermediate layers of the generative model. Such a residue is essential as it contains the expressive component deposited in the generative model from any input facial expression images. Seven public facial expression databases are employed in our experiments. With two databases (BU-4DFE and BP4D-spontaneous) for pre-training, the DeRL method has been evaluated on five databases, CK+, Oulu-CASIA, MMI, BU-3DFE, and BP4D+. The experimental results demonstrate the superior performance of the proposed method.
Huiyuan Yang, Umur A. Ciftci, Lijun Yin 0001
CVPR1
2018 Identity-Adaptive Facial Expression Recognition through Expression Regeneration Using Conditional Generative Adversarial Networks
abstract
Subject variation is a challenging issue for fa- cial expression recognition, especially when handling unseen subjects with small-scale lableled facial expression databases. Although transfer learning has been widely used to tackle the problem, the performance degrades on new data. In this paper, we present a novel approach (so-called IA-gen) to alleviate the issue of subject variations by regenerating expressions from any input facial images. First of all, we train conditional generative models to generate six prototypic facial expressions from any given query face image while keeping the identity related information unchanged. Generative Adversarial Networks are employed to train the conditional generative models, and each of them is designed to generate one of the prototypic facial expression images. Second, a regular CNN (FER-Net) is fine- tuned for expression classification. After the corresponding prototypic facial expressions are regenerated from each facial image, we output the last FC layer of FER-Net as features for both the input image and the generated images. Based on the minimum distance between the input image and the generated expression images in the feature space, the input image is classified as one of the prototypic expressions consequently. Our proposed method can not only alleviate the influence of inter-subject variations, but will also be flexible enough to integrate with any other FER CNNs for person-independent facial expression recognition. Our method has been evaluated on CK+, Oulu-CASIA, BU-3DFE and BU-4DFE databases, and the results demonstrate the effectiveness of our proposed method.
Huiyuan Yang, Zheng Zhang 0023, Lijun Yin 0001
FG1
2017 CNN based 3D facial expression recognition using masking and landmark features
abstract
Automatically recognizing facial expression is an important part for human-machine interaction. In this paper, we first review the previous studies on both 2D and 3D facial expression recognition, and then summarize the key research questions to solve in the future. Finally, we propose a 3D facial expression recognition (FER) algorithm using convolutional neural networks (CNNs) and landmark features/masks, which is invariant to pose and illumination variations due to the solely use of 3D geometric facial models without any texture information. The proposed method has been tested on two public 3D facial expression databases: BU-4DFE and BU-3DFE. The results show that the CNN model benefits from the masking, and the combination of landmark and CNN features can further improve the 3D FER accuracy.
Huiyuan Yang, Lijun Yin 0001
ACII1
2016 Multimodal Spontaneous Emotion Corpus for Human Behavior Analysis
abstract
Emotion is expressed in multiple modalities, yet most research has considered at most one or two. This stems in part from the lack of large, diverse, well-annotated, multimodal databases with which to develop and test algorithms. We present a well-annotated, multimodal, multidimensional spontaneous emotion corpus of 140 participants. Emotion inductions were highly varied. Data were acquired from a variety of sensors of the face that included high-resolution 3D dynamic imaging, high-resolution 2D video, and thermal (infrared) sensing, and contact physiological sensors that included electrical conductivity of the skin, respiration, blood pressure, and heart rate. Facial expression was annotated for both the occurrence and intensity of facial action units from 2D video by experts in the Facial Action Coding System (FACS). The corpus further includes derived features from 3D, 2D, and IR (infrared) sensors and baseline results for facial expression and action unit detection. The entire corpus will be made available to the research community.
Zheng Zhang 0023, Jeffrey M. Girard, Yue Wu 0002, Xing Zhang 0012, Peng Liu 0039, Umur A. Ciftci, Shaun J. Canavan, Michael Reale, Andrew Horowitz, Huiyuan Yang, Jeffrey F. Cohn, Lijun Yin 0001
CVPR10
2014 Self-learning PD algorithms based on approximate dynamic programming for robot motion planning
abstract
Motion planning is a key technology of the navigation and control for mobile robots. However, when considering the complexity of exterior environment and mobile robot's kinematics and dynamics, the motion planning results obtained by some traditional methods are often hard to optimize. In this paper, we propose two self-learning PD algorithms to solve motion planning for mobile robots. We firstly utilize a virtual Proportional Derivative (PD) control strategy to transform the motion planning problem into an optimization problem of the virtual control policy. Afterwards, two approximate dynamic programming algorithms, which are the Least Squares Policy Iteration (LSPI) algorithm and the Dual Heuristic Programming (DHP) algorithm, are incorporated into the virtual control strategy to tune the PD parameters automatically, namely the LSPI-PD algorithm and the DHP-PD algorithm. Simulations have been performed to validate the effectiveness of the two algorithms, where the LSPI-PD algorithm is suitable for solving problems with discrete action spaces while the DHP-PD algorithm has an advantage in solving problems with continuous action spaces.
Huiyuan Yang, Xin Xu 0001, Chuanqiang Lian
IJCNN1