EDBT 2026 Demo / reviewers in the wild / expert
Zhilei Liu
dblp:51/8175
· DBLP profile ↗
46ranked-venue papers
12as first author
20since 2021 · last 2026
0000-0003-1447-6256ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 26 · 6 first-author · 11 since 2021Artificial intelligence and machine learning · 20 · 4 first-author · 8 since 2021Computer networks · 3 · 2 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 1 since 2021Systems, architecture and hardware · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | UAU-Net: Uncertainty-aware Representation Learning and Evidential Classification for Facial Action Unit DetectionabstractFacial action unit (AU) detection remains challenging because it involves heterogeneous, AU-specific uncertainties arising at both the representation and decision stages. Recent methods have improved discriminative feature learning, but they often treat the AU representations as deterministic, overlooking uncertainty caused by visual noise, subject-dependent appearance variations, and ambiguous inter-AU relationships, all of which can substantially degrade robustness. Meanwhile, conventional point-estimation classifiers often provide poorly calibrated confidence, producing overconfident predictions, especially under the severe label imbalance typical of AU datasets. We propose UAU-Net, an Uncertainty-aware AU detection framework that explicitly models uncertainty at both stages. At the representation stage, we introduce CV-AFE, a conditional VAE (CVAE)-based AU feature extraction module that learns probabilistic AU representations by jointly estimating feature means and variances across multiple spatio-temporal scales; conditioning on AU labels further enables CV-AFE to capture uncertainty associated with inter-AU dependencies. At the decision stage, we design AB-ENN, an Asymmetric Beta Evidential Neural Network for multi-label AU detection, which parameterizes predictive uncertainty with Beta distributions and mitigates overconfidence via an asymmetric loss tailored to highly imbalanced binary labels. Extensive experiments on BP4D and DISFA show that UAU-Net achieves strong AU detection performance, and further analyses indicate that modeling uncertainty in both representation learning and evidential prediction improves robustness and reliability. Zhilei Liu |
ICMR | 2 |
| 2025 | From Bottleneck to Breakthrough: Optimizing Scheduling for Hyperscale Containerized ClustersabstractContainer orchestration platforms have become the backbone of modern private cloud infrastructure, offering flexibility, scalability, and reliability for managing containerized workloads. Among them, Kubernetes has emerged as the de facto standard, widely adopted across the industry, with many organizations operating tens or even hundreds of clusters globally. However, as infrastructure scales to ultra-large clusters (exceeding 10,000 nodes) and faces extreme provisioning scenarios—such as abrupt workload surges, irrespective of available resource headroom—scheduling latency becomes a critical bottleneck. This issue is not adequately addressed by the default Kubernetes scheduler or existing alternative solutions, which struggle to maintain low latency under such conditions. Yuquan Ren, Xinyi Song, Zhilei Liu, Caixue Lin, Wu Xiang |
SoCC | 4 |
| 2025 | NeRF-3DTalker: Neural Radiance Field with 3D Prior Aided Audio Disentanglement for Talking Head SynthesisabstractTalking head synthesis is to synthesize a lip-synchronized talking head video using audio. Recently, the capability of NeRF to enhance the realism and texture details of synthesized talking heads has attracted the attention of researchers. However, most current NeRF methods based on audio are exclusively concerned with the rendering of frontal faces. These methods are unable to generate clear talking heads in novel views. Another prevalent challenge in current 3D talking head synthesis is the difficulty in aligning acoustic and visual spaces, which often results in suboptimal lip-syncing of the generated talking heads. To address these issues, we propose Neural Radiance Field with 3D Prior Aided Audio Disentanglement for Talking Head Synthesis (NeRF-3DTalker). Specifically, the proposed method employs 3D prior information to synthesize clear talking heads with free views. Additionally, we propose a 3D Prior Aided Audio Disentanglement module, which is designed to disentangle the audio into two distinct categories: features related to 3D awarded speech movements and features related to speaking style. Moreover, to reposition the generated frames that are distant from the speaker’s motion space in the real space, we have devised a local-global Standardized Space. This method normalizes the irregular positions in the generated frames from both global and local semantic perspectives. Through comprehensive qualitative and quantitative experiments, it has been demonstrated that our NeRF3DTalker outperforms state-of-the-art in synthesizing realistic talking head videos, exhibiting superior image quality and lip synchronization. Project page: https://nerf-3dtalker.github.io/NeRF-3Dtalker/. Xiaoxing Liu, Zhilei Liu, Chongke Bi |
ICASSP | 2 |
| 2025 | Improving Deep Q Network Based on Marketing Psychology for AUV Path Planning in Unknown Marine EnvironmentsabstractAutonomous underwater vehicles (AUVs) are commonly utilized in fields like marine environmental testing because of their adaptable and independent characteristics. However, the complexity and variability of the underwater environment present challenges for AUVs. AUVs need to safely and stably navigate through complex unknown underwater environments. The existing path planning technologies for AUVs in unknown marine environments face challenges, such as slow convergence speed and low success rate. In this article, a deep Q-learning (DQN) approach grounded in marketing psychology (MDQN) is suggested to tackle the challenges mentioned above. First, to reasonably allocate the exploration and exploitation abilities in DQN, a mathematical optimizer acceleration function is developed. Additionally, to improve the environmental adaptability of AUVs and speed up the convergence of deep networks, a reward mechanism that takes into account ocean currents and the shortest path principle is suggested. Finally, the reward feedback and experience stratification mechanism are designed based on marketing psychology principles, aiming to promote a virtuous cycle between the training process of deep networks and the learning process of the agent. Various simulation experiments comprehensively demonstrate that MDQN has excellent path planning capabilities in unknown marine environments, effectively addressing the slow convergence speed and low success rate issues of DQN. Zhilei Liu, Jiaoyi Hou, Dayong Ning, Gangda Liang |
IEEE Internet Things J. | 1 |
| 2024 | NERF-AD: Neural Radiance Field With Attention-Based Disentanglement For Talking Face SynthesisabstractTalking face synthesis driven by audio is one of the current research hotspots in the fields of multidimensional signal processing and multimedia. Neural Radiance Field (NeRF) has recently been brought to this research field in order to enhance the realism and 3D effect of the generated faces. However, most existing NeRF-based methods either burden NeRF with complex learning tasks while lacking methods for supervised multimodal feature fusion, or cannot precisely map audio to the facial region related to speech movements. These reasons ultimately result in existing methods generating inaccurate lip shapes. This paper moves a portion of NeRF learning tasks ahead and proposes a talking face synthesis method via NeRF with attention-based disentanglement (NeRF-AD). In particular, an Attention-based Disentanglement module is introduced to disentangle the face into Audio-face and Identity-face using speech-related facial action unit (AU) information. To precisely regulate how audio affects the talking face, we only fuse the Audio-face with audio feature. In addition, AU information is also utilized to supervise the fusion of these two modalities. Extensive qualitative and quantitative experiments demonstrate that our NeRF-AD outperforms state-of-the-art methods in generating realistic talking face videos, including image quality and lip synchronization. To view video results, please refer to https://xiaoxingliu02.github.io/NeRF-AD/. Chongke Bi, Xiaoxing Liu, Zhilei Liu |
ICASSP | 3 |
| 2024 | An improved algorithm optimization algorithm based on RungeKutta and golden sine strategy
Mingying Li, Zhilei Liu, Hongxiang Song |
Expert Syst. Appl. | 2 |
| 2024 | A novel Elman neural network based on Gaussian kernel and improved SOA and its applications
Zhilei Liu, Dayong Ning, Jiaoyi Hou |
Expert Syst. Appl. | 1 |
| 2023 | Self-Supervised Facial Action Unit Detection with Region and Relation LearningabstractFacial action unit (AU) detection is a challenging task due to the scarcity of manual annotations. Recent works on AU detection with self-supervised learning have emerged to address this problem, aiming to learn meaningful AU representations from numerous unlabeled data. However, most existing AU detection works with self-supervised learning utilize global facial features only, while AU-related properties such as locality and relevance are not fully explored. In this paper, we propose a novel self-supervised framework for AU detection with the region and relation learning. In particular, AU related attention map is utilized to guide the model to focus more on AU-specific regions to enhance the integrity of AU local features. Meanwhile, an improved Optimal Transport (OT) algorithm is introduced to exploit the correlation characteristics among AUs. In addition, Swin Transformer is exploited to model the long-distance dependencies within each AU region during feature learning. The evaluation results on BP4D and DISFA demonstrate that our proposed method is comparable or even superior to the state-of-the-art self-supervised learning methods and supervised AU detection methods. Juan Song, Zhilei Liu |
ICASSP | 2 |
| 2023 | Teacher-Student Network for Real-World Face Super-Resolution with Progressive Embedding of Edge InformationabstractTraditional face super-resolution (FSR) methods trained on synthetic datasets usually have poor generalization ability for real-world face images. Recent work has utilized complex degradation models or training networks to simulate the real degradation process, but this limits the performance of these methods due to the domain differences that still exist between the generated low-resolution images and the real low-resolution images. Moreover, because of the existence of a domain gap, the semantic feature information of the target domain may be affected when synthetic data and real data are utilized to train super-resolution models simultaneously. In this study, a real-world face super-resolution teacher-student model is proposed, which considers the domain gap between real and synthetic data and progressively includes diverse edge information by using the recurrent network’s intermediate outputs. Extensive experiments demonstrate that our proposed approach surpasses state-of-the-art methods in obtaining high-quality face images for real-world FSR. Zhilei Liu, Chenggong Zhang |
ICIP | 1 |
| 2023 | Joint face completion and super-resolution using multi-scale feature relation learning
Zhilei Liu, Chenggong Zhang, Yunpeng Wu, Cuicui Zhang |
J. Vis. Commun. Image Represent. | 1 |
| 2022 | Face Super-Resolution with Progressive Embedding of Multi-scale Face PriorsabstractThe face super-resolution (FSR) task is to reconstruct high-resolution face images from low-resolution inputs. Re-cent works have achieved success on this task by utilizing facial priors such as facial landmarks. Most existing methods pay more attention to global shape and structure information, but less to local texture information, which makes them cannot recover local details well. In this pa-per, we propose a novel recurrent convolutional network based framework for face super-resolution, which progres-sively introduces both global shape and local texture infor-mation. We take full advantage of the intermediate outputs of the recurrent network, and landmarks information and facial action units (AUs) information are extracted in the output of the first and second steps respectively, rather than low-resolution input. Moreover, we introduced AU clas-sification results as a novel quantitative metric for facial details restoration. Extensive experiments show that our proposed method significantly outperforms state-of-the-art FSR methods in terms of image quality and facial details restoration. Chenggong Zhang, Zhilei Liu |
IJCB | 2 |
| 2022 | Talking Head Generation for Media Interaction System with Feature DisentanglementabstractThe task of talking head generation for the media interaction system is to take images and audio clips of the target face as input, and generate a realistic video of the target synchronized with the audio. Most of the existing works directly take all the information in the image as input, which causes the problem of feature information redundancy. At the same time, in the stage of feature fusion, the image and audio information is directly spliced, and the relationship between the two modal information is ignored. To solve the problems, we initially implement the disentanglement of image by introducing a loss function to separate the image into identity features and content-related features. Besides, we introduce a multi-head selfattention mechanism to learn the relationship between the two modal information of image and audio, and implement the full fusion of multi-modal information. In addition, we validate the effectiveness of our model through extensive quantitative and qualitative analysis of two datasets. Extensive experiments show the superiority of the proposed model in all aspects. Lei Zhang 0024, Zhilei Liu |
ICPADS | 3 |
| 2022 | Cross-subject Action Unit Detection with Meta Learning and Transformer-based Relation ModelingabstractFacial Action Unit (AU) detection is a crucial task for emotion analysis from facial movements. The apparent differences of different subjects sometimes mislead changes brought by AUs, resulting in inaccurate results. However, most of the existing AU detection methods based on deep learning didn't consider the identity information of different subjects. The paper proposes a meta-learning-based cross-subject AU detection model to eliminate the identity-caused differences. Besides, a transformer-based relation learning module is introduced to learn the latent relations of multiple AUs. To be specific, our proposed work is composed of two sub-tasks. The first sub-task is meta-learning-based AU local region representation learning, called MARL, which learns discriminative representation of local AU regions that incorporates the shared information of multiple subjects and eliminates identity-caused differences. The second sub-task uses the local region representation of AU of the first sub-task as input, then adds relationship learning based on the transformer encoder architecture to capture AU relationships. The entire training process is cascaded. Ablation study and visualization show that our MARL can eliminate identity-caused differences, thus obtaining a robust and generalized AU discriminative embedding representation. Our results prove that on the two public datasets BP4D and DISFA, our method is superior to the state-of-the-art technology, and the F1 score is improved by 1.3% and 1.4%, respectively. Jiyuan Cao, Zhilei Liu, Yong Zhang 0034 |
IJCNN | 2 |
| 2022 | Facial Action Unit Detection Using Attention and Relation LearningabstractAttention mechanism has recently attracted increasing attentions in the field of facial action unit (AU) detection. By finding the region of interest of each AU with the attention mechanism, AU-related local features can be captured. Most of the existing attention based AU detection works use prior knowledge to predefine fixed attentions or refine the predefined attentions within a small range, which limits their capacity to model various AUs. In this paper, we propose an end-to-end deep learning based attention and relation learning framework for AU detection with only AU labels, which has not been explored before. In particular, multi-scale features shared by each AU are learned first, and then both channel-wise and spatial attentions are adaptively learned to select and extract AU-related local features. Moreover, pixel-level relations for AUs are further captured to refine spatial attentions so as to extract more relevant local features. Without changing the network architecture, our framework can be easily extended for AU intensity estimation. Extensive experiments show that our framework (i) soundly outperforms the state-of-the-art methods for both AU detection and AU intensity estimation on the challenging BP4D, DISFA, FERA 2015, and BP4D+ benchmarks, (ii) can adaptively capture the correlated regions of each AU, and (iii) also works well under severe occlusions and large poses. Zhiwen Shao, Zhilei Liu, Jianfei Cai 0001, Yunsheng Wu, Lizhuang Ma |
IEEE Trans. Affect. Comput. | 2 |
| 2021 | Talking Head Generation with Audio and Speech Related Facial Action Units
Zhilei Liu, Jiaxing Liu 0001, Zhengxiang Yan, Longbiao Wang |
BMVC | 2 |
| 2021 | Multimodal Emotion Recognition with Capsule Graph Convolutional Based Representation FusionabstractDue to the more robust characteristics compared to unimodal, audio-video multimodal emotion recognition (MER) has attracted a lot of attention. The efficiency of representation fusion algorithm often determines the performance of MER. Although there are many fusion algorithms, information redundancy and information complementarity are usually ignored. In this paper, we propose a novel representation fusion method, Capsule Graph Convolutional Network (CapsGCN). Firstly, after unimodal representation learning, the extracted audio and video representations are distilled by capsule network and encapsulated into multimodal capsules respectively. Multimodal capsules can effectively reduce data redundancy by the dynamic routing algorithm. Secondly, the multimodal capsules with their inter-relations and intra-relations are treated as a graph structure. The graph structure is learned by Graph Convolutional Network (GCN) to get hidden representation which is a good supplement for information complementarity. Finally, the multimodal capsules and hidden relational representation learned by CapsGCN are fed to multihead self-attention to balance the contributions of source representation and relational representation. To verify the performance, visualization of representation, the results of commonly used fusion methods, and ablation studies of the proposed CapsGCN are provided. Our proposed fusion method achieves 80.83% accuracy and 80.23% F1 score on eNTERFACE05’. Jiaxing Liu 0001, Longbiao Wang, Zhilei Liu, Yahui Fu 0001, Lili Guo 0001, Jianwu Dang 0001 |
ICASSP | 4 |
| 2021 | Improved 3D Morphable Model for Facial Action Unit Synthesis
Zhilei Liu |
ICIG (3) | 2 |
| 2021 | Multi-scale Feature Relation Modeling for Facial Expression RestorationabstractFacial expression analysis in the wild is easily vulnerable to the quality of facial images, such as low resolution or occlusion. Existing facial image restoration studies have mostly failed to take full advantage of facial expression prior information, which leads to loss of information related to facial expression in the restoration results. In this paper, we propose a multi-scale feature relation modeling GAN (MFRM-GAN) for facial expression restoration by exploring the multi-scale property and relationship of facial action units. Based on the GAN model, the MFRM-GAN integrates the graph convolution network (GCN) for feature relation modeling and feature pyramid network (FPN) for multi-scale feature extraction. Extensive qualitative and quantitative experiments on BP4D and DISFA datasets demonstrate that our proposed MFRM-GAN (i) can conduct facial expression in-painting and facial image super-resolution jointly, (ii) can recover better facial expression details comparing with state-of-the-art method in both visual effect and AU detection task. Zhilei Liu, Yunpeng Wu, Cuicui Zhang |
IJCNN | 1 |
| 2021 | A Sentiment Similarity-Oriented Attention Model with Multi-task Learning for Text-Based Emotion Recognition
Yahui Fu 0001, Lili Guo 0001, Longbiao Wang, Zhilei Liu, Jiaxing Liu 0001, Jianwu Dang 0001 |
MMM (1) | 4 |
| 2021 | JÂA-Net: Joint Facial Action Unit Detection and Face Alignment Via Adaptive Attention
Zhiwen Shao, Zhilei Liu, Jianfei Cai 0001, Lizhuang Ma |
Int. J. Comput. Vis. | 2 |
| 2020 | Speech Emotion Recognition with Local-Global Aware Deep Representation LearningabstractConvolutional neural network (CNN) based deep representation learning methods for speech emotion recognition (SER) have demonstrated great success. The basic design of CNN restricts the ability to model only local information well. Capsule network (CapsNet) can overcome the shortages of CNNs to capture the shallow global features from the spectrogram, although CapsNet cannot learn the local and deep global information. In this paper, we propose a local-global aware deep representation learning system that mainly includes two modules. One module contains a multi-scale CNN, time- frequency CNN (TFCNN) to learn the local representation. In the other module, we introduce a structure with dense connections of multiple blocks to learn shallow and deep global information. Every block in this structure is a complete CapsNet improved by a new routing algorithm. The local and global representations are fed to the classifier and achieve an absolute increase of at least 4.25% than benchmarks on IEMOCAP. Jiaxing Liu 0001, Zhilei Liu, Longbiao Wang, Lili Guo 0001, Jianwu Dang 0001 |
ICASSP | 2 |
| 2020 | Teacher-Student Competition for Unsupervised Domain AdaptationabstractWith the supervision from source domain only in class-level, existing unsupervised domain adaptation (UDA) methods mainly learn the domain-invariant representations from a shared feature extractor, which causes the source-bias problem. This paper proposes an unsupervised domain adaptation approach with Teacher-Student Competition (TSC). In particular, a student network is introduced to learn the target-specific feature space, and we design a novel competition mechanism to select more credible pseudo-labels for the training of student network. We introduce a teacher network with the structure of existing conventional UDA method, and both teacher and student networks compete to provide target pseudo-labels to constrain every target sample's training in student network. Extensive experiments demonstrate that our proposed TSC framework significantly outperforms the state-of-the-art domain adaptation methods on Office-31 and ImageCLEF-DA benchmarks. Ruixin Xiao, Zhilei Liu, Baoyuan Wu |
ICPR | 2 |
| 2020 | Temporal Attention Convolutional Network for Speech Emotion Recognition with Latent Representation
Jiaxing Liu 0001, Zhilei Liu, Longbiao Wang, Yuan Gao 0040, Lili Guo 0001, Jianwu Dang 0001 |
INTERSPEECH | 2 |
| 2020 | A Feedback Mechanism for Prediction-based Anomaly Detection In Content Delivery NetworksabstractCDN (Content Delivery Network) has become an important infrastructure of the Internet. However, building an anomaly detection system to monitor and guarantee CDN service quality is non-trivial. Current anomaly detection system usually suffers from undesirable performance in terms of high rate of false positive and false negative, which consequently impacts on its practical deployment. Identifying the root cause of a false detection is critical for diagnosing and improving the performance of anomaly detection. In this paper, we propose a novel feedback mechanism for prediction-based anomaly detection in CDN . Specifically, we introduce a carefully-designed metric named Fittingscore to diagnose whether the prediction model can fit the data well. Further, a threshold adjustment mechanism is proposed to dynamically adjust the thresholds of residual errors. Extensive experiments employing a three-month real CDN dataset collected from a top ISP-operated CDN in China show our proposed method can significantly improve the performance of anomaly detection. Zhilei Liu, Tao Lin 0001, Jiyan Sun, Yanjie Hu, Yan Zhang 0014, Zhen Xu 0009 |
ISCC | 1 |
| 2020 | Speaker-Aware Speech Emotion Recognition by Fusing Amplitude and Phase Information
Lili Guo 0001, Longbiao Wang, Jianwu Dang 0001, Zhilei Liu, Haotian Guan |
MMM (1) | 4 |
| 2020 | Relation Modeling with Graph Convolutional Networks for Facial Action Unit Detection
Zhilei Liu, Jiahui Dong, Cuicui Zhang, Longbiao Wang, Jianwu Dang 0001 |
MMM (2) | 1 |
| 2020 | Region Based Adversarial Synthesis of Facial Action Units
Zhilei Liu, Yunpeng Wu |
MMM (2) | 1 |
| 2020 | Facial Expression Restoration Based on Improved Graph Convolutional Networks
Zhilei Liu, Yunpeng Wu, Cuicui Zhang |
MMM (2) | 1 |
| 2019 | Face Inpainting with Dynamic Structural Information of Facial Action Units
Zhilei Liu, Cuicui Zhang |
ICIG (3) | 2 |
| 2019 | Time-Frequency Deep Representation Learning for Speech Emotion Recognition Integrating Self-attention
Jiaxing Liu 0001, Zhilei Liu, Longbiao Wang, Lili Guo 0001, Jianwu Dang 0001 |
ICONIP (4) | 2 |
| 2019 | A Cache-Aware Approach for Dynamic Adaptive Video Streaming over HTTPabstractMore and more CDN (Content Delivery Network) or video content providers are deploying their cache nodes at the edge of networks in order to improve user perceived QoE (Quality of Experience). However, state-of-the-art ABR (Adaptive Bitrate) algorithms do not take into account the presence of a cache on video delivery path. In this paper, we firstly investigate how these ABR algorithms behave in the context of edge caching system and how caching factors such as hit ratio and access bandwidth impact on the performance of ABR algorithms. Extensive analysis results show introducing an edge cache on video delivery path cannot necessarily improve the performance of ABR algorithms in terms of average QoE. Furthermore, we propose a cache-aware ABR approach taking into account caching information including a cache hit indicator for the past chunks and the availability of the next video chunk on the cache. Experimental results validate the significant benefits of the approach which helps to make a more accurate throughput estimation and is more likely to select the video chunks stored in the cache to increase cache utilization. Yudan Liu, Tao Lin 0001, Zhilei Liu |
ISCC | 3 |
| 2019 | Conditional adversarial synthesis of 3D facial action units
Zhilei Liu, Guoxian Song, Jianfei Cai 0001, Tat-Jen Cham, Juyong Zhang |
Neurocomputing | 1 |
| 2018 | Deep Adaptive Attention for Joint Facial Action Unit Detection and Face Alignment
Zhiwen Shao, Zhilei Liu, Jianfei Cai 0001, Lizhuang Ma |
ECCV (13) | 2 |
| 2018 | Updating the Silent Speech Challenge benchmark with deep learning
Yan Ji 0002, Licheng Liu, Hongcui Wang, Zhilei Liu, Zhibin Niu, Bruce Denby |
Speech Commun. | 4 |
| 2017 | Prior-Free Dependent Motion Segmentation Using Helmholtz-Hodge Decomposition Based Object-Motion Oriented Map
Cuicui Zhang, Zhilei Liu |
J. Comput. Sci. Technol. | 2 |
| 2015 | Perception of Mandarin tones by native tibetan speakers
Wenfu Bao, Jianwu Dang 0001, Zhilei Liu |
INTERSPEECH | 4 |
| 2015 | Implicit video emotion tagging from audiences' facial expression
Shangfei Wang, Zhilei Liu, Yachen Zhu, Menghua He |
Multim. Tools Appl. | 2 |
| 2014 | Multi-label Learning with Missing LabelsabstractIn multi-label learning, each sample can be assigned to multiple class labels simultaneously. In this work, we focus on the problem of multi-label learning with missing labels (MLML), where instead of assuming a complete label assignment is provided for each sample, only partial labels are assigned with values, while the rest are missing or not provided. The positive (presence), negative (absence) and missing labels are explicitly distinguished in MLML. We formulate MLML as a transductive learning problem, where the goal is to recover the full label assignment for each sample by enforcing consistency with available label assignments and smoothness of label assignments. Along with an exact solution, we also provide an effective and efficient approximated solution. Our method shows much better performance than several state-of-the-art methods on several benchmark data sets. Baoyuan Wu, Zhilei Liu, Shangfei Wang, Bao-Gang Hu |
ICPR | 2 |
| 2014 | Exploiting multi-expression dependences for implicit multi-emotion video tagging
Shangfei Wang, Zhilei Liu, Jun Wang 0130, Zhaoyu Wang 0003 |
Image Vis. Comput. | 2 |
| 2013 | Analyses of the Differences between Posed and Spontaneous Facial ExpressionsabstractThis paper presents comprehensive analyses of the differences between posed and spontaneous expressions from visible images. First, geometric and appearance features are extracted from the difference images between apex and onset facial images. Secondly, the differences between the posed and spontaneous facial expressions are analyzed through hypothetical testing methods from three aspects: on overall samples, on samples with different genders, and on samples with different expressions. Thirdly, Bayesian networks (BNs) are used to classify posed versus spontaneous expressions from the same three aspects. Statistical analyses on the NVIE database demonstrate the importance of the geometric and appearance features for discriminating posed and spontaneous expressions. Gender effect exists on the differences between posed and spontaneous expressions. It is easier to distinguish posed happiness from spontaneous happiness than other expressions. Recognition experimental results confirm the observations of statistical analyses in most cases. Menghua He, Shangfei Wang, Zhilei Liu |
ACII | 3 |
| 2013 | Eye localization from thermal infrared images
Shangfei Wang, Zhilei Liu, Peijia Shen |
Pattern Recognit. | 2 |
| 2013 | Analyses of a Multimodal Spontaneous Facial Expression DatabaseabstractCreating a large and natural facial expression database is a prerequisite for facial expression analysis and classification. It is, however, not only time consuming but also difficult to capture an adequately large number of spontaneous facial expression images and their meanings because no standard, uniform, and exact measurements are available for database collection and annotation. Thus, comprehensive first-hand data analyses of a spontaneous expression database may provide insight for future research on database construction, expression recognition, and emotion inference. This paper presents our analyses of a multimodal spontaneous facial expression database of natural visible and infrared facial expressions (NVIE). First, the effectiveness of emotion-eliciting videos in the database collection is analyzed with the mean and variance of the subjects' self-reported data. Second, an interrater reliability analysis of raters' subjective evaluations for apex expression images and sequences is conducted using Kappa and Kendall's coefficients. Third, we propose a matching rate matrix to explore the agreements between displayed spontaneous expressions and felt affective states. Lastly, the thermal differences between the posed and spontaneous facial expressions are analyzed using a paired-samples t-test. The results of these analyses demonstrate the effectiveness of our emotion-inducing experimental design, the gender difference in emotional responses, and the coexistence of multiple emotions/expressions. Facial image sequences are more informative than apex images for both expression and emotion recognition. Labeling an expression image or sequence with multiple categories together with their intensities could be a better approach than labeling the expression image or sequence with one dominant category. The results also demonstrate both the importance of facial expressions as a means of communication to convey affective states and the diversity of the displayed manifestations of felt emotions. There are indeed some significant differences between the temperature difference data of most posed and spontaneous facial expressions, many of which are found in the forehead and cheek regions. Shangfei Wang, Zhilei Liu, Zhaoyu Wang 0003, Guobing Wu, Peijia Shen, Xufa Wang |
IEEE Trans. Affect. Comput. | 2 |
| 2012 | Posed and spontaneous expression distinguishment from infrared thermal images
Zhilei Liu, Shangfei Wang |
ICPR | 1 |
| 2011 | Emotion Recognition Using Hidden Markov Models from Facial Temperature Sequence
Zhilei Liu, Shangfei Wang |
ACII (2) | 1 |
| 2010 | Infrared Face Recognition Based on Histogram and K-Nearest Neighbor Classification
Shangfei Wang, Zhilei Liu |
ISNN (2) | 2 |
| 2010 | A Natural Visible and Infrared Facial Expression Database for Expression Recognition and Emotion InferenceabstractTo date, most facial expression analysis has been based on visible and posed expression databases. Visible images, however, are easily affected by illumination variations, while posed expressions differ in appearance and timing from natural ones. In this paper, we propose and establish a natural visible and infrared facial expression database, which contains both spontaneous and posed expressions of more than 100 subjects, recorded simultaneously by a visible and an infrared thermal camera, with illumination provided from three different directions. The posed database includes the apex expressional images with and without glasses. As an elementary assessment of the usability of our spontaneous database for expression recognition and emotion inference, we conduct visible facial expression recognition using four typical methods, including the eigenface approach [principle component analysis (PCA)], the fisherface approach [PCA + linear discriminant analysis (LDA)], the Active Appearance Model (AAM), and the AAM-based + LDA. We also use PCA and PCA+LDA to recognize expressions from infrared thermal images. In addition, we analyze the relationship between facial temperature and emotion through statistical analysis. Our database is available for research purposes. Shangfei Wang, Zhilei Liu, Siliang Lv, Yanpeng Lv, Guobing Wu, Peng Peng 0003, Fei Chen 0013, Xufa Wang |
IEEE Trans. Multim. | 2 |