Xiaohua Huang 0003

dblp:15/8828-3 · DBLP profile ↗
← Back
39ranked-venue papers
16as first author
14since 2021 · last 2026
0000-0001-8897-3517ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 21 · 10 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 6 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author
YearPublicationVenuePosition
2026 Dual-Branch Cross-Diversion Transformer With Spatial Soft Alignment for Few-Shot Surface Defect Detection
abstract
ABSTRACT High‐performance surface defect detection is essential for industrial quality inspection, requiring accurate detection and characterisation of surface defects. While deep learning‐based methods have advanced this field, challenges persist due to limited sample availability and variation in defect types. To address these challenges, we propose a new detection framework, namely, the Dual‐branch Cross‐Diversion Transformer with Spatial Soft Alignment, specifically designed for surface defect detection under data‐scarce conditions. First, the dual‐branch cross‐transformer is leveraged to address data scarcity and enhance defect detection sensitivity through a few‐shot pipeline. Furthermore, the Adaptive Activation Downsampling module is proposed to capture coarse‐grained structural features while preserving fine‐grained defect details, ensuring comprehensive surface defect characterisation. Additionally, a Cross‐Diversion Self‐Attention mechanism further improves multi‐scale feature extraction, critical for accurate detection of diverse defect types. Finally, a Spatial Soft Alignment strategy corrects spatial misalignment between detection proposals and defect categories, reducing detection uncertainty. Through extensive experiments on two benchmark industrial datasets, our proposed architecture achieves superior performance compared to state‐of‐the‐art methods, demonstrating its robustness and accuracy. These results demonstrate the effectiveness of the proposed method and its potential to advance surface defect detection techniques.
Xiaohua Huang 0003
Expert Syst. J. Knowl. Eng.1
2026 Adaptive Key Role Guided Hierarchical Relation Inference for Enhanced Group-Level Emotion Recognition
abstract
In this paper, we propose a novel hierarchical relational network, termed Key Role Guided Hierarchical Relation Inference (KR-HRI), for enhanced group-level emotion recognition (GER). Unlike existing methods that adopt a coarse-grained approach to model interactions among all individuals, our approach adaptively identifies and emphasizes key individuals who play a crucial role in conveying group-level emotions. By integrating coarse-grained relationship modeling with fine-grained key individual enhancement and leveraging global scene information, our method effectively refines discriminative feature generation while minimizing irrelevant interference. We introduce a Multi-branch Interaction Module (MIM) to dynamically fuse features from both the global scene and local individual branches using a localized mask integration strategy. This comprehensive approach enhances the interaction between global and local features, resulting in robust group-level emotion representations. Extensive experiments on three widely adopted GER datasets demonstrate that our framework consistently outperforms state-of-the-art methods, validating the effectiveness and robustness of our proposed approach.
Qing Zhu 0002, Qirong Mao, Wenlong Dong, Xiuyan Shao, Xiaohua Huang 0003, Wenming Zheng
IEEE Trans. Affect. Comput.5
2026 A Survey on Deep Learning for Group-Level Emotion Recognition
abstract
With the rapid advancement of artificial intelligence, group-level emotion recognition (GER) has emerged as an important domain in human behavior analysis. Early GER methods primarily relied on handcrafted features. However, the recent success of deep learning has shifted the focus toward neural network-based solution, enabling more effective exploitation of the rich visual and contextual cues in group images and videos. Unlike individual-level emotion recognition, GER must account for the diversity and dynamics of multiple individuals within varied social contexts. Over the past decade, numerous deep learning-based methods have been proposed, achieving substantial performance gains. This survey provides a comprehensive review of deep learning-centric review of GER, introducing a new taxonomy that spans representation learning, graph-based modeling, attention and transformer architectures, and multimodal fusion strategies. We summarize benchmark datasets, outline prevailing GER pipelines, and consolidate performance trends from recent state-of-the-art approaches. In addition, we discuss the integration of foundation models and large language model-guided multimodal reasoning into GER. Key challenges are identified, and potential research directions are proposed to support the development of robust, real-world GER systems. This work aims to serve as a pivotal reference for future research in this evolving field.
Xiaohua Huang 0003, Xiaopeng Hong, Qirong Mao, Wenming Zheng, Abhinav Dhall
IEEE Trans. Comput. Soc. Syst.1
2026 Multimodal Spatiotemporal Semisupervised Transformer Network for Video-Based Group-Level Emotion Recognition
abstract
Group-level emotion recognition (GER) has emerged as a critical research topic for identifying collective emotions in multiperson scenarios. Despite recent advancements, such as dual branch cross-attention (CA) mechanism, existing methods struggle to differentiate ambiguous emotion categories effectively. In addition, the limited size and diversity of GER datasets hinder further performance improvements. To address these challenges, this article introduces a novel approach, the multimodal spatiotemporal semisupervised transformer (MSST). First, we propose a multimodal spatiotemporal transformer to encode spatial features, capture temporal dynamics, and fuse information from three modalities effectively. Second, a semisupervised learning (SSL) strategy leverages unlabeled data, enhancing robustness against noise and outliers. Last, a two-stage classification strategy and consistency loss are introduced to improve the model’s ability to handle category ambiguity and ensure robust predictions for similar samples. Comprehensive experiments conducted on benchmark GER datasets demonstrate that MSST either considerably outperforms or achieves competitive performance compared to state-of-the-art methods, underscoring its effectiveness in advancing GER research and overcoming the limitations of existing approaches.
Xiaohua Huang 0003, Jinke Xu
IEEE Trans. Comput. Soc. Syst.1
2025 An Empirical Study of Super-Resolution on Low-Resolution Micro-Expression Recognition
Ling Zhou 0005, Mingpei Wang, Xiaohua Huang 0003, Wenming Zheng, Qirong Mao, Guoying Zhao 0001
IEEE Trans. Affect. Comput.3
2025 Towards a Robust Group-Level Emotion Recognition via Uncertainty-Aware Learning
abstract
Group-level emotion recognition (GER) is an inseparable part of human behavior analysis, aiming to recognize an overall emotion in a multi-person scene. However, the existing methods are devoted to combing diverse emotion cues while ignoring the inherent uncertainties under unconstrained environments, such as congestion and occlusion occurring within a group. Additionally, since only group-level labels are available, inconsistent emotion predictions among individuals in one group can confuse the network. In this paper, we propose an uncertainty-aware learning (UAL) method to extract more robust representations for GER. By explicitly modeling the uncertainty, we adopt stochastic embedding sourced from a Gaussian distribution instead of deterministic point embedding. It helps capture the probabilities of emotions and facilitates diverse inferences. Additionally, we adaptively assign uncertainty-sensitive scores as the fusion weights for individuals’ faces within a group. Moreover, we developed an image enhancement module to evaluate and filter samples, strengthening the model’s data-level robustness against uncertainties. The overall three-branch model, encompassing face, object, and scene components, is guided by a proportional-weighted fusion strategy and integrates the proposed uncertainty-aware method to produce the final group-level output. Experimental results demonstrate the effectiveness and generalization ability of our method across three widely used databases.
Qing Zhu 0002, Qirong Mao, Xiaohua Huang 0003, Wenming Zheng
IEEE Trans. Affect. Comput.4
2024 Micro-expression Recognition Based on Dual-Stream Spatiotemporal Transformer
Xiaohua Huang 0003, Chuangao Tang
ICPR (13)2
2023 EMC²A-Net: An Efficient Multibranch Cross-Channel Attention Network for SAR Target Classification
abstract
In recent years, convolutional neural networks (CNNs) have demonstrated significant potential for synthetic aperture radar (SAR) target recognition. SAR images possess a strong sense of granularity and contain texture features of varying scales, including speckle noise, dominant scatterers, and target contours, which are not typically considered in traditional CNN models. This article proposes two residual blocks, termed multibranch cross-channel attention (EMC2A) blocks, with multiscale receptive fields (RFs) based on a multibranch structure and designs an efficient isotopic architecture deep CNN (DCNN) called EMC2A-Net, whose structure is interpretable from a probability and mathematical statistics perspective. EMC2A blocks employ parallel dilated convolution with different dilation rates to effectively capture multiscale contextual features without significantly increasing the computational load. To further enhance the efficiency of multiscale feature fusion, this article presented a multiscale feature cross-channel attention module, known as the EMC2A module, which adopts a local multiscale feature interaction strategy without dimensionality reduction. This strategy adaptively adjusts the weights of each channel using efficient one-dimensional (1-D)-circular convolution and sigmoid function to guide attention at the global channel-wise level. Comparative results on the moving and stationary target acquisition and recognition (MSTAR) dataset demonstrate that EMC2A-Net outperforms the other available models of the same type and possesses a relatively lightweight network structure. The ablation experimental results further demonstrate that the EMC2A module significantly enhances the model’s performance by utilizing only a few parameters and appropriate cross-channel interactions.
Zhe Geng, Xiaohua Huang 0003, Qinglu Wang, Daiyin Zhu
IEEE Trans. Geosci. Remote. Sens.4
2022 Leaders and Followers Identified by Emotional Mimicry During Collaborative Learning: A Facial Expression Recognition Study on Emotional Valence
abstract
This article explores the potential of emotional mimicry in identifying the leader and follower students in collaborative learning settings. Our data include video recorded interactions of 24 high school students who worked together in groups of three during a collaborative exam. A facial emotions recognition method was used to capture participants’ facial emotions during the collaborative work. Cross-recurrence quantification analysis was applied on the detected facial emotions to see the level and direction of emotional mimicry among the dyads in the same groups. In order to validate the cross-recurrence quantification analysis results, student interactions in terms of leading or following the task were video coded. Our findings showed that the leaders and followers identified by cross-recurrence quantification analysis findings matched the leaders and followers identified by the video coding in 70 percent of the dyadic interactions across the collaborating groups. The current findings show that video-based facial emotions recognition as a method can add to collaborative learning research, especially explaining some social, and affective dynamics about it. The study further discusses the possible variables that might confound the relationship between emotional mimicry and leader-follower interactions during collaboration.
Muhterem Dindar, Sanna Järvelä, Sara Ahola, Xiaohua Huang 0003, Guoying Zhao 0001
IEEE Trans. Affect. Comput.4
2022 Analyzing Group-Level Emotion with Global Alignment Kernel based Approach
abstract
From the perspective of social science, understanding group emotion has become increasingly important for teams to considerably accomplish organizational work. Currently, automatically analyzing the perceived affect of a group of people has been received increasingly interest in affective computing community. The variability in group size makes difficulty for group-level emotion recognition to straightforwardly measure the feature distance of two group-level images. Recent works attempted to resolve the preceding problem by using feature encoding. However, the early works lack of efficiency. To alleviate this problem, this article aims to design a new method to effectively analyze the group behavior from a group-level image. Motivated by time-series kernel approaches explored in dynamic facial expression classification, this article mainly concentrates on global alignment kernel and design support vector machine with the combined global alignment kernels (SVM-CGAK) to better recognize group-level emotion. Specifically, we first propose to use global alignment kernel to explicitly measure the distance of two group-level images. For improving the performance of global alignment kernel, we use the global weight sort scheme based on their spatial relation information to sort the faces from group-level image, making an efficient data structure to the global alignment kernel. With this new global alignment kernel, we construct the backbone of SVM-CGAK, namely, support vector machine with global alignment kernel. Furthermore, considering the challenging environment, we construct two global alignment kernels based on Reisz-based Volume Local Binary Pattern and deep convolutional neural network features, respectively. Lastly, to make the robustness of group-level emotion recognition, we propose SVM-CGAK combining both global alignment kernels with multiple kernel learning approach. It can enhance the discriminative ability of each global alignment kernel. Intensive experiments are conducted on three challenging group-level emotion databases. The experimental results demonstrate that the proposed approach achieves promising performance for group-level emotion recognition compared with the recent state-of-the-art methods.
Xiaohua Huang 0003, Abhinav Dhall, Roland Göcke, Matti Pietikäinen, Guoying Zhao 0001
IEEE Trans. Affect. Comput.1
2022 Objective Class-Based Micro-Expression Recognition Under Partial Occlusion Via Region-Inspired Relation Reasoning Network
abstract
Micro-expression recognition (MER) has attracted the attention of many researchers in the past decade. However, occlusion occurs for MER in real-world scenarios. In this paper, a challenging issue in MER that is interesting but unexplored, i.e., occlusion MER, is deeply investigated. First, to research MER under real-world occlusion conditions, synthetic occluded microexpression databases are created by using various community masks. Second, to suppress the influence of occlusion, aRegion-inspiredRelationReasoningNetwork (RRRN) is proposed to model the relations between various facial regions. The RRRN consists of a backbone network, a region-inspired (RI) module and a relation reasoning (RR) module. More specifically, the backbone network aims to extract feature representations from different facial regions, the RI module is designed to compute the adaptive weight from the facial region itself based on the unobstructedness and importance of the region for suppressing the influence of occlusion using an attention mechanism, and the RR module exploits the progressive interactions among these regions by performing graph convolutions. Experiments are conducted on two tasks of MEGC 2018: the holdout-database evaluation task and the composite database evaluation task. Experimental results show that RRRN can be utilized to significantly explore the importance of facial regions and capture the cooperative complementary relationship of facial regions for MER. The results also demonstrate that RRRN outperforms the state-of-the-art approaches, especially with respect to occlusion, where RRRN is more robust.
Qirong Mao, Ling Zhou 0005, Wenming Zheng, Xiuyan Shao, Xiaohua Huang 0003
IEEE Trans. Affect. Comput.5
2021 Micro-expression action unit detection with spatial and channel attention
abstract
Action Unit (AU) detection plays an important role in facial behaviour analysis. In the literature, AU detection has extensive researches in macro-expressions. However, to the best of our knowledge, there is limited research about AU analysis for micro-expressions. In this paper, we focus on AU detection in micro-expressions. Due to the small quantity and low intensity of micro-expression databases, micro-expression AU detection becomes challenging. To alleviate these problems, in this work, we propose a novel micro-expression AU detection method by utilizing self high-order statistics of spatio-wise and channel-wise features which can be considered as spatial and channel attentions, respectively. Through such spatial attention module, we expect to utilize rich relationship information of facial regions to increase the AU detection robustness on limited micro-expression samples. In addition, considering the low intensity of micro-expression AUs, we further propose to explore high-order statistics for better capturing subtle regional changes on face to obtain more discriminative AU features. Intensive experiments show that our proposed approach outperforms the basic framework by 0.0859 on CASME II, 0.0485 on CASME, and 0.0644 on SAMM in terms of the average F1-score.
Yante Li, Xiaohua Huang 0003, Guoying Zhao 0001
Neurocomputing2
2021 Editorial for the special issue of IMAVIS on automatic face analytics for human behavior understanding
Xiaohua Huang 0003, Abhinav Dhall, Guoying Zhao 0001, Wenming Zheng, Matti Pietikäinen
Image Vis. Comput.1
2021 Joint Local and Global Information Learning With Single Apex Frame Detection for Micro-Expression Recognition
abstract
Micro-expressions (MEs) are rapid and subtle facial movements that are difficult to detect and recognize. Most recent works have attempted to recognize MEs with spatial and temporal information from video clips. According to psychological studies, the apex frame conveys the most emotional information expressed in facial expressions. However, it is not clear how the single apex frame contributes to micro-expression recognition. To alleviate that problem, this paper firstly proposes a new method to detect the apex frame by estimating pixel-level change rates in the frequency domain. With frequency information, it performs more effectively on apex frame spotting than the currently existing apex frame spotting methods based on the spatio-temporal change information. Secondly, with the apex frame, this paper proposes a joint feature learning architecture coupling local and global information to recognize MEs, because not all regions make the same contribution to ME recognition and some regions do not even contain any emotional information. More specifically, the proposed model involves the local information learned from the facial regions contributing major emotion information, and the global information learned from the whole face. Leveraging the local and global information enables our model to learn discriminative ME representations and suppress the negative influence of unrelated regions to MEs. The proposed method is extensively evaluated using CASME, CASME II, SAMM, SMIC, and composite databases. Experimental results demonstrate that our method with the detected apex frame achieves considerably promising ME recognition performance, compared with the state-of-the-art methods employing the whole ME sequence. Moreover, the results indicate that the apex frame can significantly contribute to micro-expression recognition.
Yante Li, Xiaohua Huang 0003, Guoying Zhao 0001
IEEE Trans. Image Process.2
2020 NLWSNet: a weakly supervised network for visual sentiment analysis in mislabeled web images
abstract
Large-scale datasets are driving the rapid developments of deep convolutional neural networks for visual sentiment analysis. However, the annotation of large-scale datasets is expensive and time consuming. Instead, it is easy to obtain weakly labeled web images from the Internet. However, noisy labels still lead to seriously degraded performance when we use images directly from the web for training networks. To address this drawback, we propose an end-to-end weakly supervised learning network, which is robust to mislabeled web images. Specifically, the proposed attention module automatically eliminates the distraction of those samples with incorrect labels by reducing their attention scores in the training process. On the other hand, the special-class activation map module is designed to stimulate the network by focusing on the significant regions from the samples with correct labels in a weakly supervised learning approach. Besides the process of feature learning, applying regularization to the classifier is considered to minimize the distance of those samples within the same class and maximize the distance between different class centroids. Quantitative and qualitative evaluations on well- and mislabeled web image datasets demonstrate that the proposed algorithm outperforms the related methods.
Luoyang Xue, Qirong Mao, Xiaohua Huang 0003, Jie Chen 0069
Frontiers Inf. Technol. Electron. Eng.3
2019 Discriminative Spatiotemporal Local Binary Pattern with Revisited Integral Projection for Spontaneous Facial Micro-Expression Recognition
abstract
Recently, there have been increasing interests in inferring mirco-expression from facial image sequences. Due to subtle facial movement of micro-expressions, feature extraction has become an important and critical issue for spontaneous facial micro-expression recognition. Recent works used spatiotemporal local binary pattern (STLBP) for micro-expression recognition and considered dynamic texture information to represent face images. However, they miss the shape attribute of face images. On the other hand, they extract the spatiotemporal features from the global face regions while ignore the discriminative information between two micro-expression classes. The above-mentioned problems seriously limit the application of STLBP to micro-expression recognition. In this paper, we propose a discriminative spatiotemporal local binary pattern based on an integral projection to resolve the problems of STLBP for micro-expression recognition. First, we revisit an integral projection for preserving the shape attribute of micro-expressions by using robust principal component analysis. Furthermore, a revisited integral projection is incorporated with local binary pattern across spatial and temporal domains. Specifically, we extract the novel spatiotemporal features incorporating shape attributes into spatiotemporal texture features. For increasing the discrimination of micro-expressions, we propose a new feature selection based on Laplacian method to extract the discriminative information for facial micro-expression recognition. Intensive experiments are conducted on three availably published micro-expression databases including CASME, CASME2 and SMIC databases. We compare our method with the state-of-the-art algorithms. Experimental results demonstrate that our proposed method achieves promising performance for micro-expression recognition.
Xiaohua Huang 0003, Xin Liu 0012, Guoying Zhao 0001, Xiaoyi Feng, Matti Pietikäinen
IEEE Trans. Affect. Comput.1
2018 Visual Tracking Based on Cooperative Model
abstract
In this paper, we propose a cooperative model combined the multi-task reverse sparse representation model (MTRSR) and the AdaBoost classifier, which were used to cope with the disturbing of target gradient information caused by motion blur or target serious occlusion, and a descriptive dictionary were used to estimate the weights of each candidates. First, we use the MTRSR model to get the blur kernel which were used to get the blur target template set, meanwhile the confidence of the candidates is also obtained by the reconstruction error. Then we use the HOG features of the target templates to get the descriptive dictionary to calculate the weights of the candidates, and a AdaBoost classifier is used to calculate the confidences of all candidates. Finally, the best target is retrieved by the sum of production of weight value and the two confidences. The experimental data show that the proposed algorithm can fully cope with the target's information change which were caused by motion blur and target occlusion in the complex scene, and our algorithm can further improve the accuracy and robustness in visual tracking.
Bobin Zhang, Weidong Fang 0002, Wei Chen 0036, Fangming Bi, Chaogang Tang, Xiaohua Huang 0003
FG6
2018 Can Micro-Expression be Recognized Based on Single Apex Frame?
abstract
Micro-expressions are rapid and subtle facial movements such that they are difficult to detect and recognize.Most of recent works have attempted to recognize micro-expression by using the spatial and dynamic information from the video clip.Physiological studies have demonstrated that the apex frame can convey the most emotion expressed in facial expression.It may be reasonable to use apex frame for improving micro-expression recognition.However, it is wonder how much apex frames contribute to micro-expression recognition.In this paper, we primarily focus on resolving the contribution-level by using apex frame for micro-expression recognition.Firstly, we propose a new method to detect the apex frame in frequency domain, as it is found that apex frame has very correlated relationship with the amplitude change in frequency domain.Secondly, we propose to use deep convolutional neural network (DCNN) on apex frame to recognize micro-expression.Intensive experimental results on CASME II database shows that our method has achieved considerably improvement compared with the state-of-the-art methods in micro-expression recognition.These results also demonstrate that apex frame can express the major emotion in micro-expression.
Yante Li, Xiaohua Huang 0003, Guoying Zhao 0001
ICIP2
2018 Micro-expression recognition with small sample size by transferring long-term convolutional neural network
Bing-Jun Li, Yong-Jin Liu 0001, Wen-Jing Yan, Xinyu Ou, Xiaohua Huang 0003, Xiaolan Fu
Neurocomputing6
2018 Towards Reading Hidden Emotions: A Comparative Study of Spontaneous Micro-Expression Spotting and Recognition Methods
abstract
Micro-expressions (MEs) are rapid, involuntary facial expressions which reveal emotions that people do not intend to show. Studying MEs is valuable as recognizing them has many important applications, particularly in forensic science and psychotherapy. However, analyzing spontaneous MEs is very challenging due to their short duration and low intensity. Automatic ME analysis includes two tasks: ME spotting and ME recognition. For ME spotting, previous studies have focused on posed rather than spontaneous videos. For ME recognition, the performance of previous studies is low. To address these challenges, we make the following contributions: (i) We propose the first method for spotting spontaneous MEs in long videos (by exploiting feature difference contrast). This method is training free and works on arbitrary unseen videos. (ii) We present an advanced ME recognition framework, which outperforms previous work by a large margin on two challenging spontaneous ME databases (SMIC and CASMEII). (iii) We propose the first automatic ME analysis system (MESR), which can spot and recognize MEs from spontaneous video data. Finally, we show our method outperforms humans in the ME recognition task by a large margin, and achieves comparable performance to humans at the very challenging task of spotting and then recognizing spontaneous MEs.
Xiaopeng Hong, Antti Moilanen, Xiaohua Huang 0003, Tomas Pfister, Guoying Zhao 0001, Matti Pietikäinen
IEEE Trans. Affect. Comput.4
2018 Background Subtraction Using Spatio-Temporal Group Sparsity Recovery
abstract
Background subtraction is a key step in a wide spectrum of video applications, such as object tracking and human behavior analysis. Compressive sensing-based methods, which make little specific assumptions about the background, have recently attracted wide attention in background subtraction. Within the framework of compressive sensing, background subtraction is solved as a decomposition and optimization problem, where the foreground is typically modeled as pixel-wised sparse outliers. However, in real videos, foreground pixels are often not randomly distributed, but instead, group clustered. Moreover, due to costly computational expenses, most compressive sensing-based methods are unable to process frames online. In this paper, we take into account the group properties of foreground signals in both spatial and temporal domains, and propose a greedy pursuit-based method called spatio-temporal group sparsity recovery, which prunes data residues in an iterative process, according to both sparsity and group clustering priors, rather than merely sparsity. Furthermore, a random strategy for background dictionary learning is used to handle complex background variations, while foreground-free training is not required. Finally, we propose a two-pass framework to achieve online processing. The proposed method is validated on multiple challenging video sequences. Experiments demonstrate that our approach effectively works on a wide range of complex scenarios and achieves a state-of-the-art performance with far fewer computations.
Xin Liu 0012, Jiawen Yao, Xiaopeng Hong, Xiaohua Huang 0003, Ziheng Zhou 0003, Chun Qi, Guoying Zhao 0001
IEEE Trans. Circuits Syst. Video Technol.4
2018 Domain Regeneration for Cross-Database Micro-Expression Recognition
abstract
Recently, micro-expression recognition has attracted lots of researchers' attention due to its potential value in many practical applications, e.g., lie detection. In this paper, we investigate an interesting and challenging problem in micro-expression recognition, i.e., cross-database micro-expression recognition, in which the training and testing samples come from different micro-expression databases. Under this problem setting, the consistent feature distribution between the training and testing samples originally existing in conventional micro-expression recognition would be seriously broken and hence the performance of most current well-performing micro-expression recognition methods may sharply drop. In order to overcome it, we propose a simple yet effective framework called Domain Regeneration (DR) in this paper. DR framework aims at learning a domain regenerator to regenerate the micro-expression samples from source and target databases respectively such that they can abide by the same or similar feature distributions. Thus, we are able to use the classifier learned based on the labeled source micro-expression samples to predict the label information of the unlabeled target micro-expression samples. To evaluate the proposed DR framework, we conduct extensive cross-database micro-expression recognition experiments designed based on SMIC and CASME II databases. Experimental results show that compared with recent state-of-the-art cross-database emotion recognition methods, the proposed DR framework has more promising performance.
Yuan Zong, Wenming Zheng, Xiaohua Huang 0003, Jingang Shi, Zhen Cui 0001, Guoying Zhao 0001
IEEE Trans. Image Process.3
2018 Multimodal Framework for Analyzing the Affect of a Group of People
abstract
With the advances in multimedia and the world wide web, users upload millions of images and videos everyone on social networking platforms on the Internet. From the perspective of automatic human behavior understanding, it is of interest to analyze and model the affects that are exhibited by groups of people who are participating in social events in these images. However, the analysis of the affect that is expressed by multiple people is challenging due to the varied indoor and outdoor settings. Recently, a few interesting works have investigated face-based group-level emotion recognition (GER). In this paper, we propose a multimodal framework for enhancing the affective analysis ability of GER in challenging environments. Specifically, for encoding a person's information in a group-level image, we first propose an information aggregation method for generating feature descriptions of face, upper body, and scene. Later, we revisit localized multiple kernel learning for fusing face, upper body, and scene information for GER against challenging environments. Intensive experiments are performed on two challenging group-level emotion databases (HAPPEI and GAFF) to investigate the roles of the face, upper body, scene information, and the multimodal framework. Experimental results demonstrate that the multimodal framework achieves promising performance for GER.
Xiaohua Huang 0003, Abhinav Dhall, Roland Göcke, Matti Pietikäinen, Guoying Zhao 0001
IEEE Trans. Multim.1
2018 Learning From Hierarchical Spatiotemporal Descriptors for Micro-Expression Recognition
abstract
Micro-expression recognition aims to infer genuine emotions that people try to conceal from facial video clips. It is a very challenging task because micro-expressions have a very low intensity and short duration, which makes micro-expressions difficult to observe. Recently, researchers have designed various spatiotemporal descriptors to describe micro-expressions. It is notable that for better capturing the low-intensity facial muscle movement, a fixed spatial division grid, 8× 8 for example, is commonly used to partition the facial images into a few facial blocks before extracting descriptors. However, it is hard to choose an ideal division grid for different micro-expression samples because the division grids affect the discriminative ability of spatiotemporal descriptors to distinguish micro-expressions. To address this problem, in this paper, we design a hierarchical spatial division scheme for spatiotemporal descriptor extraction. By using the proposed scheme, it would not be a problem to determine which division grid is most suitable regarding different micro-expression samples. Furthermore, we propose a kernelized group sparse learning (KGSL) model to process hierarchical scheme based spatiotemporal descriptors such that they are more effective for micro-expression recognition tasks. To evaluate the performance of the proposed micro-expression recognition method consisting of the hierarchical scheme based spatiotemporal descriptors and KGSL, extensive experiments are conducted on two public micro-expression databases: CASME II and SMIC. Compared with many recent state-of-the-art approaches, our method achieves more promising recognition results.
Yuan Zong, Xiaohua Huang 0003, Wenming Zheng, Zhen Cui 0001, Guoying Zhao 0001
IEEE Trans. Multim.2
2017 Image denoising via group sparsity residual constraint
abstract
Group sparsity has shown great potential in various low-level vision tasks (e.g, image denoising, deblurring and inpainting). In this paper, we propose a new prior model for image denoising via group sparsity residual constraint (GSRC). To enhance the performance of group sparse-based image denoising, the concept of group sparsity residual is proposed, and thus, the problem of image denoising is translated into one that reduces the group sparsity residual. To reduce the residual, we first obtain some good estimation of the group sparse coefficients of the original image by the first-pass estimation of noisy image, and then centralize the group sparse coefficients of noisy image to the estimation. Experimental results have demonstrated that the proposed method not only outperforms many state-of-the-art denoising methods such as BM3D and WNNM, but results in a faster speed.
Zhiyuan Zha, Xin Liu 0012, Ziheng Zhou 0003, Xiaohua Huang 0003, Jingang Shi, Zhenhong Shang, Lan Tang, Yechao Bai, Qiong Wang 0002, Xinggan Zhang
ICASSP4
2017 Analyzing the group sparsity based on the rank minimization methods
abstract
Sparse coding has achieved a great success in various image processing studies. However, there is not any benchmark to measure the sparsity of image patch/group because sparse discriminant conditions cannot keep unchanged. This paper analyzes the sparsity of group based on the strategy of the rank minimization. Firstly, an adaptive dictionary for each group is designed. Then, we prove that group-based sparse coding is equivalent to the rank minimization problem, and thus the sparse coefficients of each group are measured by estimating the singular values of each group. Based on that measurement, the weighted Schatten p-norm minimization (WSNM) has been found to be the closest solution to the real singular values of each group. Thus, WSNM can be equivalently transformed into a non-convex ℓp-norm minimization problem in group-based sparse coding. Experimental results on two applications: image in painting and image compressive sensing (CS) recovery show that the proposed scheme outperforms many state-of-the-art methods.
Zhiyuan Zha, Xin Liu 0012, Xiaohua Huang 0003, Henglin Shi, Yingyue Xu, Qiong Wang 0002, Lan Tang, Xinggan Zhang
ICME3
2017 Learning a Target Sample Re-Generator for Cross-Database Micro-Expression Recognition
abstract
In this paper, we investigate the cross-database micro-expression recognition problem, where the training and testing samples are from two different micro-expression databases. Under this setting, the training and testing samples would have different feature distributions and hence the performance of most existing micro-expression recognition methods may decrease greatly. To solve this problem, we propose a simple yet effective method called Target Sample Re-Generator (TSRG) in this paper. By using TSRG, we are able to re-generate the samples from target micro-expression database and the re-generated target samples would share same or similar feature distributions with the original source samples. For this reason, we can then use the classifier learned based on the labeled source samples to accurately predict the micro-expression categories of the unlabeled target samples. To evaluate the performance of the proposed TSRG method, extensive cross-database micro-expression recognition experiments designed based on SMIC and CASME II databases are conducted. Compared with recent state-of-the-art cross-database emotion recognition methods, the proposed TSRG achieves more promising results.
Yuan Zong, Xiaohua Huang 0003, Wenming Zheng, Zhen Cui 0001, Guoying Zhao 0001
ACM Multimedia2
2016 Multi-modal emotion analysis from facial expressions and electroencephalogram
Xiaohua Huang 0003, Jukka Kortelainen, Guoying Zhao 0001, Antti Moilanen, Tapio Seppänen, Matti Pietikäinen
Comput. Vis. Image Underst.1
2016 Spontaneous facial micro-expression analysis using Spatiotemporal Completed Local Quantized Patterns
Xiaohua Huang 0003, Guoying Zhao 0001, Xiaopeng Hong, Wenming Zheng, Matti Pietikäinen
Neurocomputing1
2016 Cross-Corpus Speech Emotion Recognition Based on Domain-Adaptive Least-Squares Regression
abstract
In this letter, a novel cross-corpus speech emotion recognition (SER) method using domain-adaptive least-squares regression (DaLSR) model is proposed. In this method, an additional unlabeled data set from target speech corpus is used to serve as an auxiliary data set and combined with the labeled training data set from source speech corpus for jointly training the DaLSR model. In contrast to the traditional least-squares regression (LSR) method, the major novelty of DaLSR is that it is able to handle the mismatch problem between source and target speech corpora. Hence, the proposed DaLSR method is very suitable for coping with cross-corpus SER problem. For evaluating the performance of the proposed method in dealing with the cross-corpus SER problem, we conduct extensive experiments on three emotional speech corpora and compare the results with several state-of-the-art transfer learning methods that are widely used for cross-corpus SER problem. The experimental results show that the proposed method achieves better recognition accuracies than the state-of-the-art methods.
Yuan Zong, Wenming Zheng, Tong Zhang 0021, Xiaohua Huang 0003
IEEE Signal Process. Lett.4
2015 Riesz-based Volume Local Binary Pattern and A Novel Group Expression Model for Group Happiness Intensity Analysis
abstract
Automatic emotion analysis and understanding has received much attention over the years in affective computing. Recently, there are increasing interests in inferring the emotional intensity of a group of people. For group emotional intensity analysis, feature extraction and group expression model are two critical issues. In this paper, we propose a new method to estimate the happiness intensity of a group of people in an image. Firstly, we combine the Riesz transform and the local binary pattern descriptor, named Riesz-based volume local binary pattern, which considers neighbouring changes not only in the spatial domain of a face but also along the different Riesz faces. Secondly, we exploit the continuous conditional random fields for constructing a new group expression model, which considers global and local attributes. Intensive experiments are performed on three challenging facial expression databases to evaluate the novel feature. Furthermore, experiments are conducted on the HAPPEI database to evaluate the new group expression model with the new feature. Our experimental results demonstrate the promising performance for group happiness intensity analysis.
Xiaohua Huang 0003, Abhinav Dhall, Guoying Zhao 0001, Roland Göcke, Matti Pietikäinen
BMVC1
2015 Transductive Transfer LDA with Riesz-based Volume LBP for Emotion Recognition in The Wild
abstract
In this paper, we propose the method using Transductive Transfer Linear Discriminant Analysis (TTLDA) and Riesz-based Volume Local Binary Patterns (RVLBP) for image based static facial expression recognition challenge of the Emotion Recognition in the Wild Challenge (EmotiW 2015). The task of this challenge is to assign facial expression labels to frames of some movies containing a face under the real word environment. In our method, we firstly employ a multi-scale image partition scheme to divide each face image into some image blocks and use RVLBP features extracted from each block to describe each facial image. Then, we adopt the TTLDA approach based on RVLBP to cope with the expression recognition task. The experiments on the testing data of SFEW 2.0 database, which is used for image based static facial expression challenge, demonstrate that our method achieves the accuracy of 50%. This result has a 10.87% improvement over the baseline provided by this challenge organizer.
Yuan Zong, Wenming Zheng, Xiaohua Huang 0003, Jingwei Yan, Tong Zhang 0021
ICMI3
2014 Improved Spatiotemporal Local Monogenic Binary Pattern for Emotion Recognition in The Wild
abstract
Local binary pattern from three orthogonal planes (LBP-TOP) has been widely used in emotion recognition in the wild. However, it suffers from illumination and pose changes. This paper mainly focuses on the robustness of LBP-TOP to unconstrained environment. Recent proposed method, spatiotemporal local monogenic binary pattern (STLMBP), was verified to work promisingly in different illumination conditions. Thus this paper proposes an improved spatiotemporal feature descriptor based on STLMBP. The improved descriptor uses not only magnitude and orientation, but also the phase information, which provide complementary information. In detail, the magnitude, orientation and phase images are obtained by using an effective monogenic filter, and multiple feature vectors are finally fused by multiple kernel learning. STLMBP and the proposed method are evaluated in the Acted Facial Expression in the Wild as part of the 2014 Emotion Recognition in the Wild Challenge. They achieve competitive results, with an accuracy gain of 6.35% and 7.65% above the challenge baseline (LBP-TOP) over video.
Xiaohua Huang 0003, Qiuhai He, Xiaopeng Hong, Guoying Zhao 0001, Matti Pietikäinen
ICMI1
2014 Robust Facial Expression Recognition Using Revised Canonical Correlation
abstract
The poor alignment and large variations in the temporal sale of facial expressions are two crucial problems for facial expression recognition (FER). Canonical correlation (CC) has recently received increasing attention in surveillance face recognition because of its robustness to variations of alignment. But it could not suit well to FER, because the facial expression variations and temporal information are ignored for canonical subspace. This paper proposes the revised canonical correlation method to address the two above described issues for making FER be robust to false detection or mis-alignment. Firstly, this paper presents the local binary pattern to describe the appearance features for enhancing the spatial variations of facial expression. Secondly, this paper proposes the temporal orthogonal locality preserved projection for building a canonical subspace of a video clip, where it mostly captures the motion changes of facial expressions. Then this paper presents the discriminative CC to model the low-dimensional feature space, which increases robustness to imprecise alignment and strengthens discrimination for facial expressions. Extensive experimental results on Extended Cohn-Kanade and MAHNOB-HCI databases demonstrate that the proposed method achieves the best results in recognizing facial expressions and performs robustly with ordinary on general face detection and eye detection.
Xiaohua Huang 0003, Guoying Zhao 0001, Matti Pietikäinen, Wenming Zheng
ICPR1
2013 Emotion recognition from facial images with arbitrary views
abstract
Facial expression recognition has been predominantly utilized to analyze the emotional status of human beings. In practice nearly frontal-view facial images may not be available. Therefore, a desirable property of facial expression recognition would allow the user to have any head pose. Some methods on non-frontal-view facial images were recently proposed to recognize the facial expressions by building discriminative subspace in specific views. We argue that this kind of approach ignores (1) the discrimination of inter-class samples with the same view label and (2) the closeness of intra-class samples with all view labels. This paper proposes a new method to recognize arbitrary-view facial expressions by using discriminative neighborhood preserving embedding and multi-view concepts. It first captures the discriminative property of inter-class samples. In addition, it explores the closeness of intra-class samples with arbitrary view in a low-dimensional subspace. Experimental results on BU-3DFE and Multi-PIE databases show that our approach achieves promising results for recognizing facial expressions with arbitrary views.
Xiaohua Huang 0003, Guoying Zhao 0001, Matti Pietikäinen
BMVC1
2012 Towards a dynamic expression recognition system under facial occlusion
Xiaohua Huang 0003, Guoying Zhao 0001, Wenming Zheng, Matti Pietikäinen
Pattern Recognit. Lett.1
2012 Spatiotemporal Local Monogenic Binary Patterns for Facial Expression Recognition
abstract
Feature representation is an important research topic in facial expression recognition from video sequences. In this letter, we propose to use spatiotemporal monogenic binary patterns to describe both appearance and motion information of the dynamic sequences. Firstly, we use monogenic signals analysis to extract the magnitude, the real picture and the imaginary picture of the orientation of each frame, since the magnitude can provide much appearance information and the orientation can provide complementary information. Secondly, the phase-quadrant encoding method and the local bit exclusive operator are utilized to encode the real and imaginary pictures from orientation in three orthogonal planes, and the local binary pattern operator is used to capture the texture and motion information from the magnitude through three orthogonal planes. Finally, both concatenation method and multiple kernel learning method are respectively exploited to handle the feature fusion. The experimental results on the Extended Cohn-Kanade and Oulu-CASIA facial expression databases demonstrate that the proposed methods perform better than the state-of-the-art methods, and are robust to illumination variations.
Xiaohua Huang 0003, Guoying Zhao 0001, Wenming Zheng, Matti Pietikäinen
IEEE Signal Process. Lett.1
2011 Facial expression recognition from near-infrared videos
Guoying Zhao 0001, Xiaohua Huang 0003, Matti Taini, Stan Z. Li, Matti Pietikäinen
Image Vis. Comput.2
2010 Dynamic Facial Expression Recognition Using Boosted Component-Based Spatiotemporal Features and Multi-classifier Fusion
Xiaohua Huang 0003, Guoying Zhao 0001, Matti Pietikäinen, Wenming Zheng
ACIVS (2)1