EDBT 2026 Demo / reviewers in the wild / expert
Jiabei Zeng
dblp:128/1501
· DBLP profile ↗
35ranked-venue papers
6as first author
17since 2021 · last 2026
0000-0003-3256-4524ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 24 · 4 first-author · 9 since 2021Artificial intelligence and machine learning · 22 · 4 first-author · 12 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Task-aware pre-training for facial action analysis
Xuran Sun, Jiabei Zeng, Shiguang Shan |
Frontiers Comput. Sci. | 2 |
| 2026 | Unsupervised Gaze Representation Learning by Switching FeaturesabstractIt is prevalent to leverage unlabeled data to train deep learning models when it is difficult to collect large-scale annotated datasets. However, for 3D gaze estimation, most existing unsupervised learning methods face challenges in distinguishing subtle gaze-relevant information from dominant gaze-irrelevant information. To address this issue, we propose an unsupervised learning framework to disentangle the gaze-relevant and the gaze-irrelevant information, by seeking the shared information of a pair of input images with the same gaze and with the same eye respectively. Specifically, given two images, the framework finds their shared information by first encoding the images into two latent features via two encoders and then switching part of the features before feeding them to the decoders for image reconstruction. We theoretically prove that the proposed framework is able to encode different information into different parts of the latent feature if we properly select the training image pairs and their shared information. Based on the framework, we derive Cross-Encoder and Cross-Encoder++ to learn gaze representation from the eye images and face images, respectively. Experiments on public gaze datasets demonstrate that the Cross-Encoder and Cross-Encoder++ outperform the competitive methods. The ablation study quantitatively and qualitatively shows that the gaze feature is successfully extracted. Yunjia Sun, Jiabei Zeng, Shiguang Shan, Xilin Chen 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2026 | CoSI-Gaze: Context-Spatial Integration for gaze target detection and social gaze predictionabstractUnderstanding gaze behavior is a fundamental aspect of human social perception and a challenging problem in computer vision. Social gaze encompass not only where a person is looking but also how their gaze functions within a social context to establish connections, regulate conversations, and convey meaning. Existing social gaze analysis methods either primarily focus on detecting the spatial location of gaze targets while overlooking contextual cues, or rely exclusively on semantic representations, neglecting the spatial consistency between gaze targets and social gaze patterns. To bridge this gap, we propose CoSI-Gaze(CoSI), a Context-Spatial Integration framework for joint gaze target detection and social gaze prediction. To balance the contributions of contextual information and spatial consistency, CoSI estimates the reliability of gaze target predictions and adaptively adjusts this reliability in the integration process. To evaluate the framework’s ability to understand social gaze behaviors, we introduce DyGaze, the first dataset of dyadic interactions annotated with both gaze targets and five social gaze patterns (mutual, shared, single, miss, and void). Extensive experiments demonstrate that CoSI achieves state-of-the-art performance across DyGaze and other gaze pattern prediction benchmarks. Fei Chang, Jiabei Zeng, Dongmei Jiang, Shiguang Shan |
Pattern Recognit. | 2 |
| 2026 | Facial Action Units Generation via Cross-Modality Attention Fusion and Calibrated Denoising
Chenyue Liang, Zhenliang He, Jiabei Zeng, Dongmei Jiang, Shiguang Shan |
IEEE Signal Process. Lett. | 3 |
| 2025 | Exp-VQA: Fine-grained facial expression analysis via visual question answering
Yujian Yuan, Jiabei Zeng, Shiguang Shan |
Pattern Recognit. | 2 |
| 2025 | Multi-View Facial Expressions Analysis of Autistic Children in Social PlayabstractAtypical facial expressions during interaction are among the early symptoms of autism spectrum disorder (ASD) and are included in standard diagnostic assessments. However, current methods rely on subjective human judgments, introducing bias and limiting objectivity. This paper proposes an automated framework for objective and quantitative assessment of autistic children's facial expressions during social play. Initially, we utilize four synchronized cameras to record interactions between ASD children and teachers during structured activities dominated by the teacher. To address challenges posed by head movements and occluded faces, we introduce a multi-view facial expression recognition strategy. Its effectiveness is demonstrated by experiments in real-world applications. To quantify the patterns of affect status and the dynamic complexity of facial expressions, we use the temporally accumulated distribution of the basic facial expressions and the multi-dimensional multiscale entropy of the facial expression sequence. Analysis of these features revealed significant differences between ASD and TD groups. Experimental results, derived from our quantified features, confirm conclusions drawn from previous research and experiential observations. With these facial expression features, ASD and typically developing (TD) children are accurately classified (accuracy 92.1%, precision 94.4%, 89.5% sensitivity, 94.7% specificity) in empirical experiments, suggesting the potential of our framework for improved ASD assessment. Jiabei Zeng, Yujian Yuan, Lu Qu, Fei Chang, Xuran Sun, Jinqiuyu Gong, Xuling Han, Qiaoyun Liu, Shiguang Shan, Xilin Chen 0001 |
IEEE Trans. Affect. Comput. | 1 |
| 2024 | Facial Action Unit Detection with the Semantic PromptabstractFacial action unit (AU) detection is an essential technique for fine-grained facial expression analysis. To improve the detection performance, the associations among different action units within the detection network should be exploited. In light of this, we propose to exploit the semantic corrections between AUs and improve the detection accuracy via a novel AU prompt framework. Specifically, we incorporate a pre-trained text encoder to extract the textual embeddings for AU descriptions. Then, we treat these embeddings as semantic prompts and feed them into a vision-language cross-attention module to capture the relations among AUs. The cross-attention module will adaptively aggregate the spatial features of a face image encoder, and finally generate discriminative features for each AU. Extensive experiments on BP4D, DISFA, and GFT datasets demonstrate that the proposed framework outperforms state-of-the-art methods in both within-dataset and cross-dataset settings. Chenyue Liang, Jiabei Zeng, Dongmei Jiang, Shiguang Shan |
ICME | 2 |
| 2024 | Gaze estimation with semi-supervised eye landmark detection as an auxiliary task
Yunjia Sun, Jiabei Zeng, Shiguang Shan |
Pattern Recognit. | 2 |
| 2023 | Describe Your Facial Expressions by Linking Image Encoders and Large Language Models
Yujian Yuan, Jiabei Zeng, Shiguang Shan |
BMVC | 2 |
| 2023 | Source-Free Adaptive Gaze Estimation by Uncertainty ReductionabstractGaze estimation across domains has been explored recently because the training data are usually collected under controlled conditions while the trained gaze estimators are used in nature and diverse environments. However, due to privacy and efficiency concerns, simultaneous access to annotated source data and to-be-predicted target data can be challenging. In light of this, we present an unsupervised source-free domain adaptation approach for gaze estimation, which adapts a source-trained gaze estimator to unlabeled target domains without source data. We propose the Uncertainty Reduction Gaze Adaptation (UnReGA) framework, which achieves adaptation by reducing both sample and model uncertainty. Sample uncertainty is mitigated by enhancing image quality and making them gaze-estimation-friendly, whereas model uncertainty is reduced by minimizing prediction variance on the same inputs. Extensive experiments are conducted on six cross-domain tasks, demonstrating the effectiveness of UnReGA and its components. Results show that UnReGA outperforms other state-of-the-art cross-domain gaze estimation methods under both protocols, with and without source data. The code is available at https://github.com/caixin1998/UnReGA. Jiabei Zeng, Shiguang Shan, Xilin Chen 0001 |
CVPR | 2 |
| 2023 | Gaze Pattern Recognition in Dyadic CommunicationabstractAnalyzing gaze behaviors is crucial to interpret the nature of communication. Current studies on gaze have focused primarily on the detection of a single pattern, such as the Looking-At-Each-Other pattern or the shared attention pattern. In this work, we re-define five static gaze patterns that cover all the status during a dyadic communication and propose a network to recognize these mutual exclusive gaze patterns given an image. We annotate a benchmark, called GP-Static, for the gaze pattern recognition task, on which our method experimentally outperforms other alternate solutions. Our method also achieves the state-of-art performance on other two single gaze pattern recognition tasks. The analysis of gaze patterns on preschool children demonstrates that the statistic of the proposed static gaze patterns conforms with the findings in psychology. Fei Chang, Jiabei Zeng, Qiaoyun Liu, Shiguang Shan |
ETRA | 2 |
| 2023 | Joint spatial and scale attention network for multi-view facial expression recognition
Yuanyuan Liu 0004, Jiyao Peng, Jiabei Zeng, Shiguang Shan |
Pattern Recognit. | 4 |
| 2022 | MAFW: A Large-scale, Multi-modal, Compound Affective Database for Dynamic Facial Expression Recognition in the WildabstractDynamic facial expression recognition (FER) databases provide important data support for affective computing and applications. However, most FER databases are annotated with several basic mutually exclusive emotional categories and contain only one modality, e.g., videos. The monotonous labels and modality cannot accurately imitate human emotions and fulfill applications in the real world. In this paper, we propose MAFW, a large-scale multi-modal compound affective database with 10,045 video-audio clips in the wild. Each clip is annotated with a compound emotional category and a couple of sentences that describe the subjects' affective behaviors in the clip. For the compound emotion annotation, each clip is categorized into one or more of the 11 widely-used emotions, i.e., anger, disgust, fear, happiness, neutral, sadness, surprise, contempt, anxiety, helplessness, and disappointment. To ensure high quality of the labels, we filter out the unreliable annotations by an Expectation Maximization (EM) algorithm, and then obtain 11 single-label emotion categories and 32 multi-label emotion categories. To the best of our knowledge, MAFW is the first in-the-wild multi-modal database annotated with compound emotion annotations and emotion-related captions. Additionally, we also propose a novel Transformer-based expression snippet feature learning method to recognize the compound emotions leveraging the expression-change relations among different emotions and modalities. Extensive experiments on MAFW database show the advantages of the proposed method over other state-of-the-art methods for both uni- and multi-modal FER. Our MAFW database is publicly available from https://mafw-database.github.io/MAFW. Yuanyuan Liu 0004, Chuanxu Feng, Wenbin Wang 0001, Guanghao Yin, Jiabei Zeng, Shiguang Shan |
ACM Multimedia | 6 |
| 2022 | Learning Representations for Facial Actions From Unlabeled VideosabstractFacial actions are usually encoded as anatomy-based action units (AUs), the labelling of which demands expertise and thus is time-consuming and expensive. To alleviate the labelling demand, we propose to leverage the large number of unlabelled videos by proposing a twin-cycle autoencoder (TAE) to learn discriminative representations for facial actions. TAE is inspired by the fact that facial actions are embedded in the pixel-wise displacements between two sequential face images (hereinafter, source and target) in the video. Therefore, learning the representations of facial actions can be achieved by learning the representations of the displacements. However, the displacements induced by facial actions are entangled with those induced by head motions. TAE is thus trained to disentangle the two kinds of movements by evaluating the quality of the synthesized images when either the facial actions or head pose is changed, aiming to reconstruct the target image. Experiments on AU detection show that TAE can achieve accuracy comparable to other existing AU detection methods including some supervised methods, thus validating the discriminant capacity of the representations learned by TAE. TAE's ability in decoupling the action-induced and pose-induced movements is also validated by visualizing the generated images and analyzing the facial image retrieval results qualitatively and quantitatively. Yong Li 0032, Jiabei Zeng, Shiguang Shan |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | Landmark-aware Self-supervised Eye Semantic SegmentationabstractLearning an accurate and robust eye semantic segmentation model generally requires enormous training data with delicate segmentation annotations. However, labeling the data is time-consuming and manpower-consuming. To address this issue, we propose to segment the eyes using unlabelled eye images and a weak empirical prior on the eye shape. To make the segmentation interpretable, we leverage the prior knowledge of eye shape by converting the self-supervised learned landmarks of each eye component to the segmentation maps. Specifically, we design a symmetrical auto-encoder architecture to learn disentangled representations of eye appearance and eye shape in a self-supervised manner. The eye shape is represented as the landmarks on the eyes. The proposed method encodes the eye images into the eye shapes and appearance features and then it reconstructs the image according to the eye shape and the appearance feature of another image. Since the landmarks of the training images are unknown, we require the generated landmarks' pictorial representations to have the same distribution as a known prior by minimizing an adversarial loss. Experiments on TEyeD and UnitySeg datasets demonstrate that the proposed self-supervised method is comparable with supervised ones. When the labeled data is insufficient, the proposed self-supervised method provides a better pre-trained model than other initialization methods. Jiabei Zeng, Shiguang Shan |
FG | 2 |
| 2021 | Emotion-aware Contrastive Learning for Facial Action Unit DetectionabstractCurrent AU datasets lack sufficiency and diversity because annotating facial action units (AUs) is laborious. The lack of labeled AU datasets bottlenecks the training of a discriminative AU detector. Compared with AUs, the basic emotional categories are relatively easy to annotate and they are highly correlated to AUs. To this end, we propose an Emotion-aware Contrastive Learning (EmoCo) framework to obtain representations that retain enough AU-related information. EmoCo leverages enormous and diverse facial images without AU annotations while labeled with the six universal facial expressions. EmoCo extends the prevalent self-supervised learning architecture of Momentum Contrast by simultaneously classifying the learned features into different emotional categories and distinguishing features within each emotional category in instance level. In the experiments, we train EmoCo using AffectNet dataset labeled with emotional categories. The EmoCo-learned features outperform other self-supervised learned representations in AU detection tasks on DISFA, BP4D, and GFT datasets. The EmoCo-pretrained models that fine-tuned on the AU datasets outperform most of the state-of-the-art AU detection methods. Xuran Sun, Jiabei Zeng, Shiguang Shan |
FG | 2 |
| 2021 | Cross-Encoder for Unsupervised Gaze Representation LearningabstractIn order to train 3D gaze estimators without too many annotations, we propose an unsupervised learning framework, Cross-Encoder, to leverage the unlabeled data to learn suitable representation for gaze estimation. To address the issue that the feature of gaze is always intertwined with the appearance of the eye, Cross-Encoder disentangles the features using a latent-code-swapping mechanism on eye-consistent image pairs and gaze-similar ones. Specifically, each image is encoded as a gaze feature and an eye feature. Cross-Encoder is trained to reconstruct each image in the eye-consistent pair according to its gaze feature and the other’s eye feature, but to reconstruct each image in the gaze-similar pair according to its eye feature and the other’s gaze feature. Experimental results show the validity of our work. First, using the Cross-Encoder-learned gaze representation, the gaze estimator trained with very few samples outperforms the ones using other unsupervised learning methods, under both within-dataset and cross-dataset protocol. Second, ResNet18 pretrained by Cross-Encoder is competitive with state-of-the-art gaze estimation methods. Third, ablation study shows that Cross-Encoder disentangles the gaze feature and eye feature. Yunjia Sun, Jiabei Zeng, Shiguang Shan, Xilin Chen 0001 |
ICCV | 2 |
| 2020 | Facial Expression Recognition for In-the-wild VideosabstractIn this paper, we propose a method for facial expression recognition for in-the-wild videos. Our method combines Deep Residual Network (ResNet) and Bidirectional Recurrent Neutral Network with Long-Short-Term Memory Unit (BLSTM). This method won the 2ndplace in the seven basic expression classification track of Affective Behavior Analysis in-the-wild Competition held in conjunction with the IEEE International Conference on Automatic Face and Gesture Recognition (FG) 2020, achieving 66.9% accuracy and 40.8% final metric on the test set. We also visualize the learned attention maps and analyze the importance of different regions in facial expression recognition. Jiabei Zeng, Shiguang Shan |
FG | 2 |
| 2020 | M3F: Multi-Modal Continuous Valence-Arousal Estimation in the WildabstractIn this paper, we propose a multi-modal multi-feature (M$^{3}F$) approach for in-the-wild valence-arousal estimation. In the proposed M$^{3}F$ framework, we fuse both visual features from videos and acoustic features from the audio tracks to estimate the valence and arousal. We follow a CNN-RNN paradigm, where the spatio-temporal visual features are extracted with a 3D convolutional network and/or a pretrained 2D convolutional network, and a bidirectional recurrent neural network. We evaluated the M$^{3}F$ framework on the validation set provided by the Affective Behavior Analysis in-the-wild (ABAW) Challenge, held in conjunction with the IEEE International Conference on Automatic Face and Gesture Recognition (FG) 2020, and it significantly outperforms the baseline method. Yuanhang Zhang 0001, Rulin Huang, Jiabei Zeng, Shiguang Shan |
FG | 3 |
| 2019 | Self-Supervised Representation Learning From Videos for Facial Action Unit DetectionabstractIn this paper, we aim to learn discriminative representation for facial action unit (AU) detection from large amount of videos without manual annotations. Inspired by the fact that facial actions are the movements of facial muscles, we depict the movements as the transformation between two face images in different frames and use it as the self-supervisory signal to learn the representations. However, under the uncontrolled condition, the transformation is caused by both facial actions and head motions. To remove the influence by head motions, we propose a Twin-Cycle Autoencoder (TCAE) that can disentangle the facial action related movements and the head motion related ones. Specifically, TCAE is trained to respectively change the facial actions and head poses of the source face to those of the target face. Our experiments validate TCAE's capability of decoupling the movements. Experimental results also demonstrate that the learned representation is discriminative for AU detection, where TCAE outperforms or is comparable with the state-of-the-art self-supervised learning methods and supervised AU detection methods. Yong Li 0032, Jiabei Zeng, Shiguang Shan, Xilin Chen 0001 |
CVPR | 2 |
| 2019 | Multi-Task Learning of Emotion Recognition and Facial Action Unit Detection with Adaptively Weights Sharing NetworkabstractEmotion recognition and facial action unit(AU) detection are the most two prevalent tasks in facial expression analysis. Since the two tasks are highly correlated, in this paper, we simultaneously do emotion recognition and AU detection in a multi-task learning framework to make the tasks benefit from each other. To achieve this, we propose an Adaptively Weights Sharing Network (AWS-Net) that automatically learns where and to what extent each task should borrow information from the other by placing an AWS-Unit after each layer-pair of the two tasks' networks. The proposed AWS-Net is end-to-end trainable on data that is merely annotated with emotions or AUs. Experimental results on several facial expression recognition(FER) datasets demonstrate that AWS-Net improves the performance of both single-task models(emotion recognition and AU detection) and it outperforms other state-of-the-art multi-task learning strategies in FER. Jiabei Zeng, Shiguang Shan, Xilin Chen 0001 |
ICIP | 2 |
| 2019 | Occlusion Aware Facial Expression Recognition Using CNN With Attention MechanismabstractFacial expression recognition in the wild is challenging due to various un-constrained conditions. Although existing facial expression classifiers have been almost perfect on analyzing constrained frontal faces, they fail to perform well on partially occluded faces that are common in the wild. In this paper, we propose a Convolution Neutral Network with attention mechanism (ACNN) that can perceive the occlusion regions of the face and focus on the most discriminative unoccluded regions. ACNN is an end to end learning framework. It combines the multiple representations from facial regions of interest (ROIs). Each representation is weighed via a proposed Gate Unit that computes an adaptive weight from the region itself according to the unobstructed-ness and importance. Considering different RoIs, we introduce two versions of ACNN: patch based ACNN (pACNN) and global-local based ACNN (gACNN). pACNN only pays attention to local facial patches. gACNN integrates local representations at patch-level with global representation at image-level. The proposed ACNNs are evaluated on both real and synthetic occlusions, including a self-collected facial expression dataset with real-world occlusions (FED-RO), two largest in-the-wild facial expression datasets (RAF-DB and AffectNet) and their modifications with synthesized facial occlusions. Experimental results show that ACNNs improve the recognition accuracy on both the non-occluded faces and occluded faces. Visualization results demonstrate that, compared with the CNN without Gate Unit, ACNNs are capable of shifting the attention from the occluded patches to other related but unobstructed ones. ACNNs also outperform other state-of-the-art methods on several widely used in-the-lab facial expression datasets under the cross-dataset evaluation protocol. Yong Li 0032, Jiabei Zeng, Shiguang Shan, Xilin Chen 0001 |
IEEE Trans. Image Process. | 2 |
| 2018 | Zero-Shot Facial Expression Recognition with Multi-label Label Propagation
Zijia Lu, Jiabei Zeng, Shiguang Shan, Xilin Chen 0001 |
ACCV (3) | 2 |
| 2018 | Facial Expression Recognition with Inconsistently Annotated Datasets
Jiabei Zeng, Shiguang Shan, Xilin Chen 0001 |
ECCV (13) | 1 |
| 2018 | Multi-Channel Pose-Aware Convolution Neural Networks for Multi-View Facial Expression RecognitionabstractAlthough tremendous strides have been made in facial expression recognition(FER), recognizing facial expressions in non-frontal views remains an open challenge due to the limited access to large scale training data with various poses. To make full use of the limited data, we propose a novel multi-channel pose-aware convolution neural network (MPCNN) that consists of three parts: the multi-channel feature extraction, jointly multi-scale feature fusion, and the pose-aware recognition. The feature extraction part has 3 sub-CNNs and it learns convolutional features from different features. The joint fusion part fuses multi-scale features to enhance high-level feature representation in a hierarchical way. The fused features are fed to the pose-aware recognition part that includes pose-specific recognition branches and a pose estimation sub-network. According to the estimated pose, MPCNN finally classifies the facial expression through a conditional weighted combination of the pose-specific recognition branches. MPCNN is end-to-end trainable by minimizing the joint loss of pose and expression recognition. We evaluated the proposed method on two public multi-view FER datasets (BU-3DFE and KDEF) and a FER dataset in the wild (SFEW). The experimental results demonstrate that MPCNN outperforms the state-of-the-art FER methods with both within-dataset and cross-dataset settings. Yuanyuan Liu 0004, Jiabei Zeng, Shiguang Shan, Zhuo Zheng |
FG | 2 |
| 2018 | Automatic Engagement Prediction with GAP FeatureabstractIn this paper, we propose an automatic engagement prediction method for the Engagement in the Wild sub-challenge of EmotiW 2018. We first design a novel Gaze-AU-Pose (GAP) feature taking into account the information of gaze, action units and head pose of a subject. The GAP feature is then used for the subsequent engagement level prediction. To efficiently predict the engagement level for a long-time video, we divide the long-time video into multiple overlapped video clips and extract GAP feature for each clip. A deep model consisting of a Gated Recurrent Unit (GRU) layer and a fully connected layer is used as the engagement predictor. Finally, a mean pooling layer is applied to the per-clip estimation to get the final engagement level of the whole video. Experimental results on the validation set and test set show the effectiveness of the proposed approach. In particular, our approach achieves a promising result with an MSE of 0.0724 on the test set of Engagement Prediction Challenge of EmotiW 2018.t with an MSE of 0.072391 on the test set of Engagement Prediction Challenge of EmotiW 2018. Xuesong Niu, Hu Han 0001, Jiabei Zeng, Xuran Sun, Shiguang Shan, Yan Huang 0008, Songfan Yang, Xilin Chen 0001 |
ICMI | 3 |
| 2018 | Patch-Gated CNN for Occlusion-aware Facial Expression RecognitionabstractFacial expression recognition in the wild is challenging due to various un-constrained conditions. Although existing facial expression classifiers have been almost perfect on analyzing constrained frontal faces, they fail to perform well on partially occluded faces that are common in the wild. In this paper, we propose an end-to-end trainable Patch-Gated Convolution Neutral Network (PG-CNN) that can automatically percept the occluded region of the face and focus on the most discriminative un-occluded regions. To determine the possible regions of interest on the face, PG-CNN decomposes an intermediate feature map into several patches according to the positions of related facial landmarks. Then, via a proposed Patch-Gated Unit, PG-CNN reweighs each patch by the unobstructed-ness or importance that is computed from the patch itself. The proposed PG-CNN is evaluated on two largest in-the-wild facial expression datasets (RAF-DB and AffectNet) and their modifications with synthesized facial occlusions. Experimental results show that PG-CNN improves the recognition accuracy on both the original faces and faces with synthesized occlusions. Visualization results demonstrate that, compared with the CNN without Patch-Gated Unit, PG-CNN is capable of shifting the attention from the occluded patch to other related but unobstructed ones. Experiments also show that PG-CNN outperforms other state-of-the-art methods on several widely used in-the-lab facial expression datasets under the cross-dataset evaluation protocol. Yong Li 0032, Jiabei Zeng, Shiguang Shan, Xilin Chen 0001 |
ICPR | 2 |
| 2018 | Dimensionality Reduction in Multiple Ordinal RegressionabstractSupervised dimensionality reduction (DR) plays an important role in learning systems with high-dimensional data. It projects the data into a low-dimensional subspace and keeps the projected data distinguishable in different classes. In addition to preserving the discriminant information for binary or multiple classes, some real-world applications also require keeping the preference degrees of assigning the data to multiple aspects, e.g., to keep the different intensities for co-occurring facial expressions or the product ratings in different aspects. To address this issue, we propose a novel supervised DR method for DR in multiple ordinal regression (DRMOR), whose projected subspace preserves all the ordinal information in multiple aspects or labels. We formulate this problem as a joint optimization framework to simultaneously perform DR and ordinal regression. In contrast to most existing DR methods, which are conducted independently of the subsequent classification or ordinal regression, the proposed framework fully benefits from both of the procedures. We experimentally demonstrate that the proposed DRMOR method (DRMOR-M) well preserves the ordinal information from all the aspects or labels in the learned subspace. Moreover, DRMOR-M exhibits advantages compared with representative DR or ordinal regression algorithms on three standard data sets. Jiabei Zeng, Yang Liu 0007, Biao Leng, Zhang Xiong 0001, Yiu-Ming Cheung |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2017 | 3D Object retrieval based on viewpoint segmentation
Biao Leng, Changchun Du, Jiabei Zeng, Zhang Xiong 0001 |
Multim. Syst. | 4 |
| 2016 | Confidence Preserving Machine for Facial Action Unit DetectionabstractFacial action unit (AU) detection from video has been a long-standing problem in the automated facial expression analysis. While progress has been made, accurate detection of facial AUs remains challenging due to ubiquitous sources of errors, such as inter-personal variability, pose, and low-intensity AUs. In this paper, we refer to samples causing such errors as hard samples, and the remaining as easy samples. To address learning with the hard samples, we propose the confidence preserving machine (CPM), a novel two-stage learning framework that combines multiple classifiers following an "easy-to-hard" strategy. During the training stage, CPM learns two confident classifiers. Each classifier focuses on separating easy samples of one class from all else, and thus preserves confidence on predicting each class. During the test stage, the confident classifiers provide "virtual labels" for easy test samples. Given the virtual labels, we propose a quasi-semi-supervised (QSS) learning strategy to learn a person-specific classifier. The QSS strategy employs a spatio-temporal smoothness that encourages similar predictions for samples within a spatio-temporal neighborhood. In addition, to further improve detection performance, we introduce two CPM extensions: iterative CPM that iteratively augments training samples to train the confident classifiers, and kernel CPM that kernelizes the original CPM model to promote nonlinearity. Experiments on four spontaneous data sets GFT, BP4D, DISFA, and RU-FACS illustrate the benefits of the proposed CPM models over baseline methods and the state-of-the-art semi-supervised learning and transfer learning methods. Jiabei Zeng, Wen-Sheng Chu, Fernando De la Torre, Jeffrey F. Cohn, Zhang Xiong 0001 |
IEEE Trans. Image Process. | 1 |
| 2015 | Unsupervised Synchrony Discovery in Human InteractionabstractPeople are inherently social. Social interaction plays an important and natural role in human behavior. Most computational methods focus on individuals alone rather than in social context. They also require labelled training data. We present an unsupervised approach to discover interpersonal synchrony, referred as to two or more persons preforming common actions in overlapping video frames or segments. For computational efficiency, we develop a branch-and-bound (B&B) approach that affords exhaustive search while guaranteeing a globally optimal solution. The proposed method is entirely general. It takes from two or more videos any multi-dimensional signal that can be represented as a histogram. We derive three novel bounding functions and provide efficient extensions, including multi-synchrony detection and accelerated search, using a warm-start strategy and parallelism. We evaluate the effectiveness of our approach in multiple databases, including human actions using the CMU Mocap dataset [1], spontaneous facial behaviors using group-formation task dataset [37] and parent-infant interaction dataset [28]. Wen-Sheng Chu, Jiabei Zeng, Fernando De la Torre, Jeffrey F. Cohn, Daniel S. Messinger |
ICCV | 2 |
| 2015 | Confidence Preserving Machine for Facial Action Unit DetectionabstractVaried sources of error contribute to the challenge of facial action unit detection. Previous approaches address specific and known sources. However, many sources are unknown. To address the ubiquity of error, we propose a Confident Preserving Machine (CPM) that follows an easy-to-hard classification strategy. During training, CPM learns two confident classifiers. A confident positive classifier separates easily identified positive samples from all else, a confident negative classifier does same for negative samples. During testing, CPM then learns a person-specific classifier using "virtual labels" provided by confident classifiers. This step is achieved using a quasi-semi-supervised (QSS) approach. Hard samples are typically close to the decision boundary, and the QSS approach disambiguates them using spatio-temporal constraints. To evaluate CPM, we compared it with a baseline single-margin classifier and state-of-the-art semi-supervised learning, transfer learning, and boosting methods in three datasets of spontaneous facial behavior. With few exceptions, CPM outperformed baseline and state-of-the art methods. Jiabei Zeng, Wen-Sheng Chu, Fernando De la Torre, Jeffrey F. Cohn, Zhang Xiong 0001 |
ICCV | 1 |
| 2015 | 3-D object retrieval using topic model
Jiabei Zeng, Biao Leng, Zhang Xiong 0001 |
Multim. Tools Appl. | 1 |
| 2015 | 3D Object Retrieval With Multitopic Model Combining Relevance Feedback and LDA ModelabstractView-based 3D model retrieval uses a set of views to represent each object. Discovering the complex relationship between multiple views remains challenging in 3D object retrieval. Recent progress in the latent Dirichlet allocation (LDA) model leads us to propose its use for 3D object retrieval. This LDA approach explores the hidden relationships between extracted primordial features of these views. Since LDA is limited to a fixed number of topics, we further propose a multitopic model to improve retrieval performance. We take advantage of a relevance feedback mechanism to balance the contributions of multiple topic models with specified numbers of topics. We demonstrate our improved retrieval performance over the state-of-the-art approaches. Biao Leng, Jiabei Zeng, Zhang Xiong 0001 |
IEEE Trans. Image Process. | 2 |
| 2013 | Probability tree based passenger flow prediction and its application to the Beijing subway system
Biao Leng, Jiabei Zeng, Zhang Xiong 0001, Weifeng Lv, Yueliang Wan |
Frontiers Comput. Sci. | 2 |