EDBT 2026 Demo / reviewers in the wild / expert
Lijun Yin 0001
dblp:y/LijunYin
· DBLP profile ↗
102ranked-venue papers
18as first author
15since 2021 · last 2026
0000-0002-0343-7190ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 79 · 11 first-author · 13 since 2021Artificial intelligence and machine learning · 61 · 7 first-author · 12 since 2021Human-computer interaction and ubiquitous computing · 9 · 1 first-author · 1 since 2021Security and privacy · 3 · 1 since 2021Computer networks · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | You Only Need One Stage: Novel-View Synthesis from a Single Blind Face ImageabstractWe propose a novel one-stage method, NVB-Face, for generating consistent Novel-View images directly from a single Blind Face image. Existing approaches to novel-view synthesis for objects or faces typically require a high-resolution RGB image as input. When dealing with degraded images, the conventional pipeline follows a two-stage process: first restoring the image to high resolution, then synthesizing novel views from the restored result. However, this approach is highly dependent on the quality of the restored image, often leading to inaccuracies and inconsistencies in the final output. To address this limitation, we extract single-view features directly from the blind face image and introduce a feature manipulator that transforms these features into 3D-aware, multi-view latent representations. Leveraging the powerful generative capacity of a diffusion model, our framework synthesizes high-quality, consistent novel-view face images. Experimental results show that our method significantly outperforms traditional two-stage approaches in both consistency and fidelity. Taoyue Wang, Xiang Zhang 0030, Huiyuan Yang, Lijun Yin 0001 |
AAAI | 5 |
| 2026 | ReMiX-MAE: Learning Missing-Channel Cross-Modal Representations from RGB-Only Clinical Facial Videos for Sympathetic-Mediated Pain Assessment
Nan Bi, Taoyue Wang, Lijun Yin 0001 |
FG | 3 |
| 2026 | Inter-Stance: A Dyadic Multimodal Corpus for Conversational Stance Analysis
Xiang Zhang 0030, Taoyue Wang, Nan Bi, Cody Zhou, Zoie Wang, Yuming Su, Jeffrey F. Cohn, Lijun Yin 0001 |
FG | 12 |
| 2024 | Enhancing Face Recognition in Low-Quality Images Based on Restoration and 3D Multiview GenerationabstractFace recognition in low-quality images presents a significant challenge, particularly when images or videos are captured under adverse conditions such as atmosphere turbulence due to long distance capture causing noise, blurring, and distortion in low resolution videos. This paper proposes a novel and comprehensive approach to address the challenges in atmosphere turbulence, leveraging a multi-stage process that includes weak-strong image restoration using Generative Adversarial Network (GAN) based model and Stable Diffusion (SD) based model, 3D face view generation with Neural Radiance Fields (NeRF), and adaptive face recognition on fine-augmented datasets. Our methodology shows substantial improvements in face recognition accuracy on average from 55.6% to 74.7% on the level-3 turbulence images of LFW, CFP, CALFW, as well as improvements on datasets of TinyFace, the BRIAR BRC1, and the BRIAR BGC, demonstrating the performance enhancement in extremely challenging conditions. Xiang Zhang 0030, Taoyue Wang, Lijun Yin 0001 |
IJCB | 4 |
| 2024 | Multimodal Channel-Mixing: Channel and Spatial Masked AutoEncoder on Facial Action Unit DetectionabstractRecent studies have focused on utilizing multi-modal data to develop robust models for facial Action Unit (AU) detection. However, the heterogeneity of multi-modal data poses challenges in learning effective representations. One such challenge is extracting relevant features from multiple modalities using a single feature extractor. Moreover, previous studies have not fully explored the potential of multi-modal fusion strategies. In contrast to the extensive work on late fusion, there are limited investigations on early fusion for channel information exploration. This paper presents a novel multi-modal reconstruction network, named Multimodal Channel-Mixing (MCM), as a pre-trained model to learn robust representation for facilitating multi-modal fusion. The approach follows an early fusion setup, integrating a Channel-Mixing module, where two out of five channels are randomly dropped. The dropped channels then are reconstructed from the remaining channels using masked autoencoder. This module not only reduces channel redundancy, but also facilitates multi-modal learning and reconstruction capabilities, resulting in robust feature learning. The encoder is fine-tuned on a downstream task of automatic facial action unit detection. Pre-training experiments were conducted on BP4D+, followed by fine-tuning on BP4D and DISFA to assess the effectiveness and robustness of the proposed framework. The results demonstrate that our method meets and surpasses the performance of state-of-the-art baseline methods. Xiang Zhang 0030, Huiyuan Yang, Taoyue Wang, Lijun Yin 0001 |
WACV | 5 |
| 2024 | Disagreement Matters: Exploring Internal Diversification for Redundant Attention in Generic Facial Action AnalysisabstractThis paper demonstrates the effectiveness of a diversification mechanism for building a more robust multi-attention system in generic facial action analysis. While previous multi-attention (e.g., visual attention and self-attention) research on facial expression recognition (FER) and Action Unit (AU) detection have been thoroughly studied to focus on ”external attention diversification”, where attention branches localize different facial areas, we delve into the realm of ”internal attention diversification” and explore the impact of diverse attention patterns within the same Region of Interest (RoI). Our experiments reveal that variability in attention patterns significantly impacts model performance, indicating that unconstrained multi-attention plagued by redundancy and over-parameterization, leading to sub-optimal results. To tackle this issue, we propose a compact module that guides the model to achieve self-diversified multi-attention. Our method is applied to both CNN-based and Transformer-based models, benchmarked on popular databases such as BP4D and DISFA for AU detection, as well as CK+, MMI, BU-3DFE, and BP4D+ for facial expression recognition. We also evaluate the mechanism on Self-attention and Channel-wise attention designs for improving their adaptive capabilities in multi-modal feature fusion tasks. The multi-modal evaluation is conducted on BP4D, BP4D+, and our newly developed large-scale comprehensive emotion database BP4D++, which contains well-synchronized and aligned sensor modalities, addressing the scarcity of annotations and identities in human affective computing. We plan to release the new database to the research community, fostering further advancements in this field. Zheng Zhang 0023, Xiang Zhang 0030, Taoyue Wang, Huiyuan Yang, Umur A. Ciftci, Jeffrey F. Cohn, Lijun Yin 0001 |
IEEE Trans. Affect. Comput. | 10 |
| 2024 | Spatio-Temporal Graph Analytics on Secondary Affect Data for Improving Trustworthy Emotional AIabstractEthical affective computing (AC) requires maximizing the benefits to users while minimizing its harm to obtain trust from users. This requires responsible development and deployment to ensure fairness, bias mitigation, privacy preservation, and accountability. To obtain this, we require methodologies that can quantify, visualize, analyze, and mine insights from affect data. Hence, in this paper, we propose a spatio-temporal model for representing secondary affect data from network sciences' perspective. We propose a network science-based model to represent spatio-temporal data, e.g., action units' sequences, and continuous affect reports. In particular, the proposed model captures the spatial and temporal strength of the relationship among essential variables in the data. The proposed model allows to analyze data as a whole system. We also demonstrated the use case of the model for graph analytics on secondary affect data that can assist to measure and quantify several issues that can be originated from the study setup, data recording devices, and the influences/biases that can originate from the perspective of the affect reporters. We also demonstrated the use cases of the proposed method on ethical trustworthy emotional AI via measuring biases from de-identified data and how it contributes towards ethics, transparency, value alignment, and governance. Md Taufeeq Uddin, Lijun Yin 0001, Shaun J. Canavan |
IEEE Trans. Affect. Comput. | 2 |
| 2024 | Deepfake source detection in a heart beat
Umur A. Ciftci, Ilke Demir, Lijun Yin 0001 |
Vis. Comput. | 3 |
| 2023 | ReactioNet: Learning High-order Facial Behavior from Universal Stimulus-Reaction by Dyadic Relation ReasoningabstractDiverse visual stimuli can evoke various human affective states, which are usually manifested in an individual’s muscular actions and facial expressions. In lab-controlled emotion datasets, such a critical component (i.e., stimulus) was commonly designed in a limited way, making researchers incapable of generalizing the universal correlation and causation of stimulus-reaction as well as predicting possible emotions from context, timing, and relation. In this paper, we collected a large-scale spontaneous facial behavior database ReactioNet, which contains 1.1 million coupled stimulus-reaction tuples (visual/audio/caption from both stimuli and subjects). We introduce a new facial behavior detection scenario, Dyadic Relation Reasoning (DRR), which aims to detect facial actions by reasoning their relations with stimuli. By aggregating the dyadic information, our method essentially forms a relation prototype Universal Stimulus Reaction (U-SR), which encodes the low-order and high-order relationships between stimulus agents and facial reactions. A framework with both non-graph and graph modules is further developed to evaluate DRR-based facial action unit detection, facial expression recognition, and scene classification. Specifically, to learn "what" arouses a facial reaction, the non-graph module associates and projects the fine-grained stimulus-reaction features into common subspaces using cross-domain contrastive learning. To learn "how" stimulus-reaction pairs are mutually related, the graph module adopts Graph Convolution Network to represent, converge, and infer the dyadic U-SR relation under two relation assumptions (i.e., homophily and heterophily [68]). Extensive experiments demonstrate the effectiveness of the proposed work. The new dataset will be available for the research community. Taoyue Wang, Geran Zhao, Xiang Zhang 0030, Xi Kang, Lijun Yin 0001 |
ICCV | 6 |
| 2023 | Knowledge-Spreader: Learning Semi-Supervised Facial Action Dynamics by Consistifying Knowledge GranularityabstractRecent studies on dynamic facial action unit (AU) detection have extensively relied on dense annotations. However, manual annotations are difficult, time-consuming, and costly. The canonical semi-supervised learning (SSL) methods ignore the consistency, extensibility, and adaptability of structural knowledge across spatial-temporal domains. Furthermore, the reliance on offline design and excessive parameters hinder the efficiency of the learning process. To remedy these issues, we propose a lightweight and online semi-supervised framework, a so-called Knowledge-Spreader (KS), to learn AU dynamics with sparse annotations. By formulating SSL as a Progressive Knowledge Distillation (PKD) problem, we aim to infer cross-domain information, specifically from spatial to temporal domains, by consistifying knowledge granularity within Teacher-Students Network. Specifically, KS employs sparsely annotated key-frames to learn AU dependencies as the privileged knowledge. Then, the model spreads the learned knowledge to their unlabeled neighbours by jointly applying knowledge distillation and pseudo-labeling, and completes the temporal information as the expanded knowledge. We term the progressive knowledge distillation as "Knowledge Spreading", which allows our model to learn spatial-temporal knowledge from video clips with only one label allocated. Extensive experiments demonstrate that KS achieves competitive performance as compared to the state of the arts under the circumstances of using only 2% labels on BP4D and 5% labels on DISFA. In addition, we have tested it on our newly developed large-scale comprehensive emotion database BP4D++, which contains considerable samples across well-synchronized and aligned sensor modalities for alleviating the scarcity issue of annotations and identities. Xiang Zhang 0030, Taoyue Wang, Lijun Yin 0001 |
ICCV | 4 |
| 2023 | Weakly-Supervised Text-driven Contrastive Learning for Facial Behavior UnderstandingabstractContrastive learning has shown promising potential for learning robust representations by utilizing unlabeled data. However, constructing effective positive-negative pairs for contrastive learning on facial behavior datasets remains challenging. This is because such pairs inevitably encode the subject-ID information, and the randomly constructed pairs may push similar facial images away due to the limited number of subjects in facial behavior datasets. To address this issue, we propose to utilize activity descriptions, coarse-grained information provided in some datasets, which can provide high-level semantic information about the image sequences but is often neglected in previous studies. More specifically, we introduce a two-stage Contrastive Learning with Text-Embeded framework for Facial behavior understanding (CLEF). The first stage is a weakly-supervised contrastive learning method that learns representations from positive-negative pairs constructed using coarse-grained activity information. The second stage aims to train the recognition of facial expressions or facial action units by maximizing the similarity between the image and the corresponding text label names. The proposed CLEF achieves state-of-the-art performance on three in-the-lab datasets for AU recognition and three in-the-wild datasets for facial expression recognition. Xiang Zhang 0030, Taoyue Wang, Huiyuan Yang, Lijun Yin 0001 |
ICCV | 5 |
| 2021 | Exploiting Semantic Embedding and Visual Feature for Facial Action Unit DetectionabstractRecent study on detecting facial action units (AU) has utilized auxiliary information (i.e., facial landmarks, relationship among AUs and expressions, web facial images, etc.), in order to improve the AU detection performance. As of now, no semantic information of AUs has yet been explored for such a task. As a matter of fact, AU semantic descriptions provide much more information than the binary AU labels alone, thus we propose to exploit the Semantic Embedding and Visual feature (SEV-Net) for AU detection. More specifically, AU semantic embeddings are obtained through both Intra-AU and Inter-AU attention modules, where the Intra-AU attention module captures the relation among words within each sentence that describes individual AU, and the Inter-AU attention module focuses on the relation among those sentences. The learned AU semantic embeddings are then used as guidance for the generation of attention maps through a cross-modality attention network. The generated cross-modality attention maps are further used as weights for the aggregated feature. Our proposed method is unique in that the semantic features are exploited as the first of this kind. The approach has been evaluated on three public AU-coded facial expression databases, and has achieved a superior performance than the state-of-the-art peer methods. Huiyuan Yang, Lijun Yin 0001, Jiuxiang Gu |
CVPR | 2 |
| 2021 | Your "Attention" Deserves Attention: A Self-Diversified Multi-Channel Attention for Facial Action AnalysisabstractVisual attention has been extensively studied for learning fine-grained features in both facial expression recognition (FER) and Action Unit (AU) detection. A broad range of previous research has explored how to use attention modules to localize detailed facial parts (e,g. facial action units), learn discriminative features, and learn inter-class correlation. However, few related works pay attention to the robustness of the attention module itself. Through experiments, we found neural attention maps initialized with different feature maps yield diverse representations when learning to attend the identical Region of Interest (ROI). In other words, similar to general feature learning, the representational quality of attention maps also greatly affects the performance of a model, which means unconstrained attention learning has lots of randomnesses. This uncertainty lets conventional attention learning fall into sub-optimal. In this paper, we propose a compact model to enhance the representational and focusing power of neural attention maps and learn the “inter-attention” correlation for refined attention maps, which we term the “Self-Diversified Multi-Channel Attention Network (SMA-Net)”. The proposed method is evaluated on two benchmark databases (BP4D and DISFA) for AU detection and four databases (CK+, MMI, BU-3DFE, and BP4D+) for facial expression recognition. It achieves superior performance compared to the state-of-the-art methods. Huiyuan Yang, Geran Zhao, Lijun Yin 0001 |
FG | 5 |
| 2021 | Multi-Modal Learning for AU Detection Based on Multi-Head Fused TransformersabstractMulti-modal learning has been intensified in recent years, especially for applications in facial analysis and action unit detection whilst there still exist two main challenges in terms of 1) relevant feature learning for representation and 2) efficient fusion for multi-modalities. Recently, there are a number of works have shown the effectiveness in utilizing the attention mechanism for AU detection, however, most of them are binding the region of interest (ROI) with features but rarely apply attention between features of each AU. On the other hand, the transformer, which utilizes a more efficient self-attention mechanism, has been widely used in natural language processing and computer vision tasks but is not fully explored in AU detection tasks. In this paper, we propose a novel end-to-end Multi-Head Fused Transformer (MFT) method for AU detection, which learns AU encoding features representation from different modalities by transformer encoder and fuses modalities by another fusion transformer module. Multi-head fusion attention is designed in the fusion transformer module for the effective fusion of multiple modalities. Our approach is evaluated on two public multi-modal AU databases, BP4D, and BP4D+, and the results are superior to the state-of-the-art algorithms and baseline models. We further analyze the performance of AU detection from different modalities. Xiang Zhang 0030, Lijun Yin 0001 |
FG | 2 |
| 2021 | Integrating Semantic and Temporal Relationships in Facial Action Unit DetectionabstractFacial action unit (AU) detection is a challenging task due to the variety and subtlety of individuals' facial behavior. Facial muscle characteristics such as temporal dependencies and action correlations make AU detection differ from general multi-label classification tasks, and capturing these two characteristics is the key to accurate AU detection. However, there is little work to date taking both of them into consideration concurrently. To capture the AU correlations in an image, we first disentangle the global (image) feature into multiple AU-specific features with an AU contrastive loss, and then we compute the feature for each AU by aggregating the features from the other AUs with a self-attention based transformer. Different from the original transformer, we embed the AU semantic dependency matrix into it to weakly guide the attention learning. We then weighted fuse the AU-wise features to obtain the frame-wise features. We further capture the temporal dependencies among frames by using another attention-based transformer, which achieves information aggregation from the prior frames. Extensive experiments on two benchmark datasets (i.e., BP4D and DISFA) demonstrate that the proposed framework outperforms the state-of-the-art approaches. Xiang Deng 0002, Lijun Yin 0001 |
ACM Multimedia | 4 |
| 2020 | RE-Net: A Relation Embedded Deep Model for AU Occurrence and Intensity Estimation
Huiyuan Yang, Lijun Yin 0001 |
ACCV (5) | 2 |
| 2020 | Recognizing Perceived Emotions from Facial ExpressionsabstractExpression recognition has seen an increase in research in past years, however, little work has been on recognizing perceived emotion (i.e. subject self-reporting of emotion). Considering this, we investigate the perceived emotion of subjects that perform tasks meant to elicit emotion. To facilitate this investigation, we use the BP4D+ multimodal spontaneous emotion corpus. We first statistically analyze the subject's perceived emotions across 10 tasks available in BP4D+. We show the percentage of subjects that felt specific emotions for each of the tasks. This is done across all tested subjects, as well as male and female subjects independently. Along with our statistical analysis, we also propose a 3D convolutional neural network (CNN) architecture to recognize multiple emotions felt for each task sequence. We report accuracy, Fl-binary and AUC for all subjects, as well as male and female subjects. Saurabh Hinduja, Shaun J. Canavan, Lijun Yin 0001 |
FG | 3 |
| 2020 | An EEG-Based Multi-Modal Emotion Database with Both Posed and Authentic Facial Actions for Emotion AnalysisabstractEmotion is an experience associated with a particular pattern of physiological activity along with different physiological, behavioral and cognitive changes. One behavioral change is facial expression, which has been studied extensively over the past few decades. Facial behavior varies with a person's emotion according to differences in terms of culture, personality, age, context, and environment. In recent years, physiological activities have been used to study emotional responses. A typical signal is the electroencephalogram (EEG), which measures brain activity. Most of existing EEG-based emotion analysis has overlooked the role of facial expression changes. There exits little research on the relationship between facial behavior and brain signals due to the lack of dataset measuring both EEG and facial action signals simultaneously. To address this problem, we propose to develop a new database by collecting facial expressions, action units, and EEGs simultaneously. We recorded the EEGs and face videos of both posed facial actions and spontaneous expressions from 29 participants with different ages, genders, ethnic backgrounds. Differing from existing approaches, we designed a protocol to capture the EEG signals by evoking participants' individual action units explicitly. We also investigated the relation between the EEG signals and facial action units. As a baseline, the database has been evaluated through the experiments on both posed and spontaneous emotion recognition with images alone, EEG alone, and EEG fused with images, respectively. The database will be released to the research community to advance the state of the art for automatic emotion recognition. Xiang Zhang 0030, Huiyuan Yang, Wenna Duan, Weiying Dai, Lijun Yin 0001 |
FG | 6 |
| 2020 | Set Operation Aided Network for Action Units DetectionabstractAs a large number of parameters exist in deepmodel based methods, training such models usually requires many fully AU-annotated facial images. This is true with regard to the number of frames in two widely used datasets: BP4D[31] and DISFA [18], while those frames were captured from a small number of subjects (41, 27 respectively). This is problematic, as subjects produce highly consistent facial muscle movements, adding more frames per subject would only adds more close points in the feature space, and thus the classifier does not benefit from those extra frames. Data augmentation methods can be applied to alleviate the problem to a certain degree, but they fail to augment new subjects. We propose a novel Set Operation Aided Network (SO-Net) for action units detection. Specifically, new features and the corresponding labels are generated by adding set operations to both the feature and label spaces. The generated new features can be treated as a representation of a hypothetical image. As a result, we can implicitly obtain training examples beyond what was originally observed in the dataset. Therefore, the deep model is forced to learn subject-independent features, and is generalizable to unseen subjects. SO-Net is end-to-end trainable, and can be flexibly plugged in any CNN model during training. We evaluate the proposed method on two public datasets, BP4D and DISFA. The experiment shows a state-of-the-art performance, demonstrating the effectiveness of the proposed method. Huiyuan Yang, Taoyue Wang, Lijun Yin 0001 |
FG | 3 |
| 2020 | How Do the Hearts of Deep Fakes Beat? Deep Fake Source Detection via Interpreting Residuals with Biological SignalsabstractFake portrait video generation techniques have been posing a new threat to the society with photorealistic deep fakes for political propaganda, celebrity imitation, forged evidences, and other identity related manipulations. Following these generation techniques, some detection approaches have also been proved useful due to their high classification accuracy. Nevertheless, almost no effort was spent to track down the source of deep fakes. We propose an approach not only to separate deep fakes from real videos, but also to discover the specific generative model behind a deep fake. Some pure deep learning based approaches try to classify deep fakes using CNNs where they actually learn the residuals of the generator. We believe that these residuals contain more information and we can reveal these manipulation artifacts by disentangling them with biological signals. Our key observation yields that the spatiotemporal patterns in biological signals can be conceived as a representative projection of residuals. To justify this observation, we extract PPG cells from real and fake videos and feed these to a state-of-the-art classification network for detecting the generative model per video. Our results indicate that our approach can detect fake videos with 97.29% accuracy, and the source model with 93.39% accuracy. Umur A. Ciftci, Ilke Demir, Lijun Yin 0001 |
IJCB | 3 |
| 2020 | SAT-Net: Self-Attention and Temporal Fusion for Facial Action Unit DetectionabstractResearch on facial action unit detection has shown remarkable performances by using deep spatial learning models in recent years, however, it is far from reaching its full capacity in learning due to the lack of use of temporal information of AUs across time. Since the AU occurrence in one frame is highly likely related to previous frames in a temporal sequence, exploring temporal correlation of AUs across frames becomes a key motivation of this work. In this paper, we propose a novel temporal fusion and AU-supervised self-attention network (a socalled SAT-Net) to address the AU detection problem. First of all, we input the deep features of a sequence into a convolutional LSTM network and fuse the previous temporal information into the feature map of the last frame, and continue to learn the AU occurrence. Second, considering the AU detection problem is a multi-label classification problem that individual label depends only on certain facial areas, we propose a new self-learned attention mask by focusing the detection of each AU on parts of facial areas through the learning of individual attention mask for each AU, thus increasing the AU independence without the loss of any spatial relations. Our extensive experiments show that the proposed framework achieves better results of AU detection over the state-of-the-arts on two benchmark databases (BP4D and DISFA). Zheng Zhang 0023, Lijun Yin 0001 |
ICPR | 3 |
| 2020 | Region of Interest Based Graph Convolution: A Heatmap Regression Approach for Action Unit DetectionabstractMachine vision of human facial expressions has been studied for decades, from prototypical expressions to Action Units (AUs), from hand-crafted to deep features, from multi-class to multi-label classifications. Since the widely adopted deep networks lack interpretation on learnt representations, human prior knowledge cannot be effectively imposed and examined. On the other hand, AU is a human defined concept. In order to align with this idea, a finer level of network design is desired. In this paper, we first extend the heatmaps to ROI maps, encoding the location of both positive and negative occurred AUs, then employ a well-designed backbone network to regress it. In this way, AU detection is performed in two stages, key regions localization and occurrence classification. To prompt the spatial dependency among ROIs, we utilize graph convolution for feature refinement. The decomposition of similarity matrix is supervised by AU labels. This novel framework is evaluated on two benchmark databases (BP4D and DISFA) for AU detection. The experimental results are superior to the state-of-the-art algorithms and baseline models, demonstrating the effectiveness of our proposed method. Zheng Zhang 0023, Taoyue Wang, Lijun Yin 0001 |
ACM Multimedia | 3 |
| 2020 | Adaptive Multimodal Fusion for Facial Action Units RecognitionabstractMultimodal facial action units (AU) recognition aims to build models that are capable of processing, correlating, and integrating information from multiple modalities (i.e., 2D images from a visual sensor, 3D geometry from 3D imaging, and thermal images from an infrared sensor). Although the multimodel data can provide rich information, there are two challenges that have to be addressed when learning from multimodal data: 1) the model must capture the complex cross-modal interactions in order to utilize the additional and mutual information effectively; 2) the model must be robust enough in the circumstance of unexpected data corruptions during testing, in case of a certain modality missing or being noisy. In this paper, we propose a novel A daptive M ultimodal F usion method (AMF ) for AU detection, which learns to select the most relevant feature representations from different modalities by a re-sampling procedure conditioned on a feature scoring module. The feature scoring module is designed to allow for evaluating the quality of features learned from multiple modalities. As a result, AMF is able to adaptively select more discriminative features, thus increasing the robustness to missing or corrupted modalities. In addition, to alleviate the over-fitting problem and make the model generalize better on the testing data, a cut-switch multimodal data augmentation method is designed, by which a random block is cut and switched across multiple modalities. We have conducted a thorough investigation on two public multimodal AU datasets, BP4D and BP4D+, and the results demonstrate the effectiveness of the proposed method. Ablation studies on various circumstances also show that our method remains robust to missing or noisy modalities during tests. Huiyuan Yang, Taoyue Wang, Lijun Yin 0001 |
ACM Multimedia | 3 |
| 2019 | Reconsidering the Duchenne Smile: Indicator of Positive Emotion or Artifact of Smile Intensity?abstractThe Duchenne smile hypothesis is that smiles that include eye constriction (AU6) are the product of genuine positive emotion, whereas smiles that do not are either falsified or related to negative emotion. This hypothesis has become very influential and is often used in scientific and applied settings to justify the inference that a smile is either true or false. However, empirical support for this hypothesis has been equivocal and some researchers have proposed that, rather than being a reliable indicator of positive emotion, AU6 may just be an artifact produced by intense smiles. Initial support for this proposal has been found when comparing smiles related to genuine and feigned positive emotion; however, it has not yet been examined when comparing smiles related to genuine positive and negative emotion. The current study addressed this gap in the literature by examining spontaneous smiles from 136 participants during the elicitation of amusement, embarrassment, fear, and pain (from the BP4D+ dataset). Bayesian multilevel regression models were used to quantify the associations between AU6 and self-reported amusement while controlling for smile intensity. Models were estimated to infer amusement from AU6 and to explain the intensity of AU6 using amusement. In both cases, controlling for smile intensity substantially reduced the hypothesized association, whereas the effect of smile intensity itself was quite large and reliable. These results provide further evidence that the Duchenne smile is likely an artifact of smile intensity rather than a reliable and unique indicator of genuine positive emotion. Jeffrey M. Girard, Gayatri Shandar, Zhun Liu, Jeffrey F. Cohn, Lijun Yin 0001, Louis-Philippe Morency |
ACII | 5 |
| 2019 | Cross-domain AU Detection: Domains, Learning Approaches, and MeasuresabstractFacial action unit (AU) detectors have performed well when trained and tested within the same domain. Do AU detectors transfer to new domains in which they have not been trained? To answer this question, we review literature on cross-domain transfer and conduct experiments to address limitations of prior research. We evaluate both deep and shallow approaches to AU detection (CNN and SVM, respectively) in two large, well-annotated, publicly available databases, Expanded BP4D+ and GFT. The databases differ in observational scenarios, participant characteristics, range of head pose, video resolution, and AU base rates. For both approaches and databases, performance decreased with change in domain, often to below the threshold needed for behavioral research. Decreases were not uniform, however. They were more pronounced for GFT than for Expanded BP4D+ and for shallow relative to deep learning. These findings suggest that more varied domains and deep learning approaches may be better suited for promoting generalizability. Until further improvement is realized, caution is warranted when applying AU classifiers from one domain to another. Itir Önal, Jeffrey F. Cohn, László A. Jeni, Zheng Zhang 0023, Lijun Yin 0001 |
FG | 5 |
| 2019 | Facial Action Unit Analysis through 3D Point Cloud Neural NetworksabstractFacial expression analysis on 3D data has the potential to avoid many of the difficulties heir to 2D data, such as lighting variations and non-frontal pose. In particular, analysis of 3D point cloud data (as opposed to depth maps) offers the potential for higher-resolution, pose-invariant features. Because neural networks and deep learning have proven to be very powerful tools for a wide variety of tasks in recent history, one would naturally wish to apply deep learning for expression analysis of 3D point data. However, the overwhelming majority of these methods target 2D image data, and there are only a few works that utilize 3D point data directly in a neural network for any purpose. That said, the results of these works show improvement over using other forms of data. Therefore, in this work, we experiment with recent successful architectures and propose a new architecture, Local Continuous PointNet (LCPN), for unordered 3D point cloud analysis to detect Action Units (AUs) in the BP4D-Spontaneous database. We also perform cross-database experiments on subjects from the BP4D+ database. To the best of the authors' knowledge, this is the first work that directly processes unordered 3D point clouds in a neural network for facial expression analysis. Michael Reale, Benjamin Klinghoffer, Micah Church, Hannah Szmurlo, Lijun Yin 0001 |
FG | 5 |
| 2019 | Learning Temporal Information From A Single Image For AU DetectionabstractAutomatic Facial Action Units (AUs) detection is the recognition of the facial appearance changes caused by the contraction or relaxation of one or more related facial muscles. Compared to the sequence-based methods, a decreased performance is observed for the static image-based AU detection, due to the loss of temporal information. To solve this problem, we propose a novel method that implicitly learns temporal information from a single image for AU detection by adding a hidden optical-flow layer to concatenate two Convolutional Neural Networks (CNNs) models: optical-flow net (OF-Net) and AU detection net (AU-Net). The OF-Net is designed to estimate the facial appearance changes (optical flow) from a single input image through unsupervised learning. The AU-Net accepts the estimated optical-flow as input and predicts the AU occurrence. By training both OF-Net and AU-Net jointly, our model achieves better performance than training them separately, as the AU-Net provides semantic constraints for the optical-flow learning and helps generate more meaningful optical-flow. In return, the estimated optical-flow, which reflects facial appearance changes, benefits the AU-Net. Our proposed method has been evaluated on two benchmarks: BP4D and DISFA, and the experiments show significant performance improvement as compared to the state-of-the-art methods. Huiyuan Yang, Lijun Yin 0001 |
FG | 2 |
| 2019 | Temporal Interframe Pattern Analysis for Static and Dynamic Hand Gesture RecognitionabstractHand gesture, a common non-verbal language, is being studied for Human Computer Interaction. Hand gestures can be categorized as static hand gestures and dynamic hand gestures. In recent years, effective approaches have been applied to hand gesture recognition. However, almost all of the previous works only focus on either of the two categories instead of both, and none of them has used the temporal information on the recognition of static hand gestures.In this paper, we propose a three-level scheme to utilize the temporal interframe pattern on the recognition of both static and dynamic hand gestures. The first classifier assigns a class label to each frame of the video sequence that contains both static and dynamic hand gestures. The second classifier uses the temporal pattern of the class labels to correct the errors of the first classifier and to distinguish between static and dynamic hand gestures. The third classifier is then used to recognize dynamic hand gestures. We believe that we are the first to propose such a strategy. The extensive experiments showed promising performance and demonstrated the feasibility of using temporal interframe pattern to recognize dynamic hand gestures and to correct the errors in static hand gesture recognition. Kaoning Hu, Lijun Yin 0001 |
ICIP | 2 |
| 2019 | Multi-Modality Empowered Network for Facial Action Unit DetectionabstractThis paper presents a new thermal empowered multi-task network (TEMT-Net) to improve facial action unit detection. Our primary goal is to leverage the situation that the training set has multi-modality data while the application scenario only has one modality. Thermal images are robust to illumination and face color. In the proposed multi-task framework, we utilize both modality data. Action unit detection and facial landmark detection are correlated tasks. To utilize the advantage and the correlation of different modalities and different tasks, we propose a novel thermal empowered multi-task deep neural network learning approach for action unit detection, facial landmark detection and thermal image reconstruction simultaneously. The thermal image generator and facial landmark detection provide regularization on the learned features with shared factors as the input color images. Extensive experiments are conducted on the BP4D and MMSE databases, with the comparison to the state of the art methods. The experiments show that the multi-modality framework improves the AU detection significantly. Peng Liu 0039, Zheng Zhang 0023, Huiyuan Yang, Lijun Yin 0001 |
WACV | 4 |
| 2018 | Identity-based Adversarial Training of Deep CNNs for Facial Action Unit Recognition
Zheng Zhang 0023, Shuangfei Zhai, Lijun Yin 0001 |
BMVC | 3 |
| 2018 | Facial Expression Recognition by De-Expression Residue LearningabstractA facial expression is a combination of an expressive component and a neutral component of a person. In this paper, we propose to recognize facial expressions by extracting information of the expressive component through a de-expression learning procedure, called De-expression Residue Learning (DeRL). First, a generative model is trained by cGAN. This model generates the corresponding neutral face image for any input face image. We call this procedure de-expression because the expressive information is filtered out by the generative model; however, the expressive information is still recorded in the intermediate layers. Given the neutral face image, unlike previous works using pixel-level or feature-level difference for facial expression classification, our new method learns the deposition (or residue) that remains in the intermediate layers of the generative model. Such a residue is essential as it contains the expressive component deposited in the generative model from any input facial expression images. Seven public facial expression databases are employed in our experiments. With two databases (BU-4DFE and BP4D-spontaneous) for pre-training, the DeRL method has been evaluated on five databases, CK+, Oulu-CASIA, MMI, BU-3DFE, and BP4D+. The experimental results demonstrate the superior performance of the proposed method. Huiyuan Yang, Umur A. Ciftci, Lijun Yin 0001 |
CVPR | 3 |
| 2018 | Clinical Valid Pain Database with Biomarker and Visual Information for Pain Level AnalysisabstractPain is one of the most common and distressing symptoms reported by emergency room patients. Valid and reliable assessment of pain is essential for both clinical trials and effective pain management. A major limitation of automatic pain assessment by the facial expression is the lack of clinically validated data. This work aims at collecting a pain expression database for pain intensity analysis in clinical settings. The database includes 140 color video sequences, 140 multi-sensor sequences obtained by the Kinect 2, patients' sequence-level self-report, Cyclooxygenases (COXs) level, and inducible nitric oxide synthase (iNOS) level in the blood samples. The database also includes head poses and derived facial landmarks from the 2D video. The relationships of their self-report, Cyclooxygenases (COXs), inducible nitric oxide synthase (iNOS), head pose and facial expression are analyzed. The correlation between the clinical and non-clinical pain facial expressions have been evaluated as well. Peng Liu 0039, Idris Yazgan, Sarah Olsen, Alecia Moser, Umur A. Ciftci, Saeed Bajwa, Christian Tvetenstrand, Peter Gerhardstein, Omowunmi Sadik, Lijun Yin 0001 |
FG | 10 |
| 2018 | Identity-Adaptive Facial Expression Recognition through Expression Regeneration Using Conditional Generative Adversarial NetworksabstractSubject variation is a challenging issue for fa- cial expression recognition, especially when handling unseen subjects with small-scale lableled facial expression databases. Although transfer learning has been widely used to tackle the problem, the performance degrades on new data. In this paper, we present a novel approach (so-called IA-gen) to alleviate the issue of subject variations by regenerating expressions from any input facial images. First of all, we train conditional generative models to generate six prototypic facial expressions from any given query face image while keeping the identity related information unchanged. Generative Adversarial Networks are employed to train the conditional generative models, and each of them is designed to generate one of the prototypic facial expression images. Second, a regular CNN (FER-Net) is fine- tuned for expression classification. After the corresponding prototypic facial expressions are regenerated from each facial image, we output the last FC layer of FER-Net as features for both the input image and the generated images. Based on the minimum distance between the input image and the generated expression images in the feature space, the input image is classified as one of the prototypic expressions consequently. Our proposed method can not only alleviate the influence of inter-subject variations, but will also be flexible enough to integrate with any other FER CNNs for person-independent facial expression recognition. Our method has been evaluated on CK+, Oulu-CASIA, BU-3DFE and BU-4DFE databases, and the results demonstrate the effectiveness of our proposed method. Huiyuan Yang, Zheng Zhang 0023, Lijun Yin 0001 |
FG | 3 |
| 2018 | EAC-Net: Deep Nets with Enhancing and Cropping for Facial Action Unit DetectionabstractIn this paper, we propose a deep learning based approach for facial action unit (AU) detection by enhancing and cropping regions of interest of face images. The approach is implemented by adding two novel nets (a.k.a. layers): the enhancing layers and the cropping layers, to a pretrained convolutional neural network (CNN) model. For the enhancing layers (noted as E-Net), we have designed an attention map based on facial landmark features and apply it to a pretrained neural network to conduct enhanced learning. For the cropping layers (noted as C-Net ), we crop facial regions around the detected landmarks and design individual convolutional layers to learn deeper features for each facial region. We then combine the E-Net and the C-Net to construct a so-called Enhancing and Cropping Net (EAC-Net), which can learn both features enhancing and region cropping functions effectively. The EAC-Net integrates three important elements, i.e., learning transfer, attention coding, and regions of interest processing, making our AU detection approach more efficient and more robust to facial position and orientation changes. Our approach shows a significant performance improvement over the state-of-the-art methods when tested on the BP4D and DISFA AU datasets. The EAC-Net with a slight modification also shows its potentials in estimating accurate AU intensities. We have also studied the performance of the proposed EAC-Net under two very challenging conditions: (1) faces with partial occlusion and (2) faces with large head pose variations. Experimental results show that (1) the EAC-Net learns facial AUs correlation effectively and predicts AUs reliably even with only half of a face being visible, especially for the lower half; (2) Our EAC-Net model also works well under very large head poses, which outperforms significantly a compared baseline approach. It further shows that the EAC-Net works much better without a face frontalization than with face frontalization through image warping as pre-processing, in terms of computational efficiency and AU detection accuracy. Wei Li 0077, Farnaz Abtahi, Zhigang Zhu 0001, Lijun Yin 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2017 | CNN based 3D facial expression recognition using masking and landmark featuresabstractAutomatically recognizing facial expression is an important part for human-machine interaction. In this paper, we first review the previous studies on both 2D and 3D facial expression recognition, and then summarize the key research questions to solve in the future. Finally, we propose a 3D facial expression recognition (FER) algorithm using convolutional neural networks (CNNs) and landmark features/masks, which is invariant to pose and illumination variations due to the solely use of 3D geometric facial models without any texture information. The proposed method has been tested on two public 3D facial expression databases: BU-4DFE and BU-3DFE. The results show that the CNN model benefits from the masking, and the combination of landmark and CNN features can further improve the 3D FER accuracy. Huiyuan Yang, Lijun Yin 0001 |
ACII | 2 |
| 2017 | EAC-Net: A Region-Based Deep Enhancing and Cropping Approach for Facial Action Unit DetectionabstractIn this paper, we propose a deep learning based approach for facial action unit detection by enhancing and cropping the regions of interest. The approach is implemented by adding two novel nets (layers): the enhancing layers and the cropping layers, to a pretrained CNN model. For the enhancing layers (the E-Net), we designed an attention map based on facial landmark features and applied it to a pretrained neural network to conduct enhanced learning. For the cropping layers (the C-Net), we crop facial regions around the detected landmarks and design convolutional layers to learn deeper features for each facial region. We then fuse the E-Net and the C-Net to obtain our Enhancing and Cropping (EAC) Net, which can learn both feature enhancing and region cropping functions. Our approach shows significant improvement in performance compared to the state-of-the-art methods applied to BP4D and DISFA AU datasets. Wei Li 0077, Farnaz Abtahi, Zhigang Zhu 0001, Lijun Yin 0001 |
FG | 4 |
| 2017 | FERA 2017 - Addressing Head Pose in the Third Facial Expression Recognition and Analysis ChallengeabstractThe field of Automatic Facial Expression Analysis has grown rapidly in recent years. However, despite progress in new approaches as well as benchmarking efforts, most evaluations still focus on either posed expressions, near-frontal recordings, or both. This makes it hard to tell how existing expression recognition approaches perform under conditions where faces appear in a wide range of poses (or camera views), displaying ecologically valid expressions. The main obstacle for assessing this is the availability of suitable data, and the challenge proposed here addresses this limitation. The FG 2017 Facial Expression Recognition and Analysis challenge (FERA 2017) extends FERA 2015 to the estimation of Action Units occurrence and intensity under different camera views. In this paper we present the third challenge in automatic recognition of facial expressions, to be held in conjunction with the 12th IEEE conference on Face and Gesture Recognition, May 2017, in Washington, United States. Two sub-challenges are defined: the detection of AU occurrence, and the estimation of AU intensity. In this work we outline the evaluation protocol, the data used, and the results of a baseline method for both sub-challenges. Michel F. Valstar, Enrique Sánchez-Lozano, Jeffrey F. Cohn, László A. Jeni, Jeffrey M. Girard, Zheng Zhang 0023, Lijun Yin 0001, Maja Pantic |
FG | 7 |
| 2017 | Combining gaze and demographic feature descriptors for autism classificationabstractPeople with autism suffer from social challenges and communication difficulties, which may prevent them from leading a fruitful and enjoyable life. It is imperative to diagnose and start treatments for autism as early as possible and, in order to do so, accurate methods of identifying the disorder are vital. We propose a novel method for classifying autism through the use of eye gaze and demographic feature descriptors that include a subject's age and gender. We construct feature descriptors that incorporate the subject's age and gender, as well as features based on eye gaze data. Using eye gaze information from the National Database for Autism Research, we tested our constructed feature descriptors on three different classifiers; random regression forests, C4.5 decision tree, and PART. Our proposed method for classifying autism resulted in a top classification rate of 96.2%. Shaun J. Canavan, Melanie Chen, Robert Valdez, Miles Yaeger, Huiyi Lin, Lijun Yin 0001 |
ICIP | 7 |
| 2017 | Hand gesture recognition using a skeleton-based feature representation with a random regression forestabstractIn this paper, we propose a method for automatic hand gesture recognition using a random regression forest with a novel set of feature descriptors created from skeletal data acquired from the Leap Motion Controller. The efficacy of our proposed approach is evaluated on the publicly available University of Padova Microsoft Kinect and Leap Motion dataset, as well as 24 letters of the English alphabet in American Sign Language. The letters that are dynamic (e.g. j and z) are not evaluated. Using a random regression forest to classify the features we achieve 100% accuracy on the University of Padova Microsoft Kinect and Leap Motion dataset. We also constructed an in-house dataset using the 24 static letters of the English alphabet in ASL. A classification rate of 98.36% was achieved on this dataset. We also show that our proposed method outperforms the current state of the art on the University of Padova Microsoft Kinect and Leap Motion dataset. Shaun J. Canavan, Walter Keyes, Ryan Mccormick, Julie Kunnumpurath, Tanner Hoelzel, Lijun Yin 0001 |
ICIP | 6 |
| 2017 | Spontaneous thermal facial expression analysis based on trajectory-pooled fisher vector descriptorabstractWe present a new descriptor for spontaneous facial expression recognition from videos acquired by a thermal sensor. Previous descriptors mostly compute features from RGB videos. It is difficult to process mixed and varied spontaneous expressions with a large ambiguity of facial appearances. In contrast, thermal imaging can measure autonomic activities, which are the physiological changes evoked by the autonomic nervous system regardless of the variety and ambiguity of facial appearances. This paper presents a new thermal video representation as so-called trajectory-pooled fisher vector descriptor (TFD). To get the local energy and temperature changes, we propose to use spatio-temporal orientation energy and acceleration of dense trajectory as low level features and further improve the discriminative capacity by aggregating the local feature using an improved fisher vector. The benefits of TFD in comparison with existing approaches are illustrated in two databases using different modalities: USTC-NVIE database and MMSE (a.k.a. BP4D+) database. Peng Liu 0039, Lijun Yin 0001 |
ICME | 2 |
| 2016 | Self-Adaptive Matrix Completion for Heart Rate Estimation from Face Videos under Realistic ConditionsabstractRecent studies in computer vision have shown that, while practically invisible to a human observer, skin color changes due to blood flow can be captured on face videos and, surprisingly, be used to estimate the heart rate (HR). While considerable progress has been made in the last few years, still many issues remain open. In particular, state of-the-art approaches are not robust enough to operate in natural conditions (e.g. in case of spontaneous movements, facial expressions, or illumination changes). Opposite to previous approaches that estimate the HR by processing all the skin pixels inside a fixed region of interest, we introduce a strategy to dynamically select face regions useful for robust HR estimation. Our approach, inspired by recent advances on matrix completion theory, allows us to predict the HR while simultaneously discover the best regions of the face to be used for estimation. Thorough experimental evaluation conducted on public benchmarks suggests that the proposed approach significantly outperforms state-of the-art HR estimation methods in naturalistic conditions. Sergey Tulyakov, Xavier Alameda-Pineda, Elisa Ricci 0001, Lijun Yin 0001, Jeffrey F. Cohn, Nicu Sebe |
CVPR | 4 |
| 2016 | Multimodal Spontaneous Emotion Corpus for Human Behavior AnalysisabstractEmotion is expressed in multiple modalities, yet most research has considered at most one or two. This stems in part from the lack of large, diverse, well-annotated, multimodal databases with which to develop and test algorithms. We present a well-annotated, multimodal, multidimensional spontaneous emotion corpus of 140 participants. Emotion inductions were highly varied. Data were acquired from a variety of sensors of the face that included high-resolution 3D dynamic imaging, high-resolution 2D video, and thermal (infrared) sensing, and contact physiological sensors that included electrical conductivity of the skin, respiration, blood pressure, and heart rate. Facial expression was annotated for both the occurrence and intensity of facial action units from 2D video by experts in the Facial Action Coding System (FACS). The corpus further includes derived features from 3D, 2D, and IR (infrared) sensors and baseline results for facial expression and action unit detection. The entire corpus will be made available to the research community. Zheng Zhang 0023, Jeffrey M. Girard, Yue Wu 0002, Xing Zhang 0012, Peng Liu 0039, Umur A. Ciftci, Shaun J. Canavan, Michael Reale, Andrew Horowitz, Huiyuan Yang, Jeffrey F. Cohn, Lijun Yin 0001 |
CVPR | 13 |
| 2015 | Perception driven 3D facial expression analysis based on reverse correlation and normal componentabstractResearch on automated facial expression analysis (FEA) has been focused on applying different feature extraction methods on texture space and geometric space, using holistic or local facial regions based on regular grids or facial anatomical structure. Not much work has been investigated by taking human perception into account. In this paper, we propose to study the facial expressive regions using a reverse correlation method, and further develop a novel 3D local normal component feature representation based on human perceptions. The classification image (CI) accumulated in multiple trials reveals the shape features which alter the neutral Mona Lisa portrait to positive and negative domains. The differences can be identified by both humans and machine. Based on the CI and the derived local feature regions, a novel 3D normal component based feature (3D-NLBP) is proposed to represent positive and negative expressions (e.g., happiness and sadness). This approach achieves a good performance and has been validated by testing on both high-resolution database and real-time low resolution depth map videos. Xing Zhang 0012, Zheng Zhang 0023, Lijun Yin 0001, Daniel Hipp, Peter Gerhardstein |
ACII | 3 |
| 2015 | Landmark localization on 3D/4D range data using a shape index-based statistical shape model with global and local constraints
Shaun J. Canavan, Peng Liu 0039, Xing Zhang 0012, Lijun Yin 0001 |
Comput. Vis. Image Underst. | 4 |
| 2014 | Best of Automatic Face and Gesture Recognition 2013
Stan Sclaroff, Lijun Yin 0001 |
Image Vis. Comput. | 3 |
| 2014 | BP4D-Spontaneous: a high-resolution spontaneous 3D dynamic facial expression database
Xing Zhang 0012, Lijun Yin 0001, Jeffrey F. Cohn, Shaun J. Canavan, Michael Reale, Andy Horowitz, Peng Liu 0039, Jeffrey M. Girard |
Image Vis. Comput. | 2 |
| 2013 | Multi-scale Topological Features for Hand Posture Representation and AnalysisabstractIn this paper, we propose a multi-scale topological feature representation for automatic analysis of hand posture. Such topological features have the advantage of being posture-dependent while being preserved under certain variations of illumination, rotation, personal dependency, etc. Our method studies the topology of the holes between the hand region and its convex hull. Inspired by the principle of Persistent Homology, which is the theory of computational topology for topological feature analysis over multiple scales, we construct the multi-scale Betti Numbers matrix (MSBNM) for the topological feature representation. In our experiments, we used 12 different hand postures and compared our features with three popular features (HOG, MCT, and Shape Context) on different data sets. In addition to hand postures, we also extend the feature representations to arm postures. The results demonstrate the feasibility and reliability of the proposed method. Kaoning Hu, Lijun Yin 0001 |
ICCV | 2 |
| 2013 | Fitting and tracking 3D/4D facial data using a temporal deformable shape modelabstractIn this paper, we propose a novel method for detecting and tracking landmark facial features on purely geometric 3D and 4D range models. Our proposed method involves fitting a new multi-frame constrained 3D temporal deformable shape model (TDSM) to range data sequences. We consider this a temporal based deformable model as we concatenate consecutive deformable shape models into a single model driven by the appearance of facial expressions. This allows us to simultaneously fit multiple models over a sequence of time with one TDSM. To our knowledge, it is the first work to address multiple shape models as a whole to track 3D dynamic range sequences without assistance of any texture information. The accuracy of the tracking results is evaluated by comparing the detected landmarks to the ground truth. The efficacy of the 3D feature detection and tracking over range model sequences has also been validated through an application in 3D geometric based face and expression analysis and expression sequence segmentation. We tested our method on the publicly available databases, BU-3DFE [15], BU-4DFE [16], and FRGC 2.0 [12]. We also validated our approach on our newly developed 3D dynamic spontaneous expression database [17]. Shaun J. Canavan, Xing Zhang 0012, Lijun Yin 0001 |
ICME | 3 |
| 2013 | Saliency-guided 3D head pose estimation on 3D expression modelsabstractHead pose is an important indicator of a person's attention, gestures, and communicative behavior with applications in human-computer interaction, multimedia, and vision systems. Robust head pose estimation is a prerequisite for spontaneous facial biometrics-related applications. However, most previous head pose estimation methods do not consider the facial expression and hence are more likely to be influenced by the facial expression. In this paper, we develop a saliency-guided 3D head pose estimation on 3D expression models. We address the problem of head pose estimation based on a generic model and saliency guided segmentation on a Laplacian fairing model. We propose to perform mesh Laplacian fairing to remove noise and outliers on the 3D facial model. The salient regions are detected and segmented from the model. The salient region Iterative Closest Point (ICP) then register the test face model with the generic head model. The algorithms for pose estimation are evaluated through both static and dynamic 3D facial databases. Overall, the extensive results demonstrate the effectiveness and accuracy of our approach. Peng Liu 0039, Michael Reale, Xing Zhang 0012, Lijun Yin 0001 |
ICMI | 4 |
| 2013 | Art Critic: Multisignal Vision and Speech Interaction System in a Gaming ContextabstractTrue immersion of a player within a game can only occur when the world simulated looks and behaves as close to reality as possible. This implies that the game must correctly read and understand, among other things, the player's focus, attitude toward the objects/persons in focus, gestures, and speech. In this paper, we proposed a novel system that integrates eye gaze estimation, head pose estimation, facial expression recognition, speech recognition, and text-to-speech components for use in real-time games. Both the eye gaze and head pose components utilize underlying 3-D models, and our novel head pose estimation algorithm uniquely combines scene flow with a generic head model. The facial expression recognition module uses the local binary patterns with three orthogonal planes approach on the 2-D shape index domain rather than the pixel domain, resulting in improved classification. Our system has also been extended to use a pan-tilt-zoom camera driven by the Kinect, allowing us to track a moving player. A test game, Art Critic, is also presented, which not only demonstrates the utility of our system but also provides a template for player/non-player character (NPC) interaction in a gaming context. The player alters his/her view of the 3-D world using head pose, looks at paintings/NPCs using eye gaze, and makes an evaluation based on the player's expression and speech. The NPC artist will respond with facial expression and synthetic speech based on its personality. Both qualitative and quantitative evaluations of the system are performed to illustrate the system's effectiveness. Michael Reale, Peng Liu 0039, Lijun Yin 0001, Shaun J. Canavan |
IEEE Trans. Cybern. | 3 |
| 2012 | 3D Head Pose Estimation Based on Scene Flow and Generic Head ModelabstractHead pose is an important indicator of a person's attention, gestures, and communicative behavior with applications in human computer interaction, multimedia and vision systems. In this paper, we present a novel head pose estimation system by performing head region detection using the Kinect [2], followed by face detection, feature tracking, and finally head pose estimation using an active camera. Ten feature points on the face are defined and tracked by an Active Appearance Model (AAM). We propose to use the scene flow approach to estimate the head pose from 2D video sequences. This estimation is based upon a generic 3D head model through the prior knowledge of the head shape and the geometric relationship between the 2D images and a 3D generic model. We have tested our head pose estimation algorithm with various cameras at various distances in real time. The experiments demonstrate the feasibility and advantages of our system. Peng Liu 0039, Michael Reale, Lijun Yin 0001 |
ICME | 3 |
| 2012 | A Global Parity Measure for Incomplete Point Cloud DataabstractAbstract Shapes with complex geometric and topological features such as tunnels, neighboring sheets, and cavities are susceptible to undersampling and continue to challenge existing reconstruction techniques. In this work we introduce a new measure for point clouds to determine the likely interior and exterior regions of an object. Specifically, we adapt the concept of parity to point clouds with missing data and introduce the parity map, a global measure of parity over the volume. We first examine how parity changes over the volume with respect to missing data and develop a method for extracting topologically correct interior and exterior crusts for estimating a signed distance field and performing surface reconstruction. We evaluate our approach on real scan data representing complex shapes with missing data. Our parity measure is not only able to identify highly confident interior and exterior regions but also localizes regions of missing data. Our reconstruction results are compared to existing methods and we show that our method faithfully captures the topology and geometry of complex shapes in the presence of missing data. Lee M. Seversky, Lijun Yin 0001 |
Comput. Graph. Forum | 2 |
| 2012 | Reshaping 3D facial scans for facial appearance modeling and 3D facial expression analysis
Yanhui Huang, Xing Zhang 0012, Yangyu Fan, Lijun Yin 0001, Lee M. Seversky, James Allen, Tao Lei 0002, Weijun Dong |
Image Vis. Comput. | 4 |
| 2012 | Static and dynamic 3D facial expression recognition: A comprehensive survey
Georgia Sandbach, Stefanos Zafeiriou, Maja Pantic, Lijun Yin 0001 |
Image Vis. Comput. | 4 |
| 2012 | 3D facial behaviour analysis and understanding
Stefanos Zafeiriou, Lijun Yin 0001 |
Image Vis. Comput. | 2 |
| 2011 | Reshaping 3D facial scans for facial appearance modeling and 3D facial expression analysisabstract3D face scans have been widely used for face modeling and face analysis. Due to the fact that face scans provide variable point clouds across frames, they may not capture complete facial data or miss point-to-point correspondences across various facial scans, thus causing difficulties to use such data for analysis. This paper presents an efficient approach to represent facial shapes from face scans through the reconstruction of face models based on regional information and a generic model. A hybrid approach using two vertex mapping algorithms, displacement mapping and point-to-surface mapping, and a regional blending algorithm are proposed to reconstruct the facial surface detail. The resulting models can represent individual facial shapes consistently and adaptively, establishing the facial point correspondence across individual models. The accuracy of the generated models is evaluated quantitatively. The applicability of the models is validated through the application for 3D facial expression recognition based on the databases of static 3DFE and dynamic 4DFE. A comparison with the state of the art has also been reported. Yanhui Huang, Xing Zhang 0012, Yangyu Fan, Lijun Yin 0001, Lee M. Seversky, Tao Lei 0002, Weijun Dong |
FG | 4 |
| 2011 | Recognizing face sketches by a large number of human subjects: A perception-based study for facial distinctivenessabstractUnderstanding how humans recognize face sketches drawn by artists is of significant value to both criminal investigators and researchers in computer vision, face biometrics and cognitive psychology. However, large scale experimental studies of hand-drawn face sketches are still very limited in terms of the number of artists, the number of sketches, and the number of human evaluators involved. In this paper, we reported the results of a series of psychological experiments in which 406 volunteers were asked to recognize 250 sketches drawn by 5 different artists. The primary findings are: (i) Sketch quality (artist factor) has a significant effect on human performance. Inter-artist variation as measured by the mean recognition rate can be as high as 31%; (ii) Participants showed a higher tendency to match multiple sketches to one photo than to second-guess their answers. The multi-match ratio seems correlated to the recognition rate, while second-guessing had no significant effect on human performance; (iii) For certain highly recognized faces, their rankings were very consistent using three measuring parameters: recognition rate, multi-match ratio, and second-guess ratio, suggesting that the three parameters could provide valuable information to quantify facial distinctiveness. Yong Zhang 0017, Steve Ellyson, Anthony Zone, Priyanka Gangam, John R. Sullins, Christine McCullough, Shaun J. Canavan, Lijun Yin 0001 |
FG | 8 |
| 2011 | 3D face sketch modeling and assessment for component based face recognitionabstract3D facial representations have been widely used for face recognition. There has been intensive research on geometric matching and similarity measurement on 3D range data and 3D geometric meshes of individual faces. However, little investigation has been done on geometric measurement for 3D sketch models. In this paper, we study the 3D face recognition from 3D face sketches which are derived from hand-drawn sketches and machine generated sketches. First, we have developed a 3D sketch modeling approach to create 3D facial sketch models from 2D facial sketch images. Second, we compared the 3D sketches to the existing 3D scans. Third, the 3D face similarity is measured between 3D sketches versus 3D scans, and 3D sketches versus 3D sketches based on the spatial Hidden Markov Model (HMM) classification. Experiments are conducted on both the BU-4DFE database and YSU face sketch database, resulting in a recognition rate at around 92% on average. Shaun J. Canavan, Xing Zhang 0012, Lijun Yin 0001, Yong Zhang 0017 |
IJCB | 3 |
| 2011 | Harmonic point cloud orientation
Lee M. Seversky, Matt S. Berger, Lijun Yin 0001 |
Comput. Graph. | 3 |
| 2011 | A Multi-Gesture Interaction System Using a 3-D Iris Disk Model for Gaze Estimation and an Active Appearance Model for 3-D Hand PointingabstractIn this paper, we present a vision-based human-computer interaction system, which integrates control components using multiple gestures, including eye gaze, head pose, hand pointing, and mouth motions. To track head, eye, and mouth movements, we present a two-camera system that detects the face from a fixed, wide-angle camera, estimates a rough location for the eye region using an eye detector based on topographic features, and directs another active pan-tilt-zoom camera to focus in on this eye region. We also propose a novel eye gaze estimation approach for point-of-regard (POR) tracking on a viewing screen. To allow for greater head pose freedom, we developed a new calibration approach to find the 3-D eyeball location, eyeball radius, and fovea position. Moreover, in order to get the optical axis, we create a 3-D iris disk by mapping both the iris center and iris contour points to the eyeball sphere. We then rotate the fovea accordingly and compute the final, visual axis gaze direction. This part of the system permits natural, non-intrusive, pose-invariant POR estimation from a distance without resorting to infrared or complex hardware setups. We also propose and integrate a two-camera hand pointing estimation algorithm for hand gesture tracking in 3-D from a distance. The algorithms of gaze pointing and hand finger pointing are evaluated individually, and the feasibility of the entire system is validated through two interactive information visualization applications. Michael Reale, Shaun J. Canavan, Lijun Yin 0001, Kaoning Hu, Terry Hung |
IEEE Trans. Multim. | 3 |
| 2010 | Pointing with the eyes: Gaze estimation using a static/active camera system and 3D iris disk modelabstractThe ability to capture the direction the eyes point in while the subject is a distance away from the camera offers the potential for intuitive human-computer interfaces, allowing for a greater interactivity, more intelligent HCI behavior, and increased flexibility. In this paper, we present a two-camera system that detects the face from a fixed, wide-angle camera, estimates a rough location for the eye region using an eye detector based on topographic features, and directs another active pan-tilt-zoom camera to focus in on this eye region. We also propose a novel eye gaze estimation approach for point-of-regard (PoG) tracking on a large viewing screen. To allow for greater head pose freedom, we developed a new calibration approach to find the 3D eyeball location, eyeball radius, and fovea position. Moreover, we map both the iris center and iris contour points to the eyeball sphere (creating a 3D iris disk) to get the optical axis; we then rotate the fovea accordingly and compute our final, visual axis gaze direction. We intend to integrate this gaze estimation approach with our two-camera system, permitting natural, non-intrusive, pose-invariant PoG estimation in distance and allowing user translational freedom without resorting to infrared or complex hardware setups such as stereo-cameras or “smart rooms". Michael Reale, Terry Hung, Lijun Yin 0001 |
ICME | 3 |
| 2010 | Expression-driven salient features: Bubble-based facial expression study by human and machineabstractHumans are able to recognize facial expressions of emotion from faces displaying a large set of confounding variables, including age, gender, ethnicity and other factors. Much work has been dedicated to attempts to characterize the process by which this highly developed capacity functions. In this paper, we propose to investigate local expression-driven features important to distinguishing facial expressions using a so-called `Bubbles' technique. The bubble technique is a kind of Gaussian masking to reveal information contributing to human perceptual categorization. We conducted experiments on factors from both human and machine. Observers are required to browse through the bubble-masked expression image and identify its category. By collecting responses from observers and analyzing them statistically we can find the facial features that humans employ for identifying different expressions. Humans appear to extract and use localized information specific to each expression for recognition. Additionally, we verify the findings by selecting the resulting features for expression classification using a conventional expression recognition algorithm with a public facial expression database. Xing Zhang 0012, Lijun Yin 0001, Peter Gerhardstein, Daniel Hipp |
ICME | 2 |
| 2010 | Evaluation of Multi-frame Fusion Based Face Classification Under ShadowabstractA video sequence of a head moving across a large pose angle contains much richer information than a single-view image, and hence has greater potential for identification purposes. This paper explores and evaluates the use of a multi-frame fusion method to improve face recognition in the presence of strong shadow. The dataset includes videos of 257 subjects who rotated their heads by 0° to 90°. Experiments were carried out using ten video frames per subject that were fused on the score level. The primary findings are: (i) A significant performance increase was observed, with the recognition rate being doubled from 40% using a single frame to 80% using ten frames; (ii) The performance of multi-frame fusion is strongly related to its inter-frame variation that measures its information diversity. Shaun J. Canavan, Benjamin Johnson 0004, Michael Reale, Yong Zhang 0017, Lijun Yin 0001, John R. Sullins |
ICPR | 5 |
| 2010 | Hand Pointing Estimation for Human Computer Interaction Based on Two Orthogonal-ViewsabstractHand pointing has been an intuitive gesture for human interaction with computers. Big challenges are still posted for accurate estimation of finger pointing direction in a 3D space. In this paper, we present a novel hand pointing estimation system based on two regular cameras, which includes hand region detection, hand finger estimation, two views' feature detection, and 3D pointing direction estimation. Based on the idea of binary pattern face detector, we extend the work to hand detection, in which a polar coordinate system is proposed to represent the hand region, and achieved a good result in terms of the robustness to hand orientation variation. To estimate the pointing direction, we applied an AAM based approach to detect and track 14 feature points along the hand contour from a top view and a side view. Combining two views of the hand features, the 3D pointing direction is estimated. The experiments have demonstrated the feasibility of the system. Kaoning Hu, Shaun J. Canavan, Lijun Yin 0001 |
ICPR | 3 |
| 2010 | Scalable Cage-Driven Feature Detection and Shape Correspondence for 3D Point SetsabstractWe propose an automatic deformation-driven correspondence algorithm for 3D point sets of non-rigid articulated shapes. Our approach uses simple geometric cages to embed the point set data and extract and match a coarse set of prominent features. We seek feature correspondences which lead to low-distortion deformations of the cages while satisfying the feature pairing. Our approach operates on the simplified geometric domain of the cage instead of the more complex 3D point data. Thus, it is robust to noise, partial occlusions, and insensitive to non-regular sampling. We demonstrate the potential of our approach by finding pairwise correspondences for sequences of acquired time-varying 3D scan point data. Lee M. Seversky, Lijun Yin 0001 |
ICPR | 2 |
| 2010 | Tracking Vertex Flow and Model Adaptation for Three-Dimensional Spatiotemporal Face AnalysisabstractResearch in the areas of 3-D face recognition and 3-D facial expression analysis has intensified in recent years. However, most research has been focused on 3-D static data analysis. In this paper, we investigate the facial analysis problem using dynamic 3-D face model sequences. One of the major obstacles for analyzing such data is the lack of correspondences of features due to the variable number of vertices across individual models or 3-D model sequences. In this paper, we present an effective approach for establishing vertex correspondences using a tracking-model-based approach for vertex registration, coarse-to-fine model adaptation, and vertex motion trajectory (called vertex flow) estimation. We propose to establish correspondences across frame models based on a 2-D intermediary, which is generated using conformal mapping and a generic model adaptation algorithm. Based on our newly created 3-D dynamic face database, we also propose to use a spatiotemporal hidden Markov model (ST-HMM) that incorporates 3-D surface feature characterization to learn the spatial and temporal information of faces. The advantage of using 3-D dynamic data for face recognition has been evaluated by comparing our approach to three conventional approaches: 2-D-video-based temporal HMM model, conventional 2-D-texture-based approach (e.g., Gabor-wavelet-based approach), and static 3-D-model-based approaches. To further evaluate the usefulness of vertex flow and the adapted model, we have also applied a spatial-temporal face model descriptor for facial expression classification based on dynamic 3-D model sequences. Xiaochen Chen, Matthew J. Rosato, Lijun Yin 0001 |
IEEE Trans. Syst. Man Cybern. Part A | 4 |
| 2009 | Dynamic face appearance modeling and sight direction estimation based on local region tracking and scale-space topo-representionabstractDynamic modeling of facial appearances and sight directions are demanded for HCI and multimedia applications. Traditional approaches for face tracking and eye tracking from 2D videos do not involve explicit facial modeling. In this paper, we propose to use an explicit 3D model to model the dynamic facial appearance as well as the eye shape to estimate the viewing direction. We apply active appearance models for local region tracking, and use a scale-space topographic representation for frame model instantiation. The individualized 3D models across video sequences allow us to estimate the iris viewing orientation dynamically. The proposed framework has been realized and tested in a person-independent fashion for AAM tracking and model instantiation using a single camera. Shaun J. Canavan, Lijun Yin 0001 |
ICME | 2 |
| 2008 | Facial Expression Recognition Based on 3D Dynamic Range Model Sequences
Lijun Yin 0001 |
ECCV (2) | 2 |
| 2008 | Multi-view facial expression recognitionabstractThe ability to handle multi-view facial expressions is important for computers to understand affective behavior under less constrained environment. However, most of existing methods for facial expression recognition are based on the near-frontal view face data, which are likely to fail in the non-frontal facial expression analysis. In this paper, we conduct an investigation on analyzing multi-view facial expressions. Three local patch descriptors (HoG, LBP, and SIFT) are used to extract facial features, which are the inputs to a nearest-neighbor indexing method that identifies facial expressions. We also investigate the influence of feature dimension reductions (PCA, LDA, and LPP) and classifier fusion on the recognition performance. We test our approaches on multi-view data generated from BU-3DFE 3D facial expression database that includes 100 subjects with 6 emotions and 4 intensity levels. Our extensive person-independent experiments suggest that the SIFT descriptor outperforms HoG and LBP, and LPP outperforms PCA and LDA in this application. But the classifier fusion does not show a significant advantage over SIFT-only classifier. Yuxiao Hu 0001, Zhihong Zeng, Lijun Yin 0001, Xiaozhou Wei, Thomas S. Huang |
FG | 3 |
| 2008 | Children's sensitivity to configural cues in faces undergoing rotational motionabstractOne account of face processing in childhood claims that featural cue processing dominates early in life, and more sophisticated configural cue processing skills emerge more gradually. The opposing view holds that most experimental demonstrations of diminished configural processing in early childhood reflect overall cognitive immaturity, not a true insensitivity to configural cues. This is typically assessed in discrimination tasks, via recognition accuracy with upright faces, as well as ldquoinversion effectsrdquo, relatively high performance with upright faces compared to relatively low performance with inverted faces. Larger relative inversion effects are taken to indicate proficiency in configural processing, which is specific to upright faces. In adults, the less sophisticated featural processing is assumed to proceed when configural cues are inaccessible, as in the case of an inverted face. Compared to adults, children typically have shown worse general performance with upright faces, as well as smaller inversion costs. The current study investigated the ability of motion to facilitate configural cue processing in 8 year-old children. The use of a 3D-based stimulus production system allowed us to generate facial stimuli that could be rotated in depth, which were presented in a typical discrimination task. The utility of rotational motion has seldom been explored in previous developmental tests. The current study failed to find evidence of a configural deficit in children relative to adults was not found. Additionally, the hypothesis that motion information would allow children to construct a stronger configural representation was not supported. These data are discussed in the context of task-specific encoding strategies. Gina Shroff, Peter Gerhardstein, Lijun Yin 0001 |
FG | 3 |
| 2008 | Recognizing partial facial action units based on 3D dynamic range data for facial expression recognitionabstractResearch on automatic facial expression recognition has benefited from work in psychology, specifically the Facial Action Coding System (FACS). To date, most existing approaches are primarily based on 2D images or videos. With the emergence of real-time 3D dynamic imaging technologies, however, 3D dynamic facial data is now available, thus opening up an alternative to detect facial action units in dynamic 3D space. In this paper, we investigate how to use this new modality to improve action unit (AU) detection. We select a subset of AUs from both the upper and lower parts of a facial area, apply the active appearance model (AAM) method and take the correspondence between textures and range models to track the pre-defined facial features across the 3D model sequences. A Hidden Markov Model (HMM) based classifier is employed to recognize the partial AUs. The experiments show that our 3D dynamic tracking based approach outperforms the compared 2D feature tracking based approach. The results are also comparable with the manually-picked 3D facial features based method. Finally, we extend our approach to validate the experiment for recognizing six prototypic facial expressions. Michael Reale, Lijun Yin 0001 |
FG | 3 |
| 2008 | A high-resolution 3D dynamic facial expression databaseabstractFace information processing relies on the quality of data resource. From the data modality point of view, a face database can be 2D or 3D, and static or dynamic. From the task point of view, the data can be used for research of computer based automatic face recognition, face expression recognition, face detection, or cognitive and psychological investigation. With the advancement of 3D imaging technologies, 3D dynamic facial sequences (called 4D data) have been used for face information analysis. In this paper, we focus on the modality of 3D dynamic data for the task of facial expression recognition. We present a newly created high-resolution 3D dynamic facial expression database, which is made available to the scientific research community. The database contains 606 3D facial expression sequences captured from 101 subjects of various ethnic backgrounds. The database has been validated through our facial expression recognition experiment using an HMM based 3D spatio-temporal facial descriptor. It is expected that such a database shall be used to facilitate the facial expression analysis from a static 3D space to a dynamic 3D space, with a goal of scrutinizing facial behavior at a higher level of detail in a real 3D spatio-temporal domain. Lijun Yin 0001, Xiaochen Chen, Tony Worm, Michael Reale |
FG | 1 |
| 2008 | A study of non-frontal-view facial expressions recognitionabstractThe existing methods of facial expression recognition are typically based on the near-frontal face data. The analysis of non-frontal-view facial expression is a largely unexplored research. The accessibility to a recent 3D facial expression database (BU-3DFE database) motivates us to explore an interesting question: whether non-frontal-view facial expression analysis can achieve the same as or better performance than the existing frontal-view facial expression method. Our extensive recognition experiments on data of 100 subjects with 5 yaw rotation view angles suggests that the non-frontal-view facial expression classification can outperform frontal-view facial expression recognition, given the manually labeled facial key points. Yuxiao Hu 0001, Zhihong Zeng, Lijun Yin 0001, Xiaozhou Wei, Jilin Tu, Thomas S. Huang |
ICPR | 3 |
| 2008 | Automatic pose estimation of 3D facial modelsabstractPose estimation plays an essential role in many computer vision applications, such as human computer interaction (HCI), driver attentiveness monitoring, face recognition, automatic face model editing, etc. In this paper, we propose a geometric feature based pose estimation approach based on 3D facial models. By identifying two clusters of inner eye corners of a 3D facial model, we find the nose tip with the aid of a facial reference plane. Then, a symmetry plane is generated and the pose orientation can be estimated. Our proposed algorithm is robust with respect to different persons, large pose variations, different expressions, partial facial data missing, and non-facial outliers. All computation is based on the pure 3D mesh model of a face without texture information. We tested our approach using 1,200 3D raw facial models and 1,200 corresponding clean facial models, and achieved 92.1% and 96.4% correct pose estimation rate, respectively. Lijun Yin 0001 |
ICPR | 2 |
| 2007 | Multiple-View Face Tracking For Modeling and Analysis Based On Non-Cooperative Video Imageryabstract3D face analysis has been researched intensively in recent decades. Most 3D data (so called range facial data) are obtained from 3D range imaging systems. Such data representations have been proven effective for face recognition in 3D space. However, obtaining such data requires subject cooperation in a constrained environment, which is not practical for many real applications of video surveillance. It is therefore in high demand to use regular video cameras to generate 3D face models for further classification. The goal of our research is to develop a method of tracking feature points on a face in multiple views in order to build 3D models of individual faces. We proposed a three-view based video tracking and model creation algorithm, which is based on the Active Appearance Model and a generic facial model. We will describe how to build useful individual models over time, and validate the created dynamic model sequences through the application of face recognition. Tracking multiple view fiducial points of a face in a time sequence can also be used for facial expression analysis. Our experiments demonstrated the feasibility of the proposed work. Scott Von Duhn, Lijun Yin 0001, Myung Jin Ko, Terry Hung |
CVPR | 2 |
| 2007 | Static topographic modeling for facial expression recognition and analysis
Lijun Yin 0001 |
Comput. Vis. Image Underst. | 2 |
| 2007 | Using geometric properties of topographic manifold to detect and track eyes for human-computer interactionabstractAutomatic eye detection and tracking is an important component for advanced human-computer interface design. Accurate eye localization can help develop a successful system for face recognition and emotion identification. In this article, we propose a novel approach to detect and track eyes using geometric surface features on topographic manifold of eye images. First, in the joint spatial-intensity domain, a facial image is treated as a 3D terrain surface or image topographic manifold. In particular, eye regions exhibit certain intrinsic geometric traits on this topographic manifold, namely, the pit -labeled center and hillside -like surround regions. Applying a terrain classification procedure on the topographic manifold of facial images, each location of the manifold can be labeled to generate a terrain map. We use the distribution of terrain labels to represent the eye terrain pattern. The Bhattacharyya affinity is employed to measure the distribution similarity between two topographic manifolds. Based on the Bhattacharyya kernel, a support vector machine is applied for selecting proper eye pairs from the pit-labeled candidates. Second, given detected eyes on the first frame of a video sequence, a mutual-information-based fitting function is defined to describe the similarity between two terrain surfaces of neighboring frames. By optimizing the fitting function, eye locations are updated for subsequent frames. The distinction of the proposed approach lies in that both eye detection and eye tracking are performed on the derived topographic manifold, rather than on an original-intensity image domain. The robustness of the approach is demonstrated under various imaging conditions and with different facial appearances, using both static images and video sequences without background constraints. Lijun Yin 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2006 | 3D Facial Expression Recognition Based on Primitive Surface Feature DistributionabstractThe creation of facial range models by 3D imaging systems has led to extensive work on 3D face recognition [19]. However, little work has been done to study the usefulness of such data for recognizing and understanding facial expressions. Psychological research shows that the shape of a human face, a highly mobile facial surface, is critical to facial expression perception. In this paper, we investigate the importance and usefulness of 3D facial geometric shapes to represent and recognize facial expressions using 3D facial expression range data. We propose a novel approach to extract primitive 3D facial expression features, and then apply the feature distribution to classify the prototypic facial expressions. In order to validate our proposed approach, we have conducted experiments for person-independent facial expression recognition using our newly created 3D facial expression database. We also demonstrate the advantages of our 3D geometric based approach over 2D texture based approaches in terms of various head poses. Lijun Yin 0001, Xiaozhou Wei |
CVPR (2) | 2 |
| 2006 | Efficient Hand Gesture Rendering and Decoding using a Simple Gesture LibraryabstractRecent work in hand gesture rendering and decoding has treated the two fields as separate and distinct. As the work of rendering evolves, it emphasizes exact movement replication, including more muscle and skeletal parameterization. The work in gesture decoding is largely centered on trained systems, which require large amounts of time in front of a camera rendering a gesture in order to decode movement. This paper presents a new scheme which more tightly couples the gesture rendering and decoding processes. While this scheme is simpler than existing techniques, the rendering remains natural looking, and decoding a new gesture does not require extensive training Lijun Yin 0001 |
ICME | 2 |
| 2006 | Real-time automatic 3D scene generation from natural language voice and text descriptionsabstractAutomatic scene generation using voice and text offers a unique multimedia approach to classic storytelling and human computer interaction with 3D graphics. In this paper, we present a newly developed system that generates 3D scenes from voice and text natural language input. Our system is intended to benefit non-graphics domain users and applications by providing advanced scene production through an automatic system. Scene descriptions are constructed in real-time using a method for depicting spatial relationships between and among different objects. Only the polygon representations of the objects are required for object placement. In addition, our system is robust. The system supports different quality polygon models such as those widely available on the Internet. Lee M. Seversky, Lijun Yin 0001 |
ACM Multimedia | 2 |
| 2005 | Eye Detection Under Unconstrained Background by the Terrain FeatureabstractLocating eyes in face images is an important step for automatic face analysis and recognition. In this paper, we present a novel approach for eye detection without finding the face region using topographic features. First, we argue that the eyes show certain topographic pattern if the gray-level face image is treated as a 3D terrain surface. Then a terrain map, which denotes the terrain type of each pixel, is derived from the original image by applying topographic classification approach. From the terrain map, eyes usually locate in the region around the pit-labeled pixels because of their intrinsic reflectance characteristic. At last, we construct Gaussian mixture model based probabilistic model to describe the distribution of pit-labeled candidates. The eye pair detection problem is transformed to maximize the probability that the selected candidate pair belongs to the eye space. Experiments show that our method has certain robustness to uncontrolled background Lijun Yin 0001 |
ICME | 2 |
| 2005 | Hyper-Resolution: Image detail reconstruction through parametric edges
Lijun Yin 0001, Matt T. Yourst |
Comput. Graph. | 1 |
| 2004 | On the importance of skin color for "other-race" effectabstractThe paper investigates the importance of facial skin color to account for the psychological discovery of the "other-race" effect, which is the phenomenon that other-race faces are perceived to be more alike and less discriminable than own-race faces. Skin colour is one of the important traits that could be used to distinguish people of one race from another. To understand better the role of skin color in the other-race effect, we propose to use a color transfer algorithm to simulate the effect by transplanting the skin color from one race to another. The processed face that is "made up" by another person's skin color can be used for the psychological study of race-related face identification. The principle of our color transfer method is to search the most similar pixel intensity in the face regions of two images based on the average intensity values and transfer the entire color mood of the source image to the target image. A color fine tune process is further performed on the characteristic areas to improve the result. The final psychophysical study shows that human skin color is one of the factors for the other-race effect, yet not the dominant one. Jingrong Jia, Lijun Yin 0001, Joseph P. Morrissey |
ICME | 2 |
| 2004 | Avatar-mediated face tracking and lip reading for human computer interactionabstractAdvanced human computer interaction requires automatic reading of human face in order to make the computer interact with human in the same way as human-to-human communication. We developed an automatic face tracking and lip reading system through a 3D face avatar to facilitate HCI applications in speech learning, emotional state monitoring, and non-verbal human computer interface design. The system implements a novel active face feature tracking algorithm with an uncalibrated camera. The 3D face pose is estimated and tracked by a Kalman filter-based matching process with a dynamic face model updating and constraint. The obtained facial motion parameters are transferred to an individualized 3D face avatar. As a result, a person's lip shape or expressions can be cloned to the animated 3D face avatar, by which all lip shapes from the same speech of different subjects can be easily compared and measured. This real time system targets the automatic facial expression analysis and synthesis for the next generation of HCI design. Xiaozhou Wei, Lijun Yin 0001 |
ACM Multimedia | 2 |
| 2004 | Facial expression representation and recognition based on texture augmentation and topographic maskingabstractThe variation of facial texture and surface due to the change of expression is an important cue for analyzing and modeling facial expressions. In this paper, we propose a new approach to represent the facial expression by using a so-called topographic feature. In order to capture the variation of facial surface structure, facial textures are processed by increasing the resolution. The topographical structure of human face is analyzed based on the resolution-enhanced textures. We investigate the relationship between the facial expression and its topographic features, and propose to represent the facial expression by the topographic labels. The detected topographic facial surface and the expressive regions reflect the status of facial skin movement. Based on the observation that the facial texture and its topographic features change along with facial expressions, we compare the disparity of these features between the neutral face and the expressive face to distinguish a number of universal expressions. The experiment demonstrates the feasibility of the proposed approach for facial expression representation and recognition. Lijun Yin 0001, Johnny Loi |
ACM Multimedia | 1 |
| 2004 | Generating 3D views of facial expressions from frontal face video based on topographic analysisabstractIn this paper, we report our newly developed 3D face modeling system with arbitrary expressions in a high level of detail using the topographic analysis and mesh instantiation process. Given a sequence of images of facial expressions at frontal views, we automatically generate 3D expressions at arbitrary views. Our face modeling system consists of two major components: facial surface representation using topographic analysis and generic model individualization based on labeled surface features and surface curvatures. The realism of the generated individual model is demonstrated through 3D views of facial expressions in videos. This work targets the accurate modeling of face and face expression for human computer interaction and 3D face recognition. Lijun Yin 0001, Kenneth Weiss 0001 |
ACM Multimedia | 1 |
| 2004 | Scalable edge enhancement with automatic optimization for digital radiographic images
Lijun Yin 0001, Anup Basu, Ja Kwei Chang |
Pattern Recognit. | 1 |
| 2003 | Recognizing facial expressions using active textures with wrinklesabstractThis paper explores the use of facial wrinkle textures for recognizing the facial expressions. Based on the observation of the wrinkles appearance and change along with performed expressions, we propose to extract the partial texture information in both the facial organ areas (e.g., eyes and mouth) and the facial wrinkle areas, and use the texture dissimilarity between the neutral expression and the active expression to extract the active texture for the expression representation. We present a novel method using multiple levels of detail to measure the active texture dissimilarity. The rate of change between levels is used as the rule for discriminating 6 types of universal expressions. The experiments on video sequences demonstrate the simplicity and efficiency of the proposed method for recognizing expressions with an 82.8% correct recognition rate. Lijun Yin 0001, Sergey Royt, Matt T. Yourst, Anup Basu |
ICME | 1 |
| 2003 | Automatic Lesion/Tumor Detection Using Intelligent Mesh-Based Active ContourabstractAutomatic detection of the lesion/tumor region is always of interest in medical imaging system. We address this issue in this paper for improving the accuracy and robustness as compared to the conventional methods. Active contour has been commonly used for the detection of irregular shape of region, however, it suffers the problem of the false attraction given the noisy image, and requires the correct estimation of the initial location of the object to be detected. In this paper, we present a novel method for robustly locating the object area by using a so-called intelligent mesh. With the accurate location and shape approximation in the initial stage, the object of interest is correctly detected by using a mesh-based active contour model. The correctness and robustness of the proposed algorithm are demonstrated on extracting the lesion/tumor regions of CT and Mammography images as an intensive test. Lijun Yin 0001, Sandeep Deshpande, Ja Kwei Chang |
ICTAI | 1 |
| 2002 | Color-based mouth shape tracking for synthesizing realistic facial expressionsabstractMouth shape analysis and synthesis play an important role for realistic facial expression generation. The conventional deformable template-based method requires a fairly accurate initial localization of the template because the energy minimization process only finds a local minimum. We present a new method for accurate mouth corner detection and mouth shape estimation using color information. The proposed method can deal with a variety of shapes for both open and closed positions of the mouth. Accurate mouth corner detection and mouth status determination (open and closed) increases the accuracy of template initialization. The extracted mouth shape parameters can be used for synthesizing realistic virtual facial expressions in model based coding. Experiments on video sequences demonstrate the advantage of the proposed algorithm. Lijun Yin 0001, Anup Basu |
ICIP (1) | 1 |
| 2002 | Synthesis-based scalable image enhancement for digital radiographyabstractThe ability to resolve fine picture detail is of paramount importance in a medical imaging system when it comes to view small tissues, bone structure and anatomy in X-ray images. In order to enhance diagnostic information and suppress irrelevant detail, we present a new digital X-ray imaging processing system with the property of scalability and adaptability. First, a new optimum contrast enhancement algorithm is proposed for display. The adaptive detection of the region-of-interest is developed. Second, a so called "scalable edge enhancement algorithm" is proposed to improve the image quality for showing subtle structures of the digital X-ray images. The advantage of the scheme is demonstrated by an experiments on 200 X-ray images, in which different parts of human body structures are captured. Lijun Yin 0001, Ja Kwei Chang, Anup Basu |
ICIP (2) | 1 |
| 2002 | Active Tracking and Cloning of Facial Expressions Using Spatio-Temporal InformationabstractThis paper presents a new method to analyze and synthesize facial expressions, in which a spatio-temporal gradient based method (i.e., optical flow) is exploited to estimate the movement of facial feature points. We proposed a method (called motion correlation) to improve the conventional block correlation method for obtaining motion vectors. The tracking of facial expressions under an active camera is addressed. With the motion vectors estimated, a facial expression can be cloned by adjusting the existing 3D facial model, or synthesized using different facial models. The experimental results demonstrate that the approach proposed is feasible for applications such as low bit rate video coding and face animation. Lijun Yin 0001, Anup Basu, Matt T. Yourst |
ICTAI | 1 |
| 2001 | Nose shape estimation and tracking for model-based codingabstractFeature extraction on the face plays an important role in applications of model based coding and human face recognition. Traditionally, the eyes and mouth are considered to be the most significant features contributing to different facial expressions. However, detecting and tracking the nose shape is non-trivial, and plays an equally important role as eyes and mouth for model based coding, especially for analysis and synthesis of realistic facial expressions. A feature detection method on the facial organ areas is presented. Individual templates are designed for the nostril and nose-side. First, the feature regions are limited to certain areas by using two-stage region growing methods. Second, the pre-defined templates are applied to extract the shape of the nostril and nose-side. Finally, the extracted feature shapes are exploited to guide a facial model to complete an accurate adaptation. The advantage of the proposed scheme is demonstrated by experiments on real video sequences for low bit rate video coding. Lijun Yin 0001, Anup Basu |
ICASSP | 1 |
| 2001 | Generating Realistic Facial Expressions with Wrinkles for Model-Based Coding
Lijun Yin 0001, Anup Basu |
Comput. Vis. Image Underst. | 1 |
| 2001 | Synthesizing realistic facial animations using energy minimization for model-based coding
Lijun Yin 0001, Anup Basu, Stefan Bernögger, Axel Pinz |
Pattern Recognit. | 1 |
| 2000 | Texture decomposition and correlation thresholding for realistic low-bit rate model-based codingabstractFacial texture updating and compression are crucial in achieving realistic facial animation for low-bit rate coding. We present an efficient way to update and encode facial textures in model-based coding. Differing from traditional methods which update the entire facial texture, we propose a partial texture updating method for realistic facial expression synthesis with facial wrinkles. Experiments on video sequences demonstrate the advantage of the proposed algorithm in keeping transmission cost low while producing realistic expressions. Lijun Yin 0001, Anup Basu |
ICASSP | 1 |
| 1999 | Face Model Adaptation with Active TrackingabstractA system for face model adaptation combining active tracking is presented. Input from an active camera is used for MPEG4 model based coding. First, the background is compensated considering a moving camera (tilt or pan). Second, the talking face is segmented from the compensated background using fusion of frame differences. A morphological filter is then applied to make the system less sensitive to noise. Third, Hough transform and deformable template coupled with color information are exploited to detect the facial features, e.g., eyes, mouth. Fourth, a wireframe model is adapted to the extracted face by an extended dynamic mesh. The feasibility of the proposed system is demonstrated using several real active video sequences. Lijun Yin 0001, Anup Basu |
ICIP (4) | 1 |
| 1999 | Integrating active face tracking with model based coding
Lijun Yin 0001, Anup Basu |
Pattern Recognit. Lett. | 1 |
| 1998 | Eye tracking and animation for MPEG-4 codingabstractAccurate localization and tracking of facial features are crucial for developing high quality model-based coding (MPEG-4) systems. For teleconferencing applications at very low bit rates, it is necessary to track eye and lip movements accurately over time. These movements can be coded and transmitted to a remote site, where animation techniques can be used to synthesize facial movements on a model of a face. In this paper we describe and simple heuristics which are effective in improving the results of well-known facial feature detection and tracking algorithms. Animation models are also presented, along with experimental results to demonstrate the system being developed. We focus our discussion only on the detection, tracking and modeling of eye movements. Stefan Bernögger, Lijun Yin 0001, Anup Basu, Axel Pinz |
ICPR | 2 |
| 1998 | Analysis and synthesis of facial expressions for MPEG-4 systemabstractThis paper presents a new method to analyze and synthesize facial expressions for model-based coding, which uses optical flow to estimate the movements of facial feature points and then obtains the motion vectors based on motion correlation and block correlation. With the motion vectors estimated, a facial expression can be synthesized by adjusting the existing 3-D facial model. Preliminary experimental results demonstrate that the approach proposed is feasible for very low bit rate transmission (e.g. MPEG-4 application). Lijun Yin 0001, Anup Basu |
SMC | 1 |
| 1997 | MPEG4 Face Modeling Using Fiducial PointsabstractWe present an approach for modeling a person's face for model-based coding (MPEG4) which will cater to very low bit-rate video-conferencing applications. This approach is performed entirely automatically. We utilize two views of a person's face, and the fiducial points defined on a generic facial model, to modify the generic model into a 3D individual model of a face. To deform the generated individual facial model for tracking the expression of a face over time, a new method named the layered force spreading method (LFSM) is proposed which makes the animation of facial expressions looks more natural. The feasibility of our approach is demonstrated using a real facial image. Lijun Yin 0001, Anup Basu |
ICIP (1) | 1 |
| 1996 | Constructing a 3D individualized head model from two orthogonal views
Horace Ho-Shing Ip, Lijun Yin 0001 |
Vis. Comput. | 2 |