Jeffrey F. Cohn

dblp:54/1171 · DBLP profile ↗
← Back
132ranked-venue papers
4as first author
17since 2021 · last 2026
0000-0002-9393-1116ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 93 · 1 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 64 · 1 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 27 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2
YearPublicationVenuePosition
2026 Inter-Stance: A Dyadic Multimodal Corpus for Conversational Stance Analysis
Xiang Zhang 0030, Taoyue Wang, Nan Bi, Cody Zhou, Zoie Wang, Yuming Su, Jeffrey F. Cohn, Lijun Yin 0001
FG10
2024 Expanding PyAFAR: A Novel Privacy-Preserving Infant AU Detector
abstract
We enhance PyAFAR11Code will be available on: https:\\affectanalysisgroup.github.io/PyAFARI, an open source, Python-based library for facial action unit detection by introducing a privacy-protected infant AU detector. To prevent reconstruction of the training images, we train the infant AU detector by extracting histogram of gradients (HoG) features and using an efficient Light Gradient Boosting Machine (LightGBM) classifier. Models are trained with two large, well-annotated databases. The performance of our approach is comparable to previously developed deep models that have not been released due to privacy concerns. Our models are available for use and further fine-tuning, contributing to the advancement of facial action unit detection.
Itir Önal, Saurabh Hinduja, Maneesh Bilalpur, Daniel S. Messinger, Jeffrey F. Cohn
FG5
2024 Mitigating Class Imbalance for Facial Expression Recognition Using SMOTE on Deep Features
abstract
Applications of affective computing are used in high-stake decision-making systems. Many of these systems that have applications in medicine, education, entertainment, and security user facial expression recognition (FER). Because of the significance of these applications, fair, unbiased, generalizable, and valid systems are essential. A source of error in FER can be class imbalance, which leads to bias against labels that occur less frequently. Considering this, we propose a solution to mitigate class imbalance for FER by synthesizing new deep features utilizing the Synthetic Minority Oversampling Technique (SMOTE) on deep features extracted from a neural network. We refer this to as a SMOTE layer in a neural network. More specifically, we up-sample the deep features of each minority class so that classes are balanced during training. To validate the efficacy of the proposed approach, we perform experiments on the publicly available datasets BP4D and AffectNet. We show encouraging results on BP4D, for pain recognition from facial expressions, and on AffectNet for 8-class classification of facial expressions.
Tara Nourivandi, Saurabh Hinduja, Shivam Srivastava, Jeffrey F. Cohn, Shaun J. Canavan
FG4
2024 SMURF: Statistical Modality Uniqueness and Redundancy Factorization
abstract
Multimodal late fusion is a well-performing fusion method that sums the outputs of separately processed modalities, so-called modality contributions, to create a prediction; for example, summing contributions from vision, acoustic, and language to predict affective states. In this paper, our primary goal is to improve the interpretability of what modalities contribute to the prediction in late fusion models. More specifically, we want to factorize modality contributions into what is consistently shared by at least two modalities (pairwise redundant contributions) and what the remaining modality-specific contributions are (unique contributions). Our secondary goal is to improve robustness to missing modalities by encouraging the model to learn redundant contributions. To achieve our two goals, we propose SMURF (Statistical Modality Uniqueness and Redundancy Factorization), a late fusion method that factorizes its outputs into a) unique contributions that are uncorrelated with all other modalities and b) pairwise redundant contributions that are maximally correlated between two modalities. For our primary goal, we 1) verify SMURF's factorization on a synthetic dataset, 2) ensure that its factorization does not degrade predictive performance on eight affective datasets, and 3) observe significant relationships between its factorization and human judgments on three datasets. For our secondary goal, we demonstrate that SMURF leads to more robustness to missing modalities at test time compared to three late fusion baselines.
Torsten Wörtwein, Nicholas B. Allen, Jeffrey F. Cohn, Louis-Philippe Morency
ICMI3
2024 Time to retire F1-binary score for action unit detection
abstract
Detecting action units is an important task in face analysis, especially in facial expression recognition. This is due, in part, to the idea that expressions can be decomposed into multiple action units. To evaluate systems that detect action units, F1-binary score is often used as the evaluation metric. In this paper, we argue that F1-binary score does not reliably evaluate these models due largely to class imbalance. Because of this, F1-binary score should be retired and a suitable replacement should be used. We justify this argument through a detailed evaluation of the negative influence of class imbalance on action unit detection. This includes an investigation into the influence of class imbalance in train and test sets and in new data (i.e., generalizability). We empirically show that F1-micro should be used as the replacement for F1-binary.
Saurabh Hinduja, Tara Nourivandi, Jeffrey F. Cohn, Shaun J. Canavan
Pattern Recognit. Lett.3
2024 Multimodal Prediction of Obsessive-Compulsive Disorder and Comorbid Depression Severity and Energy Delivered by Deep Brain Electrodes
abstract
To develop reliable, valid, and efficient measures of obsessive-compulsive disorder (OCD) severity, comorbid depression severity, and total electrical energy delivered (TEED) by deep brain stimulation (DBS), we trained and compared random forests regression models in a clinical trial of participants receiving DBS for refractory OCD. Six participants were recorded during open-ended interviews at pre- and post-surgery baselines and then at 3-month intervals following DBS activation. Ground-truth severity was assessed by clinical interview and self-report. Visual and auditory modalities included facial action units, head and facial landmarks, speech behavior and content, and voice acoustics. Mixed-effects random forest regression with Shapley feature reduction strongly predicted severity of OCD, comorbid depression, and total electrical energy delivered by the DBS electrodes (intraclass correlation, ICC, = 0.83, 0.87, and 0.81, respectively. When random effects were omitted from the regression, predictive power decreased to moderate for severity of OCD and comorbid depression and remained comparable for total electrical energy delivered (ICC = 0.60, 0.68, and 0.83, respectively). Multimodal measures of behavior outperformed ones from single modalities. Feature selection achieved large decreases in features and corresponding increases in prediction. The approach could contribute to closed-loop DBS that would automatically titrate DBS based on affect measures.
Saurabh Hinduja, Ali Darzi, Itir Önal, Nicole R. Provenza, Ron Gadot, Eric A. Storch, Sameer A. Sheth, Wayne K. Goodman, Jeffrey F. Cohn
IEEE Trans. Affect. Comput.9
2024 Disagreement Matters: Exploring Internal Diversification for Redundant Attention in Generic Facial Action Analysis
abstract
This paper demonstrates the effectiveness of a diversification mechanism for building a more robust multi-attention system in generic facial action analysis. While previous multi-attention (e.g., visual attention and self-attention) research on facial expression recognition (FER) and Action Unit (AU) detection have been thoroughly studied to focus on ”external attention diversification”, where attention branches localize different facial areas, we delve into the realm of ”internal attention diversification” and explore the impact of diverse attention patterns within the same Region of Interest (RoI). Our experiments reveal that variability in attention patterns significantly impacts model performance, indicating that unconstrained multi-attention plagued by redundancy and over-parameterization, leading to sub-optimal results. To tackle this issue, we propose a compact module that guides the model to achieve self-diversified multi-attention. Our method is applied to both CNN-based and Transformer-based models, benchmarked on popular databases such as BP4D and DISFA for AU detection, as well as CK+, MMI, BU-3DFE, and BP4D+ for facial expression recognition. We also evaluate the mechanism on Self-attention and Channel-wise attention designs for improving their adaptive capabilities in multi-modal feature fusion tasks. The multi-modal evaluation is conducted on BP4D, BP4D+, and our newly developed large-scale comprehensive emotion database BP4D++, which contains well-synchronized and aligned sensor modalities, addressing the scarcity of annotations and identities in human affective computing. We plan to release the new database to the research community, fostering further advancements in this field.
Zheng Zhang 0023, Xiang Zhang 0030, Taoyue Wang, Huiyuan Yang, Umur A. Ciftci, Jeffrey F. Cohn, Lijun Yin 0001
IEEE Trans. Affect. Comput.9
2023 Automated Emotional Valence Estimation in Infants with Stochastic and Strided Temporal Sampling
abstract
We propose the first automated approach to estimate the emotional valence of infants from their facial behavior. We use the state-of-the-art transformer-based video masked autoencoder (VideoMAE) that is pre-trained on a large video dataset as a backbone, and finetune it on two large, well-annotated infant video datasets (SIBSMILE and MODELING). To augment the limited data, we propose a novel video temporal augmentation method called Stochastic and Strided Temporal Sampling (SSTS). We demonstrate the effectiveness of our approach for infant valence estimation by achieving 0.671 Concordance Correlation Coefficient (CCC) on SIBSMILE and MODELING. The experiments show that SSTS remarkably accelerates the training speed by 8 times while gaining the best valence estimation performance. Lastly, we suggest that face detection and cropping (coarse registration) is a promising alternative to landmark-based registration (i.e. fine registration) in data pre-processing when accurate infant facial landmark detectors are inaccessible.
Mang Ning, Itir Önal, Daniel S. Messinger, Jeffrey F. Cohn, Albert Ali Salah
ACII4
2023 Multimodal Feature Selection for Detecting Mothers' Depression in Dyadic Interactions with their Adolescent Offspring
abstract
Depression is the most common psychological disorder, a leading cause of disability world-wide, and a major contributor to inter-generational transmission of psychopathology within families. To contribute to our understanding of depression within families and to inform modality selection and feature reduction, it is critical to identify interpretable features in developmentally appropriate contexts. Mothers with and without depression were studied. Depression was defined as history of treatment for depression and elevations in current or recent symptoms. We explored two multimodal feature selection strategies in dyadic interaction tasks of mothers with their adolescent children for depression detection. Modalities included face and head dynamics, facial action units, speech-related behavior, and verbal features. The initial feature space was vast and inter-correlated (collinear). To reduce dimensionality and gain insight into the relative contribution of each modality and feature, we explored feature selection strategies using Variance Inflation Factor (VIF) and Shapley values. On an average collinearity correction through VIF resulted in about 4 times feature reduction across unimodal and multimodal features. Collinearity correction was also found to be an optimal intermediate step prior to Shapley analysis. Shapley feature selection following VIF yielded best performance. The top 15 features obtained through Shapley achieved 78% accuracy. The most informative features came from all four modalities sampled, which supports the importance of multimodal feature selection.
Maneesh Bilalpur, Saurabh Hinduja, Laura A. Cariola, Lisa Sheeber, Nick Alien, László A. Jeni, Louis-Philippe Morency, Jeffrey F. Cohn
FG8
2023 Interpretation of Depression Detection Models via Feature Selection Methods
abstract
Given the prevalence of depression worldwide and its major impact on society, several studies employed artificial intelligence modelling to automatically detect and assess depression. However, interpretation of these models and cues are rarely discussed in detail in the AI community, but have received increased attention lately. In this study, we aim to analyse the commonly selected features using a proposed framework of several feature selection methods and their effect on the classification results, which will provide an interpretation of the depression detection model. The developed framework aggregates and selects the most promising features for modelling depression detection from 38 feature selection algorithms of different categories. Using three real-world depression datasets, 902 behavioural cues were extracted from speech behaviour, speech prosody, eye movement and head pose. To verify the generalisability of the proposed framework, we applied the entire process to depression datasets individually and when combined. The results from the proposed framework showed that speech behaviour features (e.g. pauses) are the most distinctive features of the depression detection model. From the speech prosody modality, the strongest feature groups were F0, HNR, formants, and MFCC, while for the eye activity modality they were left-right eye movement and gaze direction, and for the head modality it was yaw head movement. Modelling depression detection using the selected features (even though there are only 9 features) outperformed using all features in all the individual and combined datasets. Our feature selection framework did not only provide an interpretation of the model, but was also able to produce a higher accuracy of depression detection with a small number of features in varied datasets. This could help to reduce the processing time needed to extract features and creating the model.
Sharifa Alghowinem, Tom Gedeon, Roland Göcke, Jeffrey F. Cohn, Gordon Parker
IEEE Trans. Affect. Comput.4
2022 Ballistic Timing of Smiles is Robust to Context, Gender, Ethnicity, and National Differences
abstract
Smiles are highly variable. In some, contraction of the orbicularis oculi raises the cheeks and amplifies their intensity. In others, smile controls counteract the oblique pull of the zygomatic major, alter their shape, and decrease their intensity. Despite this variability, some features appear to be stereotypic. These features include a high correlation between the amplitude and velocity of smile onsets and same for smile offsets. The larger a smile's amplitude, the greater its velocity. This dependence is referred to as ballistic timing. In two relatively large publicly available databases (EB+ and Belfast), we tested the hypothesis of ballistic timing of smile onsets and offsets. We found high and consistent non-linear correlations between amplitude and velocity of both onsets and offsets that were robust to individual differences in persons (gender and ethnicity), context, presence or absence of the Duchenne marker, and country of residence (United States, Ireland, Peru). All$R_{2}$exceeded 0.85. The findings were highly consistent with ballistic timing. They have implications for detecting smiles that may be posed (which have been found to violate ballistic timing) and for realistic synthesis of smiles in social robots and virtual humans. Smiles that depart from ballistic timing are likely to be perceived as false or uncanny.
Maneesh Bilalpur, Saurabh Hinduja, Kenneth Goodrich, Jeffrey F. Cohn
ACII4
2022 Language Use in Mother-Adolescent Dyadic Interaction: Preliminary Results
abstract
This preliminary study applied a computer-assisted quantitative linguistic analysis to examine the effectiveness of language-based classification models to discriminate between mothers (n = 140) with and without history of treatment for depression (51% and 49%, respectively). Mothers were recorded during a problem-solving interaction with their adolescent child. Transcripts were manually annotated and analyzed using a dictionary-based, natural-language program approach (Linguistic Inquiry and Word Count). To assess the importance of linguistic features to correctly classify history of depression, we used Support Vector Machines (SVM) with interpretable features. Using linguistic features identified in the empirical literature, an initial SVM achieved nearly 63% accuracy. A second SVM using only the top 5 highest ranked SHAP features improved accuracy to 67.15%. The findings extend the existing literature base on understanding language behavior of depressed mood states, with a focus on the linguistic style of mothers with and without a history of treatment for depression and its potential impact on child development and trans-generational transmission of depression.
Laura A. Cariola, Saurabh Hinduja, Maneesh Bilalpur, Lisa Sheeber, Nicholas B. Allen, Louis-Philippe Morency, Jeffrey F. Cohn
ACII7
2022 Toward Causal Understanding of Therapist-Client Relationships: A Study of Language Modality and Social Entrainment
abstract
The relationship between a therapist and their client is one of the most critical determinants of successful therapy. The working alliance is a multifaceted concept capturing the collaborative aspect of the therapist-client relationship; a strong working alliance has been extensively linked to many positive therapeutic outcomes. Although therapy sessions are decidedly multimodal interactions, the language modality is of particular interest given its recognized relationship to similar dyadic concepts such as rapport, cooperation, and affiliation. Specifically, in this work we study language entrainment, which measures how much the therapist and client adapt toward each other’s use of language over time. Despite the growing body of work in this area, however, relatively few studies examine causal relationships between human behavior and these relationship metrics: does an individual’s perception of their partner affect how they speak, or does how they speak affect their perception? We explore these questions in this work through the use of structural equation modeling (SEM) techniques, which allow for both multilevel and temporal modeling of the relationship between the quality of the therapist-client working alliance and the participants’ language entrainment. In our first experiment, we demonstrate that these techniques perform well in comparison to other common machine learning models, with the added benefits of interpretability and causal analysis. In our second analysis, we interpret the learned models to examine the relationship between working alliance and language entrainment and address our exploratory research questions. The results reveal that a therapist’s language entrainment can have a significant impact on the client’s perception of the working alliance, and that the client’s language entrainment is a strong indicator of their perception of the working alliance. We discuss the implications of these results and consider several directions for future work in multimodality.
Alexandria K. Vail, Jeffrey M. Girard, Lauren M. Bylsma, Jeffrey F. Cohn, Jay Fournier, Holly Swartz, Louis-Philippe Morency
ICMI4
2021 Facial Action Units and Head Dynamics in Longitudinal Interviews Reveal OCD and Depression severity and DBS Energy
abstract
N euromodulation therapy, specifically Deep Brain Stimulation (DBS) of the ventral capsule/ventral striatum (VC/VS), is promising treatment for severe and intractable obsessive-compulsive disorder (OCD). To assess treatment response to DBS, reliable biomarkers are needed. We explored the hypothesis that facial action units and head dynamics in an interview context reveal severity of OCD, related depression, and DBS energy in participants undergoing DBS treatment. Participants were 5 patients (3 females, 2 males) with implanted DBS to VC/VS. They were recorded during brief open-ended interviews by a clinician at pre- and post-surgery baselines and then at 3-month intervals following activation of the DBS electrodes. Facial action units and head dynamics were assessed using AFAR (Automatic Facial Affect Recognition). OCD severity was assessed using clinical interview (YBOCS-II) and depression symptoms were assessed using participant self-report (BDI). After testing for multicollinearity and dropping highly-correlated features, a linear mixed-effects model using chi-square feature selection predicted 61% of the variation in YBOCS-II; 59% of the variation in BDI; and 37% of the variation in delivered energy by DBS to VC/VS. These findings suggest that automatically detected facial action units and head dynamics are potential biomarkers of OCD, depression se verity, and DBS energy.
Ali Darzi, Nicole R. Provenza, László A. Jeni, David A. Borton, Sameer A. Sheth, Wayne K. Goodman, Jeffrey F. Cohn
FG7
2021 Goals, Tasks, and Bonds: Toward the Computational Assessment of Therapist Versus Client Perception of Working Alliance
abstract
Early client dropout is one of the most significant challenges facing psychotherapy: recent studies suggest that at least one in five clients will leave treatment prematurely. Clients may terminate therapy for various reasons, but one of the most common causes is the lack of a strong working alliance. The concept of working alliance captures the collaborative relationship between a client and their therapist when working toward the progress and recovery of the client seeking treatment. Unfortunately, clients are often unwilling to directly express dissatisfaction in care until they have already decided to terminate therapy. On the other side, therapists may miss subtle signs of client discontent during treatment before it is too late. In this work, we demonstrate that nonverbal behavior analysis may aid in bridging this gap. The present study focuses primarily on the head gestures of both the client and therapist, contextualized within conversational turn-taking actions between the pair during psychotherapy sessions. We identify multiple behavior patterns suggestive of an individual's perspective on the working alliance; interestingly, these patterns also differ between the client and the therapist. These patterns inform the development of predictive models for self-reported ratings of working alliance, which demonstrate significant predictive power for both client and therapist ratings. Future applications of such models may stimulate preemptive intervention to strengthen a weak working alliance, whether explicitly attempting to repair the existing alliance or establishing a more suitable client-therapist pairing, to ensure that clients encounter fewer barriers to receiving the treatment they need.
Alexandria K. Vail, Jeffrey M. Girard, Lauren M. Bylsma, Jeffrey F. Cohn, Jay Fournier, Holly Swartz, Louis-Philippe Morency
FG4
2021 Human-Guided Modality Informativeness for Affective States
abstract
This paper studies the hypothesis that not all modalities are always needed to predict affective states. We explore this hypothesis in the context of recognizing three affective states that have shown a relation to a future onset of depression: positive, aggressive, and dysphoric. In particular, we investigate three important modalities for face-to-face conversations: vision, language, and acoustic modality. We first perform a human study to better understand which subset of modalities people find informative, when recognizing three affective states. As a second contribution, we explore how these human annotations can guide automatic affect recognition systems to be more interpretable while not degrading their predictive performance. Our studies show that humans can reliably annotate modality informativeness. Further, we observe that guided models significantly improve interpretability, i.e., they attend to modalities similarly to how humans rate the modality informativeness, while at the same time showing a slight increase in predictive performance.
Torsten Wörtwein, Lisa Sheeber, Nicholas B. Allen, Jeffrey F. Cohn, Louis-Philippe Morency
ICMI4
2021 Synthetic Expressions are Better Than Real for Learning to Detect Facial Actions
abstract
Critical obstacles in training classifiers to detect facial actions are the limited sizes of annotated video databases and the relatively low frequencies of occurrence of many actions. To address these problems, we propose an approach that makes use of facial expression generation. Our approach reconstructs the 3D shape of the face from each video frame, aligns the 3D mesh to a canonical view, and then trains a GAN-based network to synthesize novel images with facial action units of interest. To evaluate this approach, a deep neural network was trained on two separate datasets: One network was trained on video of synthesized facial expressions generated from FERA17; the other network was trained on unaltered video from the same database. Both networks used the same train and validation partitions and were tested on the test partition of actual video from FERA17. The network trained on synthesized facial expressions outperformed the one trained on actual facial expressions and surpassed current state-of-the-art approaches.
Koichiro Niinuma, Itir Önal, Jeffrey F. Cohn, László A. Jeni
WACV3
2020 Message from the General and Program Chairs FG 2020
abstract
Welcome to the 15th IEEE International Conference on Automatic Face and Gesture Recognition (FG 2020). FG is the premier international conference on vision-based automatic face and body behavior analysis and applications. Since its first meeting in Zurich in 1994, the conference has been held fourteen times throughout the world. This is the 15th conference.
Juan P. Wachs, Sergio Escalera, Jeffrey F. Cohn, Albert Ali Salah, Arun Ross
FG3
2020 Multimodal Interaction in Psychopathology
abstract
This paper presents an introduction to the Multimodal Interaction in Psychopathology workshop, which is held virtually in conjunction with the 22nd ACM International Conference on Multimodal Interaction on October 25th, 2020. This workshop has attracted submissions in the context of investigating multimodal interaction to reveal mechanisms and assess, monitor, and treat psychopathology. Keynote speakers from diverse disciplines present an overview of the field from different vantages and comment on future directions. Here we summarize the goals and the content of the workshop.
Itir Önal, Jeffrey F. Cohn, Hamdi Dibeklioglu
ICMI2
2019 Reconsidering the Duchenne Smile: Indicator of Positive Emotion or Artifact of Smile Intensity?
abstract
The Duchenne smile hypothesis is that smiles that include eye constriction (AU6) are the product of genuine positive emotion, whereas smiles that do not are either falsified or related to negative emotion. This hypothesis has become very influential and is often used in scientific and applied settings to justify the inference that a smile is either true or false. However, empirical support for this hypothesis has been equivocal and some researchers have proposed that, rather than being a reliable indicator of positive emotion, AU6 may just be an artifact produced by intense smiles. Initial support for this proposal has been found when comparing smiles related to genuine and feigned positive emotion; however, it has not yet been examined when comparing smiles related to genuine positive and negative emotion. The current study addressed this gap in the literature by examining spontaneous smiles from 136 participants during the elicitation of amusement, embarrassment, fear, and pain (from the BP4D+ dataset). Bayesian multilevel regression models were used to quantify the associations between AU6 and self-reported amusement while controlling for smile intensity. Models were estimated to infer amusement from AU6 and to explain the intensity of AU6 using amusement. In both cases, controlling for smile intensity substantially reduced the hypothesized association, whereas the effect of smile intensity itself was quite large and reliable. These results provide further evidence that the Duchenne smile is likely an artifact of smile intensity rather than a reliable and unique indicator of genuine positive emotion.
Jeffrey M. Girard, Gayatri Shandar, Zhun Liu, Jeffrey F. Cohn, Lijun Yin 0001, Louis-Philippe Morency
ACII4
2019 FACS3D-Net: 3D Convolution based Spatiotemporal Representation for Action Unit Detection
abstract
Most approaches to automatic facial action unit (AU) detection consider only spatial information and ignore AU dynamics. For humans, dynamics improves AU perception. Is same true for algorithms? To make use of AU dynamics, recent work in automated AU detection has proposed a sequential spatiotemporal approach: Model spatial information using a 2D CNN and then model temporal information using LSTM (Long-Short-Term Memory). Inspired by the experience of human FACS coders, we hypothesized that combining spatial and temporal information simultaneously would yield more powerful AU detection. To achieve this, we propose FACS3D-Net that simultaneously integrates 3D and 2D CNN. Evaluation was on the Expanded BP4D+ database of 200 participants. FACS3D-Net outperformed both 2D CNN and 2D CNN-LSTM approaches. Visualizations of learnt representations suggest that FACS3D-Net is consistent with the spatiotemporal dynamics attended to by human FACS coders. To the best of our knowledge, this is the first work to apply 3D CNN to the problem of AU detection.
Le Yang 0009, Itir Önal, Jeffrey F. Cohn, Zakia Hammal, Dongmei Jiang, Hichem Sahli
ACII3
2019 PAttNet: Patch-attentive deep network for action unit detection
Itir Önal, László A. Jeni, Jeffrey F. Cohn
BMVC3
2019 Unmasking the Devil in the Details: What Works for Deep Facial Action Coding?
Koichiro Niinuma, László A. Jeni, Itir Önal, Jeffrey F. Cohn
BMVC4
2019 Automated Measurement of Head Movement Synchrony during Dyadic Depression Severity Interviews
abstract
With few exceptions, most research in automated assessment of depression has considered only the patient's behavior to the exclusion of the therapist's behavior. We investigated the interpersonal coordination (synchrony) of head movement during patient-therapist clinical interviews. Our sample consisted of patients diagnosed with major depressive disorder. They were recorded in clinical interviews (Hamilton Rating Scale for Depression, HRSD) at 7-week intervals over a period of 21 weeks. For each session, patient and therapist 3D head movement was tracked from 2D videos. Head angles in the horizontal (pitch) and vertical (yaw) axes were used to measure head movement. Interpersonal coordination of head movement between patients and therapists was measured using windowed cross-correlation. Patterns of coordination in head movement were investigated using the peak picking algorithm. Changes in head movement coordination over the course of treatment were measured using a hierarchical linear model (HLM). The results indicated a strong effect for patient-therapist head movement synchrony. Within-dyad variability in head movement coordination was found to be higher than between-dyad variability, meaning that differences over time in a dyad were higher as compared to the differences between dyads. Head movement synchrony did not change over the course of treatment with change in depression severity. To the best of our knowledge, this study is the first attempt to analyze the mutual influence of patient-therapist head movement in relation to depression severity.
Shalini Bhatia, Roland Göcke, Zakia Hammal, Jeffrey F. Cohn
FG4
2019 Cross-domain AU Detection: Domains, Learning Approaches, and Measures
abstract
Facial action unit (AU) detectors have performed well when trained and tested within the same domain. Do AU detectors transfer to new domains in which they have not been trained? To answer this question, we review literature on cross-domain transfer and conduct experiments to address limitations of prior research. We evaluate both deep and shallow approaches to AU detection (CNN and SVM, respectively) in two large, well-annotated, publicly available databases, Expanded BP4D+ and GFT. The databases differ in observational scenarios, participant characteristics, range of head pose, video resolution, and AU base rates. For both approaches and databases, performance decreased with change in domain, often to below the threshold needed for behavioral research. Decreases were not uniform, however. They were more pronounced for GFT than for Expanded BP4D+ and for shallow relative to deep learning. These findings suggest that more varied domains and deep learning approaches may be better suited for promoting generalizability. Until further improvement is realized, caution is warranted when applying AU classifiers from one domain to another.
Itir Önal, Jeffrey F. Cohn, László A. Jeni, Zheng Zhang 0023, Lijun Yin 0001
FG2
2019 AFAR: A Deep Learning Based Tool for Automated Facial Affect Recognition
abstract
Automated facial affect recognition is crucial to multiple domains (e.g., health, education, entertainment). Commercial tools are available but costly and of unknown validity. Open-source ones [1] lack user-friendly GUI for use by non-programmers. For both types, evidence of domain transfer and options for retraining for use in new domains typically are lacking.
Itir Önal, László A. Jeni, Wanqiao Ding, Jeffrey F. Cohn
FG4
2019 Bag-of-Acoustic-Words for Mental Health Assessment: A Deep Autoencoding Approach
Wenchao Du, Louis-Philippe Morency, Jeffrey F. Cohn, Alan W. Black
INTERSPEECH3
2019 Learning facial action units with spatiotemporal cues and multi-label sampling
Wen-Sheng Chu, Fernando De la Torre, Jeffrey F. Cohn
Image Vis. Comput.3
2019 Editorial of Special Issue on Human Behaviour Analysis "In-the-Wild"
abstract
The papers in this special section focus on human face and body image analysis, one of the most researched objects. One of the main reasons behind this popularity lies in the numerous applications of automatic face and body gesture analysis algorithms, that span several fields such as Human-Computer and Human-Robot Interaction (facial expression/body gesture recognition for automatic analysis of affect), medicine and healthcare (detection of emotional and cognitive disorders), as well as biometrics (face recognition, gait recognition). The papers in this section focus on recent efforts towards catalysing progress in automatic analysis of human behaviour in uncontrolled, “in-the-wild” conditions. We summarize research efforts towards the development of research methodologies, database collections and benchmarks, as well as algorithms and systems for machine analysis of human behaviour, focusing on facial expressions, body gestures, speech, as well as various other sensors. We are delighted that the special issue includes authors both from academia as well as the industry.
Mihalis A. Nicolaou, Stefanos Zafeiriou, Irene Kotsia, Guoying Zhao 0001, Jeffrey F. Cohn
IEEE Trans. Affect. Comput.5
2018 Detecting Depression Severity by Interpretable Representations of Motion Dynamics
abstract
Recent breakthroughs in deep learning using automated measurement of face and head motion have made possible the first objective measurement of depression severity. While powerful, deep learning approaches lack interpretability. We developed an interpretable method of automatically measuring depression severity that uses barycentric coordinates of facial landmarks and a Lie-algebra based rotation matrix of 3D head motion. Using these representations, kinematic features are extracted, preprocessed, and encoded using Gaussian Mixture Models (GMM) and Fisher vector encoding. A multi-class SVM is used to classify the encoded facial and head movement dynamics into three levels of depression severity. The proposed approach was evaluated in adults with history of chronic depression. The method approached the classification accuracy of state-of-the-art deep learning while enabling clinically and theoretically relevant findings. The velocity and acceleration of facial movement strongly mapped onto depression severity symptoms consistent with clinical data and theory.
Anis Kacem 0001, Zakia Hammal, Mohamed Daoudi, Jeffrey F. Cohn
FG4
2018 Automated Affect Detection in Deep Brain Stimulation for Obsessive-Compulsive Disorder: A Pilot Study
abstract
Automated measurement of affective behavior in psychopathology has been limited primarily to screening and diagnosis. While useful, clinicians more often are concerned with whether patients are improving in response to treatment. Are symptoms abating, is affect becoming more positive, are unanticipated side effects emerging? When treatment includes neural implants, need for objective, repeatable biometrics tied to neurophysiology becomes especially pressing. We used automated face analysis to assess treatment response to deep brain stimulation (DBS) in two patients with intractable obsessive-compulsive disorder (OCD). One was assessed intraoperatively following implantation and activation of the DBS device. The other was assessed three months post-implantation. Both were assessed during DBS on and o conditions. Positive and negative valence were quantified using a CNN trained on normative data of 160 non-OCD participants. Thus, a secondary goal was domain transfer of the classifiers. In both contexts, DBS-on resulted in marked positive affect. In response to DBS-off, affect flattened in both contexts and alternated with increased negative affect in the outpatient setting. Mean AUC for domain transfer was 0.87. These findings suggest that parametric variation of DBS is strongly related to affective behavior and may introduce vulnerability for negative affect in the event that DBS is discontinued.
Jeffrey F. Cohn, László A. Jeni, Itir Önal, Donald Malone, Michael S. Okun, David A. Borton, Wayne K. Goodman
ICMI1
2018 Guest Editorial: The Computational Face
abstract
The papers in this special section examine the concept of automated face analysis (AFA). AFA has received special attention from the computer vision and pattern recognition communities. Research progress often gives the impression that problems such as face recognition and face detection are solved, at least for some scenarios. Several aspects of face analysis remain open problems, including the implementation of large scale face recognition/detection methods for in the wild images, emotion recognition, micro-expression analysis, and others. The community keeps making rapid progress on these topics, with continual improvement of current methods and creation of new ones that push the state-of-the-art. Applications are countless, including security and video surveillance, human computer/robot interaction, communication, entertainment, and commerce, while having an important social impact in assistive technologies for education and health. The importance of face analysis, together with the vast amount of work on the subject and the latest developments in the field, motivated us to organize a special section on this theme. The scope of the compilation comprises all aspects of face analysis from a computer vision perspective. Including, but not limited to: recognition, detection, alignment, reconstruction of faces, pose estimation of faces, gaze analysis, age, emotion, gender, and facial attributes estimation, and applications among others.
Sergio Escalera, Xavier Baró, Isabelle Guyon, Hugo Jair Escalante, Georgios Tzimiropoulos, Michel F. Valstar, Maja Pantic, Jeffrey F. Cohn, Takeo Kanade
IEEE Trans. Pattern Anal. Mach. Intell.8
2018 Viewpoint-Consistent 3D Face Alignment
abstract
Most approaches to face alignment treat the face as a 2D object, which fails to represent depth variation and is vulnerable to loss of shape consistency when the face rotates along a 3D axis. Because faces commonly rotate three dimensionally, 2D approaches are vulnerable to significant error. 3D morphable models, employed as a second step in 2D+3D approaches are robust to face rotation but are computationally too expensive for many applications, yet their ability to maintain viewpoint consistency is unknown. We present an alternative approach that estimates 3D face landmarks in a single face image. The method uses a regression forest-based algorithm that adds a third dimension to the common cascade pipeline. 3D face landmarks are estimated directly, which avoids fitting a 3D morphable model. The proposed method achieves viewpoint consistency in a computationally efficient manner that is robust to 3D face rotation. To train and test our approach, we introduce the Multi-PIE Viewpoint Consistent database. In empirical tests, the proposed method achieved simple yet effective head pose estimation and viewpoint consistency on multiple measures relative to alternative approaches.
Sergey Tulyakov, László A. Jeni, Jeffrey F. Cohn, Nicu Sebe
IEEE Trans. Pattern Anal. Mach. Intell.3
2018 Dynamic Multimodal Measurement of Depression Severity Using Deep Autoencoding
abstract
Depression is one of the most common psychiatric disorders worldwide, with over 350 million people affected. Current methods to screen for and assess depression depend almost entirely on clinical interviews and self-report scales. While useful, such measures lack objective, systematic, and efficient ways of incorporating behavioral observations that are strong indicators of depression presence and severity. Using dynamics of facial and head movement and vocalization, we trained classifiers to detect three levels of depression severity. Participants were a community sample diagnosed with major depressive disorder. They were recorded in clinical interviews (Hamilton Rating Scale for Depression, HRSD) at seven-week intervals over a period of 21 weeks. At each interview, they were scored by the HRSD as moderately to severely depressed, mildly depressed, or remitted. Logistic regression classifiers using leave-one-participant-out validation were compared for facial movement, head movement, and vocal prosody individually and in combination. Accuracy of depression severity measurement from facial movement dynamics was higher than that for head movement dynamics, and each was substantially higher than that for vocal prosody. Accuracy using all three modalities combined only marginally exceeded that of face and head combined. These findings suggest that automatic detection of depression severity from behavioral indicators in patients is feasible and that multimodal measures afford the most powerful detection.
Hamdi Dibeklioglu, Zakia Hammal, Jeffrey F. Cohn
IEEE J. Biomed. Health Informatics3
2017 Automatic action unit detection in infants using convolutional neural network
abstract
Action unit detection in infants relative to adults presents unique challenges. Jaw contour is less distinct, facial texture is reduced, and rapid and unusual facial movements are common. To detect facial action units in spontaneous behavior of infants, we propose a multi-label Convolutional Neural Network (CNN). Eighty-six infants were recorded during tasks intended to elicit enjoyment and frustration. Using an extension of FACS for infants (Baby FACS), over 230,000 frames were manually coded for ground truth. To control for chance agreement, inter-observer agreement between Baby-FACS coders was quantified using free-margin kappa. Kappa coefficients ranged from 0.79 to 0.93, which represents high agreement. The multi-label CNN achieved comparable agreement with manual coding. Kappa ranged from 0.69 to 0.93. Importantly, the CNN-based AU detection revealed the same change in findings with respect to infant expressiveness between tasks. While further research is needed, these findings suggest that automatic AU detection in infants is a viable alternative to manual coding of infant facial expression.
Zakia Hammal, Wen-Sheng Chu, Jeffrey F. Cohn, Carrie Heike, Matthew L. Speltz
ACII3
2017 Learning Spatial and Temporal Cues for Multi-Label Facial Action Unit Detection
abstract
Facial action units (AU) are the fundamental units to decode human facial expressions. At least three aspects affect performance of automated AU detection: spatial representation, temporal modeling, and AU correlation. Unlike most studies that tackle these aspects separately, we propose a hybrid network architecture to jointly model them. Specifically, spatial representations are extracted by a Convolutional Neural Network (CNN), which, as analyzed in this paper, is able to reduce person-specific biases caused by hand-crafted descriptors (e.g., HOG and Gabor). To model temporal dependencies, Long Short-Term Memory (LSTMs) are stacked on top of these representations, regardless of the lengths of input videos. The outputs of CNNs and LSTMs are further aggregated into a fusion network to produce per-frame prediction of 12 AUs. Our network naturally addresses the three issues together, and yields superior performance compared to existing methods that consider these issues independently. Extensive experiments were conducted on two large spontaneous datasets, GFT and BP4D, with more than 400,000 frames coded with 12 AUs. On both datasets, we report improvements over a standard multi-label CNN and feature-based state-of-the-art. Finally, we provide visualization of the learned AU models, which, to our best knowledge, reveal how machines see AUs for the first time.
Wen-Sheng Chu, Fernando De la Torre, Jeffrey F. Cohn
FG3
2017 Sayette Group Formation Task (GFT) Spontaneous Facial Expression Database
abstract
Despite the important role that facial expressions play in interpersonal communication and our knowledge that interpersonal behavior is influenced by social context, no currently available facial expression database includes multiple interacting participants. The Sayette Group Formation Task (GFT) database addresses the need for well-annotated video of multiple participants during unscripted interactions. The database includes 172,800 video frames from 96 participants in 32 three-person groups. To aid in the development of automated facial expression analysis systems, GFT includes expert annotations of FACS occurrence and intensity, facial landmark tracking, and baseline results for linear SVM, deep learning, active patch learning, and personalized classification. Baseline performance is quantified and compared using identical partitioning and a variety of metrics (including means and confidence intervals). The highest performance scores were found for the deep learning and active patch learning methods. Learn more at http://osf.io/7wcyz.
Jeffrey M. Girard, Wen-Sheng Chu, László A. Jeni, Jeffrey F. Cohn
FG4
2017 FERA 2017 - Addressing Head Pose in the Third Facial Expression Recognition and Analysis Challenge
abstract
The field of Automatic Facial Expression Analysis has grown rapidly in recent years. However, despite progress in new approaches as well as benchmarking efforts, most evaluations still focus on either posed expressions, near-frontal recordings, or both. This makes it hard to tell how existing expression recognition approaches perform under conditions where faces appear in a wide range of poses (or camera views), displaying ecologically valid expressions. The main obstacle for assessing this is the availability of suitable data, and the challenge proposed here addresses this limitation. The FG 2017 Facial Expression Recognition and Analysis challenge (FERA 2017) extends FERA 2015 to the estimation of Action Units occurrence and intensity under different camera views. In this paper we present the third challenge in automatic recognition of facial expressions, to be held in conjunction with the 12th IEEE conference on Face and Gesture Recognition, May 2017, in Washington, United States. Two sub-challenges are defined: the detection of AU occurrence, and the estimation of AU intensity. In this work we outline the evaluation protocol, the data used, and the results of a baseline method for both sub-challenges.
Michel F. Valstar, Enrique Sánchez-Lozano, Jeffrey F. Cohn, László A. Jeni, Jeffrey M. Girard, Zheng Zhang 0023, Lijun Yin 0001, Maja Pantic
FG3
2017 A Branch-and-Bound Framework for Unsupervised Common Event Discovery
Wen-Sheng Chu, Fernando De la Torre, Jeffrey F. Cohn, Daniel S. Messinger
Int. J. Comput. Vis.3
2017 Dense 3D face alignment from 2D video for real-time use
László A. Jeni, Jeffrey F. Cohn, Takeo Kanade
Image Vis. Comput.2
2017 Behavioral cues help predict impact of advertising on future sales
Gábor Szirtes, Javier Orozco, István Petrás, Dániel Szolgay, Ákos Utasi, Jeffrey F. Cohn
Image Vis. Comput.6
2017 Selective Transfer Machine for Personalized Facial Expression Analysis
abstract
Automatic facial action unit (AU) and expression detection from videos is a long-standing problem. The problem is challenging in part because classifiers must generalize to previously unknown subjects that differ markedly in behavior and facial morphology (e.g., heavy versus delicate brows, smooth versus deeply etched wrinkles) from those on which the classifiers are trained. While some progress has been achieved through improvements in choices of features and classifiers, the challenge occasioned by individual differences among people remains. Person-specific classifiers would be a possible solution but for a paucity of training data. Sufficient training data for person-specific classifiers typically is unavailable. This paper addresses the problem of how to personalize a generic classifier without additional labels from the test subject. We propose a transductive learning method, which we refer to as a Selective Transfer Machine (STM), to personalize a generic classifier by attenuating person-specific mismatches. STM achieves this effect by simultaneously learning a classifier and re-weighting the training samples that are most relevant to the test subject. We compared STM to both generic classifiers and cross-domain learning methods on four benchmarks: CK+ [44], GEMEP-FERA [67], RUFACS [4] and GFT [57]. STM outperformed generic classifiers in all.
Wen-Sheng Chu, Fernando De la Torre, Jeffrey F. Cohn
IEEE Trans. Pattern Anal. Mach. Intell.3
2016 Continuous Supervised Descent Method for Facial Landmark Localisation
Marc Oliu, Ciprian A. Corneanu, László A. Jeni, Jeffrey F. Cohn, Takeo Kanade, Sergio Escalera
ACCV (2)4
2016 Self-Adaptive Matrix Completion for Heart Rate Estimation from Face Videos under Realistic Conditions
abstract
Recent studies in computer vision have shown that, while practically invisible to a human observer, skin color changes due to blood flow can be captured on face videos and, surprisingly, be used to estimate the heart rate (HR). While considerable progress has been made in the last few years, still many issues remain open. In particular, state of-the-art approaches are not robust enough to operate in natural conditions (e.g. in case of spontaneous movements, facial expressions, or illumination changes). Opposite to previous approaches that estimate the HR by processing all the skin pixels inside a fixed region of interest, we introduce a strategy to dynamically select face regions useful for robust HR estimation. Our approach, inspired by recent advances on matrix completion theory, allows us to predict the HR while simultaneously discover the best regions of the face to be used for estimation. Thorough experimental evaluation conducted on public benchmarks suggests that the proposed approach significantly outperforms state-of the-art HR estimation methods in naturalistic conditions.
Sergey Tulyakov, Xavier Alameda-Pineda, Elisa Ricci 0001, Lijun Yin 0001, Jeffrey F. Cohn, Nicu Sebe
CVPR5
2016 Multimodal Spontaneous Emotion Corpus for Human Behavior Analysis
abstract
Emotion is expressed in multiple modalities, yet most research has considered at most one or two. This stems in part from the lack of large, diverse, well-annotated, multimodal databases with which to develop and test algorithms. We present a well-annotated, multimodal, multidimensional spontaneous emotion corpus of 140 participants. Emotion inductions were highly varied. Data were acquired from a variety of sensors of the face that included high-resolution 3D dynamic imaging, high-resolution 2D video, and thermal (infrared) sensing, and contact physiological sensors that included electrical conductivity of the skin, respiration, blood pressure, and heart rate. Facial expression was annotated for both the occurrence and intensity of facial action units from 2D video by experts in the Facial Action Coding System (FACS). The corpus further includes derived features from 3D, 2D, and IR (infrared) sensors and baseline results for facial expression and action unit detection. The entire corpus will be made available to the research community.
Zheng Zhang 0023, Jeffrey M. Girard, Yue Wu 0002, Xing Zhang 0012, Peng Liu 0039, Umur A. Ciftci, Shaun J. Canavan, Michael Reale, Andrew Horowitz, Huiyuan Yang, Jeffrey F. Cohn, Lijun Yin 0001
CVPR11
2016 Cross-Cultural Depression Recognition from Vocal Biomarkers
abstract
No studies have investigated cross-cultural and cross-language characteristics of depressed speech. We investigated the generalisability of a vocal biomarker-based approach to depression detection in clinical interviews recorded in three countries (Australia, the USA and Germany), two languages (German and English) and different accents (Australian and American). Several approaches to training and testing within and between datasets were evaluated. Using the same experimental protocol separately within each dataset, (cross-classification) accuracy was high.combining datasets, high accuracy was high again and consistent across language, recording environment, and culture. Training and testing between datasets, however, attenuated accuracy. These finding emphasize the importance of heterogeneous training sets for robust depression detection.
Sharifa Alghowinem, Roland Göcke, Julien Epps, Michael Wagner 0004, Jeffrey F. Cohn
INTERSPEECH5
2016 Seventh International Workshop on Human Behavior Understanding (HBU 2016)
abstract
With advances in pattern recognition and multimedia computing, it becomes possible to analyze human behavior via multimodal sensors at varying time-scales, levels of analysis, and meaning. This ability opens up far-ranging possibilities for multimedia and multimodal interaction. Research has the, potential to endow computers with the capacity to detect and understand people's actions and activities and infer their attitudes, preferences, personality, and social relationships. This workshop brings together researchers in this rapidly emerging area and especially those concerned with behavior analysis and multimedia in children.
Mohamed Chetouani, Jeffrey F. Cohn, Albert Ali Salah
ACM Multimedia2
2016 Editorial of special issue on spontaneous facial behaviour analysis
Stefanos Zafeiriou, Guoying Zhao 0001, Matti Pietikäinen, Rama Chellappa, Irene Kotsia, Jeffrey F. Cohn
Comput. Vis. Image Underst.6
2016 Cascade of Tasks for facial expression analysis
Wen-Sheng Chu, Fernando De la Torre, Jeffrey F. Cohn
Image Vis. Comput.4
2016 Survey on RGB, 3D, Thermal, and Multimodal Approaches for Facial Expression Recognition: History, Trends, and Affect-Related Applications
abstract
Facial expressions are an important way through which humans interact socially. Building a system capable of automatically recognizing facial expressions from images and video has been an intense field of study in recent years. Interpreting such expressions remains challenging and much research is needed about the way they relate to human affect. This paper presents a general overview of automatic RGB, 3D, thermal and multimodal facial expression analysis. We define a new taxonomy for the field, encompassing all steps from face detection to facial expression recognition, and describe and classify the state of the art methods accordingly. We also present the important datasets and the bench-marking of most influential methods. We conclude with a general discussion about trends, important questions and future lines of research.
Ciprian A. Corneanu, Marc Oliu, Jeffrey F. Cohn, Sergio Escalera
IEEE Trans. Pattern Anal. Mach. Intell.3
2016 Confidence Preserving Machine for Facial Action Unit Detection
abstract
Facial action unit (AU) detection from video has been a long-standing problem in the automated facial expression analysis. While progress has been made, accurate detection of facial AUs remains challenging due to ubiquitous sources of errors, such as inter-personal variability, pose, and low-intensity AUs. In this paper, we refer to samples causing such errors as hard samples, and the remaining as easy samples. To address learning with the hard samples, we propose the confidence preserving machine (CPM), a novel two-stage learning framework that combines multiple classifiers following an "easy-to-hard" strategy. During the training stage, CPM learns two confident classifiers. Each classifier focuses on separating easy samples of one class from all else, and thus preserves confidence on predicting each class. During the test stage, the confident classifiers provide "virtual labels" for easy test samples. Given the virtual labels, we propose a quasi-semi-supervised (QSS) learning strategy to learn a person-specific classifier. The QSS strategy employs a spatio-temporal smoothness that encourages similar predictions for samples within a spatio-temporal neighborhood. In addition, to further improve detection performance, we introduce two CPM extensions: iterative CPM that iteratively augments training samples to train the confident classifiers, and kernel CPM that kernelizes the original CPM model to promote nonlinearity. Experiments on four spontaneous data sets GFT, BP4D, DISFA, and RU-FACS illustrate the benefits of the proposed CPM models over baseline methods and the state-of-the-art semi-supervised learning and transfer learning methods.
Jiabei Zeng, Wen-Sheng Chu, Fernando De la Torre, Jeffrey F. Cohn, Zhang Xiong 0001
IEEE Trans. Image Process.4
2016 Joint Patch and Multi-label Learning for Facial Action Unit and Holistic Expression Recognition
abstract
Most action unit (AU) detection methods use one-versus-all classifiers without considering dependences between features or AUs. In this paper, we introduce a joint patch and multi-label learning (JPML) framework that models the structured joint dependence behind features, AUs, and their interplay. In particular, JPML leverages group sparsity to identify important facial patches, and learns a multi-label classifier constrained by the likelihood of co-occurring AUs. To describe such likelihood, we derive two AU relations, positive correlation and negative competition, by statistically analyzing more than 350,000 video frames annotated with multiple AUs. To the best of our knowledge, this is the first work that jointly addresses patch learning and multi-label learning for AU detection. In addition, we show that JPML can be extended to recognize holistic expressions by learning common and specific patches, which afford a more compact representation than the standard expression recognition methods. We evaluate JPML on three benchmark datasets CK+, BP4D, and GFT, using within-and cross-dataset scenarios. In four of five experiments, JPML achieved the highest averaged F1 scores in comparison with baseline and alternative methods that use either patch learning or multi-label learning alone.
Kaili Zhao, Wen-Sheng Chu, Fernando De la Torre, Jeffrey F. Cohn, Honggang Zhang 0002
IEEE Trans. Image Process.4
2015 What can head and facial movements convey about positive and negative affect?
abstract
We investigated whether the dynamics of head and facial movements apart from specific facial expressions communicate affect in infants. Age-appropriate tasks were used to elicit positive and negative affect in 28 ethnically diverse 12-month-old infants. 3D head and facial movements were tracked from 2D video. Strong effects were found for both head and facial movements. For head movement, angular velocity and angular acceleration of pitch, yaw, and roll were higher during negative relative to positive affect. For facial movement, displacement, velocity, and acceleration also increased during negative relative to positive affect. Our results suggest that the dynamics of head and facial movements communicate affect at ages as young as 12 months. These findings deepen our understanding of emotion communication and provide a basis for studying individual differences in emotion in socio-emotional development.
Zakia Hammal, Jeffrey F. Cohn, Carrie Heike, Matthew L. Speltz
ACII2
2015 Joint patch and multi-label learning for facial action unit detection
abstract
The face is one of the most powerful channel of nonverbal communication. The most commonly used taxonomy to describe facial behaviour is the Facial Action Coding System (FACS). FACS segments the visible effects of facial muscle activation into 30+ action units (AUs). AUs, which may occur alone and in thousands of combinations, can describe nearly all-possible facial expressions. Most existing methods for automatic AU detection treat the problem using one-vs-all classifiers and fail to exploit dependencies among AU and facial features. We introduce joint-patch and multi-label learning (JPML) to address these issues. JPML leverages group sparsity by selecting a sparse subset of facial patches while learning a multi-label classifier. In four of five comparisons on three diverse datasets, CK+, GFT, and BP4D, JPML produced the highest average F1 scores in comparison with state-of-the art.
Kaili Zhao, Wen-Sheng Chu, Fernando De la Torre, Jeffrey F. Cohn, Honggang Zhang 0002
CVPR4
2015 Unsupervised Synchrony Discovery in Human Interaction
abstract
People are inherently social. Social interaction plays an important and natural role in human behavior. Most computational methods focus on individuals alone rather than in social context. They also require labelled training data. We present an unsupervised approach to discover interpersonal synchrony, referred as to two or more persons preforming common actions in overlapping video frames or segments. For computational efficiency, we develop a branch-and-bound (B&B) approach that affords exhaustive search while guaranteeing a globally optimal solution. The proposed method is entirely general. It takes from two or more videos any multi-dimensional signal that can be represented as a histogram. We derive three novel bounding functions and provide efficient extensions, including multi-synchrony detection and accelerated search, using a warm-start strategy and parallelism. We evaluate the effectiveness of our approach in multiple databases, including human actions using the CMU Mocap dataset [1], spontaneous facial behaviors using group-formation task dataset [37] and parent-infant interaction dataset [28].
Wen-Sheng Chu, Jiabei Zeng, Fernando De la Torre, Jeffrey F. Cohn, Daniel S. Messinger
ICCV4
2015 Confidence Preserving Machine for Facial Action Unit Detection
abstract
Varied sources of error contribute to the challenge of facial action unit detection. Previous approaches address specific and known sources. However, many sources are unknown. To address the ubiquity of error, we propose a Confident Preserving Machine (CPM) that follows an easy-to-hard classification strategy. During training, CPM learns two confident classifiers. A confident positive classifier separates easily identified positive samples from all else, a confident negative classifier does same for negative samples. During testing, CPM then learns a person-specific classifier using "virtual labels" provided by confident classifiers. This step is achieved using a quasi-semi-supervised (QSS) approach. Hard samples are typically close to the decision boundary, and the QSS approach disambiguates them using spatio-temporal constraints. To evaluate CPM, we compared it with a baseline single-margin classifier and state-of-the-art semi-supervised learning, transfer learning, and boosting methods in three datasets of spontaneous facial behavior. With few exceptions, CPM outperformed baseline and state-of-the art methods.
Jiabei Zeng, Wen-Sheng Chu, Fernando De la Torre, Jeffrey F. Cohn, Zhang Xiong 0001
ICCV4
2015 Multimodal Detection of Depression in Clinical Interviews
abstract
Current methods for depression assessment depend almost entirely on clinical interview or self-report ratings. Such measures lack systematic and efficient ways of incorporating behavioral observations that are strong indicators of psychological disorder. We compared a clinical interview of depression severity with automatic measurement in 48 participants undergoing treatment for depression. Interviews were obtained at 7-week intervals on up to four occasions. Following standard cut-offs, participants at each session were classified as remitted, intermediate, or depressed. Logistic regression classifiers using leave-one-out validation were compared for facial movement dynamics, head movement dynamics, and vocal prosody individually and in combination. Accuracy (remitted versus depressed) for facial movement dynamics was higher than that for head movement dynamics; and each was substantially higher than that for vocal prosody. Accuracy for all three modalities together reached 88.93%, exceeding that for any single modality or pair of modalities. These findings suggest that automatic detection of depression from behavioral indicators is feasible and that multimodal measures afford most powerful detection.
Hamdi Dibeklioglu, Zakia Hammal, Ying Yang 0007, Jeffrey F. Cohn
ICMI4
2015 Estimating smile intensity: A better way
Jeffrey M. Girard, Jeffrey F. Cohn, Fernando De la Torre
Pattern Recognit. Lett.2
2015 Head Movement Dynamics during Play and Perturbed Mother-Infant Interaction
abstract
We investigated the dynamics of head movement in mothers and infants during an age-appropriate, well-validated emotion induction, the Still Face paradigm. In this paradigm, mothers and infants play normally for 2 minutes (Play) followed by 2 minutes in which the mothers remain unresponsive (Still Face), and then two minutes in which they resume normal behavior (Reunion). Participants were 42 ethnically diverse 4-month-old infants and their mothers. Mother and infant angular displacement and angular velocity were measured using the CSIRO head tracker. In male but not female infants, angular displacement increased from Play to Still-Face and decreased from Still Face to Reunion. Infant angular velocity was higher during Still-Face than Reunion with no differences between male and female infants. Windowed cross-correlation suggested changes in how infant and mother head movements are associated, revealing dramatic changes in direction of association. Coordination between mother and infant head movement velocity was greater during Play compared with Reunion. Together, these findings suggest that angular displacement, angular velocity and their coordination between mothers and infants are strongly related to age-appropriate emotion challenge. Attention to head movement can deepen our understanding of emotion communication.
Zakia Hammal, Jeffrey F. Cohn, Daniel S. Messinger
IEEE Trans. Affect. Comput.2
2015 Predicting Ad Liking and Purchase Intent: Large-Scale Analysis of Facial Responses to Ads
abstract
Billions of online video ads are viewed every month. We present a large-scale analysis of facial responses to video content measured over the Internet and their relationship to marketing effectiveness. We collected over 12,000 facial responses from 1,223 people to 170 ads from a range of markets and product categories. The facial responses were automatically coded frame-by-frame. Collection and coding of these 3.7 million frames would not have been feasible with traditional research methods. We show that detected expressions are sparse but that aggregate responses reveal rich emotion trajectories. By modeling the relationship between the facial responses and ad effectiveness, we show that ad liking can be predicted accurately (ROC AUC = 0.85) from webcam facial responses. Furthermore, the prediction of a change in purchase intent is possible (ROC AUC = 0.78). Ad liking is shown by eliciting expressions, particularly positive expressions. Driving purchase intent is more complex than just making viewers smile: peak positive responses that are immediately preceded by a brand appearance are more likely to be effective. The results presented here demonstrate a reliable and generalizable system for predicting ad effectiveness automatically from facial responses without a need to elicit self-report responses from the viewers. In addition we can gain insight into the structure of effective ads.
Daniel McDuff, Rana El Kaliouby, Jeffrey F. Cohn, Rosalind W. Picard
IEEE Trans. Affect. Comput.3
2014 Spatio-temporal Event Classification Using Time-Series Kernel Based Structured Sparsity
László A. Jeni, András Lörincz, Zoltán Szabó 0001, Jeffrey F. Cohn, Takeo Kanade
ECCV (4)4
2014 Jointly detecting infants' multiple facial action units expressed during spontaneous face-to-face communication
abstract
Automatic detection of spontaneous facial Action Units (AUs) in video has many applications including understanding infants' emotion-mediated interactions and development. The target AUs for detection are those essential to positive and negative emotion (i.e., AU 6, AU 12, and AU 20). Tracking and extraction of facial features is especially challenging in infants. Face shape and texture markedly differ from that in adults, jaw contour often is occult, sudden changes in pose and expression are common, and AU often occur in complex combinations. We investigate the association among AUs central to positive and negative emotion and propose a methodology for jointly detecting positively correlated facial AUs of infants during spontaneous interactions with their parents. We apply a subject-independent structural output model to (1) recognize combinations of AUs simultaneously, and (2) model the dependencies between AUs. Using this approach, we improved the reliability of automatic detection of AU 12 and AU 20 in a total 90-minute video of infant-parent interaction of 12 infants.
Nazanin Zaker, Mohammad H. Mahoor, Daniel S. Messinger, Jeffrey F. Cohn
ICIP4
2014 Dyadic Behavior Analysis in Depression Severity Assessment Interviews
abstract
Previous literature suggests that depression impacts vocal timing of both participants and clinical interviewers but is mixed with respect to acoustic features. To investigate further, 57 middle-aged adults (men and women) with Major Depression Disorder and their clinical interviewers (all women) were studied. Participants were interviewed for depression severity on up to four occasions over a 21 week period using the Hamilton Rating Scale for Depression (HRSD), which is a criterion measure for depression severity in clinical trials. Acoustic features were extracted for both participants and interviewers using COVAREP Toolbox. Missing data occurred due to missed appointments, technical problems, or insufficient vocal samples. Data from 36 participants and their interviewers met criteria and were included for analysis to compare between high and low depression severity. Acoustic features for participants varied between men and women as expected, and failed to vary with depression severity for participants. For interviewers, acoustic characteristics strongly varied with severity of the interviewee's depression. Accommodation - the tendency of interactants to adapt their communicative behavior to each other - between interviewers and interviewees was inversely related to depression severity. These findings suggest that interviewers modify their acoustic features in response to depression severity, and depression severity strongly impacts interpersonal accommodation.
Stefan Scherer, Zakia Hammal, Ying Yang 0007, Louis-Philippe Morency, Jeffrey F. Cohn
ICMI5
2014 A lp-norm MTMKL framework for simultaneous detection of multiple facial action units
abstract
Facial action unit (AU) detection is a challenging topic in computer vision and pattern recognition. Most existing approaches design classifiers to detect AUs individually or AU combinations without considering the intrinsic relations among AUs. This paper presents a novel method, lp-norm multi-task multiple kernel learning (MTMKL), that jointly learns the classifiers for detecting the absence and presence of multiple AUs. lp-norm MTMKL is an extension of the regularized multi-task learning, which learns shared kernels from a given set of base kernels among all the tasks within Support Vector Machines (SVM). Our approach has several advantages over existing methods: (1) AU detection work is transformed to a MTL problem, where given a specific frame, multiple AUs are detected simultaneously by exploiting their inter-relations; (2) lp-norm multiple kernel learning is applied to increase the discriminant power of classifiers. Our experimental results on the CK+ and DISFA databases show that the proposed method outperforms the state-of-the-art methods for AU detection.
Xiao Zhang 0003, Mohammad H. Mahoor, Seyed Mohammad Mavadati, Jeffrey F. Cohn
WACV4
2014 Nonverbal social withdrawal in depression: Evidence from manual and automatic analyses
Jeffrey M. Girard, Jeffrey F. Cohn, Mohammad H. Mahoor, Seyed Mohammad Mavadati, Zakia Hammal, Dean P. Rosenwald
Image Vis. Comput.2
2014 BP4D-Spontaneous: a high-resolution spontaneous 3D dynamic facial expression database
Xing Zhang 0012, Lijun Yin 0001, Jeffrey F. Cohn, Shaun J. Canavan, Michael Reale, Andy Horowitz, Peng Liu 0039, Jeffrey M. Girard
Image Vis. Comput.3
2014 Interpersonal Coordination of HeadMotion in Distressed Couples
abstract
In automatic emotional expression analysis, head motion has been considered mostly a nuisance variable, something to control when extracting features for action unit or expression detection. As an initial step toward understanding the contribution of head motion to emotion communication, we investigated the interpersonal coordination of rigid head motion in intimate couples with a history of interpersonal violence. Episodes of conflict and non-conflict were elicited in dyadic interaction tasks and validated using linguistic criteria. Head motion parameters were analyzed using Student's paired t-tests; actor-partner analyses to model mutual influence within couples; and windowed cross-correlation to reveal dynamics of change in direction of influence over time. Partners' RMS angular displacement for yaw and RMS angular velocity for pitch and yaw each demonstrated strong mutual influence between partners. Partners' RMS angular displacement for pitch was higher during conflict. In both conflict and non-conflict, head angular displacement and angular velocity for pitch and yaw were strongly correlated, with frequent shifts in lead-lag relationships. The overall amount of coordination between partners' head movement was more highly correlated during non-conflict compared with conflict interaction. While conflict increased head motion, it served to attenuate interpersonal coordination.
Zakia Hammal, Jeffrey F. Cohn, David Ted George
IEEE Trans. Affect. Comput.2
2014 Spatial and Temporal Linearities in Posed and Spontaneous Smiles
abstract
Creating facial animations that convey an animator’s intent is a difficult task because animation techniques are necessarily an approximation of the subtle motion of the face. Some animation techniques may result in linearization of the motion of vertices in space (blendshapes, for example), and other, simpler techniques may result in linearization of the motion in time. In this article, we consider the problem of animating smiles and explore how these simplifications in space and time affect the perceived genuineness of smiles. We create realistic animations of spontaneous and posed smiles from high-resolution motion capture data for two computer-generated characters. The motion capture data is processed to linearize the spatial or temporal properties of the original animation. Through perceptual experiments, we evaluate the genuineness of the resulting smiles. Both space and time impact the perceived genuineness. We also investigate the effect of head motion in the perception of smiles and show similar results for the impact of linearization on animations with and without head motion. Our results indicate that spontaneous smiles are more heavily affected by linearizing the spatial and temporal properties than posed smiles. Moreover, the spontaneous smiles were more affected by temporal linearization than spatial linearization. Our results are in accordance with previous research on linearities in facial animation and allow us to conclude that a model of smiles must include a nonlinear model of velocities.
Laura C. Trutoiu, Elizabeth J. Carter, Nancy S. Pollard, Jeffrey F. Cohn, Jessica K. Hodgins
ACM Trans. Appl. Percept.4
2013 Head Movement Dynamics during Normal and Perturbed Parent-Infant Interaction
abstract
We investigated the dynamics of head motion in parents and infants during an age-appropriate, well-validated emotion induction, the Face-to-Face/Still-Face procedure. Participants were 12 ethnically diverse 6-month-old infants and their mother or father. During infant gaze toward the parent, infant angular amplitude and velocity of pitch and yaw decreased from face-to-face (FF) to still-face (SF) episodes and remained lower in the following Reunion (RE). During infant gaze away from the parent, angular velocity of pitch decreased from FF to SF and remained lower in the RE. Windowed cross-correlation suggested strong bidirectional effects with frequent shifts in the direction of influence. The number of significant positive and negative peaks was higher during FF than RE. Gaze toward and away from the parent was modestly predicted by head orientation. Together, these findings suggest that head motion is strongly related to age-appropriate emotion challenge, are consistent with the hypothesis that perturbations of normal responsiveness carry-over even after the parent resumes normal responsiveness in the reunion, and that there are frequent changes in direction of influence in the postural domain.
Zakia Hammal, Jeffrey F. Cohn, Daniel S. Messinger, Whitney I. Mattson, Mohammad H. Mahoor
ACII2
2013 Facing Imbalanced Data-Recommendations for the Use of Performance Metrics
abstract
Recognizing facial action units (AUs) is important for situation analysis and automated video annotation. Previous work has emphasized face tracking and registration and the choice of features classifiers. Relatively neglected is the effect of imbalanced data for action unit detection. While the machine learning community has become aware of the problem of skewed data for training classifiers, little attention has been paid to how skew may bias performance metrics. To address this question, we conducted experiments using both simulated classifiers and three major databases that differ in size, type of FACS coding, and degree of skew. We evaluated influence of skew on both threshold metrics (Accuracy, F-score, Cohen's kappa, and Krippendorf's alpha) and rank metrics (area under the receiver operating characteristic (ROC) curve and precision-recall curve). With exception of area under the ROC curve, all were attenuated by skewed distributions, in many cases, dramatically so. While ROC was unaffected by skew, precision-recall curves suggest that ROC may mask poor performance. Our findings suggest that skew is a critical factor in evaluating performance metrics. To avoid or minimize skew-biased estimates of performance, we recommend reporting skew-normalized scores along with the obtained ones.
László A. Jeni, Jeffrey F. Cohn, Fernando De la Torre
ACII2
2013 Relative Body Parts Movement for Automatic Depression Analysis
abstract
In this paper, a human body part motion analysis based approach is proposed for depression analysis. Depression is a serious psychological disorder. The absence of an (automated) objective diagnostic aid for depression leads to a range of subjective biases in initial diagnosis and ongoing monitoring. Researchers in the affective computing community have approached the depression detection problem using facial dynamics and vocal prosody. Recent works in affective computing have shown the significance of body pose and motion in analysing the psychological state of a person. Inspired by these works, we explore a body parts motion based approach. Relative orientation and radius are computed for the body parts detected using the pictorial structures framework. A histogram of relative parts motion is computed. To analyse the motion on a holistic level, space-time interest points are computed and a bag of words framework is learnt. The two histograms are fused and a support vector machine classifier is trained. The experiments conducted on a clinical database, prove the effectiveness of the proposed method.
Jyoti Joshi, Abhinav Dhall, Roland Göcke, Jeffrey F. Cohn
ACII4
2013 Selective Transfer Machine for Personalized Facial Action Unit Detection
abstract
Automatic facial action unit (AFA) detection from video is a long-standing problem in facial expression analysis. Most approaches emphasize choices of features and classifiers. They neglect individual differences in target persons. People vary markedly in facial morphology (e.g., heavy versus delicate brows, smooth versus deeply etched wrinkles) and behavior. Individual differences can dramatically influence how well generic classifiers generalize to previously unseen persons. While a possible solution would be to train person-specific classifiers, that often is neither feasible nor theoretically compelling. The alternative that we propose is to personalize a generic classifier in an unsupervised manner (no additional labels for the test subjects are required). We introduce a transductive learning method, which we refer to Selective Transfer Machine (STM), to personalize a generic classifier by attenuating person-specific biases. STM achieves this effect by simultaneously learning a classifier and re-weighting the training samples that are most relevant to the test subject. To evaluate the effectiveness of STM, we compared STM to generic classifiers and to cross-domain learning methods in three major databases: CK+ [20], GEMEP-FERA [32] and RU-FACS [2]. STM outperformed generic classifiers in all.
Wen-Sheng Chu, Fernando De la Torre, Jeffrey F. Cohn
CVPR3
2013 Facial Action Unit Event Detection by Cascade of Tasks
abstract
Automatic facial Action Unit (AU) detection from video is a long-standing problem in facial expression analysis. AU detection is typically posed as a classification problem between frames or segments of positive examples and negative ones, where existing work emphasizes the use of different features or classifiers. In this paper, we propose a method called Cascade of Tasks (CoT) that combines the use of different tasks (i.e., frame, segment and transition) for AU event detection. We train CoT in a sequential manner embracing diversity, which ensures robustness and generalization to unseen data. In addition to conventional frame-based metrics that evaluate frames independently, we propose a new event-based metric to evaluate detection performance at event-level. We show how the CoT method consistently outperforms state-of-the-art approaches in both frame-based and event-based metrics, across three public datasets that differ in complexity: CK+, FERA and RU-FACS.
Wen-Sheng Chu, Fernando De la Torre, Jeffrey F. Cohn
ICCV4
2013 DISFA: A Spontaneous Facial Action Intensity Database
abstract
Access to well-labeled recordings of facial expression is critical to progress in automated facial expression recognition. With few exceptions, publicly available databases are limited to posed facial behavior that can differ markedly in conformation, intensity, and timing from what occurs spontaneously. To meet the need for publicly available corpora of well-labeled video, we collected, ground-truthed, and prepared for distribution the Denver intensity of spontaneous facial action database. Twenty-seven young adults were video recorded by a stereo camera while they viewed video clips intended to elicit spontaneous emotion expression. Each video frame was manually coded for presence, absence, and intensity of facial action units according to the facial action unit coding system. Action units are the smallest visibly discriminable changes in facial action; they may occur individually and in combinations to comprise more molar facial expressions. To provide a baseline for use in future research, protocols and benchmarks for automated action unit intensity measurement are reported. Details are given for accessing the database for research in computer vision, machine learning, and affective and behavioral science.
Seyed Mohammad Mavadati, Mohammad H. Mahoor, Kevin Bartlett, Philip Trinh, Jeffrey F. Cohn
IEEE Trans. Affect. Comput.5
2013 Detecting Depression Severity from Vocal Prosody
abstract
To investigate the relation between vocal prosody and change in depression severity over time, 57 participants from a clinical trial for treatment of depression were evaluated at seven-week intervals using a semi-structured clinical interview for depression severity (Hamilton Rating Scale for Depression: HRSD). All participants met criteria for Major Depressive Disorder at week 1. Using both perceptual judgments by naive listeners and quantitative analyses of vocal timing and fundamental frequency, three hypotheses were tested: 1) Naive listeners can perceive the severity of depression from vocal recordings of depressed participants and interviewers. 2) Quantitative features of vocal prosody in depressed participants reveal change in symptom severity over the course of depression. And 3) Interpersonal effects occur as well; such that vocal prosody in interviewers shows corresponding effects. These hypotheses were strongly supported. Together, participants' and interviewers' vocal prosody accounted for about 60% of variation in depression scores, and detected ordinal range of depression severity (low, mild, and moderate-to-severe) in 69% of cases (kappa = 0.53). These findings suggest that analysis of vocal prosody could be a powerful tool to assist in depression screening and monitoring over the course of depressive disorder and recovery.
Ying Yang 0007, Catherine Fairbairn, Jeffrey F. Cohn
IEEE Trans. Affect. Comput.3
2012 Automatic detection of pain intensity
abstract
Previous efforts suggest that occurrence of pain can be detected from the face. Can intensity of pain be detected as well? The Prkachin and Solomon Pain Intensity (PSPI) metric was used to classify four levels of pain intensity (none, trace, weak, and strong) in 25 participants with previous shoulder injury (McMaster-UNBC Pain Archive). Participants were recorded while they completed a series of movements of their affected and unaffected shoulders. From the video recordings, canonical normalized appearance of the face (CAPP) was extracted using active appearance modeling. To control for variation in face size, all CAPP were rescaled to 96×96 pixels. CAPP then was passed through a set of Log-Normal filters consisting of 7 frequencies and 15 orientations to extract 9216 features. To detect pain level, 4 support vector machines (SVMs) were separately trained for the automatic measurement of pain intensity on a frame-by-frame level using both 5-folds cross-validation and leave-one-subject-out cross-validation. F1 for each level of pain intensity ranged from 91% to 96% and from 40% to 67% for 5-folds and leave-one-subject-out cross-validation, respectively. Intra-class correlation, which assesses the consistency of continuous pain intensity between manual and automatic PSPI was 0.85 and 0.55 for 5-folds and leave-one-subject-out cross-validation, respectively, which suggests moderate to high consistency. These findings show that pain intensity can be reliably measured from facial expression in participants with orthopedic injury.
Zakia Hammal, Jeffrey F. Cohn
ICMI2
2012 Painful monitoring: Automatic pain monitoring using the UNBC-McMaster shoulder pain expression archive database
Patrick Lucey, Jeffrey F. Cohn, Kenneth M. Prkachin, Patricia E. Solomon, Sien W. Chew, Iain A. Matthews
Image Vis. Comput.2
2012 In the Pursuit of Effective Affective Computing: The Relationship Between Features and Registration
abstract
For facial expression recognition systems to be applicable in the real world, they need to be able to detect and track a previously unseen person's face and its facial movements accurately in realistic environments. A highly plausible solution involves performing a "dense" form of alignment, where 60-70 fiducial facial points are tracked with high accuracy. The problem is that, in practice, this type of dense alignment had so far been impossible to achieve in a generic sense, mainly due to poor reliability and robustness. Instead, many expression detection methods have opted for a "coarse" form of face alignment, followed by an application of a biologically inspired appearance descriptor such as the histogram of oriented gradients or Gabor magnitudes. Encouragingly, recent advances to a number of dense alignment algorithms have demonstrated both high reliability and accuracy for unseen subjects [e.g., constrained local models (CLMs)]. This begs the question: Aside from countering against illumination variation, what do these appearance descriptors do that standard pixel representations do not? In this paper, we show that, when close to perfect alignment is obtained, there is no real benefit in employing these different appearance-based representations (under consistent illumination conditions). In fact, when misalignment does occur, we show that these appearance descriptors do work well by encoding robustness to alignment error. For this work, we compared two popular methods for dense alignment-subject-dependent active appearance models versus subject-independent CLMs-on the task of action-unit detection. These comparisons were conducted through a battery of experiments across various publicly available data sets (i.e., CK+, Pain, M3, and GEMEP-FERA). We also report our performance in the recent 2011 Facial Expression Recognition and Analysis Challenge for the subject-independent task.
Sien W. Chew, Patrick Lucey, Simon Lucey, Jason M. Saragih, Jeffrey F. Cohn, Iain A. Matthews, Sridha Sridharan
IEEE Trans. Syst. Man Cybern. Part B5
2011 Fast-FACS: A Computer-Assisted System to Increase Speed and Reliability of Manual FACS Coding
Fernando De la Torre, Tomas Simon, Zara Ambadar, Jeffrey F. Cohn
ACII (1)4
2011 Person-independent facial expression detection using Constrained Local Models
abstract
In automatic facial expression detection, very accurate registration is desired which can be achieved via a deformable model approach where a dense mesh of 60-70 points on the face is used, such as an active appearance model (AAM). However, for applications where manually labeling frames is prohibitive, AAMs do not work well as they do not generalize well to unseen subjects. As such, a more coarse approach is taken for person-independent facial expression detection, where just a couple of key features (such as face and eyes) are tracked using a Viola-Jones type approach. The tracked image is normally post-processed to encode for shift and illumination invariance using a linear bank of filters. Recently, it was shown that this preprocessing step is of no benefit when close to ideal registration has been obtained. In this paper, we present a system based on the Constrained Local Model (CLM) method which is a generic or person-independent face alignment algorithm which gains high accuracy. We show these results against the LBP feature extraction on the CK+ and GEMEP-FERA datasets.
Sien W. Chew, Patrick Lucey, Simon Lucey, Jason M. Saragih, Jeffrey F. Cohn, Sridha Sridharan
FG5
2011 Painful data: The UNBC-McMaster shoulder pain expression archive database
abstract
A major factor hindering the deployment of a fully functional automatic facial expression detection system is the lack of representative data. A solution to this is to narrow the context of the target application, so enough data is available to build robust models so high performance can be gained. Automatic pain detection from a patient's face represents one such application. To facilitate this work, researchers at McMaster University and University of Northern British Columbia captured video of participant's faces (who were suffering from shoulder pain) while they were performing a series of active and passive range-of-motion tests to their affected and unaffected limbs on two separate occasions. Each frame of this data was AU coded by certified FACS coders, and self-report and observer measures at the sequence level were taken as well. This database is called the UNBC-McMaster Shoulder Pain Expression Archive Database. To promote and facilitate research into pain and augment current datasets, we have publicly made available a portion of this database which includes: (1) 200 video sequences containing spontaneous facial expressions, (2) 48,398 FACS coded frames, (3) associated pain frame-by-frame scores and sequence-level self-report and observer measures, and (4) 66-point AAM landmarks. This paper documents this data distribution in addition to describing baseline results of our AAM/SVM system. This data will be available for distribution in March 2011.
Patrick Lucey, Jeffrey F. Cohn, Kenneth M. Prkachin, Patricia E. Solomon, Iain A. Matthews
FG2
2011 Facial action unit recognition with sparse representation
abstract
This paper presents a novel framework for recognition of facial action unit (AU) combinations by viewing the classification as a sparse representation problem. Based on this framework, we represent a facial image exhibiting the combination of AUs as a sparse linear combination of basis constituting an overcomplete dictionary. We build an overcomplete dictionary whose main elements are mean Gabor features of AU combinations under examination. The other elements of the dictionary are randomly sampled from a distribution (e.g., Gaussian distribution) that guarantees sparse signal recovery. Afterwards, by solving L1-norm minimization, a facial image is represented as a sparse vector which is used to distinguish various AU patterns. After calculating the sparse representation, the classification problem is simply viewed as a rank maximal problem. The index of the maximal value of the sparse vector is regarded as the class label of the facial image under test. Extensive experiments on the Cohn-Kanade facial expressions database demonstrate that this sparse learning framework is promising for recognition of AU combinations.
Mohammad H. Mahoor, Mu Zhou, Kevin L. Veon, Seyed Mohammad Mavadati, Jeffrey F. Cohn
FG5
2011 Prediction-based classification for audiovisual discrimination between laughter and speech
abstract
Recent evidence in neuroscience support the theory that prediction of spatial and temporal patterns in the brain plays a key role in human actions and perception. Inspired by these findings, a system that discriminates laughter from speech by modeling the spatial and temporal relationship between audio and visual features is presented. The underlying assumption is that this relationship is different between speech and laughter. Neural networks are trained which learn the audio-to-visual and visual-to-audio feature mapping together with the time evolution of audio and visual features for both classes. Classification of a new frame / sequence is performed via prediction. All the networks produce a prediction of the expected audio / visual features and their prediction errors are combined for each class. The model which best describes the audiovisual feature relationship, i.e., results in the lowest prediction error, provides its label to the input frame / sequence. Using 4 different datasets, the proposed system is compared to standard feature-level fusion on cross-database experiments. In almost all test cases, prediction-based classification outperforms feature-level fusion. Similar conclusion are drawn when adding artificial feature-level noise to the datasets.
Stavros Petridis, Maja Pantic, Jeffrey F. Cohn
FG3
2011 Real-time avatar animation from a single image
abstract
A real time facial puppetry system is presented. Compared with existing systems, the proposed method requires no special hardware, runs in real time (23 frames-per-second), and requires only a single image of the avatar and user. The user's facial expression is captured through a real-time 3D non-rigid tracking system. Expression transfer is achieved by combining a generic expression model with synthetically generated examples that better capture person specific characteristics. Performance of the system is evaluated on avatars of real people as well as masks and cartoon characters.
Jason M. Saragih, Simon Lucey, Jeffrey F. Cohn
FG3
2011 Real-time avatar animation from a single image
abstract
A real time facial puppetry system is presented. Compared with existing systems, the proposed method requires no special hardware, runs in real time (23 frames-per-second), and requires only a single image of the avatar and user. The user's facial expression is captured through a real-time 3D non-rigid tracking system. Expression transfer is achieved by combining a generic expression model with synthetically generated examples that better capture person specific characteristics. Performance of the system is evaluated on avatars of real people as well as masks and cartoon characters.
Jason M. Saragih, Simon Lucey, Jeffrey F. Cohn
FG3
2011 Deformable Model Fitting by Regularized Landmark Mean-Shift
Jason M. Saragih, Simon Lucey, Jeffrey F. Cohn
Int. J. Comput. Vis.3
2011 Dynamic Cascades with Bidirectional Bootstrapping for Action Unit Detection in Spontaneous Facial Behavior
abstract
Automatic facial action unit detection from video is a long-standing problem in facial expression analysis. Research has focused on registration, choice of features, and classifiers. A relatively neglected problem is the choice of training images. Nearly all previous work uses one or the other of two standard approaches. One approach assigns peak frames to the positive class and frames associated with other actions to the negative class. This approach maximizes differences between positive and negative classes, but results in a large imbalance between them, especially for infrequent AUs. The other approach reduces imbalance in class membership by including all target frames from onsets to offsets in the positive class. However, because frames near onsets and offsets often differ little from those that precede them, this approach can dramatically increase false positives. We propose a novel alternative, dynamic cascades with bidirectional bootstrapping (DCBB), to select training samples. Using an iterative approach, DCBB optimally selects positive and negative samples in the training data. Using Cascade Adaboost as basic classifier, DCBB exploits the advantages of feature selection, efficiency, and robustness of Cascade Adaboost. To provide a real-world test, we used the RU-FACS (a.k.a. M3) database of nonposed behavior recorded during interviews. For most tested action units, DCBB improved AU detection relative to alternative approaches.
Yunfeng Zhu, Fernando De la Torre, Jeffrey F. Cohn, Yu-Jin Zhang
IEEE Trans. Affect. Comput.3
2011 Automatically Detecting Pain in Video Through Facial Action Units
abstract
In a clinical setting, pain is reported either through patient self-report or via an observer. Such measures are problematic as they are: 1) subjective, and 2) give no specific timing information. Coding pain as a series of facial action units (AUs) can avoid these issues as it can be used to gain an objective measure of pain on a frame-by-frame basis. Using video data from patients with shoulder injuries, in this paper, we describe an active appearance model (AAM)-based system that can automatically detect the frames in video in which a patient is in pain. This pain data set highlights the many challenges associated with spontaneous emotion detection, particularly that of expression and head movement due to the patient's reaction to pain. In this paper, we show that the AAM can deal with these movements and can achieve significant improvements in both the AU and pain detection performance compared to the current-state-of-the-art approaches which utilize similarity-normalized appearance features only.
Patrick Lucey, Jeffrey F. Cohn, Iain A. Matthews, Simon Lucey, Sridha Sridharan, Jessica Howlett, Kenneth M. Prkachin
IEEE Trans. Syst. Man Cybern. Part B2
2010 Action unit detection with segment-based SVMs
abstract
Automatic facial action unit (AU) detection from video is a long-standing problem in computer vision. Two main approaches have been pursued: (1) static modeling - typically posed as a discriminative classification problem in which each video frame is evaluated independently; (2) temporal modeling - frames are segmented into sequences and typically modeled with a variant of dynamic Bayesian networks. We propose a segment-based approach, kSeg-SVM, that incorporates benefits of both approaches and avoids their limitations. kSeg-SVM is a temporal extension of the spatial bag-of-words. kSeg-SVM is trained within a structured output SVM framework that formulates AU detection as a problem of detecting temporal events in a time series of visual features. Each segment is modeled by a variant of the BoW representation with soft assignment of the words based on similarity. Our framework has several benefits for AU detection: (1) both dependencies between features and the length of action units are modeled; (2) all possible segments of the video may be used for training; and (3) no assumptions are required about the underlying structure of the action unit events (e.g., i.i.d.). Our algorithm finds the best k-or-fewer segments that maximize the SVM score. Experimental results suggest that the proposed method outperforms state-of-the-art static methods for AU detection.
Tomas Simon, Minh Hoai, Fernando De la Torre, Jeffrey F. Cohn
CVPR4
2010 Unsupervised discovery of facial events
abstract
Automatic facial image analysis has been a long standing research problem in computer vision. A key component in facial image analysis, largely conditioning the success of subsequent algorithms (e.g. facial expression recognition), is to define a vocabulary of possible dynamic facial events. To date, that vocabulary has come from the anatomically-based Facial Action Coding System (FACS) or more subjective approaches (i.e. emotion-specified expressions). The aim of this paper is to discover facial events directly from video of naturally occurring facial behavior, without recourse to FACS or other labeling schemes. To discover facial events, we propose a temporal clustering algorithm, Aligned Cluster Analysis (ACA), and a multi-subject correspondence algorithm for matching expressions. We use a variety of video sources: posed facial behavior (Cohn-Kanade database), unscripted facial behavior (RU-FACS database) and some video in infants. Accuracy of (unsupervised) ACA approached that of a supervised version, achieved moderate intersystem agreement with FACS, and proved informative as a visualization/summarization tool.
Feng Zhou 0002, Fernando De la Torre, Jeffrey F. Cohn
CVPR3
2010 Multi-PIE
Ralph Gross, Iain A. Matthews, Jeffrey F. Cohn, Takeo Kanade, Simon Baker
Image Vis. Comput.3
2010 Non-rigid face tracking with enforced convexity and local appearance consistency constraint
Simon Lucey, Yang Wang 0001, Jason M. Saragih, Jeffrey F. Cohn
Image Vis. Comput.4
2010 Best of Automatic Face and Gesture Recognition 2008
Maja Pantic, Nicu Sebe, Jeffrey F. Cohn, Thomas S. Huang
Image Vis. Comput.3
2009 Least-squares congealing for large numbers of images
abstract
In this paper we pursue the task of aligning an ensemble of images in an unsupervised manner. This task has been commonly referred to as “congealing” in literature. A form of congealing, using a least-squares criteria, has been recently demonstrated to have desirable properties over conventional congealing. Least-squares congealing can be viewed as an extension of the Lucas & Kanade (LK) image alignment algorithm. It is well understood that the alignment performance for the LK algorithm, when aligning a single image with another, is theoretically and empirically equivalent for additive and compositional warps. In this paper we: (i) demonstrate that this equivalence does not hold for the extended case of congealing, (ii) characterize the inherent drawbacks associated with least-squares congealing when dealing with large numbers of images, and (iii) propose a novel method for circumventing these limitations through the application of an inverse-compositional strategy that maintains the attractive properties of the original method while being able to handle very large numbers of images.
Mark Cox, Sridha Sridharan, Simon Lucey, Jeffrey F. Cohn
ICCV4
2009 Face alignment through subspace constrained mean-shifts
abstract
Deformable model fitting has been actively pursued in the computer vision community for over a decade. As a result, numerous approaches have been proposed with varying degrees of success. A class of approaches that has shown substantial promise is one that makes independent predictions regarding locations of the model's landmarks, which are combined by enforcing a prior over their joint motion. A common theme in innovations to this approach is the replacement of the distribution of probable landmark locations, obtained from each local detector, with simpler parametric forms. This simplification substitutes the true objective with a smoothed version of itself, reducing sensitivity to local minima and outlying detections. In this work, a principled optimization strategy is proposed where a nonparametric representation of the landmark distributions is maximized within a hierarchy of smoothed estimates. The resulting update equations are reminiscent of mean-shift but with a subspace constraint placed on the shape's variability. This approach is shown to outperform other existing methods on the task of generic face fitting.
Jason M. Saragih, Simon Lucey, Jeffrey F. Cohn
ICCV3
2009 Deformable model fitting with a mixture of local experts
abstract
Local experts have been used to great effect for fitting deformable models to images. Typically, the best location in an image for the deformable model's landmarks are found through a locally exhaustive search using these experts. In order to achieve efficient fitting, these experts should afford an efficient evaluation, which often leads to forms with restricted discriminative capacity. In this work, a framework is proposed in which multiple simple experts can be utilized to increase the capacity of the detections overall. In particular, the use of a mixture of linear classifiers is proposed, the computational complexity of which scales linearly with the number of mixture components. The fitting objective is maximized using the expectation maximization (EM) algorithm, where approximations to the true objective are made in order to facilitate efficient and numerically stable fitting. The efficacy of the proposed approach is evaluated on the task of generic face fitting where performance improvement is observed over two existing methods.
Jason M. Saragih, Simon Lucey, Jeffrey F. Cohn
ICCV3
2009 The painful face - Pain expression recognition using active appearance models
abstract
Pain is typically assessed by patient self-report. Self-reported pain, however, is difficult to interpret and may be impaired or in some circumstances (i.e., young children and the severely ill) not even possible. To circumvent these problems behavioral scientists have identified reliable and valid facial indicators of pain. Hitherto, these methods have required manual measurement by highly skilled human observers. In this paper we explore an approach for automatically recognizing acute pain without the need for human observers. Specifically, our study was restricted to automatically detecting pain in adult patients with rotator cuff injuries. The system employed video input of the patients as they moved their affected and unaffected shoulder. Two types of ground truth were considered. Sequence-level ground truth consisted of Likert-type ratings by skilled observers. Frame-level ground truth was calculated from presence/absence and intensity of facial actions previously associated with pain. Active appearance models (AAM) were used to decouple shape and appearance in the digitized face images. Support vector machines (SVM) were compared for several representations from the AAM and of ground truth of varying granularity. We explored two questions pertinent to the construction, design and development of automatic pain detection systems. First, at what level (i.e., sequence- or frame-level) should datasets be labeled in order to obtain satisfactory automatic pain detection performance? Second, how important is it, at both levels of labeling, that we non-rigidly register the face?
Ahmed Ashraf 0001, Simon Lucey, Jeffrey F. Cohn, Tsuhan Chen, Zara Ambadar, Kenneth M. Prkachin, Patricia E. Solomon
Image Vis. Comput.3
2009 Efficient constrained local model fitting for non-rigid face alignment
Simon Lucey, Yang Wang 0001, Mark Cox, Sridha Sridharan, Jeffrey F. Cohn
Image Vis. Comput.5
2009 Visual and multimodal analysis of human spontaneous behaviour: Introduction to the Special Issue
Maja Pantic, Jeffrey F. Cohn
Image Vis. Comput.2
2008 Model-Based De-Identification of Facial Images
Ralph Gross, Latanya Sweeney, Jeffrey F. Cohn, Fernando De la Torre, Simon Baker
AMIA3
2008 Least squares congealing for unsupervised alignment of images
abstract
In this paper, we present an approach we refer to as "least squares congealing" which provides a solution to the problem of aligning an ensemble of images in an unsupervised manner. Our approach circumvents many of the limitations existing in the canonical "congealing" algorithm. Specifically, we present an algorithm that:- (i) is able to simultaneously, rather than sequentially, estimate warp parameter updates, (ii) exhibits fast convergence and (iii) requires no pre-defined step size. We present alignment results which show an improvement in performance for the removal of unwanted spatial variation when compared with the related work of Learned-Miller on two datasets, the MNIST hand written digit database and the MultiPIE face database.
Mark Cox, Sridha Sridharan, Simon Lucey, Jeffrey F. Cohn
CVPR4
2008 Enforcing convexity for improved alignment with constrained local models
abstract
Constrained local models (CLMs) have recently demonstrated good performance in non-rigid object alignment/tracking in comparison to leading holistic approaches (e.g., AAMs). A major problem hindering the development of CLMs further, for non-rigid object alignment/tracking, is how to jointly optimize the global warp update across all local search responses. Previous methods have either used general purpose optimizers (e.g., simplex methods) or graph based optimization techniques. Unfortunately, problems exist with both these approaches when applied to CLMs. In this paper, we propose a new approach for optimizing the global warp update in an efficient manner by enforcing convexity at each local patch response surface. Furthermore, we show that the classic Lucas-Kanade approach to gradient descent image alignment can be viewed as a special case of our proposed framework. Finally, we demonstrate that our approach receives improved performance for the task of non-rigid face alignment/tracking on the MultiPIE database and the UNBC-McMaster archive.
Yang Wang 0001, Simon Lucey, Jeffrey F. Cohn
CVPR3
2008 Multi-PIE
abstract
A close relationship exists between the advancement of face recognition algorithms and the availability of face databases varying factors that affect facial appearance in a controlled manner. The CMU PIE database has been very influential in advancing research in face recognition across pose and illumination. Despite its success the PIE database has several shortcomings: a limited number of subjects, a single recording session and only few expressions captured. To address these issues we collected the CMU Multi-PIE database. It contains 337 subjects, imaged under 15 view points and 19 illumination conditions in up to four recording sessions. In this paper we introduce the database and describe the recording procedure. We furthermore present results from baseline experiments using PCA and LDA classifiers to highlight similarities and differences between PIE and Multi-PIE.
Ralph Gross, Iain A. Matthews, Jeffrey F. Cohn, Takeo Kanade, Simon Baker
FG3
2008 Deformable Face Fitting with Soft Correspondence Constraints
abstract
Despite significant progress in deformable model fitting over the last decade, the problem of efficient and accurate person-independent face fitting remains a challenging problem. In this work, a reformulation of the generative fitting objective is presented, where only soft correspondences between the model and the image are enforced. This has the dual effect of improving robustness to unseen faces as well as affording fitting time which scales linearly with the model's complexity. This approach is compared with three state-of-the-art fitting methods on the problem of person independent face fitting, where it is shown to closely approach the accuracy of the currently best performing method while affording significant computational savings.
Jason M. Saragih, Simon Lucey, Jeffrey F. Cohn
FG3
2008 Comparing object alignment algorithms with appearance variation: Forward-additive vs inverse-composition
abstract
A common problem that affects object alignment algorithms is when they have to deal with objects with unseen intra-class appearance variation. Several variants based on gradient-decent algorithms, such as the Lucas-Kanade (or forward-additive) and inverse-compositional algorithms, have been proposed to deal with this issue by solving for both alignment and appearance simultaneously. In [1], Baker and Matthews showed that without appearance variation, the inverse-compositional (IC) algorithm was theoretically and empirically equivalent to the forward-additive (FA) algorithm, whilst achieving significant improvement in computational efficiency. With appearance variation, it would be intuitive that a similar benefit of the IC algorithm would be experienced over the FA counterpart. However, to date no such comparison has been performed. In this paper we remedy this situation by performing such a comparison. In this comparison we show that the two algorithms are not equivalent due to the inclusion of the appearance variation parameters. Through a number of experiments on the MultiPIE face database, we show that we can gain greater refinement using the FA algorithm due to it being a truer solution than the IC approach.
Patrick Lucey, Simon Lucey, Mark Cox, Sridha Sridharan, Jeffrey F. Cohn
MMSP5
2008 Multi-View AAM Fitting and Construction
Krishnan Ramnath, Seth Koterba, Jing Xiao 0006, Changbo Hu, Iain A. Matthews, Simon Baker, Jeffrey F. Cohn, Takeo Kanade
Int. J. Comput. Vis.7
2007 Filtered Component Analysis to Increase Robustness to Local Minima in Appearance Models
abstract
Appearance models (AM) are commonly used to model appearance and shape variation of objects in images. In particular, they have proven useful to detection, tracking, and synthesis of people's faces from video. While AM have numerous advantages relative to alternative approaches, they have at least two important drawbacks. First, they are especially prone to local minima in fitting; this problem becomes increasingly problematic as the number of parameters to estimate grows. Second, often few if any of the local minima correspond to the correct location of the model error. To address these problems, we propose filtered component analysis (FCA), an extension of traditional principal component analysis (PCA). FCA learns an optimal set of filters with which to build a multi-band representation of the object. FCA representations were found to be more robust than either grayscale or Gabor filters to problems of local minima. The effectiveness and robustness of the proposed algorithm is demonstrated in both synthetic and real data.
Fernando De la Torre, Alvaro Collet, Manuel Quero, Jeffrey F. Cohn, Takeo Kanade
CVPR4
2007 Non-Rigid Object Alignment with a Mismatch Template Based on Exhaustive Local Search
abstract
Non-rigid object alignment is especially challenging when only a single appearance template is available and target and template images fail to match. Two sources of discrepancy between target and template are changes in illumination and non-rigid motion. Because most existing methods rely on a holistic representation for the alignment process, they require multiple training images to capture appearance variance. We developed a patch-based method that requires only a single appearance template of the object. Specifically, we fit the patch-based face model to an unseen image using an exhaustive local search and constrain the local warp updates within a global warping space. Our approach is not limited to intensity values or gradients, and therefore offers a natural framework to integrate multiple local features, such as filter responses, to increase robustness to large initialization error, illumination changes and non-rigid deformations. This approach was evaluated experimentally on more than 100 subjects for multiple illumination conditions and facial expressions. In all the experiments, our patch-based method outperforms the holistic gradient descent method in terms of accuracy and robustness of feature alignment and image registration.
Yang Wang 0001, Simon Lucey, Jeffrey F. Cohn
ICCV3
2007 The painful face: pain expression recognition using active appearance models
abstract
Pain is typically assessed by patient self-report. Self-reported pain, however, is difficult to interpret and may be impaired or not even possible, as in young children or the severely ill. Behavioral scientists have identified reliable and valid facial indicators of pain. Until now they required manual measurement by highly skilled observers. We developed an approach that automatically recognizes acute pain. Adult patients with rotator cuff injury were video-recorded while a physiotherapist manipulated their affected and unaffected shoulder. Skilled observers rated pain expression from the video on a 5-point Likert-type scale. From these ratings, sequences were categorized as no-pain (rating of 0), pain (rating of 3, 4, or 5), and indeterminate (rating of 1 or 2). We explored machine learning approaches for pain-no pain classification. Active Appearance Models (AAM) were used to decouple shape and appearance parameters from the digitized face images. Support vector machines (SVM) were used with several representations from the AAM. Using a leave-one-out procedure, we achieved an equal error rate of 19% (hit rate = 81%) using canonical appearance and shape features. These findings suggest the feasibility of automatic pain detection from video.
Ahmed Ashraf 0001, Simon Lucey, Jeffrey F. Cohn, Tsuhan Chen, Zara Ambadar, Kenneth M. Prkachin, Patty Solomon, Barry-John Theobald
ICMI3
2007 Real-time expression cloning using appearance models
abstract
Active Appearance Models (AAMs) are generative parametric models commonly used to track, recognise and synthesise faces in images and video sequences. In this paper we describe a method for transferring dynamic facial gestures between subjects in real-time. The main advantages of our approach are that: 1) the mapping is computed automatically and does not require high-level semantic information describing facial expressions or visual speech gestures. 2) The mapping is simple and intuitive, allowing expressions to be transferred and rendered in real-time. 3) The mapped expression can be constrained to have the appearance of the target producing the expression, rather than the source expression imposed onto the target face. 4) Near-videorealistic talking faces for new subjects can be created without the cost of recording and processing a complete training corpus for each. Our system enables face-to-face interaction with an avatar driven by an AAM of an actual person in real-time and we show examples of arbitrary expressive speech frames cloned across different subjects.
Barry-John Theobald, Iain A. Matthews, Jeffrey F. Cohn, Steven M. Boker
ICMI3
2007 Robust Biometric Person Identification Using Automatic Classifier Fusion of Speech, Mouth, and Face Experts
abstract
Information about person identity is multimodal. Yet, most person-recognition systems limit themselves to only a single modality, such as facial appearance. With a view to exploiting the complementary nature of different modes of information and increasing pattern recognition robustness to test signal degradation, we developed a multiple expert biometric person identification system that combines information from three experts: audio, visual speech, and face. The system uses multimodal fusion in an automatic unsupervised manner, adapting to the local performance (at the transaction level) and output reliability of each of the three experts. The expert weightings are chosen automatically such that the reliability measure of the combined scores is maximized. To test system robustness to train/test mismatch, we used a broad range of acoustic babble noise and JPEG compression to degrade the audio and visual signals, respectively. Identification experiments were carried out on a 248-subject subset of the XM2VTS database. The multimodal expert system outperformed each of the single experts in all comparisons. At severe audio and visual mismatch levels tested, the audio, mouth, face, and tri-expert fusion accuracies were 16.1%, 48%, 75%, and 89.9%, respectively, representing a relative improvement of 19.9% over the best performing expert
Niall A. Fox, Ralph Gross, Jeffrey F. Cohn, Richard B. Reilly
IEEE Trans. Multim.3
2006 Foundations of human computing: facial expression and emotion
abstract
Many people believe that emotions and subjective feelings are one and the same and that a goal of human-centered computing is emotion recognition. The first belief is outdated; the second mistaken. For human-centered computing to succeed, a different way of thinking is needed.Emotions are species-typical patterns that evolved because of their value in addressing fundamental life tasks[19]. Emotions consist of multiple components that may include intentions, action tendencies, appraisals, other cognitions, central and peripheral changes in physiology, and subjective feelings. Emotions are not directly observable, but are inferred from expressive behavior, self-report, physiological indicators, and context. I focus on expressive behavior because of its coherence with other indicators and the depth of research on the facial expression of emotion in behavioral and computer science. In this paper, among the topics I include are approaches to measurement, timing or dynamics, individual differences, dyadic interaction, and inference. I propose that design and implementation of perceptual user interfaces may be better informed by considering the complexity of emotion, its various indicators, measurement, individual differences, dyadic interaction, and problems of inference.
Jeffrey F. Cohn
ICMI1
2006 Spontaneous vs. posed facial behavior: automatic analysis of brow actions
abstract
Past research on automatic facial expression analysis has focused mostly on the recognition of prototypic expressions of discrete emotions rather than on the analysis of dynamic changes over time, although the importance of temporal dynamics of facial expressions for interpretation of the observed facial behavior has been acknowledged for over 20 years. For instance, it has been shown that the temporal dynamics of spontaneous and volitional smiles are fundamentally different from each other. In this work, we argue that the same holds for the temporal dynamics of brow actions and show that velocity, duration, and order of occurrence of brow actions are highly relevant parameters for distinguishing posed from spontaneous brow actions. The proposed system for discrimination between volitional and spontaneous brow actions is based on automatic detection of Action Units (AUs) and their temporal segments (onset, apex, offset) produced by movements of the eyebrows. For each temporal segment of an activated AU, we compute a number of mid-level feature parameters including the maximal intensity, duration, and order of occurrence. We use Gentle Boost to select the most important of these parameters. The selected parameters are used further to train Relevance Vector Machines to determine per temporal segment of an activated AU whether the action was displayed spontaneously or volitionally. Finally, a probabilistic decision function determines the class (spontaneous or posed) for the entire brow action. When tested on 189 samples taken from three different sets of spontaneous and volitional facial data, we attain a 90.7% correct recognition rate.
Michel F. Valstar, Maja Pantic, Zara Ambadar, Jeffrey F. Cohn
ICMI4
2006 Meticulously Detailed Eye Region Model and Its Application to Analysis of Facial Images
abstract
We propose a system that is capable of detailed analysis of eye region images in terms of the position of the iris, degree of eyelid opening, and the shape, complexity, and texture of the eyelids. The system uses a generative eye region model that parameterizes the fine structure and motion of an eye. The structure parameters represent structural individuality of the eye, including the size and color of the iris, the width, boldness, and complexity of the eyelids, the width of the bulge below the eye, and the width of the illumination reflection on the bulge. The motion parameters represent movement of the eye, including the up-down position of the upper and lower eyelids and the 2D position of the iris. The system first registers the eye model to the input in a particular frame and individualizes it by adjusting the structure parameters. The system then tracks motion of the eye by estimating the motion parameters across the entire image sequence. Combined with image stabilization to compensate for appearance changes due to head motion, the system achieves accurate registration and motion recovery of eyes.
Tsuyoshi Moriyama, Takeo Kanade, Jing Xiao 0006, Jeffrey F. Cohn
IEEE Trans. Pattern Anal. Mach. Intell.4
2005 Multi-View AAM Fitting and Camera Calibration
abstract
In this paper, we study the relationship between multi-view active appearance model (AAM) fitting and camera calibration. In the first part of the paper we propose an algorithm to calibrate the relative orientation of a set of N > 1 cameras by fitting an AAM to sets of N images. In essence, we use the human face as a (non-rigid) calibration grid. Our algorithm calibrates a set of 2 /spl times/ 3 weak-perspective camera projection matrices, protections of the world coordinate system origin into the images, depths of the world coordinate system origin, and focal lengths. We demonstrate that the performance of this algorithm is comparable to a standard algorithm using a calibration grid. In the second part of the paper, we show how calibrating the cameras improves tile performance of multi-view AAM fitting.
Seth Koterba, Simon Baker, Iain A. Matthews, Changbo Hu, Jing Xiao 0006, Jeffrey F. Cohn, Takeo Kanade
ICCV6
2005 Affective multimodal human-computer interaction
abstract
Social and emotional intelligence are aspects of human intelligence that have been argued to be better predictors than IQ for measuring aspects of success in life, especially in social interactions, learning, and adapting to what is important. When it comes to machines, not all of them will need such skills. Yet to have machines like computers, broadcast systems, and cars, capable of adapting to their users and of anticipating their wishes, endowing them with the ability to recognize user's affective states is necessary. This article discusses the components of human affect, how they might be integrated into computers, and how far are we from realizing affective multimodal human-computer interaction.
Maja Pantic, Nicu Sebe, Jeffrey F. Cohn, Thomas S. Huang
ACM Multimedia3
2004 Fitting a Single Active Appearance Model Simultaneously to Multiple Images
abstract
Active Appearance Models (AAMs) are a well studied 2D deformable model. One recently proposed extension of AAMs to multiple images is the Coupled-View AAM. Coupled-View AAMs model the 2D shape and appearance of a face in two or more views simultaneously. The major limitation of Coupled-View AAMs, however, is that they are specific to a particular set of cameras, both in geometry and the photometric responses. In this paper, we describe how a single AAM can be fit to multiple images, captured simultaneously by cameras with arbitrary geometry and response functions. Our algorithm retains the major benefits of Coupled-View AAMs: the integration of information from multiple images into a single model, and improved fitting robustness. 1
Changbo Hu, Jing Xiao 0006, Iain A. Matthews, Simon Baker, Jeffrey F. Cohn, Takeo Kanade
BMVC5
2003 Facial asymmetry quantification for expression invariant human identification
Yanxi Liu 0001, Karen L. Schmidt, Jeffrey F. Cohn, Sinjini Mitra
Comput. Vis. Image Underst.3
2002 Individual Differences in Facial Expression: Stability over Time, Relation to Self-Reported Emotion, and Ability to Inform Person Identification
abstract
The face can communicate varied personal information including subjective emotion, communicative intent, and cognitive appraisal. Accurate interpretation by observer or computer interface depends on attention to dynamic properties of the expression, context, and knowledge of what is normative for a given individual. In two separate studies, we investigated individual differences in the base rate of positive facial expression and in specific facial action units over intervals from 4 to 12 months. Facial expression was measured using convergent measures, including facial EMG, automatic feature-point tracking, and manual FACS coding. Individual differences in facial expression were stable over time, comparable in magnitude to stability of self-reported emotion, and sufficiently strong that individuals were recognized on the basis of their facial behavior alone at rates comparable to that for a commercial face recognition system (Facelt from Identix). Facial action units convey unique information about person identity that can inform interpretation of psychological states, person recognition, and design of individuated avatars.
Jeffrey F. Cohn, Karen L. Schmidt, Ralph Gross, Paul Ekman
ICMI1
2001 Dynamics Of Facial Expression: Normative Characteristics And Individual Differences
abstract
Although the importance of facial expression in human computer interaction and in normal human interaction is widely acknowledged, there is very little data on the normative characteristics and stable individual differences for even the most common facial expressions. Dynamic characteristics of 195 spontaneous smiles from 95 individuals were measured using the facial action coding system, automated facial analysis and facial electromyography. Normative patterns observed included the characteristic timing of other facial actions with respect to action unit 12 ("smile") and a mean duration of 15.7 frames for smile onset. Stable inter-individual differences included patterns of nonverbal actions associated with individuals' smiles, and the amount of activity in the zygomaticus major muscle in two sessions recorded a year apart. These data are important in quantifying and fully describing individual differences in naturalistic human facial expression, as well as adding to our knowledge of spontaneous human smiles.
Karen L. Schmidt, Jeffrey F. Cohn
ICME2
2001 Recognizing Action Units for Facial Expression Analysis
abstract
Most automatic expression analysis systems attempt to recognize a small set of prototypic expressions, such as happiness, anger, surprise, and fear. Such prototypic expressions, however, occur rather infrequently. Human emotions and intentions are more often communicated by changes in one or a few discrete facial features. In this paper, we develop an Automatic Face Analysis (AFA) system to analyze facial expressions based on both permanent facial features (brows, eyes, mouth) and transient facial features (deepening of facial furrows) in a nearly frontal-view face image sequence. The AFA system recognizes fine-grained changes in facial expression into action units (AUs) of the Facial Action Coding System (FACS), instead of a few prototypic expressions. Multistate face and facial component models are proposed for tracking and modeling the various facial features, including lips, eyes, brows, cheeks, and furrows. During tracking, detailed parametric descriptions of the facial features are extracted. With these parameters as the inputs, a group of action units (neutral expression, six upper face AUs and 10 lower face AUs) are recognized whether they occur alone or in combinations. The system has achieved average recognition rates of 96.4 percent (95.4 percent if neutral expressions are excluded) for upper face AUs and 96.7 percent (95.6 percent with neutral expressions excluded) for lower face AUs. The generalizability of the system has been tested by using independent image databases collected and FACS-coded for ground-truth by different research teams.
Yingli Tian, Takeo Kanade, Jeffrey F. Cohn
IEEE Trans. Pattern Anal. Mach. Intell.3
2000 Recognizing Upper Face Action Units for Facial Expression Analysis
abstract
We develop an automatic system to analyze subtle changes in upper face expressions based on both permanent facial features (brows, eyes, mouth) and transient facial features (deepening of facial furrows) in a nearly frontal image sequence. Our system recognizes fine-grained changes in facial expression based on Facial Action Coding System (FACS) action units (AUs). Multi-state facial component models are proposed for tracting and modeling different facial features, including eyes, brews, cheeks, and furrows. Then we convert the results of tracking to detailed parametric descriptions of the facial features. These feature parameters are fed to a neural network which recognizes 7 upper face action units. A recognition rate of 95% is obtained for the test data that include both single action units and AU combinations.
Yingli Tian, Takeo Kanade, Jeffrey F. Cohn
CVPR3
2000 Comprehensive Database for Facial Expression Analysis
abstract
Within the past decade, significant effort has occurred in developing methods of facial expression analysis. Because most investigators have used relatively limited data sets, the generalizability of these various methods remains unknown. We describe the problem space for facial expression analysis, which includes level of description, transitions among expressions, eliciting conditions, reliability and validity of training and test data, individual differences in subjects, head orientation and scene complexity image characteristics, and relation to non-verbal behavior. We then present the CMU-Pittsburgh AU-Coded Face Expression Image Database, which currently includes 2105 digitized image sequences from 182 adult subjects of varying ethnicity, performing multiple tokens of most primary FACS action units. This database is the most comprehensive testbed to date for comparative studies of facial expression analysis.
Takeo Kanade, Yingli Tian, Jeffrey F. Cohn
FG3
2000 Dual-State Parametric Eye Tracking
abstract
Most eye trackers work well for open eyes. However blinking is a physiological necessity for humans. More over, for applications such as facial expression analysis and driver awareness systems, we need to do more than tracking of the locations of the person's eyes but obtain their detailed description. We need to recover the state of the eyes (i.e., whether they are open or closed), and the parameters of an eye model (e.g., the location and radius of the iris, and the corners and height of the eye opening). We develop a dual-state model-based system for tracking eye features that uses convergent tracking techniques and show how it can be used to detect whether the eyes are open or closed, and to recover the parameters of the eye model. Processing speed on a Pentium II 400 MHz PC is approximately 3 frames/second. In experimental tests on 500 image sequences from child and adult subjects with varying colors of skin and eye, accurate tracking results are obtained in 98% of image sequences.
Yingli Tian, Takeo Kanade, Jeffrey F. Cohn
FG3
2000 Recognizing Lower Face Action Units for Facial Expression Analysis
abstract
Most automatic expression analysis systems attempt to recognize a small set of prototypic expressions (e.g., happiness and anger). Such prototypic expressions, however, occur infrequently. Human emotions and intentions are communicated more often by changes in one or two discrete facial features. We develop an automatic system to analyze subtle changes in facial expressions based on both permanent (e.g., mouth, eye, and brow) and transient (e.g., furrows and wrinkles) facial features in a nearly frontal image sequence. Multi-state facial component models are proposed for tracking and modeling different facial features. Based on these multi-state models, and without artificial enhancement, we detect and track the facial features, including mouth, eyes, brow, cheeks, and their related wrinkles and facial furrows. Moreover we recover detailed parametric descriptions of the facial features. With these features as the inputs, 11 individual action units or action unit combinations are recognized by a neural network algorithm. A recognition rate of 96.7% is obtained. The recognition results indicate that our system can identify action units regardless of whether they occur singly or in combinations.
Yingli Tian, Takeo Kanade, Jeffrey F. Cohn
FG3
2000 Eye-State Action Unit Detection by Gabor Wavelets
Yingli Tian, Takeo Kanade, Jeffrey F. Cohn
ICMI3
2000 Image Registration Using Wavelet-Based Motion Model
Yu-Te Wu, Takeo Kanade, Ching-Chung Li, Jeffrey F. Cohn
Int. J. Comput. Vis.4
1998 Subtly Different Facial Expression Recognition and Expression Intensity Estimation
abstract
We have developed a computer vision system, including both facial feature extraction and recognition, that automatically discriminates among subtly different facial expressions. Expression classification is based on Facial Action Coding System (FACS) action units (AUs), and discrimination is performed using Hidden Markov Models (HMMs). Three methods are developed to extract facial expression information for automatic recognition. The first method is facial feature point tracking using a coarse-to-fine pyramid method. This method is sensitive to subtle feature motion and is capable of handling large displacements with sub-pixel accuracy. The second method is dense flow tracking together with principal component analysis (PCA) where the entire facial motion information per frame is compressed to a low-dimensional weight vector. The third method is high gradient component (i.e., furrow) analysis in the spatio-temporal domain, which exploits the transient variation associated with the facial expression. Upon extraction of the facial information, non-rigid facial expression is separated from the rigid head motion component, and the face images are automatically aligned and normalized using an affine transformation. This system also provides expression intensity estimation, which has significant effect on the actual meaning of the expression.
Jenn-Jier James Lien, Takeo Kanade, Jeffrey F. Cohn, Ching-Chung Li
CVPR3
1998 Feature-Point Tracking by Optical Flow Discriminates Subtle Differences in Facial Expression
Jeffrey F. Cohn, Adena J. Zlochower, Jenn-Jier James Lien, Takeo Kanade
FG1
1998 Automated Facial Expression Recognition Based on FACS Action Units
Jenn-Jier James Lien, Takeo Kanade, Jeffrey F. Cohn, Ching-Chung Li
FG3
1998 Optical Flow Estimation Using Wavelet Motion Model
abstract
A motion estimation algorithm using wavelet approximation as an optical flow model has been developed to estimate accurate dense optical flow from an image sequence. This wavelet motion model is particularly useful in estimating optical flows with large displacement. Traditional pyramid methods which use the coarse-to-fine image pyramid by image burring in estimating optical flow often produce incorrect results when the coarse-level estimates contain large errors that cannot be corrected at the subsequent finer levels. This happens when regions of low texture become flat or certain patterns result in spatial aliasing due to image blurring. Our method, in contrast, uses large-to-small full-resolution regions without blurring images, and simultaneously optimizes the coarser and finer parts of optical flow so that the large and small motion can be estimated correctly. We compare results obtained by using our method with those obtained by using one of the leading optical flow methods, the Szeliski pyramid spline-based method. The experiments include cases of small displacement (less than 4 pixels under 128/spl times/128 image size or equivalent displacement under other image sizes), and those of large displacement (10 pixels). While both methods produce comparable results when the displacements are small, our method outperforms pyramid spline-based method when the displacements are large.
Yu-Te Wu, Takeo Kanade, Jeffrey F. Cohn, Ching-Chung Li
ICCV3
1994 Quantitative description and differentiation of fundamental frequency contours
Christopher A. Moore, Jeffrey F. Cohn, Gary S. Katz
Comput. Speech Lang.2