Shaun J. Canavan

dblp:92/8116 · DBLP profile ↗
← Back
45ranked-venue papers
7as first author
17since 2021 · last 2025
0000-0002-1538-476XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 31 · 6 first-author · 9 since 2021Artificial intelligence and machine learning · 30 · 3 first-author · 11 since 2021Human-computer interaction and ubiquitous computing · 7 · 1 first-author · 4 since 2021Security and privacy · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Context-Based Screening of Autism Risk in Children
abstract
Autism Spectrum Disorder (ASD) is a neurode-velopmental condition characterized by difficulties in social communication, restricted interests, and repetitive behaviors. ASD approximately affects 1 in 36 children, however, most children are not formally diagnosed until they are four or five years old, significantly delaying early intervention. Traditional diagnostic methods for autism are often time-consuming, costly, and prone to subjective bias. Therefore, automatic methods for ASD screening and diagnosis are vital as they can help professionals to make earlier and more accurate diagnosis. Considering this, we present a context-based approach for screening autism risk in children. Our method uses a vision transformer to analyze facial data from the publicly available USF ASD dataset. The results are encouraging showing that context is important for the screening of autism in children. Empirical evaluation shows the enhanced performance of our model in accurately predicting autism risk levels, as evaluated under the RITA-T scoring criteria. We also perform a qualitative analysis using Grad-Cam visualizations, providing insights into the classification results.
Rupal Agarwal, Heather Agazzi, Shaun J. Canavan
FG3
2025 Multimodal, context-based dataset of children with Post Traumatic Stress Disorder
Saandeep Aathreya, Tara Nourivandi, Alison Salloum, Leigh J. Ruth, Eric A. Storch, Shaun J. Canavan
Pattern Recognit. Lett.6
2024 Multimodal Behavior Analysis and Impact of Culture on Affect
abstract
It has been shown that emotions are learned in a cultural way and expressions are often used to help convey these emotional states. Considering this, in this work, we investigate multimodal cultural behavior differences across 6 different cultures. More specifically, we investigate head pose, action units, and facial landmarks in British, Chinese, German, Greek, Hungarian, and Serbian cultures. Along with this, we also investigate the differences along valence and arousal dimensions for these cultures. To conduct this investigation, we evaluate the SEWA multimodal and multi-cultural dataset. We find varying differences exist that are impacted by culture, context, and modality. Based on these findings, we also perform context classification that takes into account these differences in culture. We show that incorporating culture into our pipeline improves classification performance.
Tara Nourivandi, Saandeep Aathreya, Shaun J. Canavan
ACII3
2024 FlowCon: Out-of-Distribution Detection Using Flow-Based Contrastive Learning
Saandeep Aathreya, Shaun J. Canavan
ECCV (16)2
2024 Welcome
abstract
It was our pleasure and privilege to welcome you to Istanbul for the 18th IEEE International Conference on Automatic Face and Gesture Recognition (FG 2024). We hope your experience at FG was rewarding both professionally and personally!
Hazim Kemal Ekenel, Albert Ali Salah, Arun Ross, Vitomir Struc, Lale Akarun, Xilin Chen 0001, Shaun J. Canavan
FG7
2024 Context-based Dataset for Analysis of Videos of Autistic Children
abstract
Autism affects as many as 1 in 44 youth, with many higher-functioning children not diagnosed until school-age or later. Currently, diagnosing autism is a lengthy process often delivered in varying settings (i.e., context) by a multi-disciplinary team, where the result can include subjective bias. Automatic approaches that can help professionals with diagnosis can result in earlier and quicker diagnosis. To help facilitate development of automatic approaches, we present a new, context-based dataset for analysis of videos of autistic children. The data was collected from 14 children, using the gold-standard RITA-T evaluation. Along with the dataset we also provide a baseline, context-based approach for classification of these videos. The baseline shows encouraging results that context matters for classification, and we also detail findings about which features (face, body, or gaze) are most encouraging within those contexts.
Sk Rahatul Jannat, Heather Agazzi, Shaun J. Canavan
FG3
2024 Mitigating Class Imbalance for Facial Expression Recognition Using SMOTE on Deep Features
abstract
Applications of affective computing are used in high-stake decision-making systems. Many of these systems that have applications in medicine, education, entertainment, and security user facial expression recognition (FER). Because of the significance of these applications, fair, unbiased, generalizable, and valid systems are essential. A source of error in FER can be class imbalance, which leads to bias against labels that occur less frequently. Considering this, we propose a solution to mitigate class imbalance for FER by synthesizing new deep features utilizing the Synthetic Minority Oversampling Technique (SMOTE) on deep features extracted from a neural network. We refer this to as a SMOTE layer in a neural network. More specifically, we up-sample the deep features of each minority class so that classes are balanced during training. To validate the efficacy of the proposed approach, we perform experiments on the publicly available datasets BP4D and AffectNet. We show encouraging results on BP4D, for pain recognition from facial expressions, and on AffectNet for 8-class classification of facial expressions.
Tara Nourivandi, Saurabh Hinduja, Shivam Srivastava, Jeffrey F. Cohn, Shaun J. Canavan
FG5
2024 Time to retire F1-binary score for action unit detection
abstract
Detecting action units is an important task in face analysis, especially in facial expression recognition. This is due, in part, to the idea that expressions can be decomposed into multiple action units. To evaluate systems that detect action units, F1-binary score is often used as the evaluation metric. In this paper, we argue that F1-binary score does not reliably evaluate these models due largely to class imbalance. Because of this, F1-binary score should be retired and a suitable replacement should be used. We justify this argument through a detailed evaluation of the negative influence of class imbalance on action unit detection. This includes an investigation into the influence of class imbalance in train and test sets and in new data (i.e., generalizability). We empirically show that F1-micro should be used as the replacement for F1-binary.
Saurabh Hinduja, Tara Nourivandi, Jeffrey F. Cohn, Shaun J. Canavan
Pattern Recognit. Lett.4
2024 Spatio-Temporal Graph Analytics on Secondary Affect Data for Improving Trustworthy Emotional AI
abstract
Ethical affective computing (AC) requires maximizing the benefits to users while minimizing its harm to obtain trust from users. This requires responsible development and deployment to ensure fairness, bias mitigation, privacy preservation, and accountability. To obtain this, we require methodologies that can quantify, visualize, analyze, and mine insights from affect data. Hence, in this paper, we propose a spatio-temporal model for representing secondary affect data from network sciences' perspective. We propose a network science-based model to represent spatio-temporal data, e.g., action units' sequences, and continuous affect reports. In particular, the proposed model captures the spatial and temporal strength of the relationship among essential variables in the data. The proposed model allows to analyze data as a whole system. We also demonstrated the use case of the model for graph analytics on secondary affect data that can assist to measure and quantify several issues that can be originated from the study setup, data recording devices, and the influences/biases that can originate from the perspective of the affect reporters. We also demonstrated the use cases of the proposed method on ethical trustworthy emotional AI via measuring biases from de-identified data and how it contributes towards ethics, transparency, value alignment, and governance.
Md Taufeeq Uddin, Lijun Yin 0001, Shaun J. Canavan
IEEE Trans. Affect. Comput.3
2023 Predicting Loneliness from Subject Self-Report
abstract
In this work, we propose to predict four items from the UCLA loneliness scale using subject self-report scores. Using subjective self-reporting, over 14 days, for positive and negative affect, and depression and anxiety we evaluate both subject dependent (personalized) and subject independent experimental design. We evaluate four approaches for prediction, namely random forest, support vector machine, k-nearest neighbor, and logistic regression. We find that the features (self-report) are relatively stable across all four approaches. Along with each individual self-report feature, we also evaluate the fusion of all features where they are concatenated into one feature vector. Through our experimental design, we show that UCLA loneliness items can be predicted, and that the fusion of features (positive and negative affect, and depression and anxiety) is the most encouraging way to do this prediction. We also show that the number of days, used for prediction, has a noticeable impact on the results and that personalization helps with prediction.
Liza Jivnani, Fallon Goodman, Jonathan Rottenberg, Shaun J. Canavan
ACII4
2023 Multimodal Context-Based Continuous Authentication
abstract
We present a new multimodal, context-based dataset for continuous authentication. The dataset contains 27 subjects, with an age range of [8, 72], where data has been collected across multiple sessions while the subjects are watching videos meant to elicit an emotional response. Collected data includes accelerometer data, heart rate, electrodermal activity, skin temperature, and face videos. We also propose a baseline approach for fair comparisons when using the proposed dataset. The approach uses a combination of a pretrained backbone network with supervised contrastive loss for face. Time-series features are also extracted, from the physiological signals, which are used for classification. This approach, on the proposed dataset, results in an average accuracy, precision, and recall of 76.59%, 88.90, and 53.25, respectively, on electrical signals, and 90.39%, 98.77, and 75.71, respectively on face videos.
Saandeep Aathreya, Meghna Chaudhary, Tempestt J. Neal, Shaun J. Canavan
IJCB4
2023 Cooperative Learning for Personalized Context-Aware Pain Assessment From Wearable Data
abstract
Despite the promising performance of automated pain assessment methods, current methods suffer from performance generalization due to the lack of relatively large, diverse, and annotated pain datasets. Further, the majority of current methods do not allow responsible interaction between the model and user, and do not take different internal and external factors into consideration during the model's design and development. This article aims to provide an efficient cooperative learning framework for the lack of annotated data while facilitating responsible user communication and taking individual differences into consideration during the development of pain assessment models. Our results using body and muscle movement data, collected from wearable devices, demonstrate that the proposed framework is effective in leveraging both the human and the machine to efficiently learn and predict pain.
Md Taufeeq Uddin, Ghada Zamzmi, Shaun J. Canavan
IEEE J. Biomed. Health Informatics3
2022 Quantified Facial Expressiveness for Affective Behavior Analytics
abstract
The quantified measurement of facial expressiveness is crucial to analyze human affective behavior at scale. Unfortunately, methods for expressiveness quantification at the video frame-level are largely unexplored, unlike the study of discrete expression. In this work, we propose an algorithm that quantifies facial expressiveness using a bounded, continuous expressiveness score using multimodal facial features, such as action units (AUs), landmarks, head pose, and gaze. The proposed algorithm more heavily weights AUs with high intensities and large temporal changes. The proposed algorithm can compute the expressiveness in terms of discrete expression, and can be used to perform tasks including facial behavior tracking and subjectivity quantification in context. Our results on benchmark datasets show the proposed algorithm is effective in terms of capturing temporal changes and expressiveness, measuring subjective differences in context, and extracting useful insight.
Md Taufeeq Uddin, Shaun J. Canavan
WACV2
2022 Classification of emotions using EEG activity associated with different areas of the brain
Rupal Agarwal, Marvin Andujar, Shaun J. Canavan
Pattern Recognit. Lett.3
2022 AffectiveTDA: Using Topological Data Analysis to Improve Analysis and Explainability in Affective Computing
abstract
We present an approach utilizing Topological Data Analysis to study the structure of face poses used in affective computing, i.e., the process of recognizing human emotion. The approach uses a conditional comparison of different emotions, both respective and irrespective of time, with multiple topological distance metrics, dimension reduction techniques, and face subsections (e.g., eyes, nose, mouth, etc.). The results confirm that our topology-based approach captures known patterns, distinctions between emotions, and distinctions between individuals, which is an important step towards more robust and explainable emotion recognition by machines.
Hamza Elhamdadi, Shaun J. Canavan, Paul Rosen 0001
IEEE Trans. Vis. Comput. Graph.2
2021 Investigation into Recognizing Context Over Time using Physiological Signals
abstract
In this paper, we investigate recognizing context over time using physiological signals. Using the CASE dataset we evaluate both unimodal and multimodal approaches to physiological-based context recognition, over time. For recognition, we evaluate a random forest, as well as state-of-the-art neural network. These classifiers are evaluated using accuracy, Kappa, and F1-Macro metrics. Our results suggest that the fusion of EMG signals is more accurate, at recognizing context over time, compared to the fusion of non-EMG physiological signals. Although the fusion of non-EMG has a comparatively higher accuracy, ECG data results in the highest unimodal accuracy. Considering this, we analyze how the signals are correlated, including when the are fused (i.e. multimodal). We also perform a cross-gender analysis (e.g. training on male data and testing on female data) suggesting some generalizability across gender.
Saurabh Hinduja, Shaun J. Canavan
ACII3
2021 Expression Recognition Across Age
abstract
Expression recognition is an important and growing field in AI. It has applications in fields including, but not limited to, medicine, security, and entertainment. A large portion of research, in this area, has focused on recognizing expressions of young and middle-age adults with less focus on children and elderly subjects. This focus can lead to unintentional bias across age, resulting in less accurate models. Considering this, we investigate the impact of age on expression recognition. To facilitate this investigation, we evaluate two state-of-the-art datasets, that focus on different age ranges (children and elderly), namely EmoReact and ElderReact. We propose a Siamese-network-based approach that learns the semantic similarity of expressions relative to each age. We show that the proposed approach, to expression recognition, is able to generalize across age. We show the proposed approach is comparable to or outperforms current state-of-the-art on the EmoReact and ElderReact datasets.
Sk Rahatul Jannat, Shaun J. Canavan
FG2
2020 Real-time Action Unit Intensity Detection
abstract
We present a real-time system for action unit intensity detection. We train a convolutional neural network, with images from DISFA+, to detect the intensities of 12 action units that are commonly found in the literature. Along with real-time capabilities, the system is also able to detect intensities on static images, and in the wild videos. While the focus of this work is on the detection of action unit intensities, we are able to implicitly detect the occurrence of them as well. This is done by setting all action units with an intensity greater than 0 as active. In doing this, we calculate the F1-micro and F1-binary scores for the DISFA+ dataset.
Saurabh Hinduja, Shaun J. Canavan
FG2
2020 Multimodal Fusion of Physiological Signals and Facial Action Units for Pain Recognition
abstract
In this paper, we propose a method for pain recognition by fusing physiological signals (heart rate, respiration, blood pressure, and electrodermal activity) and facial action units. We provide experimental validation that the fusion of these signals results in a positive impact to the accuracy of pain recognition, compared to using only one modality (i.e. physiological or action units). These experiments are conducted on subjects from the BP4D+ multimodal emotion corpus, and include same- and cross-gender experiments. We also investigate the correlation between the two modalities to gain further insight into applications of pain recognition. Results suggest the need for larger and more varied datasets that include physiological signals and action units that have been coded for all facial frames.
Saurabh Hinduja, Shaun J. Canavan
FG2
2020 Recognizing Perceived Emotions from Facial Expressions
abstract
Expression recognition has seen an increase in research in past years, however, little work has been on recognizing perceived emotion (i.e. subject self-reporting of emotion). Considering this, we investigate the perceived emotion of subjects that perform tasks meant to elicit emotion. To facilitate this investigation, we use the BP4D+ multimodal spontaneous emotion corpus. We first statistically analyze the subject's perceived emotions across 10 tasks available in BP4D+. We show the percentage of subjects that felt specific emotions for each of the tasks. This is done across all tested subjects, as well as male and female subjects independently. Along with our statistical analysis, we also propose a 3D convolutional neural network (CNN) architecture to recognize multiple emotions felt for each task sequence. We report accuracy, Fl-binary and AUC for all subjects, as well as male and female subjects.
Saurabh Hinduja, Shaun J. Canavan, Lijun Yin 0001
FG2
2020 Three-level Training of Multi-Head Architecture for Pain Detection
abstract
Precise pain detection is a complex task even for trained professionals. There are many occasions where self-reporting fails to capture the decisive pain measurement. Facial expressions help extract the subtle emotions which can be leveraged to detect pain levels. In this paper, we present our approach to task 1 of the FG 2020 EmoPain Challenge: Pain-related Behavior Analysis, as well as our experimental design and results. We utilize a multi-head approach with the combined features of facial action units, facial landmarks, HOG and deep features. This multimodal approach provides insight into the contribution of each of these features and their consolidated effect. To improve our regression model, we adopt a three-level architecture where we observe an increase in prediction accuracy as the levels deepen. We record results comparable to the baseline on the challenge validation set.
Saandeep Aathreya, Saurabh Hinduja, Shaun J. Canavan
FG3
2020 Mood Versus Identity: Studying the Influence of Affective States on Mobile Biometrics
abstract
Mobile device usage data such as mobile app use and acceleration measurements fluctuate often as individuals carry out their daily tasks. As these data have emerged in recent years as promising biometric identifiers, it is important to understand the many causes of these variations such that these systems can adapt without degradation in performance. In this paper, we seek to understand the impact of changes in a person's mood on the performance of a mobile biometric system using a publicly available dataset of 27 subjects. We explore the verification and identification tasks, along with mood prediction from smartphone data. We achieved an equal error rate of 3% and a d-prime value of 5.05 for the verification task, wherein experiments showed that verification is minimally influenced by an individual's mood, although negative arousal slightly degraded performance. We created a multi-class problem to study the identification task, achieving an average 83% F 1-score. Here, we observed that subjects with lower identification accuracy (95%) experienced more mood changes compared to the average. Contrasting previous claims, our findings suggest that frequent changes in mood may have little negative impact performance. Finally, positive arousal and negative valence yielded the highest area under the curve (0.67) for mood prediction. This was also the class associated with the highest average genuine and lowest average imposter scores for verification experiments, suggesting a correspondence between recognition and mood prediction tasks that applications such as sensor-enhanced mHealth apps could leverage.
Tempestt J. Neal, Shaun J. Canavan
FG2
2020 Sign Language Recognition in Virtual Reality
abstract
A real-time system for signal language recognition in virtual reality (VR) is presented in this paper. The system makes use of an egocentric view with the Vive HTC VR headset along with a Leap Motion controller. In this demo, a random forest is used to classify the 26 letters of the alphabet, in American Sign Language, from hand-crafted features extracted from the Leap Motion controller. We detail offline classification results showing the expressive power of the features used for recognition.
Jacob Schioppo, Zachary Meyer, Diego Fabiano, Shaun J. Canavan
FG4
2020 Multimodal Multilevel Fusion for Sequential Protective Behavior Detection and Pain Estimation
abstract
In this paper, we present our approach to the FG 2020 EmoPain Challenge for tasks 2 (pain estimation) and 3 (protective behavior detection) from multimodal movement data. We propose to perform sequential protective behavior detection and pain estimation using human movement information. First, we predict the existence of pain, and then use this information along with the multimodal movement data for protective behavior detection. Finally, this information is fused to estimate level of pain. In this work, we apply both early fusion (feature fusion including metadata, modalities, exercises and probabilities) and post-fusion (decision fusion). The proposed approach is encouraging, as it outperforms the baseline, with high margin for both pain estimation and protective behavior detection on the EmoPain challenge 2020 dataset.
Md Taufeeq Uddin, Shaun J. Canavan
FG2
2020 Recognizing Emotion in the Wild using Multimodal Data
abstract
In this work, we present our approach for all four tracks of the eighth Emotion Recognition in the Wild Challenge (EmotiW 2020). The four tasks are group emotion recognition, driver gaze prediction, predicting engagement in the wild, and emotion recognition using physiological signals. We explore multiple approaches including classical machine learning tools such as random forests, state of the art deep neural networks, and multiple fusion and ensemble-based approaches. We also show that similar approaches can be used across tracks as many of the features generalize well to the different problems (e.g. facial features). We detail evaluation results that are either comparable to or outperform the baseline results for both the validation and testing for most of the tracks.
Shivam Srivastava, Saandeep Aathreya, Saurabh Hinduja, Sk Rahatul Jannat, Hamza Elhamdadi, Shaun J. Canavan
ICMI6
2020 Quantified Facial Temporal-Expressiveness Dynamics for Affect Analysis
abstract
The quantification of visual affect data (e.g. face images) is essential to build and monitor automated affect modeling systems efficiently. Considering this, this work proposes quantified facial Temporal-expressiveness Dynamics (TED) to quantify the expressiveness of human faces. The proposed algorithm leverages multimodal facial features by incorporating static and dynamic information to enable accurate measurements of facial expressiveness. We show that TED can be used for high-level tasks such as summarization of unstructured visual data, and expectation from and interpretation of automated affect recognition models. To evaluate the positive impact of using TED, a case study was conducted on spontaneous pain using the UNBC-McMaster spontaneous shoulder pain dataset. Experimental results show the efficacy of using TED for quantified affect analysis.
Md Taufeeq Uddin, Shaun J. Canavan
ICPR2
2020 Gaze-based classification of autism spectrum disorder
Diego Fabiano, Shaun J. Canavan, Heather Agazzi, Saurabh Hinduja, Dmitry B. Goldgof
Pattern Recognit. Lett.2
2019 Emotion Recognition Using Fused Physiological Signals
abstract
In this paper, we propose a new representation of human emotion through the fusion of physiological signals. Using the variance of these signals, the proposed method increases the effect of signals that contribute to the recognition accuracy, while decreasing the effect of those that do not. The new representation is a powerful approach to recognizing emotions. We investigate this by comparing against emotion recognition results from non-fused physiological signals. Both the fused and non-fused signals are used to train feedforward neural networks to recognize a range of emotion. We show that the fused method outperforms each individual signal across all emotions tested. We test the efficacy of the proposed approach on two publicly available datasets, namely BP4D+ and DEAP, showing state-of-the-art results on both. To the best of our knowledge this is the first work to present emotion recognition results using physiological signals on all subjects from BP4D+.
Diego Fabiano, Shaun J. Canavan
ACII2
2019 Deformable Synthesis Model for Emotion Recognition
abstract
In this paper, we propose a deformable synthesis model that can be used to synthesize data to train deep neural networks for the task of emotion recognition. This model is created through the use of 3D facial landmarks, which are then projected to the 2D image plane for training a deep network. We show that this model can accurately recognize a range of emotions that include happiness, sadness, and fear. We test the efficacy of our proposed approach on three publicly available 3D face databases, namely BU4DFE, BP4D, and BP4D+. We show that the proposed method can accurate recognize emotion when training and testing on the same database, as well as cross-database training and testing on all 3 databases. We show the proposed method results in accurate recognition of emotion using deep neural networks outperforming current state of the art on each of the tested databases.
Diego Fabiano, Shaun J. Canavan
FG2
2019 Fusion of Hand-crafted and Deep Features for Empathy Prediction
abstract
We propose an approach to the OMG-Empathy Challenge for predicting self-annotated, continuous values of valence within the range [-1,1]. We propose the fusion of hand-crafted and deep features, extracted from both actor and listener data, to predict these valence levels. The handcrafted features include image level fusion, facial landmarks, and spectrogram features. Our proposed fusion approach can utilized in multiple parts (i.e. sub-modules), specifically utilized for the generalized track, leading to unique submissions to address the challenge problem. First, both actor and listener images are fused at the image-level. Secondly, facial landmarks from both the actor and listener are fused into one feature vector which is then used to train a random forest for prediction of continuous valence levels. Finally, we use a weighted fusion of the predicted values from both hand-crafted and deep features. We show competitive results on the OMG-Empathy challenge validation set.
Saurabh Hinduja, Md Taufeeq Uddin, Sk Rahatul Jannat, Astha Sharma, Shaun J. Canavan
FG5
2019 Multi-subspace supervised descent method for robust face alignment
abstract
Supervised Descent Method (SDM) is one of the leading cascaded regression approaches for face alignment with state-of-the-art performance and a solid theoretical basis. However, SDM is prone to local optima and likely averages conflicting descent directions. This makes SDM ineffective in covering a complex facial shape space due to large head poses and rich non-rigid face deformations. In this paper, a novel two-step framework called multi-subspace SDM (MS-SDM) is proposed to equip SDM with a stronger capability for dealing with unconstrained faces. The optimization space is first partitioned with regard to shape variations using k-means. The generated subspaces show semantic significance which highly correlates with head poses. Faces among a certain subspace also show compatible shape-appearance relationships. Then, Naive Bayes is applied to conduct robust subspace prediction by concerning about the relative proximity of each subspace to the sample. This guarantees that each sample can be allocated to the most appropriate subspace-specific regressor. The proposed method is validated on benchmark face datasets with a mobile facial tracking implementation.
Jianwen Lou, Xiaoxu Cai, Yiming Wang 0001, Hui Yu 0001, Shaun J. Canavan
Multim. Tools Appl.5
2018 Spontaneous and Non-Spontaneous 3D Facial Expression Recognition Using a Statistical Model with Global and Local Constraints
abstract
In this paper, we propose a novel method for 3D facial expression recognition based on a statistical shape model with global and local constraints. We show that the combination of the global shape of the face, along with local shape index-based information can be used to recognize a range of expressions. These expressions include happiness, sadness, surprise, embarrassment, fear, nervousness, anger, disgust, and pain. We give insights into which features are important for facial expression recognition through statistical analysis. We also show that our proposed method outperforms the current state-of-the-art methods on spontaneous and non-spontaneous facial data.
Diego Fabiano, Shaun J. Canavan
ICIP2
2017 Combining gaze and demographic feature descriptors for autism classification
abstract
People with autism suffer from social challenges and communication difficulties, which may prevent them from leading a fruitful and enjoyable life. It is imperative to diagnose and start treatments for autism as early as possible and, in order to do so, accurate methods of identifying the disorder are vital. We propose a novel method for classifying autism through the use of eye gaze and demographic feature descriptors that include a subject's age and gender. We construct feature descriptors that incorporate the subject's age and gender, as well as features based on eye gaze data. Using eye gaze information from the National Database for Autism Research, we tested our constructed feature descriptors on three different classifiers; random regression forests, C4.5 decision tree, and PART. Our proposed method for classifying autism resulted in a top classification rate of 96.2%.
Shaun J. Canavan, Melanie Chen, Robert Valdez, Miles Yaeger, Huiyi Lin, Lijun Yin 0001
ICIP1
2017 Hand gesture recognition using a skeleton-based feature representation with a random regression forest
abstract
In this paper, we propose a method for automatic hand gesture recognition using a random regression forest with a novel set of feature descriptors created from skeletal data acquired from the Leap Motion Controller. The efficacy of our proposed approach is evaluated on the publicly available University of Padova Microsoft Kinect and Leap Motion dataset, as well as 24 letters of the English alphabet in American Sign Language. The letters that are dynamic (e.g. j and z) are not evaluated. Using a random regression forest to classify the features we achieve 100% accuracy on the University of Padova Microsoft Kinect and Leap Motion dataset. We also constructed an in-house dataset using the 24 static letters of the English alphabet in ASL. A classification rate of 98.36% was achieved on this dataset. We also show that our proposed method outperforms the current state of the art on the University of Padova Microsoft Kinect and Leap Motion dataset.
Shaun J. Canavan, Walter Keyes, Ryan Mccormick, Julie Kunnumpurath, Tanner Hoelzel, Lijun Yin 0001
ICIP1
2016 Multimodal Spontaneous Emotion Corpus for Human Behavior Analysis
abstract
Emotion is expressed in multiple modalities, yet most research has considered at most one or two. This stems in part from the lack of large, diverse, well-annotated, multimodal databases with which to develop and test algorithms. We present a well-annotated, multimodal, multidimensional spontaneous emotion corpus of 140 participants. Emotion inductions were highly varied. Data were acquired from a variety of sensors of the face that included high-resolution 3D dynamic imaging, high-resolution 2D video, and thermal (infrared) sensing, and contact physiological sensors that included electrical conductivity of the skin, respiration, blood pressure, and heart rate. Facial expression was annotated for both the occurrence and intensity of facial action units from 2D video by experts in the Facial Action Coding System (FACS). The corpus further includes derived features from 3D, 2D, and IR (infrared) sensors and baseline results for facial expression and action unit detection. The entire corpus will be made available to the research community.
Zheng Zhang 0023, Jeffrey M. Girard, Yue Wu 0002, Xing Zhang 0012, Peng Liu 0039, Umur A. Ciftci, Shaun J. Canavan, Michael Reale, Andrew Horowitz, Huiyuan Yang, Jeffrey F. Cohn, Lijun Yin 0001
CVPR7
2015 Landmark localization on 3D/4D range data using a shape index-based statistical shape model with global and local constraints
Shaun J. Canavan, Peng Liu 0039, Xing Zhang 0012, Lijun Yin 0001
Comput. Vis. Image Underst.1
2014 BP4D-Spontaneous: a high-resolution spontaneous 3D dynamic facial expression database
Xing Zhang 0012, Lijun Yin 0001, Jeffrey F. Cohn, Shaun J. Canavan, Michael Reale, Andy Horowitz, Peng Liu 0039, Jeffrey M. Girard
Image Vis. Comput.4
2013 Fitting and tracking 3D/4D facial data using a temporal deformable shape model
abstract
In this paper, we propose a novel method for detecting and tracking landmark facial features on purely geometric 3D and 4D range models. Our proposed method involves fitting a new multi-frame constrained 3D temporal deformable shape model (TDSM) to range data sequences. We consider this a temporal based deformable model as we concatenate consecutive deformable shape models into a single model driven by the appearance of facial expressions. This allows us to simultaneously fit multiple models over a sequence of time with one TDSM. To our knowledge, it is the first work to address multiple shape models as a whole to track 3D dynamic range sequences without assistance of any texture information. The accuracy of the tracking results is evaluated by comparing the detected landmarks to the ground truth. The efficacy of the 3D feature detection and tracking over range model sequences has also been validated through an application in 3D geometric based face and expression analysis and expression sequence segmentation. We tested our method on the publicly available databases, BU-3DFE [15], BU-4DFE [16], and FRGC 2.0 [12]. We also validated our approach on our newly developed 3D dynamic spontaneous expression database [17].
Shaun J. Canavan, Xing Zhang 0012, Lijun Yin 0001
ICME1
2013 Art Critic: Multisignal Vision and Speech Interaction System in a Gaming Context
abstract
True immersion of a player within a game can only occur when the world simulated looks and behaves as close to reality as possible. This implies that the game must correctly read and understand, among other things, the player's focus, attitude toward the objects/persons in focus, gestures, and speech. In this paper, we proposed a novel system that integrates eye gaze estimation, head pose estimation, facial expression recognition, speech recognition, and text-to-speech components for use in real-time games. Both the eye gaze and head pose components utilize underlying 3-D models, and our novel head pose estimation algorithm uniquely combines scene flow with a generic head model. The facial expression recognition module uses the local binary patterns with three orthogonal planes approach on the 2-D shape index domain rather than the pixel domain, resulting in improved classification. Our system has also been extended to use a pan-tilt-zoom camera driven by the Kinect, allowing us to track a moving player. A test game, Art Critic, is also presented, which not only demonstrates the utility of our system but also provides a template for player/non-player character (NPC) interaction in a gaming context. The player alters his/her view of the 3-D world using head pose, looks at paintings/NPCs using eye gaze, and makes an evaluation based on the player's expression and speech. The NPC artist will respond with facial expression and synthetic speech based on its personality. Both qualitative and quantitative evaluations of the system are performed to illustrate the system's effectiveness.
Michael Reale, Peng Liu 0039, Lijun Yin 0001, Shaun J. Canavan
IEEE Trans. Cybern.4
2011 Recognizing face sketches by a large number of human subjects: A perception-based study for facial distinctiveness
abstract
Understanding how humans recognize face sketches drawn by artists is of significant value to both criminal investigators and researchers in computer vision, face biometrics and cognitive psychology. However, large scale experimental studies of hand-drawn face sketches are still very limited in terms of the number of artists, the number of sketches, and the number of human evaluators involved. In this paper, we reported the results of a series of psychological experiments in which 406 volunteers were asked to recognize 250 sketches drawn by 5 different artists. The primary findings are: (i) Sketch quality (artist factor) has a significant effect on human performance. Inter-artist variation as measured by the mean recognition rate can be as high as 31%; (ii) Participants showed a higher tendency to match multiple sketches to one photo than to second-guess their answers. The multi-match ratio seems correlated to the recognition rate, while second-guessing had no significant effect on human performance; (iii) For certain highly recognized faces, their rankings were very consistent using three measuring parameters: recognition rate, multi-match ratio, and second-guess ratio, suggesting that the three parameters could provide valuable information to quantify facial distinctiveness.
Yong Zhang 0017, Steve Ellyson, Anthony Zone, Priyanka Gangam, John R. Sullins, Christine McCullough, Shaun J. Canavan, Lijun Yin 0001
FG7
2011 3D face sketch modeling and assessment for component based face recognition
abstract
3D facial representations have been widely used for face recognition. There has been intensive research on geometric matching and similarity measurement on 3D range data and 3D geometric meshes of individual faces. However, little investigation has been done on geometric measurement for 3D sketch models. In this paper, we study the 3D face recognition from 3D face sketches which are derived from hand-drawn sketches and machine generated sketches. First, we have developed a 3D sketch modeling approach to create 3D facial sketch models from 2D facial sketch images. Second, we compared the 3D sketches to the existing 3D scans. Third, the 3D face similarity is measured between 3D sketches versus 3D scans, and 3D sketches versus 3D sketches based on the spatial Hidden Markov Model (HMM) classification. Experiments are conducted on both the BU-4DFE database and YSU face sketch database, resulting in a recognition rate at around 92% on average.
Shaun J. Canavan, Xing Zhang 0012, Lijun Yin 0001, Yong Zhang 0017
IJCB1
2011 A Multi-Gesture Interaction System Using a 3-D Iris Disk Model for Gaze Estimation and an Active Appearance Model for 3-D Hand Pointing
abstract
In this paper, we present a vision-based human-computer interaction system, which integrates control components using multiple gestures, including eye gaze, head pose, hand pointing, and mouth motions. To track head, eye, and mouth movements, we present a two-camera system that detects the face from a fixed, wide-angle camera, estimates a rough location for the eye region using an eye detector based on topographic features, and directs another active pan-tilt-zoom camera to focus in on this eye region. We also propose a novel eye gaze estimation approach for point-of-regard (POR) tracking on a viewing screen. To allow for greater head pose freedom, we developed a new calibration approach to find the 3-D eyeball location, eyeball radius, and fovea position. Moreover, in order to get the optical axis, we create a 3-D iris disk by mapping both the iris center and iris contour points to the eyeball sphere. We then rotate the fovea accordingly and compute the final, visual axis gaze direction. This part of the system permits natural, non-intrusive, pose-invariant POR estimation from a distance without resorting to infrared or complex hardware setups. We also propose and integrate a two-camera hand pointing estimation algorithm for hand gesture tracking in 3-D from a distance. The algorithms of gaze pointing and hand finger pointing are evaluated individually, and the feasibility of the entire system is validated through two interactive information visualization applications.
Michael Reale, Shaun J. Canavan, Lijun Yin 0001, Kaoning Hu, Terry Hung
IEEE Trans. Multim.2
2010 Evaluation of Multi-frame Fusion Based Face Classification Under Shadow
abstract
A video sequence of a head moving across a large pose angle contains much richer information than a single-view image, and hence has greater potential for identification purposes. This paper explores and evaluates the use of a multi-frame fusion method to improve face recognition in the presence of strong shadow. The dataset includes videos of 257 subjects who rotated their heads by 0° to 90°. Experiments were carried out using ten video frames per subject that were fused on the score level. The primary findings are: (i) A significant performance increase was observed, with the recognition rate being doubled from 40% using a single frame to 80% using ten frames; (ii) The performance of multi-frame fusion is strongly related to its inter-frame variation that measures its information diversity.
Shaun J. Canavan, Benjamin Johnson 0004, Michael Reale, Yong Zhang 0017, Lijun Yin 0001, John R. Sullins
ICPR1
2010 Hand Pointing Estimation for Human Computer Interaction Based on Two Orthogonal-Views
abstract
Hand pointing has been an intuitive gesture for human interaction with computers. Big challenges are still posted for accurate estimation of finger pointing direction in a 3D space. In this paper, we present a novel hand pointing estimation system based on two regular cameras, which includes hand region detection, hand finger estimation, two views' feature detection, and 3D pointing direction estimation. Based on the idea of binary pattern face detector, we extend the work to hand detection, in which a polar coordinate system is proposed to represent the hand region, and achieved a good result in terms of the robustness to hand orientation variation. To estimate the pointing direction, we applied an AAM based approach to detect and track 14 feature points along the hand contour from a top view and a side view. Combining two views of the hand features, the 3D pointing direction is estimated. The experiments have demonstrated the feasibility of the system.
Kaoning Hu, Shaun J. Canavan, Lijun Yin 0001
ICPR2
2009 Dynamic face appearance modeling and sight direction estimation based on local region tracking and scale-space topo-represention
abstract
Dynamic modeling of facial appearances and sight directions are demanded for HCI and multimedia applications. Traditional approaches for face tracking and eye tracking from 2D videos do not involve explicit facial modeling. In this paper, we propose to use an explicit 3D model to model the dynamic facial appearance as well as the eye shape to estimate the viewing direction. We apply active appearance models for local region tracking, and use a scale-space topographic representation for frame model instantiation. The individualized 3D models across video sequences allow us to estimate the iris viewing orientation dynamically. The proposed framework has been realized and tested in a person-independent fashion for AAM tracking and model instantiation using a single camera.
Shaun J. Canavan, Lijun Yin 0001
ICME1