Saurabh Hinduja

dblp:244/8136 · DBLP profile ↗
← Back
15ranked-venue papers
7as first author
8since 2021 · last 2024
0000-0003-1637-5950ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 6 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2024 Expanding PyAFAR: A Novel Privacy-Preserving Infant AU Detector
abstract
We enhance PyAFAR11Code will be available on: https:\\affectanalysisgroup.github.io/PyAFARI, an open source, Python-based library for facial action unit detection by introducing a privacy-protected infant AU detector. To prevent reconstruction of the training images, we train the infant AU detector by extracting histogram of gradients (HoG) features and using an efficient Light Gradient Boosting Machine (LightGBM) classifier. Models are trained with two large, well-annotated databases. The performance of our approach is comparable to previously developed deep models that have not been released due to privacy concerns. Our models are available for use and further fine-tuning, contributing to the advancement of facial action unit detection.
Itir Önal, Saurabh Hinduja, Maneesh Bilalpur, Daniel S. Messinger, Jeffrey F. Cohn
FG2
2024 Mitigating Class Imbalance for Facial Expression Recognition Using SMOTE on Deep Features
abstract
Applications of affective computing are used in high-stake decision-making systems. Many of these systems that have applications in medicine, education, entertainment, and security user facial expression recognition (FER). Because of the significance of these applications, fair, unbiased, generalizable, and valid systems are essential. A source of error in FER can be class imbalance, which leads to bias against labels that occur less frequently. Considering this, we propose a solution to mitigate class imbalance for FER by synthesizing new deep features utilizing the Synthetic Minority Oversampling Technique (SMOTE) on deep features extracted from a neural network. We refer this to as a SMOTE layer in a neural network. More specifically, we up-sample the deep features of each minority class so that classes are balanced during training. To validate the efficacy of the proposed approach, we perform experiments on the publicly available datasets BP4D and AffectNet. We show encouraging results on BP4D, for pain recognition from facial expressions, and on AffectNet for 8-class classification of facial expressions.
Tara Nourivandi, Saurabh Hinduja, Shivam Srivastava, Jeffrey F. Cohn, Shaun J. Canavan
FG2
2024 Time to retire F1-binary score for action unit detection
abstract
Detecting action units is an important task in face analysis, especially in facial expression recognition. This is due, in part, to the idea that expressions can be decomposed into multiple action units. To evaluate systems that detect action units, F1-binary score is often used as the evaluation metric. In this paper, we argue that F1-binary score does not reliably evaluate these models due largely to class imbalance. Because of this, F1-binary score should be retired and a suitable replacement should be used. We justify this argument through a detailed evaluation of the negative influence of class imbalance on action unit detection. This includes an investigation into the influence of class imbalance in train and test sets and in new data (i.e., generalizability). We empirically show that F1-micro should be used as the replacement for F1-binary.
Saurabh Hinduja, Tara Nourivandi, Jeffrey F. Cohn, Shaun J. Canavan
Pattern Recognit. Lett.1
2024 Multimodal Prediction of Obsessive-Compulsive Disorder and Comorbid Depression Severity and Energy Delivered by Deep Brain Electrodes
abstract
To develop reliable, valid, and efficient measures of obsessive-compulsive disorder (OCD) severity, comorbid depression severity, and total electrical energy delivered (TEED) by deep brain stimulation (DBS), we trained and compared random forests regression models in a clinical trial of participants receiving DBS for refractory OCD. Six participants were recorded during open-ended interviews at pre- and post-surgery baselines and then at 3-month intervals following DBS activation. Ground-truth severity was assessed by clinical interview and self-report. Visual and auditory modalities included facial action units, head and facial landmarks, speech behavior and content, and voice acoustics. Mixed-effects random forest regression with Shapley feature reduction strongly predicted severity of OCD, comorbid depression, and total electrical energy delivered by the DBS electrodes (intraclass correlation, ICC, = 0.83, 0.87, and 0.81, respectively. When random effects were omitted from the regression, predictive power decreased to moderate for severity of OCD and comorbid depression and remained comparable for total electrical energy delivered (ICC = 0.60, 0.68, and 0.83, respectively). Multimodal measures of behavior outperformed ones from single modalities. Feature selection achieved large decreases in features and corresponding increases in prediction. The approach could contribute to closed-loop DBS that would automatically titrate DBS based on affect measures.
Saurabh Hinduja, Ali Darzi, Itir Önal, Nicole R. Provenza, Ron Gadot, Eric A. Storch, Sameer A. Sheth, Wayne K. Goodman, Jeffrey F. Cohn
IEEE Trans. Affect. Comput.1
2023 Multimodal Feature Selection for Detecting Mothers' Depression in Dyadic Interactions with their Adolescent Offspring
abstract
Depression is the most common psychological disorder, a leading cause of disability world-wide, and a major contributor to inter-generational transmission of psychopathology within families. To contribute to our understanding of depression within families and to inform modality selection and feature reduction, it is critical to identify interpretable features in developmentally appropriate contexts. Mothers with and without depression were studied. Depression was defined as history of treatment for depression and elevations in current or recent symptoms. We explored two multimodal feature selection strategies in dyadic interaction tasks of mothers with their adolescent children for depression detection. Modalities included face and head dynamics, facial action units, speech-related behavior, and verbal features. The initial feature space was vast and inter-correlated (collinear). To reduce dimensionality and gain insight into the relative contribution of each modality and feature, we explored feature selection strategies using Variance Inflation Factor (VIF) and Shapley values. On an average collinearity correction through VIF resulted in about 4 times feature reduction across unimodal and multimodal features. Collinearity correction was also found to be an optimal intermediate step prior to Shapley analysis. Shapley feature selection following VIF yielded best performance. The top 15 features obtained through Shapley achieved 78% accuracy. The most informative features came from all four modalities sampled, which supports the importance of multimodal feature selection.
Maneesh Bilalpur, Saurabh Hinduja, Laura A. Cariola, Lisa Sheeber, Nick Alien, László A. Jeni, Louis-Philippe Morency, Jeffrey F. Cohn
FG2
2022 Ballistic Timing of Smiles is Robust to Context, Gender, Ethnicity, and National Differences
abstract
Smiles are highly variable. In some, contraction of the orbicularis oculi raises the cheeks and amplifies their intensity. In others, smile controls counteract the oblique pull of the zygomatic major, alter their shape, and decrease their intensity. Despite this variability, some features appear to be stereotypic. These features include a high correlation between the amplitude and velocity of smile onsets and same for smile offsets. The larger a smile's amplitude, the greater its velocity. This dependence is referred to as ballistic timing. In two relatively large publicly available databases (EB+ and Belfast), we tested the hypothesis of ballistic timing of smile onsets and offsets. We found high and consistent non-linear correlations between amplitude and velocity of both onsets and offsets that were robust to individual differences in persons (gender and ethnicity), context, presence or absence of the Duchenne marker, and country of residence (United States, Ireland, Peru). All$R_{2}$exceeded 0.85. The findings were highly consistent with ballistic timing. They have implications for detecting smiles that may be posed (which have been found to violate ballistic timing) and for realistic synthesis of smiles in social robots and virtual humans. Smiles that depart from ballistic timing are likely to be perceived as false or uncanny.
Maneesh Bilalpur, Saurabh Hinduja, Kenneth Goodrich, Jeffrey F. Cohn
ACII2
2022 Language Use in Mother-Adolescent Dyadic Interaction: Preliminary Results
abstract
This preliminary study applied a computer-assisted quantitative linguistic analysis to examine the effectiveness of language-based classification models to discriminate between mothers (n = 140) with and without history of treatment for depression (51% and 49%, respectively). Mothers were recorded during a problem-solving interaction with their adolescent child. Transcripts were manually annotated and analyzed using a dictionary-based, natural-language program approach (Linguistic Inquiry and Word Count). To assess the importance of linguistic features to correctly classify history of depression, we used Support Vector Machines (SVM) with interpretable features. Using linguistic features identified in the empirical literature, an initial SVM achieved nearly 63% accuracy. A second SVM using only the top 5 highest ranked SHAP features improved accuracy to 67.15%. The findings extend the existing literature base on understanding language behavior of depressed mood states, with a focus on the linguistic style of mothers with and without a history of treatment for depression and its potential impact on child development and trans-generational transmission of depression.
Laura A. Cariola, Saurabh Hinduja, Maneesh Bilalpur, Lisa Sheeber, Nicholas B. Allen, Louis-Philippe Morency, Jeffrey F. Cohn
ACII2
2021 Investigation into Recognizing Context Over Time using Physiological Signals
abstract
In this paper, we investigate recognizing context over time using physiological signals. Using the CASE dataset we evaluate both unimodal and multimodal approaches to physiological-based context recognition, over time. For recognition, we evaluate a random forest, as well as state-of-the-art neural network. These classifiers are evaluated using accuracy, Kappa, and F1-Macro metrics. Our results suggest that the fusion of EMG signals is more accurate, at recognizing context over time, compared to the fusion of non-EMG physiological signals. Although the fusion of non-EMG has a comparatively higher accuracy, ECG data results in the highest unimodal accuracy. Considering this, we analyze how the signals are correlated, including when the are fused (i.e. multimodal). We also perform a cross-gender analysis (e.g. training on male data and testing on female data) suggesting some generalizability across gender.
Saurabh Hinduja, Shaun J. Canavan
ACII1
2020 Real-time Action Unit Intensity Detection
abstract
We present a real-time system for action unit intensity detection. We train a convolutional neural network, with images from DISFA+, to detect the intensities of 12 action units that are commonly found in the literature. Along with real-time capabilities, the system is also able to detect intensities on static images, and in the wild videos. While the focus of this work is on the detection of action unit intensities, we are able to implicitly detect the occurrence of them as well. This is done by setting all action units with an intensity greater than 0 as active. In doing this, we calculate the F1-micro and F1-binary scores for the DISFA+ dataset.
Saurabh Hinduja, Shaun J. Canavan
FG1
2020 Multimodal Fusion of Physiological Signals and Facial Action Units for Pain Recognition
abstract
In this paper, we propose a method for pain recognition by fusing physiological signals (heart rate, respiration, blood pressure, and electrodermal activity) and facial action units. We provide experimental validation that the fusion of these signals results in a positive impact to the accuracy of pain recognition, compared to using only one modality (i.e. physiological or action units). These experiments are conducted on subjects from the BP4D+ multimodal emotion corpus, and include same- and cross-gender experiments. We also investigate the correlation between the two modalities to gain further insight into applications of pain recognition. Results suggest the need for larger and more varied datasets that include physiological signals and action units that have been coded for all facial frames.
Saurabh Hinduja, Shaun J. Canavan
FG1
2020 Recognizing Perceived Emotions from Facial Expressions
abstract
Expression recognition has seen an increase in research in past years, however, little work has been on recognizing perceived emotion (i.e. subject self-reporting of emotion). Considering this, we investigate the perceived emotion of subjects that perform tasks meant to elicit emotion. To facilitate this investigation, we use the BP4D+ multimodal spontaneous emotion corpus. We first statistically analyze the subject's perceived emotions across 10 tasks available in BP4D+. We show the percentage of subjects that felt specific emotions for each of the tasks. This is done across all tested subjects, as well as male and female subjects independently. Along with our statistical analysis, we also propose a 3D convolutional neural network (CNN) architecture to recognize multiple emotions felt for each task sequence. We report accuracy, Fl-binary and AUC for all subjects, as well as male and female subjects.
Saurabh Hinduja, Shaun J. Canavan, Lijun Yin 0001
FG1
2020 Three-level Training of Multi-Head Architecture for Pain Detection
abstract
Precise pain detection is a complex task even for trained professionals. There are many occasions where self-reporting fails to capture the decisive pain measurement. Facial expressions help extract the subtle emotions which can be leveraged to detect pain levels. In this paper, we present our approach to task 1 of the FG 2020 EmoPain Challenge: Pain-related Behavior Analysis, as well as our experimental design and results. We utilize a multi-head approach with the combined features of facial action units, facial landmarks, HOG and deep features. This multimodal approach provides insight into the contribution of each of these features and their consolidated effect. To improve our regression model, we adopt a three-level architecture where we observe an increase in prediction accuracy as the levels deepen. We record results comparable to the baseline on the challenge validation set.
Saandeep Aathreya, Saurabh Hinduja, Shaun J. Canavan
FG2
2020 Recognizing Emotion in the Wild using Multimodal Data
abstract
In this work, we present our approach for all four tracks of the eighth Emotion Recognition in the Wild Challenge (EmotiW 2020). The four tasks are group emotion recognition, driver gaze prediction, predicting engagement in the wild, and emotion recognition using physiological signals. We explore multiple approaches including classical machine learning tools such as random forests, state of the art deep neural networks, and multiple fusion and ensemble-based approaches. We also show that similar approaches can be used across tracks as many of the features generalize well to the different problems (e.g. facial features). We detail evaluation results that are either comparable to or outperform the baseline results for both the validation and testing for most of the tracks.
Shivam Srivastava, Saandeep Aathreya, Saurabh Hinduja, Sk Rahatul Jannat, Hamza Elhamdadi, Shaun J. Canavan
ICMI3
2020 Gaze-based classification of autism spectrum disorder
Diego Fabiano, Shaun J. Canavan, Heather Agazzi, Saurabh Hinduja, Dmitry B. Goldgof
Pattern Recognit. Lett.4
2019 Fusion of Hand-crafted and Deep Features for Empathy Prediction
abstract
We propose an approach to the OMG-Empathy Challenge for predicting self-annotated, continuous values of valence within the range [-1,1]. We propose the fusion of hand-crafted and deep features, extracted from both actor and listener data, to predict these valence levels. The handcrafted features include image level fusion, facial landmarks, and spectrogram features. Our proposed fusion approach can utilized in multiple parts (i.e. sub-modules), specifically utilized for the generalized track, leading to unique submissions to address the challenge problem. First, both actor and listener images are fused at the image-level. Secondly, facial landmarks from both the actor and listener are fused into one feature vector which is then used to train a random forest for prediction of continuous valence levels. Finally, we use a weighted fusion of the predicted values from both hand-crafted and deep features. We show competitive results on the OMG-Empathy challenge validation set.
Saurabh Hinduja, Md Taufeeq Uddin, Sk Rahatul Jannat, Astha Sharma, Shaun J. Canavan
FG1