Mohammed E. Hoque 0001

dblp:26/1236 · also Ehsan Hoque 0001, Mohammed (Ehsan) Hoque, Mohammed Ehsan Hoque · DBLP profile ↗
← Back
61ranked-venue papers
13as first author
18since 2021 · last 2026
0000-0003-4781-4733ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 35 · 9 first-author · 6 since 2021Artificial intelligence and machine learning · 26 · 6 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 PULSAR: Graph-Based Positive Unlabeled Learning with Multi-Stream Adaptive Convolutions for Parkinson's Disease Recognition
abstract
Timely diagnosis of movement disorders like Parkinson’s Disease (PD) improves quality of life. However, access to clinical diagnosis is limited in low-income countries. Here, we present PULSAR, a novel method for classifying individuals with or without PD from webcam-recorded videos of the finger-tapping task used in the Movement Disorder Society—Unified Parkinson’s Disease Rating Scale (MDS-UPDRS). PULSAR was trained and evaluated on data from 382 participants, including 183 self-reported PD patients. We used an adaptive graph convolutional neural network to dynamically learn task-specific spatio-temporal edges and enhanced it with a multi-stream convolution model to capture critical features like finger joint locations, tapping velocity, and acceleration of tapping. As video labels are self-reported, some non-PD labels may be undiagnosed cases. To address this, we used Positive Unlabeled (PU) Learning, which outperformed traditional supervised learning. PULSAR achieved 80.95% accuracy on the validation set and 71.29% mean accuracy (2.49% standard deviation) on an independent test set. We hope PULSAR can aid in accessible PD screening and that these techniques may extend to assessing disorders like ataxia and Huntington’s disease.
Md. Zarif Ul Alam, Asif Azad, Md. Saiful Islam 0013, Mohammed E. Hoque 0001, Mohammad Saifur Rahman 0001
ACM Trans. Comput. Heal.4
2026 Exploring the Role of Randomization on Belief Rigidity in Online Social Networks
abstract
People often stick to their existing beliefs, ignoring contradicting evidence or only interacting with those who reinforce their views. This tendency, referred to as being rigid in one's beliefs, is often driven by one's emotional connection to their beliefs and worsened by social media platforms which promote highly personalized content to maximize user engagement. As increased belief rigidity can negatively impact decisions, it is crucial to study and design interventions to reduce belief rigidity in online social networks. We design an experimental framework to passively infer belief rigidity and empirically quantify the effects of introducing randomness into the network structure as an intervention strategy. We recruit 163 participants into two conditions: one emulating traditional online platforms and the other more random. Our results suggest that individuals' beliefs are influenced by peer opinions, regardless of whether those opinions are similar to or differ from their own. Moreover, people are more likely to incorporate peers with (slightly) differing opinions into their networks when the recommendations are more randomized. Our findings highlight the need for future exploration into how design interventions can shape people's psycho-social functionalities in online social platforms.
Adiba Mahbub Proma, Neeley Pate, Raiyan Abdul Baten, Sifeng Chen, James Druckman, Gourab Ghoshal, Mohammed E. Hoque 0001
IEEE Trans. Affect. Comput.7
2025 Accessible, At-Home Detection of Parkinson's Disease via Multi-Task Video Analysis
abstract
Limited accessibility to neurological care leads to under-diagnosed Parkinson's Disease (PD), preventing early intervention. Existing AI-based PD detection methods primarily focus on unimodal analysis of motor or speech tasks, overlooking the multifaceted nature of the disease. To address this, we introduce a large-scale, multi-task video dataset of 1102 sessions (each containing videos of finger tapping, facial expression, and speech tasks captured via webcam) from 845 participants (272 with PD). We propose a novel Uncertainty-calibrated Fusion Network (UFNet) that leverages this multimodal data to enhance diagnostic accuracy. UFNet employs independent task-specific networks, trained with Monte Carlo Dropout for uncertainty quantification, followed by self-attended fusion of features, with attention weights dynamically adjusted based on task-specific uncertainties. We randomly split the participants into training (60%), validation (20%), and test (20%) sets to ensure patient-centered evaluation. UFNet significantly outperformed single-task models in terms of accuracy, area under the ROC curve (AUROC), and sensitivity while maintaining non-inferior specificity. Withholding uncertain predictions further boosted the performance, achieving 88.0 +- 0.3% accuracy, 93.0 +- 0.2% AUROC, 79.3 +- 0.9% sensitivity, and 92.6 +- 0.3% specificity, at the expense of not being able to predict for 2.3 +- 0.3% data (+- denotes 95% confidence interval). Further analysis suggests that the trained model does not exhibit any detectable bias across sex and ethnic subgroups and is most effective for individuals aged between 50 and 80. By merely requiring a webcam and microphone, our approach facilitates accessible home-based PD screening, especially in regions with limited healthcare resources.
Md. Saiful Islam 0013, T. M. Tariq Adnan, Jan Freyberg, Sangwu Lee, Abdelrahman Abdelkader, Meghan Pawlik, Catherine Schwartz, Karen Jaffe, Ruth B. Schneider, Earl Ray Dorsey, Mohammed E. Hoque 0001
AAAI11
2025 Personalizing LLM Responses to Combat Political Misinformation
abstract
Despite various efforts to tackle online misinformation, people inevitably encounter and engage with it, especially on social media platforms.Recent advances in LLMs present an opportunity to develop personalized interventions to address misinformed beliefs, and potentially offer more effective approaches than existing nontailored methods.In this paper, we design and evaluate personalized LLM agent that can consider users' demographics and personalities to tailor responses to mitigate misinformed beliefs.Our pipeline is grounded in facts through an external Retrieval Augmented Generation (RAG) knowledge base and is able to generate diverse output as a result of the personalization, with an average cosine similarity of 0.538.Our pipeline scores an average rating of 3.99 out of 5 when evaluated by a GPT-4o-mini LLM judge for response persuasiveness.Our methods can be adapted to design similar personalized agents in other domains.
Adiba Mahbub Proma, Neeley Pate, James Druckman, Gourab Ghoshal, Mohammed E. Hoque 0001
UMAP5
2025 'Poker with Play Money': Exploring Psychotherapist Training with Virtual Patients
abstract
Role-play exercises are widely utilized for training across a variety of domains; however, they have many shortcomings, including low availability, resource intensity, and lack of diversity. Large language model-driven virtual agents offer a potential avenue to mitigate these limitations and offer lower-risk role-play. The implications, however, of shifting this human-human collaboration to human-agent collaboration are still largely unexplored. In this work we focus on the context of psychotherapy, as psychotherapists-in-training extensively engage in role-play exercises with peers and/or supervisors to practice the interpersonal and therapeutic skills required for effective treatment. We provide a case study of a realistic ''virtual patient'' system for mental health training, evaluated by trained psychotherapists in comparison to their previous experiences with both real role-play partners and real patients. Our qualitative, reflexive analysis generated three themes and thirteen subthemes regarding key interpersonal skills of psychotherapy, the utility of the system compared to traditional role-play techniques, and factors which impacted psychotherapist-perceived ''humanness'' of the virtual patient. Although psychotherapists were optimistic about the system's potential to bolster therapeutic skills, this utility was impacted by the extent to which the virtual patient was perceived as human-like. We leverage the Computers Are Social Actors framework to discuss human-virtual-patient collaboration for practicing rapport, and discuss challenges of prototyping novel human-AI systems for clinical contexts which require a high degree of unpredictability. We pull from the ''SEEK'' three-factor theory of anthropomorphism to stress the importance of adequately representing a variety of cultural communities within mental health AI systems, in alignment with decolonial computing.
Cynthia M. Baseman, Masum Hasan, Nathaniel Swinger, Sheila A. M. Rauch, Mohammed E. Hoque 0001, Rosa I. Arriaga
Proc. ACM Hum. Comput. Interact.5
2024 Getting on the Right Foot: Using Observational and Quantitative Methods to Evaluate Movement Disorders
abstract
Currently doctors rely on tools such as the Unified Parkinson’s Disease Rating Scale (MDS-UDPRS) and the Scale for the Assessment and Rating of Ataxia (SARA) to make diagnoses for movement disorders based on clinical observations of a patient’s motor movement. Observation-based assessments however are inherently subjective and can differ by person. Moreover, different movement disorders show overlapping symptoms, challenging neurologists to make a correct diagnosis based on eyesight alone. In this work, we create an intelligent interface to highlight movements and gestures that are indicative of a movement disorder to observing doctors. First, we analyzed the walking patterns of 43 participants with Parkinson’s Disease (PD), 60 participants with ataxia, and 52 participants with no movement disorder to find ten metrics that can be used to distinguish PD from ataxia. Next, we designed an interface that provides context to the gestures that are relevant to a movement disorder diagnosis. Finally, we surveyed two neurologists (one who specializes in PD and the other who specializes in ataxia) on how useful this interface is for making a diagnosis. Our results not only showcase additional metrics that can be used to evaluate movement disorders quantitatively but also outline steps to be taken when designing an interface for these kinds of diagnostic tasks.
James Spann, Sarah A Chen, Tetsuo Ashizawa, Mohammed E. Hoque 0001
IUI4
2024 Managing Emotional Dialogue for a Virtual Cancer Patient: A Schema-Guided Approach
abstract
In this paper, we describe a general-purpose dialogue management framework used to design SOPHIE (Standardized Online Patient for Healthcare Interaction Education). SOPHIE simulates a virtual standardized cancer patient that allows physicians to practice skills such as empathy and patient empowerment in end-of-life communication. To provide the user with an opportunity to practice these skills, SOPHIE must produce a natural, emotionally appropriate conversation, yet handle topic shifts and open-ended questions from the user. To accomplish this, our approach to dialogue management loosely followsschemas– explicit representations of thetypicalflows of dialogue in end-of-life communication – while also using flexible pattern-driven methods for interpretation and generation. We conduct a crowdsourced evaluation of conversations between medical students and SOPHIE. Our agent is judged to produce responses that are natural, emotionally appropriate, and consistent with her role as a cancer patient. Furthermore, it significantly outperforms an end-to-end neural model fine-tuned on a human standardized patient corpus, attesting to the advantages of a schema-guided approach in this domain. However, the system is currently limited in its ability to generate responses that are judged to demonstrate deep understanding of the user, suggesting that future work should place focus on integrating this framework with robust natural language understanding and commonsense reasoning methods.
Benjamin Kane, Catherine Giugno, Lenhart K. Schubert, Kurtis Haut, Caleb Wohn, Mohammed E. Hoque 0001
IEEE Trans. Affect. Comput.6
2023 Validating a virtual human and automated feedback system for training doctor-patient communication skills
abstract
Effective communication between a clinician and their patient is critical for delivering healthcare maximizing outcomes. Unfortunately, traditional communication training approaches that use human standardized patients and expert coaches are difficult to scale. Here, we present the development and validation of a scalable, easily accessible, digital tool known as the Standardized Online Patient for Health Interaction Education (SOPHIE) for practicing and receiving feedback on doctor-patient communication skills. SOPHIE was validated by conducting an experiment with 30 medical professionals. We found that participants who underwent SOPHIE condition performed significantly better than the control in overall communication, aggregate scores, empowering the patient, and showing empathy (p < 0.05 in all cases). The results presented in this paper provide early evidence of scalable and accessible virtual humans, chatbot technologies, and automated feedback generation being transformative for medical communication training by supplementing existing resources.
Kurtis Haut, Caleb Wohn, Benjamin Kane, Thomas Carroll, Catherine Giugno, Ronald M. Epstein, Lenhart K. Schubert, Mohammed E. Hoque 0001
ACII9
2023 Novel Computational Linguistic Measures, Dialogue System and the Development of SOPHIE: Standardized Online Patient for Healthcare Interaction Education
abstract
In this article, we describe the iterative participatory design of SOPHIE, an online virtual patient for feedback-based practice of sensitive patient-physician conversations, and discuss an initial qualitative evaluation of the system by professional end users. The design of SOPHIE was motivated from a computational linguistic analysis of the transcripts of 383 patient-physician conversations from an essential office visit of late stage cancer patients with their oncologists. We developed methods for the automatic detection of two behavioral paradigms, lecturing and positive language usage patterns (sentiment trajectory of conversation), that are shown to be significantly associated with patient prognosis understanding. These automated metrics associated with effective communication were incorporated into SOPHIE, and a pilot user study identified that SOPHIE was favorably reviewed by a user group of practicing physicians.
Mohammad Rafayet Ali, Taylan K. Sen, Benjamin Kane, Shagun Bose, Thomas M. Carroll, Ronald M. Epstein, Lenhart K. Schubert, Mohammed E. Hoque 0001
IEEE Trans. Affect. Comput.8
2023 DBATES: Dataset for Discerning Benefits of Audio, Textual, and Facial Expression Features in Competitive Debate Speeches
abstract
In this article, we present a database of multimodal communication features extracted from debate speeches in the 2019 North American Universities Debate Championships (NAUDC). Feature sets were extracted from the visual (facial expression, gaze, and head pose), audio (PRAAT), and textual (word sentiment and linguistic category) modalities of raw video recordings of competitive collegiate debaters (N=716 6-minute recordings from 140 unique debaters). Each speech has an associated competition debate score (range: 67-96) from experienced judges as well as competitor demographic and per-round reflection surveys. We observe the fully multimodal model performs best in comparison to models trained on various compositions of individual modalities. We also find that the weights of some features (such as the expression ofjoyand the use of the word ”we”) change in direction between the aforementioned models. We use these results to highlight the value of a multimodal dataset for studying competitive, collegiate debate.
Taylan K. Sen, Gazi Naven, Luke Gerstner, Daryl Bagley, Raiyan Abdul Baten, Wasifur Rahman, Md. Kamrul Hasan 0003, Kurtis Haut, Abdullah Al Mamun 0002, Samiha Samrose, Anne Solbu, R. Eric Barnes, Mark G. Frank, Mohammed E. Hoque 0001
IEEE Trans. Affect. Comput.14
2022 Assistive Video Filters for People with Parkinson's Disease to Remove Tremors and Adjust Voice
abstract
COVID ushered in the widespread use of videoconferencing and it's here to stay. In virtual communication, we can alter everything from our appearance, voice and backgrounds. Most of these changes are fun gimmicks, but what if we could leverage these filtering technologies to a life-changing assistive technology? We propose the idea of developing assistive video filters for people with Parkinson's disease (PwP) that will remove involuntary tremors and smooth the stuttering in their voice. We surveyed 177 PwP and 107 people from the general public, and we personally interviewed 52 PwP as well as 3 health care professionals. We find overwhelming statistical evidence that these filters would fulfill a demonstrated communication need for PwP and that the general public also approves of a video filter that could assist with communication for PwP. To test the feasibility of our concept, we developed a filter prototype to remove physical tremors and tested it on two PwP. Although this paper focuses on PwP as a use case, we hope this work encourages others to ethically develop filtering technologies to help individuals with other movement disorders, eye-contact impairment and stuttering in computer-mediated conversations.
Kurtis Haut, Adira Blumenthal, Sarah Atterbury, Xiaofei Zhou 0004, Wasifur Rahman, Emanuela Natali, Mohammad Rafayet Ali, Mohammed E. Hoque 0001
ACII8
2022 Demographic Feature Isolation for Bias Research using Deepfakes
abstract
This paper explores the complexity of what constitutes the demographic features of race and how race is perceived. "Race" is composed of a variety of factors including skin tone, facial features, and accent. Isolating these interrelated race features is a difficult problem and failure to do so properly can easily invite confounding factors. Here we propose a novel method to isolate features of race by using AI-based technology and measure the impact these modifications have on an outcome variable of interest; i.e., perceived credibility. We used videos from a deception dataset for which the ground-truth is known and create three conditions: 1) a Black vs White CycleGAN image condition; 2) an original vs deepfake video condition; 3) an original vs deepfake still frame condition. We crowd-sourced 1736 responses to measure how credibility was influenced by changing the perceived race. We found that it is possible to alter perceived race through modifying demographically visual features alone. However, we did not find any statistically significant differences for credibility across our experiments based on these changes. Our findings help quantify intuitions from prior research that the relationship between racial perception and credibility is more complex than visual features alone. Our presented deepfake framework could be incorporated to precisely measure the impact of a wider range of demographic features (such as gender or age) due to the fine-grained isolation and control that was previously impossible in a lab setting.
Kurtis Haut, Caleb Wohn, Victor Antony, Aidan Goldfarb, Melissa Welsh, Dillanie Sumanthiran, Md. Rafayet Ali, Mohammed E. Hoque 0001
ACM Multimedia8
2022 MIA: Motivational Interviewing Agent for Improving Conversational Skills in Remote Group Discussions
abstract
Since online discussion platforms can limit the perception of social cues, effective collaboration over videochat requires additional attention to conversational skills. However, self-affirmation and defensive bias theories indicate that feedback may appear confrontational, especially when users are not motivated to incorporate them. We develop a feedback chatbot that employs Motivational Interviewing (MI), a directive counseling method that encourages commitment to behavior change, with the end goal of improving the user's conversational skills. We conduct a within-subject study with 21 participants in 8 teams to evaluate our MI-agent 'MIA' and a non-MI-agent 'Roboto'. After interacting with an agent, participants are tasked with conversing over videochat to evaluate candidate résumés for a job circular. Our quantitative evaluation shows that the MI-agent effectively motivates users, improves their conversational skills, and is likable. Through a qualitative lens, we present the strategies and the cautions needed to fulfill individual and team goals during group discussions. Our findings reveal the potential of the MI technique to improve collaboration and provide examples of conversational tactics important for optimal discussion outcomes.
Samiha Samrose, Mohammed E. Hoque 0001
Proc. ACM Hum. Comput. Interact.2
2022 Are You Really Looking at Me? A Feature-Extraction Framework for Estimating Interpersonal Eye Gaze From Conventional Video
abstract
Despite a revolution in the pervasiveness of video cameras in our daily lives, one of the most meaningful forms of nonverbal affective communication, interpersonal eye gaze, i.e., eye gaze relative to a conversation partner, is not available from common video. We introduce the Interpersonal-Calibrating Eye-gaze Encoder (ICE), which automatically extracts interpersonal gaze from video recordings without specialized hardware and without prior knowledge of participant locations. Leveraging the intuition that individuals spend a large portion of a conversation looking at each other enables the ICE dynamic clustering algorithm to extract interpersonal gaze. We validate ICE in both video chat using an objective metric with an infrared gaze tracker (F1 = 0.846, N = 8), as well as in face-to-face communication with expert-rated evaluations of eye contact (r = 0.37, N = 170). We then use ICE to analyze behavior in two different, yet important affective communication domains: interrogation-based deception detection, and communication skill assessment in speed dating. We find that honest witnesses break interpersonal gaze contact and look down more often than deceptive witnesses when answering questions (p = 0.004, d = 0.79). In predicting expert communication skill ratings in speed dating videos, we demonstrate that interpersonal gaze alone has more predictive power than facial expressions.
Taylan K. Sen, Kurtis Haut, Mohammad Rafayet Ali, Mohammed E. Hoque 0001
IEEE Trans. Affect. Comput.5
2022 Discourse Behavior of Older Adults Interacting with a Dialogue Agent Competent in Multiple Topics
abstract
We present a conversational agent designed to provide realistic conversational practice to older adults at risk of isolation or social anxiety, and show the results of a content analysis on a corpus of data collected from experiments with elderly patients interacting with our system. The conversational agent, represented by a virtual avatar, is designed to hold multiple sessions of casual conversation with older adults. Throughout each interaction, the system analyzes the prosodic and nonverbal behavior of users and provides feedback to the user in the form of periodic comments and suggestions on how to improve. Our avatar is unique in its ability to hold natural dialogues on a wide range of everyday topics—27 topics in three groups, developed in collaboration with a team of gerontologists. The three groups vary in “degrees of intimacy,” and as such in degrees of cognitive difficulty for the user. After collecting data from nine participants who interacted with the avatar for seven to nine sessions over a period of 3 to 4 weeks, we present results concerning dialogue behavior and inferred sentiment of the users. Analysis of the dialogues reveals correlations such as greater elaborateness for more difficult topics, increasing elaborateness with successive sessions, stronger sentiments in topics concerned with life goals rather than routine activities, and stronger self-disclosure for more intimate topics. In addition to their intrinsic interest, these results also reflect positively on the sophistication and practical applicability of our dialogue system.
Seyedeh Zahra Razavi, Lenhart K. Schubert, Kimberly Van Orden, Mohammad Rafayet Ali, Benjamin Kane, Mohammed E. Hoque 0001
ACM Trans. Interact. Intell. Syst.6
2021 Humor Knowledge Enriched Transformer for Understanding Multimodal Humor
abstract
Recognizing humor from a video utterance requires understanding the verbal and non-verbal components as well as incorporating the appropriate context and external knowledge. In this paper, we propose Humor Knowledge enriched Transformer (HKT) that can capture the gist of a multimodal humorous expression by integrating the preceding context and external knowledge. We incorporate humor centric external knowledge into the model by capturing the ambiguity and sentiment present in the language. We encode all the language, acoustic, vision, and humor centric features separately using Transformer based encoders, followed by a cross attention layer to exchange information among them. Our model achieves 77.36% and 79.41% accuracy in humorous punchline detection on UR-FUNNY and MUStaRD datasets -- achieving a new state-of-the-art on both datasets with the margin of 4.93% and 2.94% respectively. Furthermore, we demonstrate that our model can capture interpretable, humor-inducing patterns from all modalities.
Md. Kamrul Hasan 0003, Sangwu Lee, Wasifur Rahman, Amir Zadeh 0001, Rada Mihalcea, Louis-Philippe Morency, Mohammed E. Hoque 0001
AAAI7
2021 Hitting your MARQ: Multimodal ARgument Quality Assessment in Long Debate Video
abstract
The combination of gestures, intonations, and textual content plays a key role in argument delivery.However, the current literature mostly considers textual content while assessing the quality of an argument, and is limited to datasets containing short sequences (18-48 words).In this paper, we study argument quality assessment in a multimodal context, and experiment on DBATES, a publicly available dataset of long debate videos.First, we propose a set of interpretable debate-centric features such as clarity, content variation, body movement cues, and pauses, inspired by theories of argumentation quality.Second, we design the Multimodal ARgument Quality assessor (MARQ) -a hierarchical neural network model that summarizes the multimodal signals on long sequences and enriches the multimodal embedding with debate-centric features.Our proposed MARQ model achieves an accuracy of 81.91% on the argument quality prediction task and outperforms established baseline models with an error rate reduction of 22.7%.Through ablation studies, we demonstrate the importance of multimodal cues in modeling argument quality.
Md. Kamrul Hasan 0003, James Spann, Masum Hasan, Md. Saiful Islam 0013, Kurtis Haut, Rada Mihalcea, Mohammed E. Hoque 0001
EMNLP (1)7
2021 Analyzing Head Pose in Remotely Collected Videos of People with Parkinson's Disease
abstract
We developed an intelligent web interface that guides users to perform several Parkinson’s disease (PD) motion assessment tests in front of their webcam. After gathering data from 329 participants (N = 199 with PD, N = 130 without PD), we developed a methodology for measuring head motion randomness based on the frequency distribution of the motion. We found PD is associated with significantly higher randomness in side-to-side head motion as measured by the variance and number of large frequency components compared to the age-matched non-PD control group (p = 0.001, d = 0.13). Additionally, in participants taking levodopa (N = 151), the most common drug to treat Parkinson’s, the degree of random side-to-side head motion was found to follow an exponential-decay activity model following the time of the last dose taken (r = −0.404, p = 6e-5). A logistic regression model for classifying PD vs. non-PD groups identified that higher frequency components are more associated with PD. Our findings could potentially be useful toward objectively quantifying differences in head motions that may be due to either PD or PD medications.
Mohammad Rafayet Ali, Taylan K. Sen, Qianyi Li, Raina Langevin, Taylor Myers, Earl Ray Dorsey, Saloni Sharma, Mohammed E. Hoque 0001
ACM Trans. Comput. Heal.8
2020 Integrating Multimodal Information in Large Pretrained Transformers
abstract
Recent Transformer-based contextual word representations, including BERT and XLNet, have shown state-of-the-art performance in multiple disciplines within NLP. Fine-tuning the trained contextual models on task-specific datasets has been the key to achieving superior performance downstream. While fine-tuning these pre-trained models is straight-forward for lexical applications (applications with only language modality), it is not trivial for multimodal language (a growing area in NLP focused on modeling face-to-face communication). Pre-trained models don't have the necessary components to accept two extra modalities of vision and acoustic. In this paper, we proposed an attachment to BERT and XLNet called Multimodal Adaptation Gate (MAG). MAG allows BERT and XLNet to accept multimodal nonverbal data during fine-tuning. It does so by generating a shift to internal representation of BERT and XLNet; a shift that is conditioned on the visual and acoustic modalities. In our experiments, we study the commonly used CMU-MOSI and CMU-MOSEI datasets for multimodal sentiment analysis. Fine-tuning MAG-BERT and MAG-XLNet significantly boosts the sentiment analysis performance over previous baselines as well as language-only fine-tuning of BERT and XLNet. On the CMU-MOSI dataset, MAG-XLNet achieves human-level multimodal sentiment analysis performance for the first time in the NLP community.
Wasifur Rahman, Md. Kamrul Hasan 0003, Sangwu Lee, Amir Zadeh 0001, Chengfeng Mao, Louis-Philippe Morency, Mohammed E. Hoque 0001
ACL7
2020 Spatio-Temporal Attention and Magnification for Classification of Parkinson's Disease from Videos Collected via the Internet
abstract
We present an automated framework for detecting Parkinson's disease (PD) from videos collected through a scalable online platform. We analyzed 1380 videos of age-matched participants performing four standard motor tasks from the MDS-UPDRS. Our proposed framework leverages multiple deep neural networks to temporally and spatially segment the videos as well as magnify relevant motions. Frequency domain representations of the resulting data are then classified using supervised learning. Overall, the proposed framework achieves an accuracy of 82.5% when discriminating between those with PD and those without, and 61.8% when discriminating between those with PD with treatment, with PD without treatment, and those without PD. These results increased up to 91.8% and 73.5%, respectively, when combining the predictions of multiple models. To understand the contributions of each part of our framework we perform systematic ablation studies. We also compare between motion features based on pixel, phase-based and deep learning-based representations. This work demonstrates the possibility of identifying PD cues in challenging real-life settings with inexpensive webcams.
Mohammad Rafayet Ali, Javier Hernandez, Earl Ray Dorsey, Mohammed E. Hoque 0001, Daniel McDuff
FG4
2020 Leveraging Shared and Divergent Facial Expression Behavior Between Genders in Deception Detection
abstract
While facial expression behavior has been understood as mostly universal between the genders, recent research has highlighted important differences, including the expression of smiles and surprise. Despite such gender differences, studies involving facial expression often have limited sample sizes such that splitting the data set in half to train separate male and female models has been untenable. In order to leverage gender divergent complexity in facial expression models while also using a full dataset to train shared behaviors, we developed GAHL: the Gender-Augmented Hyper-Linear model. GAHL selectively increases non-linear model complexity with regards to gender divergent features. Using both simulated data and data from a study of facial expressions during deception (N=80, >6 hours), we demonstrate that when the facial expression data set size is in the range of N<; 7 5, GAHL outperforms several mainstream machine learning models including logistic regression, decision tree, and SVM with polynomial and radial basis function kernels.
Gazi Naven, Taylan K. Sen, Luke Gerstner, Kurtis Haut, Melissa Wen, Mohammed E. Hoque 0001
FG6
2020 A Virtual Conversational Agent for Teens with Autism Spectrum Disorder: Experimental Results and Design Lessons
abstract
We present the design of an online social skills development interface for teenagers with autism spectrum disorder (ASD). The interface is intended to enable private conversation practice anywhere, anytime using a web-browser. Users converse informally with a virtual agent, receiving feedback on nonverbal cues in realtime, and summary feedback. The prototype was developed in consultation with an expert UX designer, two psychologists, and a pediatrician. Using the data from 47 individuals, feedback and dialogue generation were automated using a hidden Markov model and a schema-driven dialogue manager capable of handling multi-topic conversations. We conducted a study with nine high-functioning ASD teenagers. Through a thematic analysis of post-experiment interviews, identified several key design considerations, notably:
Mohammad Rafayet Ali, Seyedeh Zahra Razavi, Raina Langevin, Abdullah Al Mamun 0002, Benjamin Kane, Reza Rawassizadeh, Lenhart K. Schubert, Mohammed E. Hoque 0001
IVA8
2019 What Computers Can Teach Us About Doctor-Patient Communication: Leveraging Gender Differences in Cancer Care
abstract
Advanced cancer patients sometimes spend their final days in unnecessary distress while receiving aggressive cancer treatment that is unlikely to work. Part of this problem stems from patients having incorrect understanding of their prognosis. Although studies have identified that effective doctor-patient communication is associated with better patient outcomes, most cancer patients misunderstand their prognosis. We applied computational language analysis tools (word category and language sentiment) to identify gender-specific communication characteristics associated with improved patient prognosis understanding. Analysis of 382 conversations between oncologists and patients identified that for female doctors, discussing feelings, using positive sentiment language, and speaking in shorter turns were strongly associated with better patient prognosis understanding. For male doctors, allowing patients to speak more, discussing the future, and not focusing heavily on religion or death were important. Synchrony between the doctors and patients usage of positive sentiment language was shown to be relevant only for female doctors.
Mohammad Rafayet Ali, Taylan K. Sen, Viet-Duy Nguyen, Mohammed E. Hoque 0001, Ronald M. Epstein, Reza Rawassizadeh, Paul Duberstein
ACII4
2019 Upskilling Together: How Peer-interaction Influences Speaking-skills Development Online
abstract
We explore the characteristics and values of online peer-interactions in developing a fundamental soft-skill such as speaking. 60 participants recorded speech videos on 5 job-interview prompts and exchanged comments and performance ratings with their peers. We find that both (i) receiving suggestions for improvement ('tips') and (ii) having access to peers with better average ratings than one's own correspond to performance improvement (p<; 0.001 for both). Using linguistic features (e.g., emotions, personality, sentiment), we are able to classify tips from non-tip comments (AUC 0.89). Linguistic features from the received comments and the average ratings of one's peers incrementally improve the prediction of one's future ratings, showing the simultaneous importance of the two peer-learning sources. Qualitative analysis reveals dyadic and community-level peer-influence factors: context-driven feedback, first-hand demonstration, empathetic support, acknowledgment, opinion diversity, sense of community and comfort in interaction. These insights inform the building of intelligent human-machine symbiosis systems for speaking-skills development.
Raiyan Abdul Baten, Famous Clark, Mohammed E. Hoque 0001
ACII3
2019 Facial Expression Based Imagination Index and a Transfer Learning Approach to Detect Deception
abstract
In this paper, we introduce a framework to automatically distinguish between facial expression sequences associated with imagining vs. remembering while answering a question. Our experiment includes a baseline and relevant questioning technique in the context of deception with 220 participants (20 hours long). Baseline questioning includes participants being separately asked to remember and imagine an arbitrary experience. During the relevant questioning, participants were prompted to either lie or tell the truth about a certain task. We trained a neural network model on the baseline data and achieved an accuracy of 60% on classifying imagining vs. remembering, whereas human performance for this task is 51%. Relevant questioning included a set of questions, each of which became an independent response segment. Using a transfer learning approach, we use the pretrained model from the baseline to obtain an imagination probability score for each relevant response segment. We define this individual probability per response as the Imagination Index. We apply the imagination indices as a feature vector to classify the whole relevant section as truth vs. bluff with an accuracy of 70%, significantly outperforming the human performance of 52%.
Md. Kamrul Hasan 0003, Wasifur Rahman, Luke Gerstner, Taylan K. Sen, Sangwu Lee, Kurtis Haut, Mohammed E. Hoque 0001
ACII7
2019 LIWC into the Eyes: Using Facial Features to Contextualize Linguistic Analysis in Multimodal Communication
abstract
This paper demonstrates that analyzing language patterns in light of their associated facial expressions elicits significant differences between deceptive and truthful communication. Facial Action Units (AU) were analyzed in video recordings (1.2M frames) of 151 dyadic conversations following an interrogation protocol, in which one of the participants is known to be either lying or telling the truth. Linguistic features were extracted from the transcripts using Linguistic Inquiry and Word Count (LIWC) dictionary. Our framework extracted facial-feature contexts automatically corresponding to high and low intensities of AU occurrences. This helped us dive deeper into answers corresponding to the video segments where the witnesses kept their eyes wide open (high intensity of AU05-upper lid raise). We found that in these segments, deceivers used significantly fewer `Seeing', `Perceptual' and `Cognitive' words and their answers were significantly shorter than truth-tellers.
Md. Kamrul Hasan 0003, Taylan K. Sen, Raiyan Abdul Baten, Kurtis Haut, Mohammed E. Hoque 0001
ACII6
2019 Visual Cues for Disrespectful Conversation Analysis
abstract
Toxic, abusive, or disrespectful behavior analysis is a non-trivial problem previously addressed mostly from the language perspective. In this paper, we present a novel video dataset containing disrespect and non-disrespect labels, and introduce such behavior analysis by using visual cues. The dataset is collected from YouTube news show videos of two-party conversations, in which a host and a guest interact through teleconferencing. Each video is labeled by three trained raters to identify disrespect expressed through face and gesture, voice, and language. By resolving confounding factors, we generate the corresponding pairwise samples of non-disrespect. To particularly show the influence of visual cues in disrespectful interactions, we present 222 labeled clips (duration=974.41(s), mean duration=4.39(s)). We extract and analyze the facial action units (AVs) prevalent in disrespectful behavior. Our result shows statistically significant differences after Bonferroni correction for Inner Brow raise (AV01), Lip Corner Depressor (AV15), and Chin Raiser (AV17). For prediction, we build two classifiers using logistic regression and linear Support Vector Machine with 62.61 % and 61.48 % accuracy, respectively. For an in-depth analysis of overall face and gesture features, we conduct a qualitative analysis using theme extraction. Our qualitative analysis provides further insights on leveraging synchronous and asynchronous features, along with combining text and audio data with visual cues to better detect disrespect.
Samiha Samrose, Wenyi Chu, Carolina He, Yuebai Gao, Syeda Sarah Shahrin, Mohammed E. Hoque 0001
ACII7
2019 UR-FUNNY: A Multimodal Language Dataset for Understanding Humor
abstract
Md Kamrul Hasan, Wasifur Rahman, AmirAli Bagher Zadeh, Jianyuan Zhong, Md Iftekhar Tanveer, Louis-Philippe Morency, Mohammed (Ehsan) Hoque. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Md. Kamrul Hasan 0003, Wasifur Rahman, Amir Zadeh 0001, Jianyuan Zhong, Md. Iftekhar Tanveer, Louis-Philippe Morency, Mohammed E. Hoque 0001
EMNLP/IJCNLP (1)7
2018 Understanding Social Interpersonal Interaction via Synchronization Templates of Facial Events
abstract
Automatic facial expression analysis in inter-personal communication is challenging. Not only because conversation partners' facial expressions mutually influence each other, but also because no correct interpretation of facial expressions is possible without taking social context into account. In this paper, we propose a probabilistic framework to model interactional synchronization between conversation partners based on their facial expressions. Interactional synchronization manifests temporal dynamics of conversation partners' mutual influence. In particular, the model allows us to discover a set of common and unique facial synchronization templates directly from natural interpersonal interaction without recourse to any predefined labeling schemes. The facial synchronization templates represent periodical facial event coordinations shared by multiple conversation pairs in a specific social context. We test our model on two different dyadic conversations of negotiation and job-interview. Based on the discovered facial event coordination, we are able to predict their conversation outcomes with higher accuracy than HMMs and GMMs.
Rui Li 0002, Jared Curhan, Mohammed E. Hoque 0001
AAAI3
2018 Awe the Audience: How the Narrative Trajectories Affect Audience Perception in Public Speaking
abstract
Telling a great story often involves a deliberate alteration of emotions. In this paper, we objectively measure and analyze the narrative trajectories of stories in public speaking and their impact on subjective ratings. We conduct the analysis using the transcripts of over 2000 TED talks and estimate potential audience response using over 5 million spontaneous annotations from the viewers. We use IBM Watson Tone Analyzer to extract sentence-wise emotion, language, and social scores. Our study indicates that it is possible to predict (with AUC as high as 0.88) the subjective ratings of the audience by analyzing the narrative trajectories. Additionally, we find that some trajectories (for example, a flat trajectory of joy) correlate well with some specific ratings (e.g. "Longwinded') assigned by the viewers. Such an association could be useful in forecasting audience responses using objective analysis.
Md. Iftekhar Tanveer, Samiha Samrose, Raiyan Abdul Baten, Mohammed E. Hoque 0001
CHI4
2018 The What, When, and Why of Facial Expressions: An Objective Analysis of Conversational Skills in Speed-Dating Videos
abstract
In this paper, we demonstrate the importance of combinations of facial expressions and their timing, in explaining a person's conversational skills in a series of brief non-romantic conversations. Video recordings of 365 four-minute conversations before and after a randomized intervention were analyzed in which facial action units (AUs) were examined over different time segments. Male subjects (N=47) were evaluated in their conversation skills using the Conversational Skills Rating Scale (CSRS). A linear regression model was used to compare the importance of AU features from different time segments in predicting CSRS ratings. In the first minute of conversation, CSRS ratings were best predicted by activity levels in action units associated with speaking (Lips part, AU25). In the last minute of conversation, affective indicators associated with expressions of laughter (Jaw Drop, AU26) and warmth (Happy faces) emerged as the most important. These findings suggest that feedback on nonverbal skills must dynamically account for shifting goals of conversation.
Mohammad Rafayet Ali, Taylan K. Sen, Dev Crasta, Viet-Duy Nguyen, Ronald D. Rogge, Mohammed E. Hoque 0001
FG6
2018 Say CHEESE: Common Human Emotional Expression Set Encoder and Its Application to Analyze Deceptive Communication
abstract
In this paper we introduce the Common Human Emotional Expression Set Encoder (CHEESE) framework for objectively determining which, if any, subsets of the facial action units associated with smiling are well represented by a small finite set of clusters according to an information theoretic metric. Smile-related AUs (6,7,10,12,14) in over 1.3M frames of facial expressions from 151 pairs of individuals playing a communication game involving deception were analyzed with CHEESE. The combination of AU6 (cheek raiser) and AU12 (lip corner puller) are shown to cluster well into five different types of expression. Liars showed high intensity AU6 and AU12 more often compared to honest speakers. Additionally, interrogators were found to express a higher frequency of low intensity AU6 with high intensity AU12 (i.e. polite smiles) when they were being lied to, suggesting that deception analysis should be done in consideration of both the message sender's and the receiver's facial expressions.
Taylan K. Sen, Md. Kamrul Hasan 0003, Mohammed E. Hoque 0001
FG5
2018 Analyzing the Impact of Gender on the Automation of Feedback for Public Speaking
abstract
This paper explores gender differences in the evaluation of male and female speakers' affective features in public speaking. We analyzed 260 two-minute behavioral videos (200 of females and 60 of males), collected from an online public speaking practice tool. We adopted a linear regression model that utilized facial and prosodic features, including facial action units (AU), word count, pitch, and volume, to automatically assess speaker performance. The model was evaluated against ratings from 2 expert speakers from Toastmasters, an international public speaking club, on speaker performance. Our feature analysis suggests that certain combinations of features are correlated with higher ratings only in males, such as the combined increase of speech rate and vocal pitch variation. Moreover, our clustering analysis suggests that exhibiting certain negative emotions correlates with higher ratings for males but not for females, illustrating the impact of gender in generating effective feedback on public speaking.
Astha Singhal, Mohammad Rafayet Ali, Raiyan Abdul Baten, Chigusa Kurumada, Elizabeth West Marvin, Mohammed E. Hoque 0001
FG6
2018 Aging and Engaging: A Social Conversational Skills Training Program for Older Adults
abstract
We developed 'Aging and Engaging, a web-based intelligent interface, to improve communication skills among older adults. The interface allows users to practice conversations with a virtual assistant and receive feedback on eye contact, speaking volume, smiling, and valence of speech content. Feedback is generated automatically by analyzing the temporal properties of the conversation using the hidden Markov model. The interface was designed with the assistance of an expert advisory panel that works with geriatric patients, as well as a focus group of 12 older adults. To evaluate its effectiveness, we conducted a study with 25 older adults, each of whom participated in four conversations. Participants' response times to questions, as well as the amount of positive feedback, increased gradually through these interactions, as assessed by human judges. Participants found the feedback useful, easy to interpret, and fairly accurate, and expressed their interest in using the system at home. We plan to enroll subjects with difficulties in social communication; have them use the system over time at home in a randomized, controlled study; and measure any changes in their behavior.
Mohammad Rafayet Ali, Kimberly Van Orden, Kimberly Parkhurst, Viet-Duy Nguyen, Paul Duberstein, Mohammed E. Hoque 0001
IUI7
2018 Automated Analysis and Prediction of Job Interview Performance
abstract
We present a computational framework for automatically quantifying verbal and nonverbal behaviors in the context of job interviews. The proposed framework is trained by analyzing the videos of 138 interview sessions with 69 internship-seeking undergraduates at the Massachusetts Institute of Technology (MIT). Our automated analysis includes facial expressions (e.g., smiles, head gestures, facial tracking points), language (e.g., word counts, topic modeling), and prosodic information (e.g., pitch, intonation, and pauses) of the interviewees. The ground truth labels are derived by taking a weighted average over the ratings of nine independent judges. Our framework can automatically predict the ratings for interview traits such as excitement, friendliness, and engagement with correlation coefficients of 0.70 or higher, and can quantify the relative importance of prosody, language, and facial expressions. By analyzing the relative feature weights learned by the regression models, our framework recommends to speak more fluently, use fewer filler words, speak as “we” (versus “I”), use more unique words, and smile more. We also find that the students who were rated highly while answering the first interview question were also rated highly overall (i.e., first impression matters). Finally, our MIT Interview dataset is available to other researchers to further validate and expand our findings.
Iftekhar Naim, Md. Iftekhar Tanveer, Daniel Gildea, Mohammed E. Hoque 0001
IEEE Trans. Affect. Comput.4
2017 Automated video interview judgment on a large-sized corpus collected online
abstract
Online video-based job interviews are becoming very popular in the screening of potential employees. In this study, we collected a corpus of 1891 monologue job interview videos (63 hours in duration) from 260 online workers. These videos were annotated for personality traits and hiring recommendation score by experts from a major assessment company. We proposed a unified method of automatic analysis that consists of using clustering to convert continuous audio/video analysis output to discrete pseudoword documents, and then applying modern text classification methods to process speech content, prosody and facial expressions. Our experiments showed that using what the interviewees say (i.e., spoken text), we can predict their personality traits such as openness, conscientiousness, extraversion, agreeableness, and emotional stability with an F-measure of 0.8 or better, while we get an F-measure of 0.6 in predicting hiring recommendation score. Prosody and facial expressions added limited usefulness on interview judgments and need further investigation.
Lei Chen 0004, Chee Wee Leong, Blair Lehman, Gary Feng, Mohammed E. Hoque 0001
ACII6
2017 Modeling doctor-patient communication with affective text analysis
abstract
We present a method of automatic analysis of doctor-patient communication and present findings after applying this methodology in a post hoc study of communication between oncologists and their cancer patients (N=122). We analyzed several features of each participant in the conversation including the number of words spoken, the average positive/negative sentiment expressed, the number of questions asked, and the word diversity (unique word count). We found that the number of words spoken by the doctor is correlated with the highest doctor communication ability ratings made by patients. We additionally found that unsupervised clustering of conversation features into “styles” identified that certain styles are associated with higher communication ratings. Two well-defined styles emerged when clustering based on doctor word diversity and doctor sentiment: a high word diversity-neutral sentiment style, which was associated with higher ratings, and a low word diversity-positive sentiment style with lower average ratings. Machine learning models were trained to automatically predict whether a doctor-patient interaction will be rated high or not with a best-performing 71% test set accuracy.
Taylan K. Sen, Mohammad Rafayet Ali, Mohammed E. Hoque 0001, Ronald M. Epstein, Paul Duberstein
ACII3
2016 ROC comment: automated descriptive and subjective captioning of behavioral videos
abstract
We present an automated interface, ROC Comment, for generating natural language comments on behavioral videos. We focus on the domain of public speaking, which many people consider their greatest fear. We collect a dataset of 196 public speaking videos from 49 individuals and gather 12,173 comments, generated by more than 500 independent human judges. We then train a k-Nearest-Neighbor (k-NN) based model by extracting prosodic (e.g., volume) and facial (e.g., smiles) features. Given a new video, we extract features and select the closest comments using k-NN model. We further filter the comments by clustering them using DBScan, and eliminating the outliers. Evaluation of our system with 30 participants conclude that while the generated comments are helpful, there is room for improvement in further personalizing them. Our model has been deployed online, allowing individuals to upload their videos and receive open-ended and interpretative comments. Our system is available at http://tinyurl.com/roccomment.
Mohammad Rafayet Ali, Facundo Ciancio, Iftekhar Naim, Mohammed E. Hoque 0001
UbiComp5
2016 AutoManner: An Automated Interface for Making Public Speakers Aware of Their Mannerisms
abstract
Many individuals exhibit unconscious body movements called mannerisms while speaking. These repeated changes often distract the audience when not relevant to the verbal context. We present an intelligent interface that can automatically extract human gestures using Microsoft Kinect to make speakers aware of their mannerisms. We use a sparsity-based algorithm, Shift Invariant Sparse Coding, to automatically extract the patterns of body movements. These patterns are displayed in an interface with subtle question and answer-based feedback scheme that draws attention to the speaker's body language. Our formal evaluation with 27 participants shows that the users became aware of their body language after using the system. In addition, when independent observers annotated the accuracy of the algorithm for every extracted pattern, we find that the patterns extracted by our algorithm is significantly (p<0.001) more accurate than just random selection. This represents a strong evidence that the algorithm is able to extract human-interpretable body movement patterns. An interactive demo of AutoManner is available at http://tinyurl.com/AutoManner.
Md. Iftekhar Tanveer, Kezhen Chen, Zoe Tiet, Mohammed E. Hoque 0001
IUI5
2016 The LISSA Virtual Human and ASD Teens: An Overview of Initial Experiments
Seyedeh Zahra Razavi, Mohammad Rafayet Ali, Tristram H. Smith, Lenhart K. Schubert, Mohammed E. Hoque 0001
IVA5
2015 LISSA - Live Interactive Social Skill Assistance
abstract
We present LISSA — Live Interactive Social Skill Assistance — a web-based system that helps people practice their conversational skills by having short conversations with a human like virtual agent and receiving real-time feedback on their nonverbal behavior. In this paper, we describe the development of an interface for these features and examine the viability of real time feedback using a Wizard of Oz prototype. We then evaluated our system using a speed-dating study design. We invited 47 undergraduate male students to interact with staff and randomly assigned them to intervention with LISSA or a self-help control group. Results suggested participants who practiced with the LISSA system were rated as significantly better in nodding when compared to a self-help control group and marginally better in eye contact and gesturing. The system usability and surveys showed that participants found the feedback provided by the system useful, unobtrusive, and easy to understand.
Mohammad Rafayet Ali, Dev Crasta, Agustin Baretto, Joshua Pachter, Ronald D. Rogge, Mohammed E. Hoque 0001
ACII7
2015 Exploring temporal patterns in classifying frustrated and delighted smiles (Extended abstract)
abstract
We created two experimental situations to elicit two affective states: frustration and delight. In the first experiment, participants were asked to recall situations while expressing either frustration or delight. The second experiment tried to elicit these states naturally with a frustrating experience and a delightful video. There were two significant differences between the acted and natural occurrences of the expressions. First, the acted instances were much easier for the computer to classify. Second, in 90 percent of the acted cases, participants did not smile when frustrated. In 90 percent of the natural cases, participants smiled during the frustrating interaction, despite self-reporting significant frustration with the experience. As a follow up study, we develop an automated system to distinguish between naturally occurring spontaneous smiles under frustrating and delightful stimuli by exploring their temporal patterns, given video of both. We extracted local and global features related to human smile dynamics. Next, we evaluated and compared two variants of Support Vector Machines (SVM), Hidden Markov Models (HMM), and Hidden-state Conditional Random Fields (HCRF) for binary classification. While human classification of the smile videos under frustrating stimuli was below chance, a dynamic SVM classifier obtained an accuracy of 92 percent in distinguishing smiles under frustrating and delighted stimuli.
Mohammed E. Hoque 0001, Daniel McDuff, Rosalind W. Picard
ACII1
2015 ROC speak: semi-automated personalized feedback on nonverbal behavior from recorded videos
abstract
We present a framework that couples computer algorithms with human intelligence in order to automatically sense and interpret nonverbal behavior. The framework is cloud-enabled and ubiquitously available via a web browser, and has been validated in the context of public speaking. The system automatically captures audio and video data in-browser through the user's webcam, and then analyzes the data for smiles, movement, and volume modulation. Our framework allows users to opt in and receive subjective feedback from Mechanical Turk workers ("Turkers"). Our system synthesizes the Turkers' interpretations, ratings, and comment rankings with the machine-sensed data and enables users to interact with, explore, and visualize personalized and presentational feedback. Our results provide quantitative and qualitative evidence in support of our proposed synthesized feedback, relative to video-only playback with impersonal tips. Our interface can be seen here: http://tinyurl.com/feedback-ui (Supported in Google Chrome.)
Michelle Fung, Yina Jin, RuJie Zhao, Mohammed E. Hoque 0001
UbiComp4
2015 Vowel shapes: an open-source, interactive tool to assist singers with learning vowels
abstract
The mastery of vowel production is central to developing vocal technique and may be influenced by language, musical context or a coach's direction. Currently, students learn through verbal descriptions, demonstration of correct vowel sounds, and customized exercises. Vowel Shapes is an interactive practice tool that automatically captures and visualizes vowel sounds in real time to assist singers in correctly producing target vowels. The system may be used during a lesson or as a practice tool when an instructor is not present. Our system's design was informed by iterative evaluations with 14 students and their vocal professor from the Eastman School of Music, University of Rochester. Results from an exploratory evaluation of the system with 10 students indicated that 70% of the participants improved their time to reach an instructor-defined target. 90% of the students in the evaluation would use this system during practice sessions.
Cynthia Ryan, Katherine Ciesinski, Mohammed E. Hoque 0001
UbiComp3
2015 Rhema: A Real-Time In-Situ Intelligent Interface to Help People with Public Speaking
abstract
A large number of people rate public speaking as their top fear. What if these individuals were given an intelligent interface that provides live feedback on their speaking skills? In this paper, we present Rhema, an intelligent user interface for Google Glass to help people with public speaking. The interface automatically detects the speaker's volume and speaking rate in real time and provides feedback during the actual delivery of speech. While designing the interface, we experimented with two different strategies of information delivery: 1) Continuous streams of information, and 2) Sparse delivery of recommendation. We evaluated our interface with 30 native English speakers. Each participant presented three speeches (avg. duration 3 minutes) with 2 different feedback strategies (continuous, sparse) and a baseline (no feeback) in a random order. The participants were significantly more pleased (p < 0.05) with their speech while using the sparse feedback strategy over the continuous one and no feedback.
Md. Iftekhar Tanveer, Emy Lin, Mohammed E. Hoque 0001
IUI3
2015 Unsupervised Extraction of Human-Interpretable Nonverbal Behavioral Cues in a Public Speaking Scenario
abstract
We present a framework for unsupervised detection of nonverbal behavioral cues---hand gestures, pose, body movements, etc.---from a collection of motion capture (MoCap) sequences in a public speaking setting. We extract the cues by solving a sparse and shift-invariant dictionary learning problem, known as shift-invariant sparse coding. We find that the extracted behavioral cues are human-interpretable in the context of public speaking. Our technique can be applied to automatically identify the common patterns of body movements and the time-instances of their occurrences, minimizing time and efforts needed for manual detection and coding of nonverbal human behaviors.
Md. Iftekhar Tanveer, Ji Liu 0002, Mohammed E. Hoque 0001
ACM Multimedia3
2014 M.I.D.A.S. touch: magnetic interactive device for alternative sight through touch
abstract
This paper describes the development of a working prototype of a novel free-space haptic human-computer interface called MIDAS Touch. MIDAS Touch works by applying physical forces to a user's finger through the production of a dynamic magnetic field. The magnetic field strength is adjusted in real-time based on the user's movement of his/her finger. A user's hand/finger motion in the real world is mapped to movement of a virtual finger in a virtual world through the use of a Leap Motion 3D tracking sensor. As a person's virtual finger collides with objects in the virtual world, the magnetic field strength is varied. In this demo, we present a case of MIDAS Touch coupled to a standard PC as a computer drawing viewer and drawing application for helping individuals with visual impairment feel what they or others have drawn.
Taylan K. Sen, Morgan W. Sinko, Alex T. Wilson, Mohammed E. Hoque 0001
ASSETS4
2014 A google glass app to help the blind in small talk
abstract
In this paper we present a wearable prototype that can automatically recognize affective cues such as number of people present, their age and gender distributions given an image. We customize the prototype in the context of helping people with visual impairments to better navigate social scenarios. Running an experiment to validate this technology in social scenarios remains part of our future work.
Md. Iftekhar Tanveer, Mohammed E. Hoque 0001
ASSETS2
2013 Automated Coach to Practice Conversations
abstract
We present a real-time system including a 3D character that can converse, capture, analyze and interpret subtle and multidimensional human nonverbal behaviors for possible applications such as job interviews, public speaking, or even automated speech therapy. The system works in a personal computer and senses nonverbal data from video (i.e., facial expressions) and audio (i.e., speech recognition and prosody analysis) using a standard web cam. We contextualized the development and evaluation of our system as a training scenario for job interviews. Using user-centered design and iterations, we determine how the nonverbal data could be presented to the user in an intuitive and educational manner. We tested efficacy of the system in context of job interviews with 90 MIT undergraduate students. Our results suggest that the participants who used our system to improve their interview skills were perceived to be better candidates by human judges. Participants reported that the most useful feature was being given feedback on their speaking rate, and overall they reported strong agreement that would consider using this system again for self-reflection.
Mohammed E. Hoque 0001, Rosalind W. Picard
ACII1
2013 MACH: my automated conversation coach
abstract
MACH--My Automated Conversation coacH--is a novel system that provides ubiquitous access to social skills training. The system includes a virtual agent that reads facial expressions, speech, and prosody and responds with verbal and nonverbal behaviors in real time. This paper presents an application of MACH in the context of training for job interviews. During the training, MACH asks interview questions, automatically mimics certain behavior issued by the user, and exhibit appropriate nonverbal behaviors. Following the interaction, MACH provides visual feedback on the user's performance. The development of this application draws on data from 28 interview sessions, involving employment-seeking students and career counselors. The effectiveness of MACH was assessed through a weeklong trial with 90 MIT undergraduates. Students who interacted with MACH were rated by human experts to have improved in overall interview performance, while the ratings of students in control groups did not improve. Post-experiment interviews indicate that participants found the interview experience informative about their behaviors and expressed interest in using MACH in the future.
Mohammed E. Hoque 0001, Matthieu Courgeon, Jean-Claude Martin, Bilge Mutlu, Rosalind W. Picard
UbiComp1
2012 Mood meter: counting smiles in the wild
abstract
In this study, we created and evaluated a computer vision based system that automatically encouraged, recognized and counted smiles on a college campus. During a ten-week installation, passersby were able to interact with the system at four public locations. The aggregated data was displayed in real time in various intuitive and interactive formats on a public website. We found privacy to be one of the main design constraints, and transparency to be the best strategy to gain participants' acceptance. In a survey (with 300 responses), participants reported that the system made them smile more than they expected, and it made them and others around them feel momentarily better. Quantitative analysis of the interactions revealed periodic patterns (e.g., more smiles during the weekends) and strong correlation with campus events (e.g., fewer smiles during exams, most smiles the day after graduation), reflecting the emotional responses of a large community.
Javier Hernandez, Mohammed E. Hoque 0001, Will Drevo, Rosalind W. Picard
UbiComp2
2012 My automated conversation helper (MACH): helping people improve social skills
abstract
Ever been in a situation where you didn't get that job despite being the deserving candidate? What went wrong? Psychology literature suggests that the most important skill towards making an impression during interviews is your interpersonal/social skills. Is it possible for people to improve their social skills (e.g., vary the voice intonation, and pauses appropriately; use social smiles, when appropriate; and maintain eye contact) through a computerized intervention? In this thesis, I propose to develop an autonomous and Automated Conversation Helper (3D virtual character) that can play the role of the interviewer, allowing participants to practice their social skills in context of job interviews. The Automated Conversation Helper is being developed with the ability to "see" (facial expression processing), "hear" (speech recognition and prosody analysis) and "respond" (speech and behavior synthesis) in real-time and provide live feedback on participant's non-verbal behavior.
Mohammed E. Hoque 0001
ICMI1
2012 Exploring Temporal Patterns in Classifying Frustrated and Delighted Smiles
abstract
We create two experimental situations to elicit two affective states: frustration, and delight. In the first experiment, participants were asked to recall situations while expressing either delight or frustration, while the second experiment tried to elicit these states naturally through a frustrating experience and through a delightful video. There were two significant differences in the nature of the acted versus natural occurrences of expressions. First, the acted instances were much easier for the computer to classify. Second, in 90 percent of the acted cases, participants did not smile when frustrated, whereas in 90 percent of the natural cases, participants smiled during the frustrating interaction, despite self-reporting significant frustration with the experience. As a follow up study, we develop an automated system to distinguish between naturally occurring spontaneous smiles under frustrating and delightful stimuli by exploring their temporal patterns given video of both. We extracted local and global features related to human smile dynamics. Next, we evaluated and compared two variants of Support Vector Machine (SVM), Hidden Markov Models (HMM), and Hidden-state Conditional Random Fields (HCRF) for binary classification. While human classification of the smile videos under frustrating stimuli was below chance, an accuracy of 92 percent distinguishing smiles under frustrating and delighted stimuli was obtained using a dynamic SVM classifier.
Mohammed E. Hoque 0001, Daniel McDuff, Rosalind W. Picard
IEEE Trans. Affect. Comput.1
2011 Machine Learning for Affective Computing
Mohammed E. Hoque 0001, Daniel McDuff, Louis-Philippe Morency, Rosalind W. Picard
ACII (2)1
2011 Are You Friendly or Just Polite? - Analysis of Smiles in Spontaneous Face-to-Face Interactions
Mohammed E. Hoque 0001, Louis-Philippe Morency, Rosalind W. Picard
ACII (1)1
2011 Acted vs. natural frustration and delight: Many people smile in natural frustration
abstract
This work is part of research to build a system to combine facial and prosodic information to recognize commonly occurring user states such as delight and frustration. We create two experimental situations to elicit two emotional states: the first involves recalling situations while expressing either delight or frustration; the second experiment tries to elicit these states directly through a frustrating experience and through a delightful video. We find two significant differences in the nature of the acted vs. natural occurrences of expressions. First, the acted ones are much easier for the computer to recognize. Second, in 90% of the acted cases, participants did not smile when frustrated, whereas in 90% of the natural cases, participants smiled during the frustrating interaction, despite self-reporting significant frustration with the experience. This paper begins to explore the differences in the patterns of smiling that are seen under natural frustration and delight conditions, to see if there might be something measurably different about the smiles in these two cases, which could ultimately improve the performance of classifiers applied to natural expressions.
Mohammed E. Hoque 0001, Rosalind W. Picard
FG1
2009 Exploring speech therapy games with children on the autism spectrum
abstract
Individuals on the autism spectrum often have difficulties producing intelligible speech with either high or low speech rate, and atypical pitch and/or amplitude affect.In this study, we present a novel intervention towards customizing speech enabled games to help them produce intelligible speech.In this approach, we clinically and computationally identify the areas of speech production difficulties of our participants.We provide an interactive and customized interface for the participants to meaningfully manipulate the prosodic aspects of their speech.Over the course of 12 months, we have conducted several pilots to set up the experimental design, developed a suite of games and audio processing algorithms for prosodic analysis of speech.Preliminary results demonstrate our intervention being engaging and effective for our participants.
Mohammed E. Hoque 0001, Joseph K. Lane, Rana El Kaliouby, Matthew S. Goodwin, Rosalind W. Picard
INTERSPEECH1
2009 When Human Coders (and Machines) Disagree on the Meaning of Facial Affect in Spontaneous Videos
Mohammed E. Hoque 0001, Rana El Kaliouby, Rosalind W. Picard
IVA1
2008 Analysis of speech properties of neurotypicals and individuals diagnosed with autism and down
abstract
Many individuals diagnosed with autism and Down syndrome have difficulties producing intelligible speech. Systematic analysis of their voice parameters could lead to better understanding of the specific challenges they face in achieving proper speech production. In this study, 100 minutes of speech data from natural conversations between neurotypicals and individuals diagnosed with autism/Down-syndrome was used. Analyzing their voice parameters indicated new findings across a variety of speech parameters. An immediate extension of this work would be to customize this technology allowing participants to visualize and control their speech parameters in real time and get live feedback.
Mohammed E. Hoque 0001
ASSETS1
2007 What Speech Tells Us About Discourse: The Role of Prosodic and Discourse Features in Speech Act Classification
abstract
This paper explores the relative importance of discourse features, prosodic features and their fusion in robust classification of speech acts. Five different feature selection algorithms were used to select set of features to improve the robustness of the classification. The results showed that the ensemble-based classifiers performed best in the classification of 12 speech acts using subsets of both prosodic and discourse features.
Mohammed E. Hoque 0001, Mohammad S. Sorower, Mohammed Yeasin, Max M. Louwerse
IJCNN1
2006 Robust Recognition of Emotion from Speech
Mohammed E. Hoque 0001, Mohammed Yeasin, Max M. Louwerse
IVA1