EDBT 2026 Demo / reviewers in the wild / expert
Theodora Chaspari
dblp:59/10649
· DBLP profile ↗
54ranked-venue papers
10as first author
26since 2021 · last 2026
0000-0002-7603-8633ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 23 · 8 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 19 · 2 first-author · 11 since 2021Artificial intelligence and machine learning · 14 · 3 first-author · 6 since 2021Computer networks · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | User judgment of an AI model is biased by its description: A study in a job interview training context
Sharon Lynn Chu Yew Yee, Marcin Karcz, Amal Hashky, Neha Rani, Theodora Chaspari, Winfred Arthur Jr., Eric D. Ragan |
Int. J. Hum. Comput. Stud. | 5 |
| 2025 | Leveraging Pre-Trained Transformers and Facial Embeddings for Multimodal Hirability Prediction in Job Interviews
Eric Fithian, Theodora Chaspari |
ICMI | 2 |
| 2025 | Investigating the Reasoning Abilities of Large Language Models for Understanding Spoken Language in Interpersonal Interactions
Pranjal Aggarwal, Ghritachi Mahajani, Pavan Kumar Malasani, Vaibhav Jamadagni, Caroline J. Wendt, Ehsanul Haque Nirjhar, Theodora Chaspari |
INTERSPEECH | 7 |
| 2025 | Assessing the feasibility of Large Language Models for detecting micro-behaviors in team interactions during space missions
Ankush Raut, Projna Paromita, Sydney R. Begerowski, Suzanne T. Bell, Theodora Chaspari |
INTERSPEECH | 5 |
| 2025 | Modeling Gold Standard Moment-to-Moment Ratings of Perception of Stress From Audio RecordingsabstractEnabling continuous and unobtrusive monitoring of stress is essential for delivering personalized stress interventions at opportune moments. To achieve automatic stress detection on a time-continuous basis, reliable moment-to-moment ratings of stress are required. However, the current research lacks a large-scale multimodal dataset that provides time-continuous ratings of perceived stress. Existing datasets mainly consist of single-valued self-reported ratings obtained after the stress-inducing task or rely on audio-visual recordings to capture moment-to-moment ratings from multiple annotators. The collection of time-continuous ratings of stress based solely on audio recordings has not been extensively explored. In this paper, we introduce an updated version of the publicly available VerBIO dataset that contains moment-to-moment ratings of perceived stress from multiple annotators. These annotators rated their perception of stress by listening to participants who had conducted a public speaking task. Time-continuous ratings of stress are obtained from four annotators using 22 hours of audio recordings from 339 public speaking sessions performed by 53 individuals. These time-continuous ratings of stress perception were obtained from the annotators solely based on speech, without incorporating the visual modality as an expressive marker. We examine the reliability of the annotation scheme employed in this study and investigate the factors contributing to the observed variation in perceived stress among annotators. Next, we introduce an annotation fusion technique based on expectation-maximization to obtain a reliable gold standard rating by aggregating the ratings from multiple annotators. Results indicate that the proposed annotation fusion technique yields aggregated ratings that can be estimated more reliably using acoustic features compared to the ratings yielded from conventional annotation fusion techniques. The newly generated annotations are publicly available within the proposed updated version of the existing VerBIO dataset, facilitating research in the field of continuous stress detection. Ehsanul Haque Nirjhar, Theodora Chaspari |
IEEE Trans. Affect. Comput. | 2 |
| 2024 | A Linguistic Analysis of the Impact of Team Interactions on Team Performance During Space Exploration MissionsabstractThis paper examines the impact of team interaction on team performance via linguistic content analysis. Linguistic markers expressed during the interaction are associated with self-assessed team performance via a linear mixed-effect model (LME) in the presence or absence of subtle interactions called “micro-behaviors”. We applied the LME analysis to longitudinal data collected from 9 teams participating in a 45-day simulated spaceflight mission in a ground analog. More specifically we used data collected from team interaction battery (TIB) tasks teams performed during missions which comprise around on average 1.5 hours of conversation data for 5 days per team. The same data was used to predict self-assessed team performance using a deep learning model based on a graph neural network (GNN) to capture interaction patterns between members and a recurrent neural network (RNN) to capture the history of conversation among speakers. Results obtained via an ablation study suggested that the proposed model significantly outperforms the baseline models that modeled either the group interaction between team members or the sequential nature of the data. The study indicates that interaction dynamics among team members are important for the automatic assessment of team performance, thereby informing the development of personalized interventions with concrete and tailored action items to mitigate team performance degradation. Projna Paromita, Theodora Chaspari |
ACII | 2 |
| 2024 | Perception of Stress: A Comparative Multimodal Analysis of Time-Continuous Stress Ratings from Self and ObserversabstractTime-continuous ratings of stress are necessary for designing robust stress detection algorithms that operate in real-time. Common methods for obtaining these ratings in the field of affective computing are through self-reports or by employing multiple external observers. However, limited research has explored the association between these two methods, as well as their respective relation with multimodal bio-behavioral features. Using a mock job interview as a stress inducing task, this paper investigates time-continuous ratings of stress from self-reports and external observers. By analyzing the data from 223 question/answer exchanges from 31 participants, results suggest that observer ratings display low correlation with self ratings (r = 0.145, p < 0.05) and this degree of association varies depending on the inter-rater reliability of external observers. Findings also indicate that multimodal bio-behavioral features show higher association with observer ratings compared to self ratings, and therefore, machine learning models based on this multimodal data can estimate observer ratings (CCC = 0.4688 ± 0.247) better than self ratings (CCC = 0.2172 ± 0.205). Ehsanul Haque Nirjhar, Winfred Arthur, Theodora Chaspari |
ICMI | 3 |
| 2024 | A multimodal analysis of environmental stress experienced by older adults during outdoor walking trips: Implications for designing new intelligent technologies to enhance walkability in low-income Latino communitiesabstractNeighborhood walkability has a significant influence on older adults’ physical and mental health. These effects are amplified in underserved communities (e.g., low-income groups, ethnic minorities) which are often associated with worsening pedestrian infrastructure and safety concerns. This paper investigates environmental stressors linked with decreased walkability of older adults from a low-income Latino community, and how these are associated with physiological, physical, environmental, and sociological variables. 68 older adults were recruited from a primarily Hispanic neighborhood, and each collected two-weeks of multimodal data using wearable and smartphone devices. The data included location, acceleration, and physiological data, such as heart rate and electrodermal activity, from participants’ outdoor walking trips. Environmental stressors participants encountered during each walking trip were self-reported through a mobile application. The first part of this paper discusses unique challenges faced when working with this under-studied population and strategies used to address these challenges while maintaining scientific rigor. The second part of the paper describes results from the preliminary analysis employing linear mixed models (LMM) and machine learning classifiers to examine potential associations between self-reported and objectively-measured stress levels among participants, as well as the effect of environmental, sociological, and individual variables on physiological stress responses while walking. Findings from this study support new avenues for engaging with and gaining deeper insights into a unique and often overlooked population while laying the groundwork for developing new computational models for quantifying environmental stress using wearable and smartphone devices. Raquel Yupanqui, John Sohn, Yoojun Kim, Raquel Flores, Hanwool Lee, Youngjib Ham, Chanam Lee, Theodora Chaspari |
ICMI | 10 |
| 2023 | A Knowledge-Driven Vowel-Based Approach of Depression Classification from Speech Using Data AugmentationabstractWe propose a novel explainable machine learning (ML) model that identifies depression from speech, by modeling the temporal dependencies across utterances and utilizing the spectrotemporal information at the vowel level. Our method first models the variable-length utterances at the local-level into a fixed-size vowel-based embedding using a convolutional neural network with a spatial pyramid pooling layer (vowel CNN). Following that, the depression is classified at the global-level from a group of vowel CNN embeddings that serve as the input to another 1D CNN (depression CNN). Different data augmentation methods are designed for both the training of vowel CNN and depression CNN. We investigate the performance of the proposed system at various temporal granularities when modeling short, medium, and long analysis windows, corresponding to 10, 21, and 42 utterances, respectively. The proposed method reaches comparable performance with previous state-of-the-art approaches and depicts explainable properties with respect to the depression outcome. The findings from this work may benefit clinicians by providing additional intuitions during joint human-ML decision-making tasks. Kexin Feng, Theodora Chaspari |
ICASSP | 2 |
| 2023 | Toward Privacy-Enhancing Ambulatory-Based Well-Being Monitoring: Investigating User Re-Identification Risk in Multimodal DataabstractThe sensitivity of data collected via ambulatory monitoring, which regularly involve the recording of speech signals and sensor information, can cause strong privacy concerns. We investigate user re-identification risk in a corpus of such data collected to observe the interplay between behavior, physiology, and well-being of healthcare workers in their daily life. We then develop a user anonymization approach that preserves well-being information (i.e., anxiety), but eliminates user identify (ID) information. We formulate this via an auto-encoder that learns a transformed version of the original feature set in an adversarial manner so that it minimizes the anxiety estimation loss and maximizes the user classification loss. Results indicate that the original features bear a large user re-identification risk, while also having a good ability to classify a user’s anxiety. After removing the most prone features to user re-identification from the original feature set, the user classification accuracy decreases, while the anxiety classification performance is preserved. The final features transformed via the auto-encoder further reduce evidence of user ID and preserve anxiety classification ability. Findings from this study can contribute to the design privacy-aware bio-behavioral models that can be used for responsible ambulatory monitoring in healthcare and beyond. Ravi Pranjal, Ranjana Seshadri, Rakesh Kumar Sanath Kumar Kadaba, Tiantian Feng, Shri Narayanan, Theodora Chaspari |
ICASSP | 6 |
| 2023 | EMSAssist: An End-to-End Mobile Voice Assistant at the Edge for Emergency Medical ServicesabstractAccurate and prompt delivery of Emergency Medical Services (EMS) is critical in emergency incidents, e.g., man-made or natural disaster areas. However, quickly selecting the correct EMS protocol(s) (which dictate the medical procedures to be administered to patients) in complex medical scenarios, remains a key, demanding task for Emergency Medical Technicians (EMT). In this paper, we present EMSAssist, the first end-to-end mobile voice assistant at the edge for EMS. EMSAssist consists of three major components that address technical challenges present in state-of-the-art solutions: 1) For the first time, EMSAssist proposes and applies a few-sample fine-tuning technique in medical speech recognition task, that achieves a faster and more accurate speech transcription on our EMS audio dataset, when compared to Google Cloud Speech-to-Text; 2) A WordPiece tokenizer helps boosting the end-to-end EMS protocol selection accuracy by retrieving useful information from incorrect transcriptions; 3) A novel data customization framework that enables our data-driven EMSMobileBERT model to become the new state-of-the-art for EMS protocol selection. Extensive end-to-end evaluation results at the edge show EMSAssist can more accurately select EMS protocols (Top-5 accuracy above 96%) for EMTs, with end-to-end latencies of around 4.2 seconds. Liuyi Jin, Tian Liu 0006, Amran Haroon, Radu Stoleru, Michael Middleton, Ziwei Zhu 0001, Theodora Chaspari |
MobiSys | 7 |
| 2023 | Demo: EMSAssist - An End-to-End Mobile Voice Assistant at the Edge for Emergency Medical ServicesabstractWe present EMSAssist, the first end-to-end mobile voice assistant for emergency medical services (EMS). EMSAssist allows Emergency Medical Technicians (EMT) to verbally describe patients' signs and symptoms and uses EMTs' voice input to recommend top-5 EMS protocols. Through this demo, we allow the attendees to evaluate EMSAssist through a pair of Google Glass and a mobile phone. Both mobile devices can collect users' voices as input and output top-5 recommended protocols. A companion youtube video of using EMSAssist on the Google Glass is provided: https://www.youtube.com/watch?v=bj7aQJKf4aE Liuyi Jin, Tian Liu 0006, Amran Haroon, Radu Stoleru, Michael Middleton, Ziwei Zhu 0001, Theodora Chaspari |
MobiSys | 7 |
| 2023 | An Engineering View on Emotions and Speech: From Analysis and Predictive Models to Responsible Human-Centered ApplicationsabstractThe substantial growth of Internet-of-Things technology and the ubiquity of smartphone devices has increased the public and industry focus on speech emotion recognition (SER) technologies. Yet, conceptual, technical, and societal challenges restrict the wide adoption of these technologies in various domains, including, healthcare, and education. These challenges are amplified when automated emotion recognition systems are called to function “in-the-wild” due to the inherent complexity and subjectivity of human emotion, the difficulty of obtaining reliable labels at high temporal resolution, and the diverse contextual and environmental factors that confound the expression of emotion in real life. In addition, societal and ethical challenges hamper the wide acceptance and adoption of these technologies, with the public raising questions about user privacy, fairness, and explainability. This article briefly reviews the history of affective speech processing, provides an overview of current state-of-the-art approaches to SER, and discusses algorithmic approaches to render these technologies accessible to all, maximizing their benefits and leading to responsible human-centered computing applications. Chi-Chun Lee, Theodora Chaspari, Emily Mower Provost, Shri Narayanan |
Proc. IEEE | 2 |
| 2023 | Few-Shot Learning in Emotion Recognition of Spontaneous Speech Using a Siamese Neural Network With Adaptive Sample Pair FormationabstractSpeech-based machine learning (ML) has been heralded as a promising solution for tracking prosodic and spectrotemporal patterns in real-life that are indicative of emotional changes, providing a valuable window into one's cognitive and mental state. Yet, the scarcity of labelled data in ambulatory studies prevents the reliable training of ML models, which usually rely on “data-hungry” distribution-based learning. Leveraging the abundance of labelled speech data from acted emotions, this paper proposes a few-shot learning approach for automatically recognizing emotion in spontaneous speech from a small number of labelled samples. Few-shot learning is implemented via a metric learning approach through a siamese neural network, which models the relative distance between samples rather than relying on learning absolute patterns of the corresponding distributions of each emotion. Results indicate the feasibility of the proposed metric learning in recognizing emotions from spontaneous speech in four datasets, even with a small amount of labelled samples. They further demonstrate superior performance of the proposed metric learning compared to commonly used adaptation methods, including network fine-tuning and adversarial learning. Findings from this work provide a foundation for the ambulatory tracking of human emotion in spontaneous speech contributing to the real-life assessment of mental health degradation. Kexin Feng, Theodora Chaspari |
IEEE Trans. Affect. Comput. | 2 |
| 2023 | Editorial: Special Issue on Unobtrusive Physiological Measurement Methods for Affective ApplicationsabstractIn The formative years of Affective Computing [1], from the late 1990s and into the early 2000s, a significant fraction of research attention was focused on the development of methods forunobtrusive physiological measurement. It quickly became obvious that wiring people with electrodes and strapping cumbersome hardware to their bodies was not only restricting the types of experiments that could be performed but also was not conducive to unbiased observations. For instance, subjects with fingers wrapped with electrodermal activity (EDA) and photoplethysmography (PPG) sensors could hardly type, drive or sleep comfortably. Hence, there was a need for more elegant and scalable physiological measurement methods [2]. Ioannis Pavlidis, Theodora Chaspari, Daniel McDuff |
IEEE Trans. Affect. Comput. | 2 |
| 2022 | Investigating the Interplay Between Self-Reported and Bio-Behavioral Measures of Stress: A Pilot Study of Civilian Job Interviews with Military VeteransabstractTransitioning from the military to the civilian lifestyle, especially for military veterans who decide to pursue careers in the civilian workforce, is often a difficult experience. The job interview, a task in which the interviewees meet and discuss their skills and career goals with strangers in a position of authority, is the first step of assimilation into the civilian workplace, which might cause them to experience nervousness or anxiety. This feeling of excessive stress may compromise the interviewee's performance, therefore potentially impeding their successful transition to the workforce. Intelligent interview training technologies would benefit from automated stress detection systems that could assist interviewees in better understanding causes and antecedents of stressors during their interaction with the interviewer. This paper examines self-reported and bio-behavioral measures of stress experienced during mock job interviews conducted with 24 U.S. military veterans. Self-reported measures were captured via a global measure of stress reported by the participant at the conclusion of the interview, and a continuous moment-to-moment annotation of stress resulting from the retrospective inspection of the interview video recording. Bio-behavioral indices of stress include physiological reactivity measures captured via electrodermal activity and electrocardiogram signals, as well as acoustic measures extracted from speech. Results indicate that physiological reactivity measures exhibit moderate-to-strong correlation with self-reported measures of stress, and can be thus used to estimate the self-reported stress measures. Augmenting the feature space with demographic and psychological traits can further improve the accurate detection of stress during the interviews. Ehsanul Haque Nirjhar, Ellen Hagen, Neha Rani, Sharon Lynn Chu Yew Yee, Winfred Arthur, Amir H. Behzadan, Theodora Chaspari |
ACII | 8 |
| 2022 | Multimodal Affect and Aesthetic ExperienceabstractThe term “aesthetic experience” corresponds to the inner state of a person exposed to the form and content of artistic objects. Quantifying and interpreting the aesthetic experience of people in different contexts can contribute towards (a) creating context and (b) better understanding people’s affective reactions to different aesthetic stimuli. Focusing on different types of artistic content, such as movies, music, literature, urban art, ancient artwork, and modern interactive technology, the goal of this workshop is to enhance the interdisciplinary collaboration among researchers coming from the following domains: affective computing, aesthetics, human-robot/computer interaction, digital archaeology and art, culture, addictive games. Theodoros Kostoulas, Michal Muszynski, Leimin Tian, Edgar Roman-Rangel, Theodora Chaspari, Panos Amelidis |
ICMI | 5 |
| 2022 | Evaluating Just-In-Time Vibrotactile Feedback for Communication AnxietyabstractWrist-worn vibrotactile feedback has been heralded as a promising intervention for reducing state anxiety during stressor events. However, current work has focused on the continuous delivery of the vibrotactile stimulus, which entails the risk of habituation to the potentially relieving effects of the feedback. This paper examines the just-in-time administration of vibrotactile feedback during a public speaking task in an effort to reduce communication apprehension. We evaluate two types of vibrotactile feedback delivery mechanisms compared to a control in a between-subjects design – one that delivers stimulus over random time points and one that delivers stimulus during moments of heightened physiological reactivity, as determined by changes in electrodermal activity. The results from these interventions indicate that vibrotactile feedback administered during high physiological arousal improves stress-related physiological measures (e.g., heart rate) and self-reported stress annotations early on in the intervention, and contributes to increased vocal stability during the public speaking task, but these effects diminish over time. Delivering the vibrotactile feedback over random points in time appears to worsen stress-related measures overall. Jason Raether, Ehsanul Haque Nirjhar, Theodora Chaspari |
ICMI | 3 |
| 2022 | Exploring Individual Differences of Public Speaking Anxiety in Real-Life and Virtual PresentationsabstractPublic speaking is a vital skill for making good impressions, effectively exchanging ideas, and influencing others. Yet, public speaking anxiety (PSA) ranks as a top social phobia. Recent advancements in wearable devices and ubiquitous virtual reality (VR) interfaces can help measure and mitigate PSA. This research quantifies PSA through bio-behavioral markers related to individuals’ physiological and acoustic characteristics. The effect of virtual reality (VR) training on alleviating PSA is measured through self-reported and bio-behavioral indices. Psychological (e.g., general trait anxiety, personality) and demographic (e.g., age, gender, highest education, native language) traits are examined as moderating factors between bio-behavioral indices and PSA, as well as moderating factors for measuring the VR effectiveness in mitigating PSA. These measures are also used as clustering criteria for stratifying participants in group-based models of PSA. Results indicate the significance of such traits to modeling PSA with the proposed group-based models yielding Spearman’s correlation of 0.55 ($p<0.05$) between the actual and predicted outcome. Results further demonstrate that systematic exposure to public speaking in VR can alleviate PSA in terms of both self-reported ($p<0.05$) and physiological ($p<0.05$) indices. Findings from this study will enable researchers to better understand antecedents and causes of PSA and lay the foundation for personalized adaptive feedback for PSA interventions. Megha Yadav, Ehsanul Haque Nirjhar, Kexin Feng, Amir H. Behzadan, Theodora Chaspari |
IEEE Trans. Affect. Comput. | 6 |
| 2022 | Predicting the Macronutrient Composition of Mixed Meals From Dietary Biomarkers in BloodabstractDiet monitoring is an essential intervention component for a number of diseases, from type 2 diabetes to cardiovascular diseases. However, current methods for diet monitoring are burdensome and often inaccurate. In prior work, we showed that continuous glucose monitors (CGMs) may be used to predict meal macronutrients (e.g., carbohydrates, protein, fat) by analyzing the shape of the post-prandial glucose response. In this study, we examine a number of additional dietary biomarkers in blood by their ability to improve macronutrient prediction, compared to using CGMs alone. For this purpose, we conducted a nutritional study where (n = 10) participants consumed nine different mixed meals with varied but known macronutrient amounts, and we analyzed the concentration of 33 dietary biomarkers (including amino acids, insulin, triglycerides, and glucose) at various times post-prandially. Then, we built machine learning models to predict macronutrient amounts from (1) individual biomarkers and (2) their combinations. We find that the additional blood biomarkers provide complementary information, and more importantly, achieve lower normalized root mean squared error (NRMSE) for the three macronutrients (carbohydrates: 22.9%; protein: 23.4%; fat: 32.3%) than CGMs alone (carbohydrates: 28.9%, t(18) =1.64, p =0.060; protein: 46.4%, t(18) =5.38, p 0.001; fat: 40.0%, t(18) =2.09, p =0.025). Our main conclusion is that augmenting CGMs to measure these additional dietary biomarkers improves macronutrient prediction performance, and may ultimately lead to the development of automated methods to monitor nutritional intake. This work is significant to biomedical research as it provides a potential solution to the long-standing problem of diet monitoring, facilitating new interventions for a number of diseases. Anurag Das, Bobak Mortazavi, Seyedhooman Sajjadi, Theodora Chaspari, Laura Ruebush, Nicolaas E. P. Deutz, Gerard L. Coté, Ricardo Gutierrez-Osuna |
IEEE J. Biomed. Health Informatics | 4 |
| 2021 | A Sparse Coding Approach to Automatic Diet Monitoring with Continuous Glucose MonitorsabstractMeasuring dietary intake is a major challenge in the management of chronic diseases. Current methods rely on self-report measures, which are cumbersome to obtain and often unreliable. This article presents an approach to estimate dietary intake automatically by analyzing the post-prandial glucose response (PPGR) of a meal, as measured with continuous glucose monitors. In particular, we propose a sparse-coding technique that can be used to estimate the amounts of macronutrients (carbohydrates, protein, fat) in a meal from the meal’s PPGR. We use Lasso regularization to represent the PPGR of a new meal as a sparse combination of PPGRs in a dictionary, then combine the sparse weights with the macronutrient amounts in the dictionary’s meals to estimate the macronutrients in the new meal. We evaluate the approach on a dataset containing nine standardized meals and their corresponding PPGRs, consumed by fifteen participants. The proposed technique consistently outperforms two baseline systems based on ridge regression and nearest-neighbors, in terms of correlation and normalized root mean square error of the predictions. Anurag Das, Seyedhooman Sajjadi, Bobak Mortazavi, Theodora Chaspari, Projna Paromita, Laura Ruebush, Nicolaas E. P. Deutz, Ricardo Gutierrez-Osuna |
ICASSP | 4 |
| 2021 | Towards The Development of Subject-Independent Inverse Metabolic ModelsabstractDiet monitoring is an important component of interventions in type 2 diabetes, but is time intensive and often inaccurate. To address this issue, we describe an approach to monitor diet automatically, by analyzing fluctuations in glucose after a meal is consumed. In particular, we evaluate three standardization techniques (baseline correction, feature normalization, and model personalization) that can be used to compensate for the large individual differences that exist in food metabolism. Then, we build machine learning models to predict the amounts of macronutrients in a meal from the associated glucose responses. We evaluate the approach on a dataset containing glucose responses for 15 participants who consumed 9 meals. Three techniques improve the accuracy of the models: subtracting the baseline glucose, performing z-score normalization, and scaling the amount of macronutrients by each individuals’ body mass index. Seyedhooman Sajjadi, Anurag Das, Ricardo Gutierrez-Osuna, Theodora Chaspari, Projna Paromita, Laura Ruebush, Nicolaas E. P. Deutz, Bobak Mortazavi |
ICASSP | 4 |
| 2021 | Workshop on Multimodal Affect and Aesthetic ExperienceabstractThe term “aesthetic experience” corresponds to inner states of individuals exposed to art. Investigating form, content, and aesthetic values of artistic objects, indoor and outdoor spaces, urban areas, and modern interactive technology is essential to improve social behaviour, quality of life, and health of humans in the long term. Quantifying and interpreting the aesthetic experience of art receivers in different contexts can contribute towards (a) creating art and (b) better understanding humans’ affective reactions to aesthetic stimuli. Focusing on different types of artistic content, such as movies, music, urban art, ancient artwork, and modern interactive technology, the goal of the Second International Workshop on Multimodal Affect and Aesthetic Experience is to enhance the interdisciplinary collaboration among researchers from the following domains: affective computing, aesthetics, human-robot interaction, and digital archaeology and art. Michal Muszynski, Edgar Roman-Rangel, Leimin Tian, Theodoros Kostoulas, Theodora Chaspari, Panos Amelidis |
ICMI | 5 |
| 2021 | Knowledge- and Data-Driven Models of Multimodal Trajectories of Public Speaking Anxiety in Real and Virtual SettingsabstractPublic speaking skills are essential to professional success. Yet, public speaking anxiety (PSA) is considered one of the most common social phobias. Understanding PSA can help communication experts identify effective ways to treat this communication-based disorder. Existing works on PSA rely on self-reports and aggregate multimodal measures which do not capture the temporal variation in PSA. This paper examines temporal trajectories of acoustic and physiological measures throughout the public speaking encounter with real and virtual audiences, and aims to model those in both knowledge- and data-driven ways. Knowledge-driven models leverage theoretically-grounded patterns through fitting interpretable parametric functions to the corresponding signals. Data-driven models consider the functional nature of multimodal signals via functional principal component analysis. Results indicate that the parameters of the proposed models can successfully estimate individuals’ trait anxiety in both real-life and virtual reality settings, and suggest that models trained on data obtained in virtual public speaking stimuli are able to estimate levels of PSA in real-life. Ehsanul Haque Nirjhar, Amir H. Behzadan, Theodora Chaspari |
ICMI | 3 |
| 2021 | Investigating Trust in Human-Machine Learning Collaboration: A Pilot Study on Estimating Public Anxiety from SpeechabstractTrust is a key element in the development of effective collaborative relationships between humans and increasingly complex artificial intelligence (AI) systems. Here, we examine trust in AI in the context of a human-AI partnership that involves a joint decision making task for estimating levels of public speaking anxiety based on speech signals. The AI system is comprised of an explainable machine learning (ML) algorithm, that takes acoustic characteristics as input and outputs the estimate of public speaking anxiety levels, a local explanation about the most important features that contributed to the decision of each speech sample, and a global explanation about the most important features for the data overall. We analyze interactions between AI and human annotators with background in psychological sciences, and measure trust over time via the annotators’ agreement with the AI model and the annotators’ self-reports. We further examine factors of trust that are related to the characteristics of the human annotator and the ML algorithm. Results indicate that trust in AI depends on the openness level of the annotator and the importance level of input features. Findings from this study can provide guidelines to designing solutions that properly calibrate human trust in AI in collaborative human-AI tasks. Abdullah Aman Tutul, Ehsanul Haque Nirjhar, Theodora Chaspari |
ICMI | 3 |
| 2021 | Assessment of Daily Routine Uniformity in a Smart Home Environment Using Hierarchical ClusteringabstractThe gradual decline in routine patterns is a major symptom of early-stage dementia, therefore an unobtrusive real-life assessment of the elder's routine can potentially be of significant clinical importance. This article focuses on the assessment of changes in a person's daily routine using longitudinal data recorded from a network of nonintrusive motion sensors in a smart home environment. In this article, we propose to identify repeating patterns in a person's daily routine over the span of multiple days using hierarchical clustering algorithms, which provide an effective way to mitigate noise artifacts and confounding factors that contribute to the momentary variability of the sensor data. We have evaluated our proposed algorithm on both synthetic and real-world data recorded in the span of 50-100 days from four elderly adults. Our results indicate that the proposed hierarchical clustering approach can more reliably capture the gradual change in the degree of routineness compared to baseline approaches that measure the similarity between two consecutive days or capture variations in the occurrence of recognized activities. Prakhar Mohan, Bogyeong Lee, Theodora Chaspari, Changbum R. Ahn |
IEEE J. Biomed. Health Informatics | 3 |
| 2020 | Exploring Bio-Behavioral Signal Trajectories of State Anxiety During Public SpeakingabstractPublic speaking anxiety (PSA) is among the top social phobias in the world. Quantifying PSA in a reliable and unobtrusive manner can lay the foundation toward personalized and inexpensive technology-based interventions. Existing work for quantifying PSA often relies on self-reported measures and statistical aggregates of bio-behavioral indices, such as physiology and speech. Such aggregated bio-behavioral indices are not able to capture time-based trajectories of PSA variation, that can be very useful for better understanding and reliably predicting moments of anxiety. We tackle this problem by introducing temporal parametric models to quantify bio-behavioral trajectories of PSA throughout a public speaking encounter. Using data from 55 participants in a real-life public speaking task, the parameters of the proposed models are found to be significantly correlated with individuals' trait characteristics of general and communication-based anxiety, outperforming aggregate mean bio-behavioral measures. Ehsanul Haque Nirjhar, Amir H. Behzadan, Theodora Chaspari |
ICASSP | 3 |
| 2020 | Predicting the Effectiveness of Systematic Desensitization Through Virtual Reality for Mitigating Public Speaking AnxietyabstractPublic speaking is central to socialization in casual, professional, or academic settings. Yet, public speaking anxiety (PSA) is known to impact a considerable portion of the general population. This paper utilizes bio-behavioral indices captured from wearable devices to quantify the effectiveness of systematic exposure to virtual reality (VR) audiences for mitigating PSA. The effect of separate bio-behavioral features and demographic factors is studied, as well as the amount of necessary data from the VR sessions that can yield a reliable predictive model of the VR training effectiveness. Results indicate that acoustic and physiological reactivity during the VR exposure can reliably predict change in PSA before and after the training. With the addition of demographic features, both acoustic and physiological feature sets achieve improvements in performance. Finally, using bio-behavioral data from six to eight VR sessions can yield reliable prediction of PSA change. Findings of this study will enable researchers to better understand how bio-behavioral factors indicate improvements in PSA with VR training. Margaret von Ebers, Ehsanul Haque Nirjhar, Amir H. Behzadan, Theodora Chaspari |
ICMI | 4 |
| 2020 | Multimodal Affect and Aesthetic ExperienceabstractThe term 'aesthetic experience' corresponds to the inner state of a person exposed to form and content of artistic objects. Exploring certain aesthetic values of artistic objects, as well as interpreting the aesthetic experience of people when exposed to art can contribute towards understanding (a) art and (b) people's affective reactions to artwork. Focusing on different types of artistic content, such as movies, music, urban art and other artwork, the goal of this workshop is to enhance the interdisciplinary collaboration between affective computing and aesthetics researchers. Theodoros Kostoulas, Michal Muszynski, Theodora Chaspari, Panos Amelidis |
ICMI | 3 |
| 2020 | Preserving Privacy in Image-based Emotion Recognition through User AnonymizationabstractThe large amount of data captured by ambulatory sensing devices can afford us insights into longitudinal behavioral patterns, which can be linked to emotional, psychological, and cognitive outcomes. Yet, the sensitivity of behavioral data, which regularly involve speech signals and facial images, can cause strong privacy concerns, such as the leaking of the user identity. We examine the interplay between emotion-specific and user identity-specific information in image-based emotion recognition systems. We further study a user anonymization approach that preserves emotion-specific information, but eliminates user-dependent information from the convolutional kernel of convolutional neural networks (CNN), therefore reducing user re-identification risks. We formulate an adversarial learning problem implemented with a multitask CNN, that minimizes emotion classification and maximizes user identification loss. The proposed system is evaluated on three datasets achieving moderate to high emotion recognition and poor user identity recognition performance. The resulting image transformation obtained by the convolutional layer is visually inspected, attesting to the efficacy of the proposed system in preserving emotion-specific information. Implications from this study can inform the design of privacy-aware emotion recognition systems that preserve facets of human behavior, while concealing the identity of the user, and can be used in ambulatory monitoring applications related to health, well-being, and education. Vansh Narula, Kexin Feng, Theodora Chaspari |
ICMI | 3 |
| 2020 | Saliency detection analysis of collective physiological responses of pedestrians to evaluate neighborhood built environments
Megha Yadav, Theodora Chaspari, Changbum R. Ahn |
Adv. Eng. Informatics | 3 |
| 2020 | Sub-Population Specific Models of Couples' ConflictabstractInterpersonal conflict between couples is a significant source of stress with long-lasting effects on partners’ physical and psychological health. Motivated by findings in psychological science, we study how couples with distinct relationship functioning characteristics experience conflict in real life. We propose sub-population specific machine learning models using hierarchical and adaptive learning frameworks to automatically detect interpersonal conflict through the ambulatory monitoring of couples’ physiological signals, audio samples, and linguistic indices. Results indicate that the proposed models outperform a general model learned for the entire population and separate models independently trained on each sub-population, providing a foundation toward personalized health applications. Krit Gupta, Aditya Gujral, Theodora Chaspari, Adela C. Timmons, Sohyun C. Han, Yehsong Kim, Sarah Barrett, Stassja Sichko, Gayla Margolin |
ACM Trans. Internet Techn. | 3 |
| 2019 | Virtual reality interfaces and population-specific models to mitigate public speaking anxietyabstractPublic speaking is key to effectively exchanging ideas, persuading others, and making a tangible impact. Yet, public speaking anxiety (PSA) ranks as a top social phobia among many people. This paper leverages bio-behavioural indices captured from wearable devices and virtual reality (VR) interfaces to quantify PSA. The significance of individual-specific factors, such as general trait anxiety and personality, as well as contextual factors, such as age, gender, highest education, and native language, in moderating the association between bio-behavioral indices and PSA is further examined through group-based machine learning models. Results highlight the importance of including such factors for detecting PSA with the proposed group-based PSA models yielding Spearman's correlation of 0.55(p <; 0.05) between the actual and predicted state-based anxiety scores. This work further analyzes whether systematic exposure to public speaking tasks in the VR environment can help alleviate PSA. Results indicate that systematic exposure to public speaking in VR can alleviate PSA in terms of both self-reported (p <; 0.05) and physiological (p <; 0.05) indices. Findings of this study will enable researchers to better understand antedecedents and causes of PSA contributing to behavioral interventions using VR. Megha Yadav, Kexin Feng, Theodora Chaspari, Amir H. Behzadan |
ACII | 4 |
| 2019 | An Attention-aware Bidirectional Multi-residual Recurrent Neural Network (Abmrnn): A Study about Better Short-term Text ClassificationabstractLong Short-Term Memory (LSTM) has been proven an efficient way to model sequential data, because of its ability to overcome the gradient diminishing problem during training. However, due to the limited memory capacity in LSTM cells, LSTM is weak in capturing long-time dependency in sequential data. To address this challenge, we propose an Attention-aware Bidirectional Multi-residual Recurrent Neural Network (ABMRNN) to overcome the deficiency. Our model considers both past and future information at every time step with omniscient attention based on LSTM. In addition to that, the multi-residual mechanism has been leveraged in our model which aims to model the relationship between current time step with further distant time steps instead of a just previous time step. The results of experiments show that our model achieves state-of-the-art performance in classification tasks. Ye Wang 0006, Xinxiang Zhang, Theodora Chaspari, Yoonsuck Choe, Mi Lu |
ICASSP | 4 |
| 2019 | Exploring Transfer Learning between Scripted and Spontaneous Speech for Emotion RecognitionabstractInternet of Things technologies yield large amounts of real-life speech data related to human emotions. Yet, labelled data of human emotion from spontaneous speech are extremely limited due to the difficulties in the annotation of such large volumes of audio samples. A potential way to address this limitation is to augment emotion models of spontaneous speech with fully annotated data collected using scripted scenarios. We investigate whether and to what extent knowledge related to speech emotional content can be transferred between datasets of scripted and spontaneous speech. We implement transfer learning through: (1) a feed-forward neural network trained on the source data and whose last layers are fine-tuned based on the target data; and (2) a progressive neural network retaining a pool of pre-trained models and learning lateral connections between source and target task. We explore the effectiveness of the proposed approach using four publicly available datasets of emotional speech. Our results indicate that transfer learning can effectively leverage corpora of scripted data to improve emotion recognition performance for spontaneous speech. Theodora Chaspari |
ICMI | 2 |
| 2018 | Human-Habitat for Health (H3): Human-habitat Multimodal Interaction for Promoting Health and Well-being in the Internet of Things EraabstractThis paper presents an introduction to the "Human-Habitat for Health (H3): Human-habitat multimodal interaction for promoting health and well-being in the Internet of Things era" workshop, which was held at the 20th ACM International Conference on Multimodal Interaction on October 16th, 2018, in Boulder, CO, USA. The main theme of the workshop focused on the effect of the physical or virtual environment on individual's behavior, well-being, and health. The H3 workshop included keynote speeches that provided an overview and future directions of the field, as well as presentations including position papers and research contributions. The workshop brought together experts from academia and industry spanning a set of multi-disciplinary fields, including computer science, speech and spoken language understanding, construction science, life-sciences, health sciences, and psychology, to discuss their respective views and identify synergistic and converging research directions and solutions. Theodora Chaspari, Angeliki Metallinou, Leah I. Stein Duker, Amir H. Behzadan |
ICMI | 1 |
| 2018 | Population-specific Detection of Couples' Interpersonal Conflict using Multi-task LearningabstractThe inherent diversity of human behavior limits the capabilities of general large-scale machine learning systems, that usually require ample amounts of data to provide robust descriptors of the outcomes of interest. Motivated by this challenge, personalized and population-specific models comprise a promising line of work for representing human behavior, since they can make decisions for clusters of people with common characteristics, reducing the amount of data needed for training. We propose a multi-task learning (MTL) framework for developing population-specific models of interpersonal conflict between couples using ambulatory sensor and mobile data from real-life interactions. The criteria for population clustering include global indices related to couples' relationship quality and attachment style, person-specific factors of partners' positivity, negativity, and stress levels, as well as fluctuating factors of daily emotional arousal obtained from acoustic and physiological indices. Population-specific information is incorporated through a MTL feed-forward neural network (FF-NN), whose first layers capture the common information across all data samples, while its last layers are specific to the unique characteristics of each population. Our results indicate that the proposed MTL FF-NN trained solely on the sensor-based acoustic, linguistic, and physiological modalities provides unweighted and weighted F1-scores of 0.51 and 0.75, respectively, outperforming the corresponding baselines of a single general FF-NN trained on the entire dataset and separate FF-NNs trained on each population cluster individually. These demonstrate the feasibility of such ambulatory systems for detecting real-life behaviors and possibly intervening upon them, and highlights the importance of taking into account the inherent diversity of different populations from the general pool of data. Aditya Gujral, Theodora Chaspari, Adela C. Timmons, Yehsong Kim, Sarah Barrett, Gayla Margolin |
ICMI | 2 |
| 2018 | Automated ergonomic risk monitoring using body-mounted sensors and machine learning
Nipun D. Nath, Theodora Chaspari, Amir H. Behzadan |
Adv. Eng. Informatics | 2 |
| 2017 | Exploring sparse representation measures of physiological synchrony for romantic couplesabstractQuantifying the inherent coordination between interacting individuals can afford us new insights into their emotions, communicative intent, and relationship quality. We propose a novel framework to capture the physiological synchrony between romantic partners through sparse representation techniques and appropriately designed parametric dictionaries that take into account the characteristic structure of the considered signals. Physiological synchrony is operationalized as the similarity of co-occurring electrodermal activity (EDA) streams captured through the distance in the corresponding parametric representation space, as well as through the joint signal representation errors. Results indicate that the proposed sparse EDA synchrony measures (SESM)-evaluated on two datasets of couples' interactions-differ across tasks of various emotional intensity and are associated with the partners' attachment style. These results provide a foundation towards designing novel descriptors of interaction and physiological linkage between individuals for emerging affective computing applications. Theodora Chaspari, Adela C. Timmons, Brian R. Baucom, Laura Perrone, Katherine J. W. Baucom, Panayiotis G. Georgiou, Gayla Margolin, Shri Narayanan |
ACII | 1 |
| 2017 | A knowledge-driven framework for ECG representation and interpretation for wearable applicationsabstractThe increasing use of wearable technology creates the need for reliable signal representations with low storage and transmission cost, as well as interpretable models that can be used to translate signals into meaningful constructs. We propose a knowledge-driven sparse representation of the electrocardiogram (ECG) that takes into account the characteristic structure of the corresponding signal through the use of appropriately designed parametric dictionaries containing Hermite and amplitude-modulated sinusoidal atoms for the P, T waves and QRS complex, respectively. We further demonstrate how these atoms can be used to automatically interpret the ECG morphology through the QRS detection and beat classification. Our results indicate relative errors of the order of 10-2, compression rates 10 times smaller than the actual signal, as well as reliable QRS detection (93%) and beat classification (78%). These are discussed in terms of developing efficient and reliable wearable ECG applications. Ramasubramanian Balasubramanian, Theodora Chaspari, Shri Narayanan |
ICASSP | 2 |
| 2017 | Quantifying regulation mechanisms in dating couples through a dynamical systems model of acoustic and physiological arousalabstractNegative emotional arousal during conflict has been related to negative outcomes in romantic relationships and degraded quality of family life. Despite its extensive study in psychology, it is still challenging to quantify emotional arousal in a meaningful way with objective indices beyond traditionally-used self-reported scores. We examine the association of acoustic and physiological arousal between dating couples through speech prosodic patterns and Electrodermal Activity (EDA) features. We use a dynamical systems model (DSM) approach to capture the interplay of arousal indices within and between people. The DSM parameters reflect the amount of self-regulation with respect to the acoustic and physiological cues within a person, the degree of cross-regulation between the two modalities, as well as the within-couple co-regulation. Our results through statistical analysis and classification experiments indicate a significant association between the estimated system parameters and the participants' self-reported relationship satisfaction measures. This is consistent with previous findings and can help towards better understanding regulation mechanisms and escalation effects of emotional arousal during couples' discussions. Theodora Chaspari, Sohyun C. Han, Daniel Bone, Adela C. Timmons, Laura Perrone, Gayla Margolin, Shri Narayanan |
ICASSP | 1 |
| 2016 | Pathological speech processing: State-of-the-art, current challenges, and future directionsabstractThe study of speech pathology involves evaluation and treatment of speech production related disorders affecting phonation, fluency, intonation and aeromechanical components of respiration. Recently, speech pathology has garnered special interest amongst machine learning and signal processing (ML-SP) scientists. This growth in interest is led by advances in novel data collection technology, data science, speech processing and computational modeling. These in turn have enabled scientists in better understanding both the causes and effects of pathological speech conditions. In this paper, we review the application of machine learning and signal processing techniques to speech pathology and specifically focus on three different aspects. First, we list challenges such as controlling subjectivity in pathological speech assessments and patient variability in the application of ML-SP tools to the domain. Second, we discuss feature design methods and machine learning algorithms using a combination of domain knowledge and data driven methods. Finally, we present some case studies related to analysis of pathological speech and discuss their design. Rahul Gupta 0001, Theodora Chaspari, Jangwon Kim, Naveen Kumar 0004, Daniel Bone, Shri Narayanan |
ICASSP | 2 |
| 2016 | An Acoustic Analysis of Child-Child and Child-Robot Interactions for Understanding Engagement during Speech-Controlled Computer Games
Theodora Chaspari, Jill Fain Lehman |
INTERSPEECH | 1 |
| 2015 | Quantifying EDA synchrony through joint sparse representation: A case-study of couples' interactionsabstractThe co-variation degree between individuals in their physiological signals can reveal insights about the quality of their interaction as well as their personal characteristics. In an effort to capture the amount of synchrony between Electrodermal Activity (EDA) streams occurring in parallel during dyadic interactions, we propose Sparse EDA Synchrony Measure (SESM), an index derived from the joint sparse representation of EDA ensembles. Sparse decomposition is performed using Simultaneous Orthogonal Matching Pursuit (SOMP) from a knowledge-driven dictionary of tonic and phasic atoms, capturing the slow-varying trends and high-frequency signal fluctuations, respectively. At each iteration the atom having the maximum average correlation with the residuals is selected. We compute SESM as the negative natural logarithm of the joint reconstruction error and evaluate it with data from interactions of married and young dating couples participating in tasks of varying emotional intensity. Through statistical analysis and multiple linear regression experiments, our results indicate that SESM depicts significant differences across tasks in both datasets considered and can be associated to individuals' attachment-related characteristics. Theodora Chaspari, Brian R. Baucom, Adela C. Timmons, Andreas Tsiartas, Larissa Borofsky Del Piero, Katherine J. W. Baucom, Panayiotis G. Georgiou, Gayla Margolin, Shri Narayanan |
ICASSP | 1 |
| 2015 | Analysis and modeling of the role of laughter in motivational interviewing based psychotherapy conversations
Rahul Gupta 0001, Theodora Chaspari, Panayiotis G. Georgiou, David C. Atkins, Shri Narayanan |
INTERSPEECH | 2 |
| 2014 | A non-homogeneous poisson process model of Skin Conductance Responses integrated with observed regulatory behaviors for Autism interventionabstractEarly intervention in individuals with Autism Spectrum Disorder (ASD) can improve core and associated symptoms and facilitate skills that increase social opportunities. However, determining effective intervention success in this population, and the mechanisms that produce it, is currently restricted to observable behavior. The need of therapy assessment metrics beyond traditional behavioral criteria, led to the use of physiological signals for capturing child-therapist internal dynamics during an intervention session. Internal physiological states were measured through Electrodermal Activity (EDA) and modeled in relation to observed self- and co-regulatory behaviors. A common measure of EDA, Skin Conductance Response (SCR), was the primary signal of interest and assumed to form a non-homogeneous Poisson Process whose rate function is determined by observed regulatory behaviors. Through likelihood and residual goodness of fit analysis, statistical tests and classification tasks, our results indicate that SCR changes and observable behavior in child-therapist dyads are temporally associated and the estimated model parameters can be linked to the types of regulation stimuli. Theodora Chaspari, Matthew S. Goodwin, Oliver Wilder-Smith, Amanda Gulsrud, Charlotte A. Mucchetti, Connie Kasari, Shri Narayanan |
ICASSP | 1 |
| 2013 | Using physiology and language cues for modeling verbal response latencies of children with ASDabstractSignal-derived measures can provide effective ways towards quantifying human behavior. Verbal Response Latencies (VRLs) of children with Autism Spectrum Disorders (ASD) during conversational interactions are able to convey valuable information about their cognitive and social skills. Motivated by the inherent gap between the external behavior and inner affective state of children with ASD, we study their VRLs in relation to their explicit but also implicit behavioral cues. Explicit cues include the children's language use, while implicit cues are based on physiological signals. Using these cues, we perform classification and regression tasks to predict the duration type (short/long) and value of VRLs of children with ASD while they interacted with an Embodied Conversational Agent (ECA) and their parents. Since parents are active participants in these triadic interactions, we also take into account their linguistic and physiological behaviors. Our results suggest an association between VRLs and these externalized and internalized signal information streams, providing complementary views of the same problem. Theodora Chaspari, Daniel Bone, James Gibson, Chi-Chun Lee, Shri Narayanan |
ICASSP | 1 |
| 2013 | Classifying language-related developmental disorders from speech cues: the promise and the potential confoundsabstractSpeech and spoken language cues offer a valuable means to measure and model human behavior. Computational models of speech behavior have the potential to support health care through assistive technologies, informed intervention, and effi-cient long-term monitoring. The Interspeech 2013 Autism Sub-Challenge addresses two developmental disorders that manifest in speech: autism spectrum disorders and specific language im-pairment. We present classification results with an analysis on the development set including a discussion of potential con-founds in the data such as recording condition differences. We hence propose study of features within these domains that may inform realistic separability between groups as well as have the potential to be used for behavioral intervention and monitoring. We investigate template-based prosodic and formant modeling as well as goodness of pronunciation modeling, reporting above chance classification accuracies. Index Terms: autism spectrum disorders, intonation, specific language impairment, goodness of pronunciation Daniel Bone, Theodora Chaspari, Kartik Audhkhasi, James Gibson, Andreas Tsiartas, Maarten Van Segbroeck, Ming Li 0026, Sungbok Lee, Shri Narayanan |
INTERSPEECH | 2 |
| 2013 | Acoustic-prosodic, turn-taking, and language cues in child-psychologist interactions for varying social demandabstractImpaired social communication and social reciprocity are the primary phenotypic distinctions between autism spectrum dis-orders (ASD) and other developmental disorders. We investi-gate quantitative conversational cues in child-psychologist in-teractions using acoustic-prosodic, turn-taking, and language features. Results indicate the conversational quality degraded for children with higher ASD severity, as the child exhibited difficulties conversing and the psychologist varied her speech and language strategies to engage the child. When interacting with children with increasing ASD severity, the psychologist exhibited higher prosodic variability, increased pausing, more speech, atypical voice quality, and less use of conventional con-versational cue such as assents and non-fluencies. Children with increasing ASD severity spoke less, spoke slower, responded later, had more variable prosody, and used personal pronouns, affect language, and fillers less often. We also investigated the predictive power of features from interaction subtasks with varying social demands placed on the child. We found that acoustic prosodic and turn-taking features were more predictive during higher social demand tasks, and that the most predictive features vary with context of interaction. We also observed that psychologist language features may be robust to the amount of speech in a subtask, showing significance even when the child is participating in minimal-speech, low social-demand tasks. Index Terms: autism spectrum disorders, atypical prosody, so-cial reciprocity, turn-taking, language cues Daniel Bone, Chi-Chun Lee, Theodora Chaspari, Matthew Black, Marian E. Williams, Sungbok Lee, Pat Levitt, Shri Narayanan |
INTERSPEECH | 3 |
| 2013 | Analyzing the structure of parent-moderated narratives from children with ASD using an entity-based approachabstractStorytelling is a commonly used technique for rating linguistic and communicative abilities of children with Autism Spectrum Disorders (ASD). It highlights their language use beyond sentence-level production, and their ability to cohesively link events into a plot, including incorporating social context. A key scenario of interest we consider is spoken narrative creation in interactive settings, where confederates such as parents can offer scaffolding to their children’s narratives by eliciting answers with appropriate questions, shaping the structure of the resulting narrative. We analyze the structure of children’s stories narrated with the help of their parents using entity-based feature-level patterns in order to see how there are influenced by the parents’ narrative elicitation techniques. The frequency distribution and evolution of entities -meaning the co-referent people, objects and ideas- can capture the main axis of the story plot. Our results indicate that the type of questions the parents ask can be reflected in the entity-based features of a narrative, affecting its underlying structure and coherence. Index Terms: Narrative Structure, Coherence, Text Entities, Autism Spectrum Disorders Theodora Chaspari, Emily Mower Provost, Shri Narayanan |
INTERSPEECH | 1 |
| 2013 | Multi-band long-term signal variability features for robust voice activity detectionabstractIn this paper, we propose robust features for the problem of voice activity detection (VAD). In particular, we extend the long term signal variability (LTSV) feature to accommodate multiple spectral bands. The motivation of the multi-band approach stems from the non-uniform frequency scale of speech phonemes and noise characteristics. Our analysis shows that the multi-band approach offers advantages over the single band LTSV for voice activity detection. In terms of classification accuracy, we show 0.3%-61.2% relative improvement over the best accuracy of the baselines considered for 7 out 8 different noisy channels. Experimental results, and error analysis, are reported on the DARPA RATS corpora of noisy speech. Index Terms: noisy speech data, voice activity detection, robust feature extraction Andreas Tsiartas, Theodora Chaspari, Athanasios Katsamanis, Prasanta Kumar Ghosh, Ming Li 0026, Maarten Van Segbroeck, Alexandros Potamianos, Shri Narayanan |
INTERSPEECH | 2 |
| 2012 | An acoustic analysis of shared enjoyment in ECA interactions of children with autismabstractThe quality of shared enjoyment in interactions is a key aspect related to Autism Spectrum Disorders (ASD). This paper discusses two types of enjoyment: the first refers to humorous events and is associated with one's positive affective state and the second is used to facilitate social interactions between people. These types of shared enjoyment are objectively specified by their proximity to a voiced and unvoiced laughter instance, respectively. The goal of this work is to study the acoustic differences of areas surrounding the two kinds of shared enjoyment instances, called “social zones”, using data collected from children with autism, and their parents, interacting with an Embodied Conversational Agent (ECA). A classification task was performed to predict whether a “social zone” surrounds a voiced or an unvoiced laughter instance. Our results indicate that humorous events are more easily recognized than events acting as social facilitators and that related speech patterns vary more across children compared to other interlocutors. Theodora Chaspari, Emily Mower Provost, Athanasios Katsamanis, Shri Narayanan |
ICASSP | 1 |
| 2012 | Interplay between verbal response latency and physiology of children with autism during ECA interactions
Theodora Chaspari, Chi-Chun Lee, Shri Narayanan |
INTERSPEECH | 1 |
| 2011 | Analyzing the Nature of ECA Interactions in Children with AutismabstractEmbodied conversational agents (ECA) offer platforms for the collection of structured interaction and communication data. This paper discusses the data collected from the Rachel system, an ECA developed at the University of Southern California, for interactions with children with autism. Two dyads each com-posed of a child with autism and his parent participated in an experiment with two modes: interactions with and without the ECA present. The goal of this work is to assess the naturalness of the data recorded in the ECA interaction. This analysis was carried out using a classification framework with a prediction variable of the presence or absence of the ECA in the inter-action. The results demonstrate that it is possible to estimate whether or not a parent is interacting with the ECA using their speech data. However, it is not generally possible to do so for the child suggesting that the Rachel system is eliciting commu-nication data that is similar to that elicited through interactions between the child and his parent. Index Terms: Embodied conversational agent, multimodal in-terface, audio-video recording, autism, children’s speech Emily Mower Provost, Chi-Chun Lee, James Gibson, Theodora Chaspari, Marian E. Williams, Shri Narayanan |
INTERSPEECH | 4 |