VLDB 2026 Research / reviewers in the wild / expert
Ehsanul Haque Nirjhar
dblp:213/7960
· DBLP profile ↗
10ranked-venue papers
5as first author
8since 2021 · last 2025
0000-0002-2345-895XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 6 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Investigating the Reasoning Abilities of Large Language Models for Understanding Spoken Language in Interpersonal Interactions
Pranjal Aggarwal, Ghritachi Mahajani, Pavan Kumar Malasani, Vaibhav Jamadagni, Caroline J. Wendt, Ehsanul Haque Nirjhar, Theodora Chaspari |
INTERSPEECH | 6 |
| 2025 | Modeling Gold Standard Moment-to-Moment Ratings of Perception of Stress From Audio RecordingsabstractEnabling continuous and unobtrusive monitoring of stress is essential for delivering personalized stress interventions at opportune moments. To achieve automatic stress detection on a time-continuous basis, reliable moment-to-moment ratings of stress are required. However, the current research lacks a large-scale multimodal dataset that provides time-continuous ratings of perceived stress. Existing datasets mainly consist of single-valued self-reported ratings obtained after the stress-inducing task or rely on audio-visual recordings to capture moment-to-moment ratings from multiple annotators. The collection of time-continuous ratings of stress based solely on audio recordings has not been extensively explored. In this paper, we introduce an updated version of the publicly available VerBIO dataset that contains moment-to-moment ratings of perceived stress from multiple annotators. These annotators rated their perception of stress by listening to participants who had conducted a public speaking task. Time-continuous ratings of stress are obtained from four annotators using 22 hours of audio recordings from 339 public speaking sessions performed by 53 individuals. These time-continuous ratings of stress perception were obtained from the annotators solely based on speech, without incorporating the visual modality as an expressive marker. We examine the reliability of the annotation scheme employed in this study and investigate the factors contributing to the observed variation in perceived stress among annotators. Next, we introduce an annotation fusion technique based on expectation-maximization to obtain a reliable gold standard rating by aggregating the ratings from multiple annotators. Results indicate that the proposed annotation fusion technique yields aggregated ratings that can be estimated more reliably using acoustic features compared to the ratings yielded from conventional annotation fusion techniques. The newly generated annotations are publicly available within the proposed updated version of the existing VerBIO dataset, facilitating research in the field of continuous stress detection. Ehsanul Haque Nirjhar, Theodora Chaspari |
IEEE Trans. Affect. Comput. | 1 |
| 2024 | Perception of Stress: A Comparative Multimodal Analysis of Time-Continuous Stress Ratings from Self and ObserversabstractTime-continuous ratings of stress are necessary for designing robust stress detection algorithms that operate in real-time. Common methods for obtaining these ratings in the field of affective computing are through self-reports or by employing multiple external observers. However, limited research has explored the association between these two methods, as well as their respective relation with multimodal bio-behavioral features. Using a mock job interview as a stress inducing task, this paper investigates time-continuous ratings of stress from self-reports and external observers. By analyzing the data from 223 question/answer exchanges from 31 participants, results suggest that observer ratings display low correlation with self ratings (r = 0.145, p < 0.05) and this degree of association varies depending on the inter-rater reliability of external observers. Findings also indicate that multimodal bio-behavioral features show higher association with observer ratings compared to self ratings, and therefore, machine learning models based on this multimodal data can estimate observer ratings (CCC = 0.4688 ± 0.247) better than self ratings (CCC = 0.2172 ± 0.205). Ehsanul Haque Nirjhar, Winfred Arthur, Theodora Chaspari |
ICMI | 1 |
| 2022 | Investigating the Interplay Between Self-Reported and Bio-Behavioral Measures of Stress: A Pilot Study of Civilian Job Interviews with Military VeteransabstractTransitioning from the military to the civilian lifestyle, especially for military veterans who decide to pursue careers in the civilian workforce, is often a difficult experience. The job interview, a task in which the interviewees meet and discuss their skills and career goals with strangers in a position of authority, is the first step of assimilation into the civilian workplace, which might cause them to experience nervousness or anxiety. This feeling of excessive stress may compromise the interviewee's performance, therefore potentially impeding their successful transition to the workforce. Intelligent interview training technologies would benefit from automated stress detection systems that could assist interviewees in better understanding causes and antecedents of stressors during their interaction with the interviewer. This paper examines self-reported and bio-behavioral measures of stress experienced during mock job interviews conducted with 24 U.S. military veterans. Self-reported measures were captured via a global measure of stress reported by the participant at the conclusion of the interview, and a continuous moment-to-moment annotation of stress resulting from the retrospective inspection of the interview video recording. Bio-behavioral indices of stress include physiological reactivity measures captured via electrodermal activity and electrocardiogram signals, as well as acoustic measures extracted from speech. Results indicate that physiological reactivity measures exhibit moderate-to-strong correlation with self-reported measures of stress, and can be thus used to estimate the self-reported stress measures. Augmenting the feature space with demographic and psychological traits can further improve the accurate detection of stress during the interviews. Ehsanul Haque Nirjhar, Ellen Hagen, Neha Rani, Sharon Lynn Chu Yew Yee, Winfred Arthur, Amir H. Behzadan, Theodora Chaspari |
ACII | 1 |
| 2022 | Evaluating Just-In-Time Vibrotactile Feedback for Communication AnxietyabstractWrist-worn vibrotactile feedback has been heralded as a promising intervention for reducing state anxiety during stressor events. However, current work has focused on the continuous delivery of the vibrotactile stimulus, which entails the risk of habituation to the potentially relieving effects of the feedback. This paper examines the just-in-time administration of vibrotactile feedback during a public speaking task in an effort to reduce communication apprehension. We evaluate two types of vibrotactile feedback delivery mechanisms compared to a control in a between-subjects design – one that delivers stimulus over random time points and one that delivers stimulus during moments of heightened physiological reactivity, as determined by changes in electrodermal activity. The results from these interventions indicate that vibrotactile feedback administered during high physiological arousal improves stress-related physiological measures (e.g., heart rate) and self-reported stress annotations early on in the intervention, and contributes to increased vocal stability during the public speaking task, but these effects diminish over time. Delivering the vibrotactile feedback over random points in time appears to worsen stress-related measures overall. Jason Raether, Ehsanul Haque Nirjhar, Theodora Chaspari |
ICMI | 2 |
| 2022 | Exploring Individual Differences of Public Speaking Anxiety in Real-Life and Virtual PresentationsabstractPublic speaking is a vital skill for making good impressions, effectively exchanging ideas, and influencing others. Yet, public speaking anxiety (PSA) ranks as a top social phobia. Recent advancements in wearable devices and ubiquitous virtual reality (VR) interfaces can help measure and mitigate PSA. This research quantifies PSA through bio-behavioral markers related to individuals’ physiological and acoustic characteristics. The effect of virtual reality (VR) training on alleviating PSA is measured through self-reported and bio-behavioral indices. Psychological (e.g., general trait anxiety, personality) and demographic (e.g., age, gender, highest education, native language) traits are examined as moderating factors between bio-behavioral indices and PSA, as well as moderating factors for measuring the VR effectiveness in mitigating PSA. These measures are also used as clustering criteria for stratifying participants in group-based models of PSA. Results indicate the significance of such traits to modeling PSA with the proposed group-based models yielding Spearman’s correlation of 0.55 ($p<0.05$) between the actual and predicted outcome. Results further demonstrate that systematic exposure to public speaking in VR can alleviate PSA in terms of both self-reported ($p<0.05$) and physiological ($p<0.05$) indices. Findings from this study will enable researchers to better understand antecedents and causes of PSA and lay the foundation for personalized adaptive feedback for PSA interventions. Megha Yadav, Ehsanul Haque Nirjhar, Kexin Feng, Amir H. Behzadan, Theodora Chaspari |
IEEE Trans. Affect. Comput. | 3 |
| 2021 | Knowledge- and Data-Driven Models of Multimodal Trajectories of Public Speaking Anxiety in Real and Virtual SettingsabstractPublic speaking skills are essential to professional success. Yet, public speaking anxiety (PSA) is considered one of the most common social phobias. Understanding PSA can help communication experts identify effective ways to treat this communication-based disorder. Existing works on PSA rely on self-reports and aggregate multimodal measures which do not capture the temporal variation in PSA. This paper examines temporal trajectories of acoustic and physiological measures throughout the public speaking encounter with real and virtual audiences, and aims to model those in both knowledge- and data-driven ways. Knowledge-driven models leverage theoretically-grounded patterns through fitting interpretable parametric functions to the corresponding signals. Data-driven models consider the functional nature of multimodal signals via functional principal component analysis. Results indicate that the parameters of the proposed models can successfully estimate individuals’ trait anxiety in both real-life and virtual reality settings, and suggest that models trained on data obtained in virtual public speaking stimuli are able to estimate levels of PSA in real-life. Ehsanul Haque Nirjhar, Amir H. Behzadan, Theodora Chaspari |
ICMI | 1 |
| 2021 | Investigating Trust in Human-Machine Learning Collaboration: A Pilot Study on Estimating Public Anxiety from SpeechabstractTrust is a key element in the development of effective collaborative relationships between humans and increasingly complex artificial intelligence (AI) systems. Here, we examine trust in AI in the context of a human-AI partnership that involves a joint decision making task for estimating levels of public speaking anxiety based on speech signals. The AI system is comprised of an explainable machine learning (ML) algorithm, that takes acoustic characteristics as input and outputs the estimate of public speaking anxiety levels, a local explanation about the most important features that contributed to the decision of each speech sample, and a global explanation about the most important features for the data overall. We analyze interactions between AI and human annotators with background in psychological sciences, and measure trust over time via the annotators’ agreement with the AI model and the annotators’ self-reports. We further examine factors of trust that are related to the characteristics of the human annotator and the ML algorithm. Results indicate that trust in AI depends on the openness level of the annotator and the importance level of input features. Findings from this study can provide guidelines to designing solutions that properly calibrate human trust in AI in collaborative human-AI tasks. Abdullah Aman Tutul, Ehsanul Haque Nirjhar, Theodora Chaspari |
ICMI | 2 |
| 2020 | Exploring Bio-Behavioral Signal Trajectories of State Anxiety During Public SpeakingabstractPublic speaking anxiety (PSA) is among the top social phobias in the world. Quantifying PSA in a reliable and unobtrusive manner can lay the foundation toward personalized and inexpensive technology-based interventions. Existing work for quantifying PSA often relies on self-reported measures and statistical aggregates of bio-behavioral indices, such as physiology and speech. Such aggregated bio-behavioral indices are not able to capture time-based trajectories of PSA variation, that can be very useful for better understanding and reliably predicting moments of anxiety. We tackle this problem by introducing temporal parametric models to quantify bio-behavioral trajectories of PSA throughout a public speaking encounter. Using data from 55 participants in a real-life public speaking task, the parameters of the proposed models are found to be significantly correlated with individuals' trait characteristics of general and communication-based anxiety, outperforming aggregate mean bio-behavioral measures. Ehsanul Haque Nirjhar, Amir H. Behzadan, Theodora Chaspari |
ICASSP | 1 |
| 2020 | Predicting the Effectiveness of Systematic Desensitization Through Virtual Reality for Mitigating Public Speaking AnxietyabstractPublic speaking is central to socialization in casual, professional, or academic settings. Yet, public speaking anxiety (PSA) is known to impact a considerable portion of the general population. This paper utilizes bio-behavioral indices captured from wearable devices to quantify the effectiveness of systematic exposure to virtual reality (VR) audiences for mitigating PSA. The effect of separate bio-behavioral features and demographic factors is studied, as well as the amount of necessary data from the VR sessions that can yield a reliable predictive model of the VR training effectiveness. Results indicate that acoustic and physiological reactivity during the VR exposure can reliably predict change in PSA before and after the training. With the addition of demographic features, both acoustic and physiological feature sets achieve improvements in performance. Finally, using bio-behavioral data from six to eight VR sessions can yield reliable prediction of PSA change. Findings of this study will enable researchers to better understand how bio-behavioral factors indicate improvements in PSA with VR training. Margaret von Ebers, Ehsanul Haque Nirjhar, Amir H. Behzadan, Theodora Chaspari |
ICMI | 2 |