Hanna Drimalla

dblp:188/1779 · DBLP profile ↗
← Back
7ranked-venue papers
1as first author
6since 2021 · last 2025
0000-0003-3783-7237ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 4 · 4 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Improving Autism Detection with Multimodal Behavioral Analysis
William Saakyan, Matthias Norden, Lola Eversmann, Simon Kirsch, Muyu Lin, Simón Guendelman, Isabel Dziobek, Hanna Drimalla
MICCAI (4)8
2025 Explainable AI for Audio and Visual Affective Computing: A Scoping Review
abstract
Affective computing often relies on audiovisual data to identify affective states from non-verbal signals, such as facial expressions and vocal cues. Since automatic affect recognition can be used in sensitive applications, such as healthcare and education, it is crucial to understand how models arrive at their decisions. Interpretability of machine learning models is the goal of the emerging research area of Explainable AI (explainable AI (XAI)). This scoping review aims to survey the field of audiovisual affective machine learning to identify how XAI is applied in this domain. We first provide an overview of XAI concepts relevant to affective computing. Next, following the recommended PRISMA guidelines, we perform a literature search in the ACM, IEEE, Web of Science and PubMed databases. After systematically reviewing 1190 articles, a final set of 65 papers is included in our analysis. We quantitatively summarize the scope, methods and evaluation of the XAI techniques used in the identified papers. Our findings show encouraging developments for using XAI to explain models in audiovisual affective computing, yet only a limited set of methods are used in the reviewed works. Following a critical discussion, we provide recommendations for incorporating interpretability in future work for affective machine learning.
David Johnson 0009, Olya Hakobyan, Jonas Paletschek, Hanna Drimalla
IEEE Trans. Affect. Comput.4
2024 A Paradigm to Investigate Social Signals of Understanding and Their Susceptibility to Stress
abstract
In human-machine explanation interactions, such as tutoring systems or customer support chatbots, it is important for the machine explainer to infer the human user's understanding. Nonverbal signals play an important role for expressing mental states like understanding and confusion in these interactions. However, an individual's expressions may vary depending on other factors. In cases where these factors are unknown, machine learning methods that infer understanding from nonverbal cues become unreliable. Stress for example has been shown to affect human expression, but it is not clear from the current research how stress affects the expression of understanding. To address this gap, we design a paradigm that induces understanding and con-fusion through game rule explanations. During the explanations, self-perceived understanding and confusion are annotated by the participants. A stress condition is also introduced to enable the investigation of changes in the expression of social signals under stress. We conducted a study to validate the stress induction and participants reported a statistically significant increase in stress during the stress condition compared to the neutral control condition. Additionally, feedback from participants shows that the paradigm is effective in inducing understanding and confusion. This paradigm paves the way for further studies investigating social signals of understanding to improve human-machine explanation interactions for varying contexts. Index Terms-Understanding, Nonverbal Social Signals, Stress Induction, Explanation, Machine Learning Bias.
Jonas Paletschek, Jan Bleimling, David Johnson 0009, Hanna Drimalla
ACII4
2024 Introducing the "Simulated Interaction Task for Children" (Kids-SIT): Recording and Analyzing Social Interaction Behavior
abstract
Figure 1: General procedure of the Kids-SIT research tool.A child participant follows a fully automated conversation scenario with pre-recorded videos of an actress.The actress initiates a dialogue on meal preferences, talking about her and asking the participant on their favorite and disliked foods one after the other.The actress empathically waits for responses and automatically continues the conversation in the following.For the whole time, the participant behavior is video recorded through the front camera.
Matthias Norden, William Saakyan, Nadine Vietmeier, Simone Kirst, Isabel Dziobek, Julia Asbrand, Hanna Drimalla
MUM7
2023 On Scalable and Interpretable Autism Detection from Social Interaction Behavior
abstract
Autism Spectrum Condition (ASC) is characterized by social interaction difficulties that can be challenging to assess objectively in the diagnostic process. In this paper, we evaluate the capability of using videos of a standardized social interaction to differentiate non-verbal behaviors of individuals with and without ASC. We collected a large video dataset consisting of 164 participants with ASC (n = 83) and neurotypical individuals (n = 81) who completed the computer-based Simulated Interaction Task (SIT) in different studies including lab and home settings. To classify individuals with and without ASC, we trained uni-and multimodal machine learning models based on different modalities such as facial expressions, gaze behavior, head pose and voice features. Our results indicate that a multimodal late fusion approach achieved the highest accuracy (74%). In the unimodal setting, classification based on facial expressions (accuracy 73%) and voice features (accuracy 70%) were most effective. An explainability analysis of the most relevant features for the facial expression model indicated that features from all emotional parts as well as from both the speaking and listening part of the interaction were informative. Based on our results, we developed a scalable online version of the SIT to collect diverse data on a large scale for the development of machine learning models that can differentiate between different clinical conditions. Our study highlights the potential of machine learning on videos of standardized social interactions in supporting clinical diagnosis and the objective and effective measurement of differences in social interaction behavior.
William Saakyan, Matthias Norden, Lola Herrmann, Simon Kirsch, Muyu Lin, Simón Guendelman, Isabel Dziobek, Hanna Drimalla
ACII8
2022 Automatic Detection of Subjective, Annotated and Physiological Stress Responses from Video Data
abstract
Machine-learning-based stress detection systems differ with respect to the ground truth used for training the algorithms. It is unclear how models trained on different facets of the stress reaction (e.g., biological, psychological, social) can be compared, interpreted and applied. In this study, we investigate the influence of the stress label on the performance of machine learning models trained on either vocal characteristics or facial expressions extracted from videos. We collected videos from 40 male participants while being exposed to the Trier Social Stress Test (TSST) and assessed self-reported, live observed, video-annotated and neuro-endocrinological stress levels. We train three standard machine learning models to separately predict different stress labels using either voice or facial cues. Analyzing the relationships of different stress facets we found that observers' annotations were significantly positively associated (live vs. video annotated,$\boldsymbol{\rho}_{\mathbf{s}}\ =\ . 53$). Similarly, the neuro-endocrinological stress indices correlated with each other (cortisol vs. sAA,$\boldsymbol{\rho}_{\mathbf{s}}$=.39). Machine learning experiments resulted in predictions that were positively associated with panel-annotated stress levels showing significantly stronger correlations in voice-based models$(\boldsymbol{\rho}_{\mathbf{s}=}.54\ \mathbf{v}\mathbf{s}. \boldsymbol{\rho}_{\mathbf{S}}=.30)$. Predictions of self-reported stress were positively related to ground truth values for face-based$(\boldsymbol{\rho}_{\mathbf{s}}$=.24) but not for voice-based models. There was no evidence for successful predictions of video-annotations or endocrinological stress levels in both settings. We provide evidence that machine learning models trained on different stress assessments perform differently and should be interpreted and applied accordingly. Implications and recommendations for future work on video-based stress detection are discussed.
Matthias Norden, Oliver T. Wolf, Lennart Lehmann, Katja Langer, Christoph Lippert, Hanna Drimalla
ACII6
2018 Detecting Autism by Analyzing a Simulated Social Interaction
Hanna Drimalla, Niels Landwehr, Irina Baskow, Behnoush Behnia, Stefan Roepke, Isabel Dziobek, Tobias Scheffer
ECML/PKDD (1)1