EDBT 2026 Demo / reviewers in the wild / expert
Michael Neumann 0001
dblp:82/7836-1
· DBLP profile ↗
25ranked-venue papers
9as first author
18since 2021 · last 2025
0000-0002-1023-9061ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 21 · 9 first-author · 16 since 2021Artificial intelligence and machine learning · 20 · 6 first-author · 16 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Multimodal Speech-Based Biomarkers Outperform the ALS Functional Rating Scale in Predicting Individual Disease Progression in ALS
Hardik Kothare, Michael Neumann 0001, Vikram Ramanarayanan |
INTERSPEECH | 2 |
| 2025 | An Exploration of Interpretable Deep Learning Models for the Assessment of Mild Cognitive ImpairmentabstractEarly diagnosis and intervention are crucial for mild cognitive impairment (MCI), as MCI often progresses to more severe neurodegenerative conditions. In this study, we explore utilizing deep learning for MCI detection without loosing the interpretability provided by feature-based approaches. We used a dataset consisting of 90 MCI patients and 91 controls collected via a remote assessment platform and analyzed the participants' spontaneous speech responses to the Patient Report of Problems (PROP) which asks patients to report their most bothersome general health problems. The proposed deep neural network, which features a bottleneck layer including 13 interpretable symptom domains, achieved an AUC of 0.62, thereby outperforming a set of feature-based classifiers while ensuring interpretability due to the bottleneck layer. We further illustrated the model's interpretability by examining how the predicted PROP domains influence final predictions using Shapley values. Emma C. L. Leschly, Oliver Roesler, Michael Neumann 0001, Jackson Liscombe, Abhishek Hosamath, Lakshmi Arbatti, Line Harder Clemmensen, Melanie Ganz-Benjaminsen, Vikram Ramanarayanan |
INTERSPEECH | 3 |
| 2025 | Multimodal Speech, Language and Orofacial Analysis for Remote Assessment of Positive, Negative and Cognitive Symptoms in Schizophrenia
Michael Neumann 0001, Hardik Kothare, Beverly Insel, Anzalee Khan, Danyah Nadim, Jean-Pierre Lindenmayer, Vikram Ramanarayanan |
INTERSPEECH | 1 |
| 2025 | Accessible Real-time Eye-gaze Tracking for Neurocognitive Health Assessment: A Multimodal Web-based Approach
Daniel Tisdale, Jackson Liscombe, David Pautler, Michael Neumann 0001, Vikram Ramanarayanan |
INTERSPEECH | 4 |
| 2024 | How Consistent are Speech-Based Biomarkers in Remote Tracking of ALS Disease Progression Across Languages? A Case Study of English and DutchabstractPrevious work has demonstrated the utility of speech-based digital biomarkers for remotely tracking longitudinal progression in people with Amyotrophic Lateral Sclerosis (pALS). Here, we investigate the responsiveness of these biomarkers across languages for consistency. We collected audiovisual data using a cloud-based multimodal dialogue platform, where pALS interacted with a virtual guide to perform several speaking exercises. We automatically extracted speech, linguistic and orofacial metrics from 143 English-speaking pALS (36 bulbar onset, 107 non-bulbar onset) and 26 Dutch-speaking pALS (10 bulbar, 16 non-bulbar onset). We used growth curve models to estimate the trajectory of these metrics over time. We observe that for most of these metrics, English-speaking pALS and Dutch-speaking pALS follow similar trajectories, i.e. the slopes are not statistically different from each other, demonstrating the potential of such speech-based biomarkers for remote monitoring across languages. Hardik Kothare, Michael Neumann 0001, Cathy Zhang, Jackson Liscombe, Jordi W. J. van Unnik, Lianne C. M. Botman, Leonard H. van den Berg, Ruben P. A van Eijk, Vikram Ramanarayanan |
INTERSPEECH | 2 |
| 2024 | Multimodal Digital Biomarkers for Longitudinal Tracking of Speech Impairment Severity in ALS: An Investigation of Clinically Important Differences
Michael Neumann 0001, Hardik Kothare, Jackson Liscombe, Emma C. L. Leschly, Oliver Roesler, Vikram Ramanarayanan |
INTERSPEECH | 1 |
| 2024 | Towards Scalable Remote Assessment of Mild Cognitive Impairment Via Multimodal Dialog
Oliver Roesler, Jackson Liscombe, Michael Neumann 0001, Hardik Kothare, Abhishek Hosamath, Lakshmi Arbatti, Doug Habberstad, Christiane Suendermann-Oeft, Meredith Bartlett, Cathy Zhang, Nikhil Sukhdev, Kolja Wilms, Anusha Badathala, Sandrine Istas, Steve Ruhmel, Bryan Hansen, Madeline Hannan, David Henley, Arthur W. Wallace, Ira Shoulson, David Suendermann-Oeft, Vikram Ramanarayanan |
INTERSPEECH | 3 |
| 2023 | Responsiveness, Sensitivity and Clinical Utility of Timing-Related Speech Biomarkers for Remote Monitoring of ALS Disease Progressionabstract= 94). We further evaluated the sensitivity of speech metrics in tracking disease progression in pALS while their ALSFRS-R speech score remained unchanged at 3 out of a total possible score of 4. We observed that timing-related speech metrics showed significant longitudinal changes even after accounting for learning effects. The findings of this study have the potential to inform disease prognosis and functional outcomes of clinical trials. Hardik Kothare, Michael Neumann 0001, Jackson Liscombe, Jordan R. Green, Vikram Ramanarayanan |
INTERSPEECH | 2 |
| 2023 | A Multimodal Investigation of Speech, Text, Cognitive and Facial Video Features for Characterizing Depression With and Without Medication
Michael Neumann 0001, Hardik Kothare, Doug Habberstad, Vikram Ramanarayanan |
INTERSPEECH | 1 |
| 2023 | Combining Multiple Multimodal Speech Features into an Interpretable Index Score for Capturing Disease Progression in Amyotrophic Lateral SclerosisabstractMultiple speech biomarkers have been shown to carry useful information regarding Amyotrophic Lateral Sclerosis (ALS) pathology. We propose a two-step framework to compute optimal linear combinations (indexes) of these biomarkers that are more discriminative and noise-robust than the individual markers, which is important for clinical care and pharmaceutical trial applications. First, we use a hierarchical clustering based method to select representative speech metrics from a dataset comprising 143 people with ALS and 135 age- and sex-matched healthy controls. Second, we analyze three methods of index computation that optimize linear discriminability, Youden Index, and sparsity of logistic regression model weights, respectively, and evaluate their performance with 5-fold cross validation. We find that the proposed indexes are generally more discriminative of bulbar vs non-bulbar onset in ALS than their individual component metrics as well as an equally-weighted baseline. Michael Neumann 0001, Hardik Kothare, Vikram Ramanarayanan |
INTERSPEECH | 1 |
| 2023 | When Words Speak Just as Loudly as Actions: Virtual Agent Based Remote Health Assessment Integrating What Patients Say with What They Do
Vikram Ramanarayanan, David Pautler, Lakshmi Arbatti, Abhishek Hosamath, Michael Neumann 0001, Hardik Kothare, Oliver Roesler, Jackson Liscombe, Andrew Cornish, Doug Habberstad, Vanessa Richter, David Suendermann-Oeft, Ira Shoulson |
INTERSPEECH | 5 |
| 2023 | Remote Assessment for ALS using Multimodal Dialog Agents: Data Quality, Feasibility and Task ComplianceabstractWe investigate the feasibility, task compliance and audiovisual data quality of a multimodal dialog-based solution for remote assessment of Amyotrophic Lateral Sclerosis (ALS). 53 people with ALS and 52 healthy controls interacted with Tina, a cloud-based conversational agent, in performing speech tasks designed to probe various aspects of motor speech function while their audio and video was recorded. We rated a total of 250 recordings for audio/video quality and participant task compliance, along with the relative frequency of different issues observed. We observed excellent compliance (98%) and audio (95.2%) and visual quality rates (84.8%), resulting in an overall yield of 80.8% recordings that were both compliant and of high quality. Furthermore, recording quality and compliance were not affected by level of speech severity and did not differ significantly across end devices. These findings support the utility of dialog systems for remote monitoring of speech in ALS. Vanessa Richter, Michael Neumann 0001, Jordan R. Green, Brian Richburg, Oliver Roesler, Hardik Kothare, Vikram Ramanarayanan |
INTERSPEECH | 2 |
| 2023 | Ethical Awareness in Paralinguistics: A Taxonomy of ApplicationsabstractSince the end of the last century, the automatic processing of paralinguistics has been investigated widely and put into practice in many applications, on wearables, smartphones, and computers. In this contribution, we address ethical awareness for paralinguistic applications, by establishing taxonomies for data representations, system designs for and a typology of applications, and users/test sets and subject areas. These are related to an “ethical grid” consisting of the most relevant ethical cornerstones, based on principalism. The characteristics of and the interdependencies between these taxonomies are described and exemplified. This makes it possible to assess more or less critical “ethical constellations.” To the best of our knowledge, this is the first attempt of its kind. Anton Batliner, Michael Neumann 0001, Felix Burkhardt, Alice Baird, Sarina Meyer, Ngoc Thang Vu, Björn W. Schuller |
Int. J. Hum. Comput. Interact. | 2 |
| 2022 | Statistical and clinical utility of multimodal dialogue-based speech and facial metrics for Parkinson's disease assessment
Hardik Kothare, Michael Neumann 0001, Jackson Liscombe, Oliver Roesler, William Burke, Andrew Exner, Sandy Snyder, Andrew Cornish, Doug Habberstad, David Pautler, David Suendermann-Oeft, Jessica Huber, Vikram Ramanarayanan |
INTERSPEECH | 2 |
| 2021 | "It seemed like an annoying woman": On the Perception and Ethical Considerations of Affective Language in Text-Based Conversational AgentsabstractPrevious research has found that task-oriented conversational agents are perceived more positively by users when they provide information in an empathetic manner compared to a plain, emotionless information exchange.However, users' perception and ethical considerations related to a dialog systems' response language style have received comparatively little attention in the field of human-computer interaction.To bridge this gap, we explored these ethical implications through a scenario-based user study.127 participants interacted with one of three variants of an affective, task-oriented conversational agent, each variant providing responses in a different language style.After the interaction, participants filled out a survey about their feelings during the experiment and their perception of various aspects of the chatbot.Based on statistical and qualitative analysis of the responses, we found language style played an important role in how humanlike participants perceived a dialog agent as well as how likable.Language style also had a direct effect on how users perceived the use of personal pronouns 'I' and 'You' and how they projected gender onto the chatbot.Finally, we identify and discuss ethical implications.In particular we focus on what factors/stereotypes influenced participants' impressions of gender, and what trade-offs a more human-like chatbot brings. Lindsey Vanderlyn, Gianna Weber, Michael Neumann 0001, Dirk Väth, Sarina Meyer, Ngoc Thang Vu |
CoNLL | 3 |
| 2021 | Investigating the Interplay Between Affective, Phonatory and Motoric Subsystems in Autism Spectrum Disorder Using a Multimodal Dialogue AgentabstractAbstract We explore the utility of an on-demand multimodal conversational platform in extracting speech and facial metrics in children with Autism Spectrum Disorder (ASD). We investigate the extent to which these metrics correlate with objective clinical measures, particularly as they pertain to the interplay be-tween the affective, phonatory and motoric subsystems. 22 participants diagnosed with ASD engaged with a virtual agent in conversational affect production tasks designed to elicit facial and vocal affect. We found significant correlations between vocal pitch and loudness extracted by our platform during these tasks and accuracy in recognition of facial and vocal affect, as-sessed via the Diagnostic Analysis of Nonverbal Accuracy-2 (DANVA-2) neuropsychological task. We also found significant correlations between jaw kinematic metrics extracted using our platform and motor speed of the dominant hand assessed via a standardised neuropsychological finger tapping task. These findings offer preliminary evidence for the usefulness of these audiovisual analytic metrics and could help us better model the interplay between different physiological subsystems in individuals with ASD. Hardik Kothare, Vikram Ramanarayanan, Oliver Roesler, Michael Neumann 0001, Jackson Liscombe, William Burke, Andrew Cornish, Doug Habberstad, Alaa Sakallah, Sara Markuson, Seemran Kansara, Afik Faerman, Yasmine Bensidi-Slimane, Laura Fry, Saige Portera, David Suendermann-Oeft, David Pautler, Carly Demopoulos |
Interspeech | 4 |
| 2021 | Investigating the Utility of Multimodal Conversational Technology and Audiovisual Analytic Measures for the Assessment and Monitoring of Amyotrophic Lateral Sclerosis at ScaleabstractWe propose a cloud-based multimodal dialog platform for the remote assessment and monitoring of Amyotrophic Lateral Sclerosis (ALS) at scale. This paper presents our vision, technology setup, and an initial investigation of the efficacy of the various acoustic and visual speech metrics automatically extracted by the platform. 82 healthy controls and 54 people with ALS (pALS) were instructed to interact with the platform and completed a battery of speaking tasks designed to probe the acoustic, articulatory, phonatory, and respiratory aspects of their speech. We find that multiple acoustic (rate, duration, voicing) and visual (higher order statistics of the jaw and lip) speech metrics show statistically significant differences between controls, bulbar symptomatic and bulbar pre-symptomatic patients. We report on the sensitivity and specificity of these metrics using five-fold cross-validation. We further conducted a LASSO-LARS regression analysis to uncover the relative contributions of various acoustic and visual features in predicting the severity of patients' ALS (as measured by their self-reported ALSFRS-R scores). Our results provide encouraging evidence of the utility of automatically extracted audiovisual analytics for scalable remote patient assessment and monitoring in ALS. Michael Neumann 0001, Oliver Roesler, Jackson Liscombe, Hardik Kothare, David Suendermann-Oeft, David Pautler, Indu Navar, Aria Anvar, Jochen Kumm, Raquel Norel, Ernest Fraenkel, Alexander V. Sherman, James D. Berry, Gary L. Pattee, Jun Wang 0037, Jordan R. Green, Vikram Ramanarayanan |
Interspeech | 1 |
| 2021 | Investigations on audiovisual emotion recognition in noisy conditionsabstractIn this paper we explore audiovisual emotion recognition under noisy acoustic conditions with a focus on speech features. We attempt to answer the following research questions: (i) How does speech emotion recognition perform on noisy data? and (ii) To what extend does a multimodal approach improve the accuracy and compensate for potential performance degradation at different noise levels? We present an analytical investigation on two emotion datasets with superimposed noise at different signal-to-noise ratios, comparing three types of acoustic features. Visual features are incorporated with a hybrid fusion approach: The first neural network layers are separate modality-specific ones, followed by at least one shared layer before the final prediction. The results show a significant performance decrease when a model trained on clean audio is applied to noisy data and that the addition of visual features alleviates this effect. Michael Neumann 0001, Ngoc Thang Vu |
SLT | 1 |
| 2020 | ClaVis: An Interactive Visual Comparison System for ClassifiersabstractWe propose ClaVis, a visual analytics system for comparative analysis of classification models. ClaVis allows users to visually compare the performance and behavior of tens to hundreds of classifiers trained with different hyperparameter configurations. Our approach is plugin-based and classifier-agnostic and allows users to add their own datasets and classifier implementations. It provides multiple visualizations, including a multivariate ranking, a similarity map, a scatterplot that reveals correlations between parameters and scores, and a training history chart. We demonstrate the effectivity of our approach in multiple case studies for training classification models in the domain of natural language processing. Frank Heyen, Tanja Munz-Körner, Michael Neumann 0001, Daniel Ortega, Ngoc Thang Vu, Daniel Weiskopf, Michael Sedlmair |
AVI | 3 |
| 2020 | Toward Remote Patient Monitoring of Speech, Video, Cognitive and Respiratory Biomarkers Using Multimodal Dialog Technology
Vikram Ramanarayanan, Oliver Roesler, Michael Neumann 0001, David Pautler, Doug Habberstad, Andrew Cornish, Hardik Kothare, Vignesh Murali, Jackson Liscombe, Dirk Schnelle-Walka, Patrick L. Lange, David Suendermann-Oeft |
INTERSPEECH | 3 |
| 2019 | Improving Speech Emotion Recognition with Unsupervised Representation Learning on Unlabeled SpeechabstractIn this paper we present our findings on how representation learning on large unlabeled speech corpora can be beneficially utilized for speech emotion recognition (SER). Prior work on representation learning for SER mostly focused on the relatively small emotional speech datasets without making use of additional unlabeled speech data. We show that integrating representations learnt by an unsupervised autoencoder into a CNN-based emotion classifier improves the recognition accuracy. To gain insights about what those models learn, we analyze visualizations of the different representations using t-distributed neighbor embeddings (t-SNE). We evaluate our approach on IEMOCAP and MSP-IMPROV by means of within- and cross-corpus testing. Michael Neumann 0001, Ngoc Thang Vu |
ICASSP | 1 |
| 2019 | CycleGAN-Based Emotion Style Transfer as Data Augmentation for Speech Emotion Recognition
Fang Bao, Michael Neumann 0001, Ngoc Thang Vu |
INTERSPEECH | 2 |
| 2019 | NEMSI: A Multimodal Dialog System for Screening of Neurological or Mental ConditionsabstractWe present NEMSI, a cloud-based multimodal dialog system designed to have naturalistic interactions with individuals for the purpose of screening neurological or mental conditions. The system has been used by thousands of people capturing audio and video responses to open-ended questions and structured health surveys. David Suendermann-Oeft, Amanda Robinson 0002, Andrew Cornish, Doug Habberstad, David Pautler, Dirk Schnelle-Walka, Franziska Haller, Jackson Liscombe, Michael Neumann 0001, Mike Merrill, Oliver Roesler, Renko Geffarth |
IVA | 9 |
| 2018 | CRoss-lingual and Multilingual Speech Emotion Recognition on English and FrenchabstractResearch on multilingual speech emotion recognition faces the problem that most available speech corpora differ from each other in important ways, such as annotation methods or interaction scenarios. These inconsistencies complicate building a multilingual system. We present results for cross-lingual and multilingual emotion recognition on English and French speech data with similar characteristics in terms of interaction (human-human conversations). Further, we explore the possibility of fine-tuning a pre-trained cross-lingual model with only a small number of samples from the target language, which is of great interest for low-resource languages. To gain more insights in what is learned by the deployed convolutional neural network, we perform an analysis on the attention mechanism inside the network. Michael Neumann 0001, Ngoc Thang Vu |
ICASSP | 1 |
| 2017 | Attentive Convolutional Neural Network Based Speech Emotion Recognition: A Study on the Impact of Input Features, Signal Length, and Acted SpeechabstractSpeech emotion recognition is an important and challenging task in the realm of human-computer interaction. Prior work proposed a variety of models and feature sets for training a system. In this work, we conduct extensive experiments using an attentive convolutional neural network with multi-view learning objective function. We compare system performance using different lengths of the input signal, different types of acoustic features and different types of emotion speech (improvised/scripted). Our experimental results on the Interactive Emotional Motion Capture (IEMOCAP) database reveal that the recognition performance strongly depends on the type of speech data independent of the choice of input features. Furthermore, we achieved state-of-the-art results on the improvised speech data of IEMOCAP. Michael Neumann 0001, Ngoc Thang Vu |
INTERSPEECH | 1 |