VLDB 2026 Research / reviewers in the wild / expert
Maria Schuster
dblp:02/336
· DBLP profile ↗
27ranked-venue papers
1as first author
16since 2021 · last 2024
0000-0001-9122-7478ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 24 · 1 first-author · 14 since 2021Artificial intelligence and machine learning · 20 · 12 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Towards Interpretability of Automatic Phoneme Analysis in Cleft Lip and Palate SpeechabstractCleft Lip and Palate ranks among the most common congenital abnormalities and significantly influences speech articulation, resulting in varying phonemic impacts. In a clinical context, a detailed diagnosis is carried out by time-consuming perceptual evaluations. We use perceptual ratings of different articulatory modifications on phoneme-level as ground-truth and propose a system based on wav2vec 2.0, trained to the downstream task of classifying phonemic criteria as a multi-class and multi-label problem. The system is trained for detection on utterance level, without the usage of phoneme labels. To gain a clearer understanding of which areas of the speech signal have the greatest impact on classification, we assess the extent to which our system aligns with expert ratings at the phoneme level. Additionally, we examine which specific phonemes play a decisive role in determining the final classification of the labeled criteria. The results show that salient phonemes marked by experts contribute remarkably greater to the classification of the correct class using feature relevance explanation methods. To the best of our knowledge, this is the first study incorporating various utterance-level articulatory modifications classification and phoneme-level interpretation, offering a more comprehensive understanding for potential clinical applications. Ilja Baumann, Dominik Wagner 0002, Maria Schuster, Elmar Nöth, Tobias Bocklet |
ICASSP | 3 |
| 2024 | Contrastive Learning Approach for Assessment of Phonological Precision in Patients with Tongue Cancer Using MRI DataabstractMagnetic Resonance Imaging (MRI) allows analyzing speech production by capturing high-resolution images of the dynamic processes in the vocal tract. In clinical applications, combining MRI with synchronized speech recordings leads to improved patient outcomes, especially if a phonological-based approach is used for assessment. However, when audio signals are unavailable, the recognition accuracy of sounds is decreased when using only MRI data. We propose a contrastive learning approach to improve the detection of phonological classes from MRI data when acoustic signals are not available at inference time. We demonstrate that frame-wise recognition of phonological classes improves from an f1 of 0.74 to 0.85 when the contrastive loss approach is implemented. Furthermore, we show the utility of our approach in the clinical application of using such phonological classes to assess speech disorders in patients with tongue cancer, yielding promising results in the recognition task. Tomás Arias-Vergara, Paula Andrea Pérez-Toro, Xiaofeng Liu 0001, Fangxu Xing, Maureen Stone 0001, Jiachen Zhuo, Jerry L. Prince, Maria Schuster, Elmar Nöth, Jonghye Woo, Andreas K. Maier |
INTERSPEECH | 8 |
| 2024 | Towards Self-Attention Understanding for Automatic Articulatory Processes Analysis in Cleft Lip and Palate SpeechabstractCleft lip and palate (CLP) speech presents unique challenges for automatic phoneme analysis due to its distinct acoustic characteristics and articulatory anomalies. We perform phoneme analysis in CLP speech using a pre-trained wav2vec 2.0 model with a multi-head self-attention classification module to capture long-range dependencies within the speech signal, thereby enabling better contextual understanding of phoneme sequences. We demonstrate the effectiveness of our approach in the classification of various articulatory processes in CLP speech. Furthermore, we investigate the interpretability of self-attention to gain insights into the model’s understanding of CLP speech characteristics. Our findings highlight the potential of the selfattention mechanisms for improving automatic phoneme analysis in CLP speech, paving the way for enhanced diagnostics, adding interpretability for therapists and affected patients. Ilja Baumann, Dominik Wagner 0002, Maria Schuster, Korbinian Riedhammer, Elmar Nöth, Tobias Bocklet |
INTERSPEECH | 3 |
| 2024 | Multilingual Speech and Language Analysis for the Assessment of Mild Cognitive Impairment: Outcomes from the Taukadial Challenge
Paula Andrea Pérez-Toro, Tomás Arias-Vergara, Philipp Klumpp, Tobias Weise, Maria Schuster, Elmar Nöth, Juan Rafael Orozco-Arroyave, Andreas K. Maier |
INTERSPEECH | 5 |
| 2024 | Speaker- and Text-Independent Estimation of Articulatory Movements and Phoneme Alignments from Speech
Tobias Weise, Philipp Klumpp, Kubilay Can Demir, Paula Andrea Pérez-Toro, Maria Schuster, Elmar Nöth, Björn Heismann, Andreas K. Maier, Seung-Hee Yang |
INTERSPEECH | 5 |
| 2023 | Transferring Quantified Emotion Knowledge for the Detection of Depression in Alzheimer's Disease Using ForestnetsabstractProgressive loss of memory is the most known symptom of Alzheimer’s Disease (AD); however, it also affects other cognitive skills and leads to depression symptoms. This paper presents a transfer learning strategy for automatically detecting AD and depression in AD patients using acoustic information and ForestNet, an artificial neural network that allows computing the contribution of a set of features to a model’s decision. The methodology consists of training ForestNet with a dataset commonly used for emotion recognition; then, we fine-tune the pre-trained model to detect AD and depression in AD. We trained the models with several acoustic features commonly used for emotion and AD applications. Unweighted average recalls of up to 0.87 were achieved to classify the disease and up to 0.82 to detect depression in AD. Our results indicate that the information obtained from the Arousal Valence plane may be suitable for discriminating and analyzing depression in AD. Paula Andrea Pérez-Toro, Dalia Rodríguez-Salas, Tomás Arias-Vergara, Sebastian P. Bayerl, Philipp Klumpp, Korbinian Riedhammer, Maria Schuster, Elmar Nöth, Andreas K. Maier, Juan Rafael Orozco-Arroyave |
ICASSP | 7 |
| 2023 | Measuring Phonological Precision in Children with Cleft Lip and Palate
Tomás Arias-Vergara, Elizabeth Londoño-Mora, Paula Andrea Pérez-Toro, Maria Schuster, Elmar Nöth, Juan Rafael Orozco-Arroyave, Andreas K. Maier |
INTERSPEECH | 4 |
| 2023 | Automatic Assessment of Alzheimer's across Three Languages Using Speech and Language Features
Paula Andrea Pérez-Toro, Tomás Arias-Vergara, Franziska Braun, Florian Hönig, Carlos Tobon 0001, David Aguillón, Francisco Lopera, Liliana Hincapié-Henao, Maria Schuster, Korbinian Riedhammer, Andreas K. Maier, Elmar Nöth, Juan Rafael Orozco-Arroyave |
INTERSPEECH | 9 |
| 2022 | Alzheimer's Detection from English to Spanish Using Acoustic and Linguistic Embeddings
Paula Andrea Pérez-Toro, Philipp Klumpp, Abner Hernandez, Tomas Arias, Patricia Lillo, Andrea Slachevsky, Adolfo M. García, Maria Schuster, Andreas K. Maier, Elmar Nöth, Juan Rafael Orozco-Arroyave |
INTERSPEECH | 8 |
| 2022 | CoachLea: an Android Application to Evaluate the Speech Production and Perception of Children with Hearing Loss
P. Schäfer, Paula Andrea Pérez-Toro, Philipp Klumpp, Juan Rafael Orozco-Arroyave, Elmar Nöth, Andreas K. Maier, A. Abad, Maria Schuster, Tomás Arias-Vergara |
INTERSPEECH | 8 |
| 2022 | Disentangled Latent Speech Representation for Automatic Pathological Intelligibility AssessmentabstractSpeech intelligibility assessment plays an important role in the therapy of patients suffering from pathological speech disorders.Automatic and objective measures are desirable to assist therapists in their traditionally subjective and labor-intensive assessments.In this work, we investigate a novel approach for obtaining such a measure using the divergence in disentangled latent speech representations of a parallel utterance pair, obtained from a healthy reference and a pathological speaker.Experiments on an English database of Cerebral Palsy patients, using all available utterances per speaker, show high and significant correlation values (R = -0.9)with subjective intelligibility measures, while having only minimal deviation (±0.01) across four different reference speaker pairs.We also demonstrate the robustness of the proposed method (R = -0.89deviating ±0.02 over 1000 iterations) by considering a significantly smaller amount of utterances per speaker.Our results are among the first to show that disentangled speech representations can be used for automatic pathological speech intelligibility assessment, resulting in a reference speaker pair invariant method, applicable in scenarios with only few utterances available. Tobias Weise, Philipp Klumpp, Andreas K. Maier, Elmar Nöth, Björn Heismann, Maria Schuster, Seung-Hee Yang |
INTERSPEECH | 6 |
| 2022 | Depression assessment in people with Parkinson's disease: The combination of acoustic features and natural language processing
Paula Andrea Pérez-Toro, Tomás Arias-Vergara, Philipp Klumpp, Juan Camilo Vásquez-Correa, Maria Schuster, Elmar Nöth, Juan Rafael Orozco-Arroyave |
Speech Commun. | 5 |
| 2021 | Acoustic and Linguistic Analyses to Assess Early-Onset and Genetic Alzheimer's DiseaseabstractThe PSEN1-E280A or Paisa mutation is responsible for most of Early-Onset Alzheimer’s (EOA) disease cases in Colombia. It affects a large kindred of over 5000 members that present the same phenotype. The most common symptoms are related to language disorders, where speech fluency is also affected due to the difficulty to access semantic information intentionally. This study proposes the use of acoustic and linguistic methods to extract features from speech recordings and their transcriptions to discriminate people with conditions related to the Paisa mutation. We consider state-of-the-art word-embedding methods like Word2Vec and Bidirectional Encoder Representations from Transformer to process the transcripts. The speech signals are modeled by using traditional acoustic features and speaker embeddings. To the best of our knowledge, this is the first study focused on evaluating genetic Alzheimer’s and EOA using acoustics and linguistics. Paula Andrea Pérez-Toro, Juan Camilo Vásquez-Correa, Tomás Arias-Vergara, Philipp Klumpp, M. Sierra-Castrillón, M. E. Roldán-López, David Aguillón, Liliana Hincapié-Henao, Carlos Tobon 0001, Tobias Bocklet, Maria Schuster, Juan Rafael Orozco-Arroyave, Elmar Nöth |
ICASSP | 11 |
| 2021 | Influence of the Interviewer on the Automatic Assessment of Alzheimer's Disease in the Context of the ADReSSo ChallengeabstractAlzheimer’s Disease (AD) results from the progressive loss of neurons in the hippocampus, which affects the capability to produce coherent language. It affects lexical, grammatical, and semantic processes as well as speech fluency. This paper considers the analyses of speech and language for the assessment of AD in the context of the Alzheimer’s Dementia Recognition through Spontaneous Speech (ADReSSo) 2021 challenge. We propose to extract acoustic features such as X-vectors, prosody, and emotional embeddings as well as linguistic features such as perplexity, and word-embeddings. The data consist of speech recordings from AD patients and healthy controls. The transcriptions are obtained using a commercial automatic speech recognition system. We outperform baseline results on the test set, both for the classification and the Mini-Mental State Examination (MMSE) prediction. We achieved a classification accuracy of 80% and an RMSE of 4.56 in the regression. Additionally, we found strong evidence for the influence of the interviewer on classification results. In cross-validation on the training set, we get classification results of 85% accuracy using the combined speech of the interviewer and the participant. Using interviewer speech only we still get an accuracy of 78%. Thus, we provide strong evidence for interviewer influence on classification results. Paula Andrea Pérez-Toro, Sebastian P. Bayerl, Tomás Arias-Vergara, Juan Camilo Vásquez-Correa, Philipp Klumpp, Maria Schuster, Elmar Nöth, Juan Rafael Orozco-Arroyave, Korbinian Riedhammer |
Interspeech | 6 |
| 2021 | Multi-channel spectrograms for speech processing applications using deep learning methodsabstractAbstract Time–frequency representations of the speech signals provide dynamic information about how the frequency component changes with time. In order to process this information, deep learning models with convolution layers can be used to obtain feature maps. In many speech processing applications, the time–frequency representations are obtained by applying the short-time Fourier transform and using single-channel input tensors to feed the models. However, this may limit the potential of convolutional networks to learn different representations of the audio signal. In this paper, we propose a methodology to combine three different time–frequency representations of the signals by computing continuous wavelet transform, Mel-spectrograms, and Gammatone spectrograms and combining then into 3D-channel spectrograms to analyze speech in two different applications: (1) automatic detection of speech deficits in cochlear implant users and (2) phoneme class recognition to extract phone-attribute features. For this, two different deep learning-based models are considered: convolutional neural networks and recurrent neural networks with convolution layers. Tomás Arias-Vergara, Philipp Klumpp, Juan Camilo Vásquez-Correa, Elmar Nöth, Juan Rafael Orozco-Arroyave, Maria Schuster |
Pattern Anal. Appl. | 6 |
| 2021 | Transfer learning helps to improve the accuracy to classify patients with different speech disorders in different languages
Juan Camilo Vásquez-Correa, Cristian D. Ríos-Urrego, Tomás Arias-Vergara, Maria Schuster, Jan Rusz, Elmar Nöth, Juan Rafael Orozco-Arroyave |
Pattern Recognit. Lett. | 4 |
| 2020 | Parallel Representation Learning for the Classification of Pathological Speech: Studies on Parkinson's Disease and Cleft Lip and Palate
Juan Camilo Vásquez-Correa, Tomás Arias-Vergara, Maria Schuster, Juan Rafael Orozco-Arroyave, Elmar Nöth |
Speech Commun. | 3 |
| 2019 | Multi-channel Convolutional Neural Networks for Automatic Detection of Speech Deficits in Cochlear Implant Users
Tomás Arias-Vergara, Juan Camilo Vásquez-Correa, Sandra Gollwitzer, Juan Rafael Orozco-Arroyave, Maria Schuster, Elmar Nöth |
CIARP | 5 |
| 2019 | Convolutional Neural Networks and a Transfer Learning Strategy to Classify Parkinson's Disease from Speech in Three Different Languages
Juan Camilo Vásquez-Correa, Tomás Arias-Vergara, Cristian D. Ríos-Urrego, Maria Schuster, Jan Rusz, Juan Rafael Orozco-Arroyave, Elmar Nöth |
CIARP | 4 |
| 2019 | Phone-Attribute Posteriors to Evaluate the Speech of Cochlear Implant UsersabstractPeople with pre- and postlingual onset of deafness, i.e, age of occurrence of hearing loss, often present speech production\nproblems even after hearing rehabilitation by cochlear implantation. In this paper, the speech of 20 prelinguals (aged between 18 to 71 years old), 20 postlinguals (aged between 33 to 78 years old) and 20 healthy control (aged between 31 to 62 years old) German native speakers are analyzed considering phone-attribute features extracted with pre-trained Deep Neural Networks. Speech signals are analyzed with reference to the manner of articulation of consonants according to 5 groups: nasals, sibilants, fricatives, voiced-stops, and voiceless-stops. According to the results, it is possible to detect alterations in the consonant production of CI users when compared with healthy speakers. A comprehensive evaluation of speech changes of CI users will help in the rehabilitation after deafening. Tomás Arias-Vergara, Juan Rafael Orozco-Arroyave, Milos Cernak, Sandra Gollwitzer, Maria Schuster, Elmar Nöth |
INTERSPEECH | 5 |
| 2016 | A generalized procedure for analyzing sustained and dynamic vocal fold vibrations from laryngeal high-speed videos using phonovibrograms
Jakob Unger, Maria Schuster, Dietmar J. Hecker, Bernhard Schick, Jörg Lohscheller |
Artif. Intell. Medicine | 2 |
| 2009 | A microphone-independent visualization technique for speech disordersabstractIn this paper we introduce a novel method for the visualization of speech disorders. We demonstrate the method with disordered speech and a control group. However, both groups were recorded using two different microphones. The projection of the patient data using a single microphone yields significant correlations between the coordinates on the map and certain criteria of the disorder which were perceptually rated. However, projection of data from multiple microphones reduces this correlation. Usually, the acoustical mismatch between the microphones is greater than the mismatch between the speakers, i.e., not the disorders but the microphones form clusters in the visualization. Based on an extension of the Sammon mapping, we are able to create a map which projects the same speakers onto the same position even if multiple microphones are used. Furthermore, our method also restores the correlation between the map coordinates and the perceptual assessment. Index Terms: visualization, robustness, speech processing. Andreas K. Maier, Stefan Wenhardt, Tino Haderlein, Maria Schuster, Elmar Nöth |
INTERSPEECH | 4 |
| 2009 | PEAKS - A system for the automatic evaluation of voice and speech disorders
Andreas K. Maier, Tino Haderlein, Ulrich Eysholdt, Frank Rosanowski, Anton Batliner, Maria Schuster, Elmar Nöth |
Speech Commun. | 6 |
| 2008 | Automatic evaluation of characteristic speech disorders in children with cleft lip and palateabstractAbstract This paper discusses the automatic evaluation of speech of chil-dren with cleft lip and palate (CLP). CLP speech shows specialcharacteristics such as hypernasality, backing, and weakeningof plosives. In total ve criteria were subjectively assessed byan experienced speech expert on the phone level. This subjec-tive evaluation was used as a gold standard to train a classi-cation system. The automatic system achieves recognition re-sults on frame, phone, and word level of up to 75.8% CL. Onspeaker level signicant and high correlations between the sub-jective evaluation and the automatic system of up to 0.89 areobtained. Index Terms : pathologic speech, speech assessment, pronun-ciation scoring, children’s speech 1. Introduction Cleft Lip and Palate (CLP) is the most common malformationof the head. It constitutes almost two-thirds of the major facialdefects and almost 80% of all orofacial clefts [1]. Its prevalencediffers in different populations from 1 in 400 to 500 newborns inAsians to 1 in 1500 to 2000 in African Americans. The preva-lence in Caucasians is 1 in 750 to 900 births [2, 3].In clinical practice, articulation disorders are mainly eval-uated by subjective tools. The simplest method is the audi-tive perception, mostly performed by a speech therapist. Pre-vious studies have shown that experience is an important fac-tor that inuences the subjective estimation of speech disorderswhich leads to inaccurate evaluation by persons with only fewyears of experience as speech therapist [4]. Until now, objectivemeans exist only for quantitative measurements of nasal emis-sions [5, 6, 7] and for the detection of secondary voice disorders[8]. But other specic articulation disorders in CLP cannot besufciently quantied.In this paper, we present a new technical procedure for themeasurement and evaluation of specic speech disorders andcompare the results obtained with subjective ratings of an expe-rienced speech therapist. Andreas K. Maier, Florian Hönig, Christian Hacker, Maria Schuster, Elmar Nöth |
INTERSPEECH | 4 |
| 2007 | Towards robust automatic evaluation of pathologic telephone speechabstractFor many aspects of speech therapy an objective evaluation of the intelligibility of a patient's speech is needed. We investigate the evaluation of the intelligibility of speech by means of automatic speech recognition. Previous studies have shown that measures like word accuracy are consistent with human experts' ratings. To ease the patient's burden, it is highly desirable to conduct the assessment via phone. However, the telephone channel influences the quality of the speech signal which negatively affects the results. To reduce inaccuracies, we propose a combination of two speech recognizers. Experiments on two sets of pathological speech show that the combination results in consistent improvements in the correlation between the automatic evaluation and the ratings by human experts. Furthermore, the approach leads to reductions of 10% and 25% of the maximum error of the intelligibility measure. Korbinian Riedhammer, Georg Stemmer, Tino Haderlein, Maria Schuster, Frank Rosanowski, Elmar Nöth, Andreas K. Maier |
ASRU | 4 |
| 2007 | Automatic scoring of the intelligibility in patients with cancer of the oral cavityabstractAfter surgical treatment of cancer of the oral cavity patients often suffer from functional restrictions such as speech disorders.In this paper we present a novel approach to assess the outcome of the treatment w.r.t. the intelligibility of the patient using the result of an automatic speech recognition system.The word recognition rate was taken as intelligibility score.Compared to four speech experts this method yields results that are as good as the best speech expert compared to the other experts.The correlation between our system and the mean opinion of the experts is .92.Furthermore we show that our system has better performance than the average expert and is more reliable. Andreas K. Maier, Maria Schuster, Anton Batliner, Elmar Nöth, Emeka Nkenke |
INTERSPEECH | 2 |
| 2005 | Can you Understand him? Let's Look at his Word Accuracy - Automatic Evaluation of Tracheoesophageal SpeechabstractTracheoesophageal (TE) speech is a possibility to restore the ability to speak after laryngectomy. TE speech often shows low intelligibility. An objective means to determine and quantify the intelligibility does not exist until now and an automation of this procedure is desirable. We used a speech recognizer trained on normal, non-pathologic voices. We compared intelligibility scores for TE speech from five experienced raters with the word accuracy (WA) of our speech recognizer. A correlation coefficient of -0.84 shows that WA can be a good indicator of intelligibility for pathologic voices. An outlook for future work is presented. Maria Schuster, Elmar Nöth, Tino Haderlein, Stefan Steidl, Anton Batliner, Frank Rosanowski |
ICASSP (1) | 1 |