Philipp Klumpp

dblp:204/6913 · DBLP profile ↗
← Back
22ranked-venue papers
6as first author
16since 2021 · last 2024
0000-0002-7531-1693ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 18 · 4 first-author · 12 since 2021Artificial intelligence and machine learning · 17 · 6 first-author · 12 since 2021
YearPublicationVenuePosition
2024 Multilingual Speech and Language Analysis for the Assessment of Mild Cognitive Impairment: Outcomes from the Taukadial Challenge
Paula Andrea Pérez-Toro, Tomás Arias-Vergara, Philipp Klumpp, Tobias Weise, Maria Schuster, Elmar Nöth, Juan Rafael Orozco-Arroyave, Andreas K. Maier
INTERSPEECH3
2024 Speaker- and Text-Independent Estimation of Articulatory Movements and Phoneme Alignments from Speech
Tobias Weise, Philipp Klumpp, Kubilay Can Demir, Paula Andrea Pérez-Toro, Maria Schuster, Elmar Nöth, Björn Heismann, Andreas K. Maier, Seung-Hee Yang
INTERSPEECH2
2023 Transferring Quantified Emotion Knowledge for the Detection of Depression in Alzheimer's Disease Using Forestnets
abstract
Progressive loss of memory is the most known symptom of Alzheimer’s Disease (AD); however, it also affects other cognitive skills and leads to depression symptoms. This paper presents a transfer learning strategy for automatically detecting AD and depression in AD patients using acoustic information and ForestNet, an artificial neural network that allows computing the contribution of a set of features to a model’s decision. The methodology consists of training ForestNet with a dataset commonly used for emotion recognition; then, we fine-tune the pre-trained model to detect AD and depression in AD. We trained the models with several acoustic features commonly used for emotion and AD applications. Unweighted average recalls of up to 0.87 were achieved to classify the disease and up to 0.82 to detect depression in AD. Our results indicate that the information obtained from the Arousal Valence plane may be suitable for discriminating and analyzing depression in AD.
Paula Andrea Pérez-Toro, Dalia Rodríguez-Salas, Tomás Arias-Vergara, Sebastian P. Bayerl, Philipp Klumpp, Korbinian Riedhammer, Maria Schuster, Elmar Nöth, Andreas K. Maier, Juan Rafael Orozco-Arroyave
ICASSP5
2022 Wav2vec behind the Scenes: How end2end Models learn Phonetics
Teena tom Dieck, Paula Andrea Pérez-Toro, Tomas Arias, Elmar Nöth, Philipp Klumpp
INTERSPEECH5
2022 Alzheimer's Detection from English to Spanish Using Acoustic and Linguistic Embeddings
Paula Andrea Pérez-Toro, Philipp Klumpp, Abner Hernandez, Tomas Arias, Patricia Lillo, Andrea Slachevsky, Adolfo M. García, Maria Schuster, Andreas K. Maier, Elmar Nöth, Juan Rafael Orozco-Arroyave
INTERSPEECH2
2022 CoachLea: an Android Application to Evaluate the Speech Production and Perception of Children with Hearing Loss
P. Schäfer, Paula Andrea Pérez-Toro, Philipp Klumpp, Juan Rafael Orozco-Arroyave, Elmar Nöth, Andreas K. Maier, A. Abad, Maria Schuster, Tomás Arias-Vergara
INTERSPEECH3
2022 Disentangled Latent Speech Representation for Automatic Pathological Intelligibility Assessment
abstract
Speech intelligibility assessment plays an important role in the therapy of patients suffering from pathological speech disorders.Automatic and objective measures are desirable to assist therapists in their traditionally subjective and labor-intensive assessments.In this work, we investigate a novel approach for obtaining such a measure using the divergence in disentangled latent speech representations of a parallel utterance pair, obtained from a healthy reference and a pathological speaker.Experiments on an English database of Cerebral Palsy patients, using all available utterances per speaker, show high and significant correlation values (R = -0.9)with subjective intelligibility measures, while having only minimal deviation (±0.01) across four different reference speaker pairs.We also demonstrate the robustness of the proposed method (R = -0.89deviating ±0.02 over 1000 iterations) by considering a significantly smaller amount of utterances per speaker.Our results are among the first to show that disentangled speech representations can be used for automatic pathological speech intelligibility assessment, resulting in a reference speaker pair invariant method, applicable in scenarios with only few utterances available.
Tobias Weise, Philipp Klumpp, Andreas K. Maier, Elmar Nöth, Björn Heismann, Maria Schuster, Seung-Hee Yang
INTERSPEECH2
2022 Common Phone: A Multilingual Dataset for Robust Acoustic Modelling
abstract
Current state of the art acoustic models can easily comprise more than 100 million parameters. This growing complexity demands larger training datasets to maintain a decent generalization of the final decision function. An ideal dataset is not necessarily large in size, but large with respect to the amount of unique speakers, utilized hardware and varying recording conditions. This enables a machine learning model to explore as much of the domain-specific input space as possible during parameter estimation. This work introduces Common Phone, a gender-balanced, multilingual corpus recorded from more than 76.000 contributors via Mozilla’s Common Voice project. It comprises around 116 hours of speech enriched with automatically generated phonetic segmentation. A Wav2Vec 2.0 acoustic model was trained with the Common Phone to perform phonetic symbol recognition and validate the quality of the generated phonetic annotation. The architecture achieved a PER of 18.1 % on the entire test set, computed with all 101 unique phonetic symbols, showing slight differences between the individual languages. We conclude that Common Phone provides sufficient variability and reliable phonetic annotation to help bridging the gap between research and application of acoustic models.
Philipp Klumpp, Tomás Arias-Vergara, Paula Andrea Pérez-Toro, Elmar Nöth, Juan Rafael Orozco-Arroyave
LREC1
2022 The phonetic footprint of Parkinson's disease
Philipp Klumpp, Tomás Arias-Vergara, Juan Camilo Vásquez-Correa, Paula Andrea Pérez-Toro, Juan Rafael Orozco-Arroyave, Anton Batliner, Elmar Nöth
Comput. Speech Lang.1
2022 Depression assessment in people with Parkinson's disease: The combination of acoustic features and natural language processing
Paula Andrea Pérez-Toro, Tomás Arias-Vergara, Philipp Klumpp, Juan Camilo Vásquez-Correa, Maria Schuster, Elmar Nöth, Juan Rafael Orozco-Arroyave
Speech Commun.3
2021 Acoustic and Linguistic Analyses to Assess Early-Onset and Genetic Alzheimer's Disease
abstract
The PSEN1-E280A or Paisa mutation is responsible for most of Early-Onset Alzheimer’s (EOA) disease cases in Colombia. It affects a large kindred of over 5000 members that present the same phenotype. The most common symptoms are related to language disorders, where speech fluency is also affected due to the difficulty to access semantic information intentionally. This study proposes the use of acoustic and linguistic methods to extract features from speech recordings and their transcriptions to discriminate people with conditions related to the Paisa mutation. We consider state-of-the-art word-embedding methods like Word2Vec and Bidirectional Encoder Representations from Transformer to process the transcripts. The speech signals are modeled by using traditional acoustic features and speaker embeddings. To the best of our knowledge, this is the first study focused on evaluating genetic Alzheimer’s and EOA using acoustics and linguistics.
Paula Andrea Pérez-Toro, Juan Camilo Vásquez-Correa, Tomás Arias-Vergara, Philipp Klumpp, M. Sierra-Castrillón, M. E. Roldán-López, David Aguillón, Liliana Hincapié-Henao, Carlos Tobon 0001, Tobias Bocklet, Maria Schuster, Juan Rafael Orozco-Arroyave, Elmar Nöth
ICASSP4
2021 End-2-End Modeling of Speech and Gait from Patients with Parkinson's Disease: Comparison Between High Quality Vs. Smartphone Data
abstract
Parkinson’s disease is a neurodegenerative disorder characterized by the presence of different motor impairments. Speech and gait signals have been analyzed to detect the presence of the disease and the severity in patients. However, most studies have been performed in controlled conditions using high quality data, which make those studies not suitable for a continuous at-home evaluation of the state of the patients. The developed technology should be evaluated in more realistic scenarios, for instance using smartphone data. We propose the use of state-of-the-art deep learning techniques to evaluate the speech and gait symptoms of patients. The proposed methods are evaluated in two scenarios to cover both high quality and smartphone data. The results indicate that it is possible to classify patients and healthy subjects with accuracies over 92% in both scenarios. The proposed methods are also promising to evaluate the severity of the speech symptoms and the global motor state of the patients.
Juan Camilo Vásquez-Correa, Tomás Arias-Vergara, Philipp Klumpp, Paula Andrea Pérez-Toro, Juan Rafael Orozco-Arroyave, Elmar Nöth
ICASSP3
2021 The Phonetic Footprint of Covid-19?
abstract
Against the background of the ongoing pandemic, this year’s Computational Paralinguistics Challenge featured a classification problem to detect Covid-19 from speech recordings. The presented approach is based on a phonetic analysis of speech samples, thus it enabled us not only to discriminate between Covid and non-Covid samples, but also to better understand how the condition influenced an individual’s speech signal. Our deep acoustic model was trained with datasets collected exclusively from healthy speakers. It served as a tool for segmentation and feature extraction on the samples from the challenge dataset. Distinct patterns were found in the embeddings of phonetic classes that have their place of articulation deep inside the vocal tract. We observed profound differences in classification results for development and test splits, similar to the baseline method. We concluded that, based on our phonetic findings, it was safe to assume that our classifier was able to reliably detect a pathological condition located in the respiratory tract. However, we found no evidence to claim that the system was able to discriminate between Covid-19 and other respiratory diseases.
Philipp Klumpp, Tobias Bocklet, Tomás Arias-Vergara, Juan Camilo Vásquez-Correa, Paula Andrea Pérez-Toro, Sebastian P. Bayerl, Juan Rafael Orozco-Arroyave, Elmar Nöth
Interspeech1
2021 Influence of the Interviewer on the Automatic Assessment of Alzheimer's Disease in the Context of the ADReSSo Challenge
abstract
Alzheimer’s Disease (AD) results from the progressive loss of neurons in the hippocampus, which affects the capability to produce coherent language. It affects lexical, grammatical, and semantic processes as well as speech fluency. This paper considers the analyses of speech and language for the assessment of AD in the context of the Alzheimer’s Dementia Recognition through Spontaneous Speech (ADReSSo) 2021 challenge. We propose to extract acoustic features such as X-vectors, prosody, and emotional embeddings as well as linguistic features such as perplexity, and word-embeddings. The data consist of speech recordings from AD patients and healthy controls. The transcriptions are obtained using a commercial automatic speech recognition system. We outperform baseline results on the test set, both for the classification and the Mini-Mental State Examination (MMSE) prediction. We achieved a classification accuracy of 80% and an RMSE of 4.56 in the regression. Additionally, we found strong evidence for the influence of the interviewer on classification results. In cross-validation on the training set, we get classification results of 85% accuracy using the combined speech of the interviewer and the participant. Using interviewer speech only we still get an accuracy of 78%. Thus, we provide strong evidence for interviewer influence on classification results.
Paula Andrea Pérez-Toro, Sebastian P. Bayerl, Tomás Arias-Vergara, Juan Camilo Vásquez-Correa, Philipp Klumpp, Maria Schuster, Elmar Nöth, Juan Rafael Orozco-Arroyave, Korbinian Riedhammer
Interspeech5
2021 Wearables-based multi-task gait and activity segmentation using recurrent neural networks
Christine Martindale, Vincent Christlein, Philipp Klumpp, Björn M. Eskofier
Neurocomputing3
2021 Multi-channel spectrograms for speech processing applications using deep learning methods
abstract
Abstract Time–frequency representations of the speech signals provide dynamic information about how the frequency component changes with time. In order to process this information, deep learning models with convolution layers can be used to obtain feature maps. In many speech processing applications, the time–frequency representations are obtained by applying the short-time Fourier transform and using single-channel input tensors to feed the models. However, this may limit the potential of convolutional networks to learn different representations of the audio signal. In this paper, we propose a methodology to combine three different time–frequency representations of the signals by computing continuous wavelet transform, Mel-spectrograms, and Gammatone spectrograms and combining then into 3D-channel spectrograms to analyze speech in two different applications: (1) automatic detection of speech deficits in cochlear implant users and (2) phoneme class recognition to extract phone-attribute features. For this, two different deep learning-based models are considered: convolutional neural networks and recurrent neural networks with convolution layers.
Tomás Arias-Vergara, Philipp Klumpp, Juan Camilo Vásquez-Correa, Elmar Nöth, Juan Rafael Orozco-Arroyave, Maria Schuster
Pattern Anal. Appl.2
2020 Surgical Mask Detection with Deep Recurrent Phonetic Models
Philipp Klumpp, Tomás Arias-Vergara, Juan Camilo Vásquez-Correa, Paula Andrea Pérez-Toro, Florian Hönig, Elmar Nöth, Juan Rafael Orozco-Arroyave
INTERSPEECH1
2019 Feature Space Visualization with Spatial Similarity Maps for Pathological Speech Data
Philipp Klumpp, Juan Camilo Vásquez-Correa, Tino Haderlein, Elmar Nöth
INTERSPEECH1
2019 Apkinson: A Mobile Solution for Multimodal Assessment of Patients with Parkinson's Disease
Juan Camilo Vásquez-Correa, Tomás Arias-Vergara, Philipp Klumpp, M. Strauss, Arne Küderle, Nils Roth, Sebastian P. Bayerl, Nicanor García, Paula Andrea Pérez-Toro, L. Felipe Parra-Gallego, Cristian D. Ríos-Urrego, Daniel Escobar-Grisales, Juan Rafael Orozco-Arroyave, Björn M. Eskofier, Elmar Nöth
INTERSPEECH3
2019 Phonet: A Tool Based on Gated Recurrent Neural Networks to Extract Phonological Posteriors from Speech
Juan Camilo Vásquez-Correa, Philipp Klumpp, Juan Rafael Orozco-Arroyave, Elmar Nöth
INTERSPEECH2
2018 Unobtrusive Monitoring of Speech Impairments of Parkinson'S Disease Patients Through Mobile Devices
abstract
Parkinson's disease (PD) produces several speech impairments in the patients. Automatic classification of PD patients is performed considering speech recordings collected in noncontrolled acoustic conditions during normal phone calls in a unobtrusive way. A speech enhancement algorithm is applied to improve the quality of the signals. Two different classification approaches are considered: the classification of PD patients and healthy speakers and a multi-class experiment to classify patients in several stages of the disease. According to the results it is possible to classify PD patients and healthy controls with a AUe of up to 0.87. This work is a step forward to the development of telemonitoring systems to assess the speech of the patients.
Tomás Arias-Vergara, Juan Camilo Vásquez-Correa, Juan Rafael Orozco-Arroyave, Philipp Klumpp, Elmar Nöth
ICASSP4
2017 Apkinson - A Mobile Monitoring Solution for Parkinson's Disease
Philipp Klumpp, Thomas Janu, Tomás Arias-Vergara, Juan Camilo Vásquez-Correa, Juan Rafael Orozco-Arroyave, Elmar Nöth
INTERSPEECH1