Juan Camilo Vásquez-Correa

dblp:172/6510 · DBLP profile ↗
← Back
39ranked-venue papers
15as first author
15since 2021 · last 2026
0000-0003-4946-9232ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 30 · 13 first-author · 8 since 2021Artificial intelligence and machine learning · 26 · 9 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Do modern speech LLMs and re-scoring techniques improve bilingual ASR performance for Basque and Spanish in domain-specific contexts?
abstract
This paper presents an extended evaluation of Vicomtech’s automatic speech recognition (ASR) systems developed for the Albayzín 2024 Bilingual Basque-Spanish Speech-to-Text (BBS-S2T) Challenge, a task focused on transcribing bilingual parliamentary recordings featuring frequent intra- and inter-sentential code-switching between Basque and Spanish. These recordings, drawn from Basque Parliament plenary sessions, pose significant challenges due to the abrupt language alternations, the limited availability of digital resources for Basque, and the absence of contextual and speaker information. The study incorporates additional analysis of state-of-the-art ASR architectures, namely Phi4-multimodal and CrisperWhisper, fine-tuned on the challenge dataset. Furthermore, the systems were evaluated on a complementary benchmark to assess model robustness. A detailed comparison of automatic hypothesis selection techniques, including both traditional n -gram and large language model (LLM)-based approaches, is also provided. Results demonstrate that optimal word error rate (WER) does not always correlate with the most accurate transcriptions, highlighting the complexity of evaluating ASR performance in code-switching scenarios. • Leading ASR systems have been assessed on a challenging bilingual domain. • Fusion and re-scoring frameworks were applied using statistical and large language models. • External language models prioritised meaning over transcription accuracy. • An extensive and dedicated analysis had been performed to further evaluate recognition errors. • Optimal WER may not reflect the top quality transcriptions in specific applications.
Ander González-Docasal, Juan Camilo Vásquez-Correa, Haritz Arzelus, Aitor Álvarez 0001, Santiago Andres Moreno-Acevedo
Comput. Speech Lang.2
2024 Representation learning strategies to model pathological speech: Effect of multiple spectral resolutions
abstract
This paper considers a representation learning strategy to model speech signals from patients with Parkinson’s disease, with the goal of predicting the presence of the disease, and evaluating the level of degradation of a patient’s speech. In particular, we propose a novel fusion strategy that combines wideband and narrowband spectral resolutions using a representation learning strategy based on autoencoders , called the multi-spectral autoencoder. The proposed model is able to classify the speech from Parkinson’s disease patients with accuracy up to 97%. The proposed model is also able to assess the dysarthria severity of Parkinson’s disease patients with a Spearman correlation up to 0.79. These results outperform those observed in literature where the same problem was addressed with the same corpus.
Gabriel F. Miller, Juan Camilo Vásquez-Correa, Juan Rafael Orozco-Arroyave, Elmar Nöth
Comput. Speech Lang.2
2024 Evaluation of effectiveness in conversations between humans and chatbots using parallel convolutional neural networks with multiple temporal resolutions
abstract
Abstract Chatbots enable the automation of several components in customer service and allow the support of multiple users. Despite their multiple advantages, due to the large amount of conversations generated by a chatbot, it is difficult to determine whether customer requests are well-addressed. For practical reasons, chatbot’s effectiveness is evaluated manually based upon a small sample (randomly chosen) of conversations or through self-reported user satisfaction. This procedure does not guarantee the correct evaluation of the service because the sample is generally not large enough and self-reports might be influenced by different external factors not directly associated to the chatbot’s functioning. This study proposes a methodology for automatic evaluation of chatbot effectiveness in real production environments. The analysis considers convolutional neural networks adapted for natural language processing, using two parallel convolutional layers to evaluate questions and answers independently. The proposed model also incorporates filters to extract features with multiple temporal resolution. This methodology is tested upon real conversations of chatbots that provide service to two different companies. The results are compared to baseline models based on classical techniques with different pre-trained word embedding models. According to our results, the proposed approach provides accuracies between 78.95% and 80.18%, which outperforms the best result of the baseline models by 2.9%.
Daniel Escobar-Grisales, Juan Camilo Vásquez-Correa, Juan Rafael Orozco-Arroyave
Multim. Tools Appl.2
2023 User State Modeling Based on the Arousal-Valence Plane: Applications in Customer Satisfaction and Health-Care
abstract
The acoustic analysis helps to discriminate emotions according to non-verbal information, while linguistics aims to capture verbal information from written sources. Acoustic and linguistic analyses can be addressed for different applications, where information related to emotions, mood, or affect are involved. The Arousal-Valence plane is commonly used to model emotional states in a multidimensional space. This study proposes a methodology focused on modeling the user’s state based on the Arousal-Valence plane in different scenarios. Acoustic and linguistic information are used as input to feed different deep learning architectures mainly based on convolutional and recurrent neural networks, which are trained to model the Arousal-Valence plane. The proposed approach is used for the evaluation of customer satisfaction in call-centers and for health-care applications in the assessment of depression in Parkinson’s disease and the discrimination of Alzheimer’s disease. F-scores of up to 0.89 are obtained for customer satisfaction, of up to 0.82 for depression in Parkinson’s patients, and of up to 0.80 for Alzheimer’s patients. The proposed approach confirms that there is information embedded in the Arousal-Valence plane that can be used for different purposes.
Paula Andrea Pérez-Toro, Juan Camilo Vásquez-Correa, Tobias Bocklet, Elmar Nöth, Juan Rafael Orozco-Arroyave
IEEE Trans. Affect. Comput.2
2022 The phonetic footprint of Parkinson's disease
Philipp Klumpp, Tomás Arias-Vergara, Juan Camilo Vásquez-Correa, Paula Andrea Pérez-Toro, Juan Rafael Orozco-Arroyave, Anton Batliner, Elmar Nöth
Comput. Speech Lang.3
2022 Empirical Mode Decomposition articulation feature extraction on Parkinson's Diadochokinesia
Alice Rueda, Juan Camilo Vásquez-Correa, Juan Rafael Orozco-Arroyave, Elmar Nöth, Sridhar Krishnan 0001
Comput. Speech Lang.2
2022 Depression assessment in people with Parkinson's disease: The combination of acoustic features and natural language processing
Paula Andrea Pérez-Toro, Tomás Arias-Vergara, Philipp Klumpp, Juan Camilo Vásquez-Correa, Maria Schuster, Elmar Nöth, Juan Rafael Orozco-Arroyave
Speech Commun.4
2021 Colombian Dialect Recognition Based on Information Extracted from Speech and Text Signals
abstract
Dialect recognition is useful in many industrial sectors, par-ticularly with the aim of allowing a better interaction between customers and providers. The core idea is to improve or customize marketing and customer service strategies, de-pending on the geographic location, birthplace and culture. This study proposes different models to automatically dis-criminate between two Colombian dialects: “Antioqueño” and “Bogotano”, to the best of our knowledge this is the first work of Colombian dialect recognition based on real conver-sations from customer service centers. The proposed strategy consists of independent analyses, using information from speech recordings and their corresponding transliterations. On the one hand, classical approaches are used to model speech including prosody features, Mel frequency cepstral coefficients and the mean Hilbert envelope coefficients. For text models, Word2Vec and bidirectional encoding represen-tations from transformer embeddings are considered. On the other hand, a deep learning approach is applied by considering convolutional neural networks, which are trained using spectrograms and embedding matrices for speech and text, respectively. The implemented deep learning models seem to be more promising than the classical ones for the addressed problem. Further experiments will be considered to validate this claim in a wider spectrum of methods.
Daniel Escobar-Grisales, Cristian D. Ríos-Urrego, Diego Alexander Lopez-Santander, Jeferson David Gallo-Aristizábal, Juan Camilo Vásquez-Correa, Elmar Nöth, Juan Rafael Orozco-Arroyave
ASRU5
2021 Acoustic and Linguistic Analyses to Assess Early-Onset and Genetic Alzheimer's Disease
abstract
The PSEN1-E280A or Paisa mutation is responsible for most of Early-Onset Alzheimer’s (EOA) disease cases in Colombia. It affects a large kindred of over 5000 members that present the same phenotype. The most common symptoms are related to language disorders, where speech fluency is also affected due to the difficulty to access semantic information intentionally. This study proposes the use of acoustic and linguistic methods to extract features from speech recordings and their transcriptions to discriminate people with conditions related to the Paisa mutation. We consider state-of-the-art word-embedding methods like Word2Vec and Bidirectional Encoder Representations from Transformer to process the transcripts. The speech signals are modeled by using traditional acoustic features and speaker embeddings. To the best of our knowledge, this is the first study focused on evaluating genetic Alzheimer’s and EOA using acoustics and linguistics.
Paula Andrea Pérez-Toro, Juan Camilo Vásquez-Correa, Tomás Arias-Vergara, Philipp Klumpp, M. Sierra-Castrillón, M. E. Roldán-López, David Aguillón, Liliana Hincapié-Henao, Carlos Tobon 0001, Tobias Bocklet, Maria Schuster, Juan Rafael Orozco-Arroyave, Elmar Nöth
ICASSP2
2021 End-2-End Modeling of Speech and Gait from Patients with Parkinson's Disease: Comparison Between High Quality Vs. Smartphone Data
abstract
Parkinson’s disease is a neurodegenerative disorder characterized by the presence of different motor impairments. Speech and gait signals have been analyzed to detect the presence of the disease and the severity in patients. However, most studies have been performed in controlled conditions using high quality data, which make those studies not suitable for a continuous at-home evaluation of the state of the patients. The developed technology should be evaluated in more realistic scenarios, for instance using smartphone data. We propose the use of state-of-the-art deep learning techniques to evaluate the speech and gait symptoms of patients. The proposed methods are evaluated in two scenarios to cover both high quality and smartphone data. The results indicate that it is possible to classify patients and healthy subjects with accuracies over 92% in both scenarios. The proposed methods are also promising to evaluate the severity of the speech symptoms and the global motor state of the patients.
Juan Camilo Vásquez-Correa, Tomás Arias-Vergara, Philipp Klumpp, Paula Andrea Pérez-Toro, Juan Rafael Orozco-Arroyave, Elmar Nöth
ICASSP1
2021 The Phonetic Footprint of Covid-19?
abstract
Against the background of the ongoing pandemic, this year’s Computational Paralinguistics Challenge featured a classification problem to detect Covid-19 from speech recordings. The presented approach is based on a phonetic analysis of speech samples, thus it enabled us not only to discriminate between Covid and non-Covid samples, but also to better understand how the condition influenced an individual’s speech signal. Our deep acoustic model was trained with datasets collected exclusively from healthy speakers. It served as a tool for segmentation and feature extraction on the samples from the challenge dataset. Distinct patterns were found in the embeddings of phonetic classes that have their place of articulation deep inside the vocal tract. We observed profound differences in classification results for development and test splits, similar to the baseline method. We concluded that, based on our phonetic findings, it was safe to assume that our classifier was able to reliably detect a pathological condition located in the respiratory tract. However, we found no evidence to claim that the system was able to discriminate between Covid-19 and other respiratory diseases.
Philipp Klumpp, Tobias Bocklet, Tomás Arias-Vergara, Juan Camilo Vásquez-Correa, Paula Andrea Pérez-Toro, Sebastian P. Bayerl, Juan Rafael Orozco-Arroyave, Elmar Nöth
Interspeech4
2021 Influence of the Interviewer on the Automatic Assessment of Alzheimer's Disease in the Context of the ADReSSo Challenge
abstract
Alzheimer’s Disease (AD) results from the progressive loss of neurons in the hippocampus, which affects the capability to produce coherent language. It affects lexical, grammatical, and semantic processes as well as speech fluency. This paper considers the analyses of speech and language for the assessment of AD in the context of the Alzheimer’s Dementia Recognition through Spontaneous Speech (ADReSSo) 2021 challenge. We propose to extract acoustic features such as X-vectors, prosody, and emotional embeddings as well as linguistic features such as perplexity, and word-embeddings. The data consist of speech recordings from AD patients and healthy controls. The transcriptions are obtained using a commercial automatic speech recognition system. We outperform baseline results on the test set, both for the classification and the Mini-Mental State Examination (MMSE) prediction. We achieved a classification accuracy of 80% and an RMSE of 4.56 in the regression. Additionally, we found strong evidence for the influence of the interviewer on classification results. In cross-validation on the training set, we get classification results of 85% accuracy using the combined speech of the interviewer and the participant. Using interviewer speech only we still get an accuracy of 78%. Thus, we provide strong evidence for interviewer influence on classification results.
Paula Andrea Pérez-Toro, Sebastian P. Bayerl, Tomás Arias-Vergara, Juan Camilo Vásquez-Correa, Philipp Klumpp, Maria Schuster, Elmar Nöth, Juan Rafael Orozco-Arroyave, Korbinian Riedhammer
Interspeech4
2021 On Modeling Glottal Source Information for Phonation Assessment in Parkinson's Disease
abstract
Parkinson's disease produces several motor symptoms, including different speech impairments that are known as hypokinetic dysarthria. Symptoms associated to dysarthria affect different dimensions of speech such as phonation, articulation, prosody, and intelligibility. Studies in the literature have mainly focused on the analysis of articulation and prosody because they seem to be the most prominent symptoms associated to dysarthria severity. However, phonation impairments also play a significant role to evaluate the global speech severity of Parkinson's patients. This paper proposes an extensive comparison of different methods to automatically evaluate the severity of specific phonation impairments in Parkinson's patients. The considered models include the computation of perturbation and glottal-based features, in addition to features extracted from a zero frequency filtered signals. We consider as well end-to-end models based on 1D CNNs, which are trained to learn features from the raw speech waveform, reconstructed glottal signals, and zero-frequency filtered signals. The results indicate that it is possible to automatically classify between speakers with low versus high phonation severity due to the presence of dysarthria and at the same time to evaluate the severity of the phonation impairments on a continuous scale, posed as a regression problem.
Juan Camilo Vásquez-Correa, Julian Fritsch, Juan Rafael Orozco-Arroyave, Elmar Nöth, Mathew Magimai-Doss
Interspeech1
2021 Multi-channel spectrograms for speech processing applications using deep learning methods
abstract
Abstract Time–frequency representations of the speech signals provide dynamic information about how the frequency component changes with time. In order to process this information, deep learning models with convolution layers can be used to obtain feature maps. In many speech processing applications, the time–frequency representations are obtained by applying the short-time Fourier transform and using single-channel input tensors to feed the models. However, this may limit the potential of convolutional networks to learn different representations of the audio signal. In this paper, we propose a methodology to combine three different time–frequency representations of the signals by computing continuous wavelet transform, Mel-spectrograms, and Gammatone spectrograms and combining then into 3D-channel spectrograms to analyze speech in two different applications: (1) automatic detection of speech deficits in cochlear implant users and (2) phoneme class recognition to extract phone-attribute features. For this, two different deep learning-based models are considered: convolutional neural networks and recurrent neural networks with convolution layers.
Tomás Arias-Vergara, Philipp Klumpp, Juan Camilo Vásquez-Correa, Elmar Nöth, Juan Rafael Orozco-Arroyave, Maria Schuster
Pattern Anal. Appl.3
2021 Transfer learning helps to improve the accuracy to classify patients with different speech disorders in different languages
Juan Camilo Vásquez-Correa, Cristian D. Ríos-Urrego, Tomás Arias-Vergara, Maria Schuster, Jan Rusz, Elmar Nöth, Juan Rafael Orozco-Arroyave
Pattern Recognit. Lett.1
2020 Comparison of User Models Based on GMM-UBM and I-Vectors for Speech, Handwriting, and Gait Assessment of Parkinson's Disease Patients
abstract
Parkinson's disease is a neurodegenerative disorder characterized by the presence of different motor impairments. Information from speech, handwriting, and gait signals have been considered to evaluate the neurological state of the patients. On the other hand, user models based on Gaussian mixture models - universal background models (GMMUBM) and i-vectors are considered the state-of-the-art in biometric applications like speaker verification because they are able to model specific speaker traits. This study introduces the use of GMM-UBM and i-vectors to evaluate the neurological state of Parkinson's patients using information from speech, handwriting, and gait. The results show the importance of different feature sets from each type of signal in the assessment of the neurological state of the patients.
Juan Camilo Vásquez-Correa, Tobias Bocklet, Juan Rafael Orozco-Arroyave, Elmar Nöth
ICASSP1
2020 Surgical Mask Detection with Deep Recurrent Phonetic Models
Philipp Klumpp, Tomás Arias-Vergara, Juan Camilo Vásquez-Correa, Paula Andrea Pérez-Toro, Florian Hönig, Elmar Nöth, Juan Rafael Orozco-Arroyave
INTERSPEECH3
2020 Parallel Representation Learning for the Classification of Pathological Speech: Studies on Parkinson's Disease and Cleft Lip and Palate
Juan Camilo Vásquez-Correa, Tomás Arias-Vergara, Maria Schuster, Juan Rafael Orozco-Arroyave, Elmar Nöth
Speech Commun.1
2019 Multi-channel Convolutional Neural Networks for Automatic Detection of Speech Deficits in Cochlear Implant Users
Tomás Arias-Vergara, Juan Camilo Vásquez-Correa, Sandra Gollwitzer, Juan Rafael Orozco-Arroyave, Maria Schuster, Elmar Nöth
CIARP2
2019 Convolutional Neural Networks and a Transfer Learning Strategy to Classify Parkinson's Disease from Speech in Three Different Languages
Juan Camilo Vásquez-Correa, Tomás Arias-Vergara, Cristian D. Ríos-Urrego, Maria Schuster, Jan Rusz, Juan Rafael Orozco-Arroyave, Elmar Nöth
CIARP1
2019 Articulation and Empirical Mode Decomposition Features in Diadochokinetic Exercises for the Speech Assessment of Parkinson's Disease Patients
Juan Camilo Vásquez-Correa, Cristian D. Ríos-Urrego, Alice Rueda, Juan Rafael Orozco-Arroyave, Sri Krishnan, Elmar Nöth
CIARP1
2019 Feature Space Visualization with Spatial Similarity Maps for Pathological Speech Data
Philipp Klumpp, Juan Camilo Vásquez-Correa, Tino Haderlein, Elmar Nöth
INTERSPEECH2
2019 Feature Representation of Pathophysiology of Parkinsonian Dysarthria
Alice Rueda, Juan Camilo Vásquez-Correa, Cristian D. Ríos-Urrego, Juan Rafael Orozco-Arroyave, Sridhar Krishnan 0001, Elmar Nöth
INTERSPEECH2
2019 Apkinson: A Mobile Solution for Multimodal Assessment of Patients with Parkinson's Disease
Juan Camilo Vásquez-Correa, Tomás Arias-Vergara, Philipp Klumpp, M. Strauss, Arne Küderle, Nils Roth, Sebastian P. Bayerl, Nicanor García, Paula Andrea Pérez-Toro, L. Felipe Parra-Gallego, Cristian D. Ríos-Urrego, Daniel Escobar-Grisales, Juan Rafael Orozco-Arroyave, Björn M. Eskofier, Elmar Nöth
INTERSPEECH1
2019 Phonet: A Tool Based on Gated Recurrent Neural Networks to Extract Phonological Posteriors from Speech
Juan Camilo Vásquez-Correa, Philipp Klumpp, Juan Rafael Orozco-Arroyave, Elmar Nöth
INTERSPEECH1
2019 Multimodal Assessment of Parkinson's Disease: A Deep Learning Approach
abstract
Parkinson's disease is a neurodegenerative disorder characterized by a variety of motor symptoms. Particularly, difficulties to start/stop movements have been observed in patients. From a technical/diagnostic point of view, these movement changes can be assessed by modeling the transitions between voiced and unvoiced segments in speech, the movement when the patient starts or stops a new stroke in handwriting, or the movement when the patient starts or stops the walking process. This study proposes a methodology to model such difficulties to start or to stop movements considering information from speech, handwriting, and gait. We used those transitions to train convolutional neural networks to classify patients and healthy subjects. The neurological state of the patients was also evaluated according to different stages of the disease (initial, intermediate, and advanced). In addition, we evaluated the robustness of the proposed approach when considering speech signals in three different languages: Spanish, German, and Czech. According to the results, the fusion of information from the three modalities is highly accurate to classify patients and healthy subjects, and it shows to be suitable to assess the neurological state of the patients in several stages of the disease. We also aimed to interpret the feature maps obtained from the deep learning architectures with respect to the presence or absence of the disease and the neurological state of the patients. As far as we know, this is one of the first works that considers multimodal information to assess Parkinson's disease following a deep learning approach.
Juan Camilo Vásquez-Correa, Tomás Arias-Vergara, Juan Rafael Orozco-Arroyave, Björn M. Eskofier, Jochen Klucken, Elmar Nöth
IEEE J. Biomed. Health Informatics1
2018 Unobtrusive Monitoring of Speech Impairments of Parkinson'S Disease Patients Through Mobile Devices
abstract
Parkinson's disease (PD) produces several speech impairments in the patients. Automatic classification of PD patients is performed considering speech recordings collected in noncontrolled acoustic conditions during normal phone calls in a unobtrusive way. A speech enhancement algorithm is applied to improve the quality of the signals. Two different classification approaches are considered: the classification of PD patients and healthy speakers and a multi-class experiment to classify patients in several stages of the disease. According to the results it is possible to classify PD patients and healthy controls with a AUe of up to 0.87. This work is a step forward to the development of telemonitoring systems to assess the speech of the patients.
Tomás Arias-Vergara, Juan Camilo Vásquez-Correa, Juan Rafael Orozco-Arroyave, Philipp Klumpp, Elmar Nöth
ICASSP2
2018 Multimodal I-vectors to Detect and Evaluate Parkinson's Disease
Nicanor García, Juan Camilo Vásquez-Correa, Juan Rafael Orozco-Arroyave, Elmar Nöth
INTERSPEECH2
2018 A Multitask Learning Approach to Assess the Dysarthria Severity in Patients with Parkinson's Disease
Juan Camilo Vásquez-Correa, Tomás Arias-Vergara, Juan Rafael Orozco-Arroyave, Elmar Nöth
INTERSPEECH1
2018 Speaker models for monitoring Parkinson's disease progression considering different communication channels and acoustic conditions
abstract
Symptoms of Parkinson's disease vary from patient to patient. Additionally, the progression of those symptoms also differs among patients. Most of the studies on the analysis of speech of people with Parkinson's disease do not consider such an individual variation. This paper presents a methodology for the automatic and individual monitoring of speech disorders developed by PD patients. The neurological state and dysarthria level of the patients are evaluated. The proposed system is based on individual speaker models which are created for each patient. Two different models are evaluated, the classical GMM–UBM and the i–vectors approach. These two methods are compared with respect to a baseline found with a traditional Support Vector Regressor. Different speech aspects (phonation, articulation, and prosody) are considered to model recordings of spontaneous speech and a read text. A multi-aspect coefficient is proposed with the aim of incorporating information from all of these speech aspects into a single measure. Two different scenarios are considered to assess a set with seven PD patients: (1) the longitudinal test set which consists of speech recordings captured in five recording sessions distributed from 2012 to 2016, and (2) the at-home test set which consists of speech recordings captured in the home of the same seven patients during 4 months (one day per month, four times per day). The UBM is trained with the recordings of 100 speakers (50 with Parkinson's disease and 50 healthy speakers) captured with controlled acoustic conditions and a professional audio-setting. With the aim of evaluating the suitability of the proposed approaches and the possibility of extending this kind of systems to remotely assess the speech of the patients, a total of five different communication channels (sound-proof booth, Skype®, Hangouts®, mobile phone, and land-line) are considered to train and test the system. Due to the reduced number of recording sessions in the longitudinal test set, the experiments that involved this set are evaluated with the Pearson's correlation. The experiments with the at-home test set are evaluated with the Spearman's correlation. The results estimating the dysarthria level of the patients in the at-home test set indicate a correlation of 0.55 with a modified version of the Frenchay Dysarthria Assessment scale when the GMM-UBM model is applied upon the Skype® recordings. The results in the longitudinal test set indicate a correlation of 0.77 using a model based on i-vectors with recordings captured in the sound-proof-booth. The evaluation of the neurological state of the patients in the longitudinal test set shows correlations of up to 0.55 with the Movement Disorder Society - Unified Parkinson's Disease Rating Scale also using models based on i-vectors created with Skype® recordings. These results suggest that the i–vector approach is suitable when the acoustic conditions among recording sessions differ (longitudinal test set). The GMM-UBM approach seems to be more suitable when the acoustic conditions do not change a lot among recording sessions (at-home test set). Particularly, the best results were obtained with the Skype® calls, which can be explained due to several preprocessing stages that this codec applies to the audio signals. In general, the results suggest that the proposed approaches are suitable for tele-monitoring the dysarthria level and the neurological state of PD patients.
Tomás Arias-Vergara, Juan Camilo Vásquez-Correa, Juan Rafael Orozco-Arroyave, Elmar Nöth
Speech Commun.2
2017 On the impact of non-modal phonation on phonological features
abstract
Different modes of vibration of the vocal folds contribute significantly to the voice quality. The neutral mode phonation, often used in a modal voice, is one against which the other modes can be contrastively described, also called non-modal phonations. This paper investigates the impact of non-modal phonation on phonological posteriors, the probabilities of phonological features inferred from the speech signal using a deep learning approach. Five different non-modal phonations are considered: falsetto, creaky, harshness, tense and breathiness. The impact of such non-modal phonation on phonological features, the Sound Patterns of English (SPE), is investigated in both speech analysis and synthesis tasks. We found that breathy and tense phonation impact the SPE features less, creaky phonation impacts the features moderately, and harsh and falsetto phonation impact the phonological features the most. We also report invariant and the most different SPE features impacted by non-modal phonation.
Milos Cernak, Elmar Nöth, Frank Rudzicz, Heidi Christensen, Juan Rafael Orozco-Arroyave, Raman Arora, Tobias Bocklet, Hamid R. Chinaei, Julius Hannink, Phani S. Nidadavolu, Juan Camilo Vásquez-Correa, Maria Yancheva, Alyssa Vann, Nikolai Vogler
ICASSP11
2017 Multi-view representation learning via gcca for multimodal analysis of Parkinson's disease
abstract
Information from different bio-signals such as speech, handwriting, and gait have been used to monitor the state of Parkinson's disease (PD) patients, however, all the multimodal bio-signals may not always be available. We propose a method based on multi-view representation learning via generalized canonical correlation analysis (GCCA) for learning a representation of features extracted from handwriting and gait that can be used as a complement to speech-based features. Three different problems are addressed: classification of PD patients vs. healthy controls, prediction of the neurological state of PD patients according to the UPDRS score, and the prediction of a modified version of the Frenchay dysarthria assessment (m-FDA). According to the results, the proposed approach is suitable to improve the results in the addressed problems, specially in the prediction of the UPDRS, and m-FDA scores.
Juan Camilo Vásquez-Correa, Juan Rafael Orozco-Arroyave, Raman Arora, Elmar Nöth, Najim Dehak, Heidi Christensen, Frank Rudzicz, Tobias Bocklet, Milos Cernak, Hamid R. Chinaei, Julius Hannink, Phani S. Nidadavolu, Maria Yancheva, Alyssa Vann, Nikolai Vogler
ICASSP1
2017 Effect of acoustic conditions on algorithms to detect Parkinson's disease from speech
abstract
Automatic detection of Parkinson's disease (PD) from speech is a basic step towards computer-aided tools supporting the diagnosis and monitoring of the disease. Although several methods have been proposed, their applicability to real-world situations is still unclear. In particular, the effect of acoustic conditions is not well understood. In this paper, the effects on the accuracy of five different methods to detect PD from speech are evaluated. Among the considered conditions, background noise produces the worst effect, while dynamic compression or some speech codecs can even have a marginal positive impact. We also consider, for the first time in this context, the problem of mismatches, i.e., when train/test acoustic conditions are different, and observe a high negative impact on all considered methods. Overall, this study is a step forward in performing a continuous monitoring of the neurological state of the patients in non-controlled acoustic conditions.
Juan Camilo Vásquez-Correa, Joan Serrà, Juan Rafael Orozco-Arroyave, Jesús Francisco Vargas-Bonilla, Elmar Nöth
ICASSP1
2017 Apkinson - A Mobile Monitoring Solution for Parkinson's Disease
Philipp Klumpp, Thomas Janu, Tomás Arias-Vergara, Juan Camilo Vásquez-Correa, Juan Rafael Orozco-Arroyave, Elmar Nöth
INTERSPEECH4
2017 Convolutional Neural Network to Model Articulation Impairments in Patients with Parkinson's Disease
Juan Camilo Vásquez-Correa, Juan Rafael Orozco-Arroyave, Elmar Nöth
INTERSPEECH1
2017 Characterisation of voice quality of Parkinson's disease using differential phonological posterior features
Milos Cernak, Juan Rafael Orozco-Arroyave, Frank Rudzicz, Heidi Christensen, Juan Camilo Vásquez-Correa, Elmar Nöth
Comput. Speech Lang.5
2016 Towards an automatic monitoring of the neurological state of Parkinson's patients from speech
abstract
The suitability of articulation measures and speech intelligibility is evaluated to estimate the neurological state of patients with Parkinson's disease (PD). A set of measures recently introduced to model the articulatory capability of PD patients is considered. Additionally, the speech intelligibility in terms of the word accuracy obtained from the Google® speech recognizer is included. Recordings of patients in three different languages are considered: Spanish, German, and Czech. Additionally, the proposed approach is tested on data recently used in the INTERSPEECH 2015 Computational Paralinguistics Challenge. According to the results, it is possible to estimate the neurological state of PD patients from speech with a Spearman's correlation of up to 0.72 with respect to the evaluations performed by neurologist experts.
Juan Rafael Orozco-Arroyave, Juan Camilo Vásquez-Correa, Florian Hönig, Julián D. Arias-Londoño, Jesús Francisco Vargas-Bonilla, Sabine Skodda, Jan Rusz, Elmar Nöth
ICASSP2
2016 Parkinson's Disease Progression Assessment from Speech Using GMM-UBM
Tomás Arias-Vergara, Juan Camilo Vásquez-Correa, Juan Rafael Orozco-Arroyave, Jesús Francisco Vargas-Bonilla, Elmar Nöth
INTERSPEECH2
2015 Automatic detection of parkinson's disease from continuous speech recorded in non-controlled noise conditions
Juan Camilo Vásquez-Correa, Tomás Arias-Vergara, Juan Rafael Orozco-Arroyave, Jesús Francisco Vargas-Bonilla, Julián D. Arias-Londoño, Elmar Nöth
INTERSPEECH1