Jan Rusz

dblp:50/10325 · DBLP profile ↗
← Back
20ranked-venue papers
0as first author
7since 2021 · last 2023
0000-0002-1036-3054ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 17 · 6 since 2021Artificial intelligence and machine learning · 14 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2023 Federated Learning for Secure Development of AI Models for Parkinson's Disease Detection Using Speech from Different Languages
abstract
Parkinson's disease (PD) is a neurological disorder impacting a person's speech. Among automatic PD assessment methods, deep learning models have gained particular interest. Recently, the community has explored cross-pathology and cross-language models which can improve diagnostic accuracy even further. However, strict patient data privacy regulations largely prevent institutions from sharing patient speech data with each other. In this paper, we employ federated learning (FL) for PD detection using speech signals from 3 real-world language corpora of German, Spanish, and Czech, each from a separate institution. Our results indicate that the FL model outperforms all the local models in terms of diagnostic accuracy, while not performing very differently from the model based on centrally combined training sets, with the advantage of not requiring any data sharing among collaborators. This will simplify inter-institutional collaborations, resulting in enhancement of patient outcomes.
Soroosh Tayebi Arasteh, Cristian D. Ríos-Urrego, Elmar Nöth, Andreas K. Maier, Seung-Hee Yang, Jan Rusz, Juan Rafael Orozco-Arroyave
INTERSPEECH6
2023 Which aspects of motor speech disorder are captured by Mel Frequency Cepstral Coefficients? Evidence from the change in STN-DBS conditions in Parkinson's disease
abstract
One of the most popular speech parametrizations for dysarthria has been Mel Frequency Cepstral Coefficients (MFCCs). Although the MFCCs ability to capture vocal tract characteristics is known, the reflected dysarthria aspects are primarily undisclosed. Thus, we investigated the relationship between key acoustic variables in Parkinson's disease (PD) and the MFCCs. 23 PD patients were recruited with ON and OFF conditions of Deep Brain Stimulation of the Subthalamic Nucleus (STN-DBS) and examined via a reading passage. The changes in dysarthria aspects were compared to changes in a global MFCC measure and individual MFCCs. A similarity was found in 2nd to 3rd MFCCs changes and voice quality. Changes in 4th to 9th MFCCs reflected articulation clarity. The global MFCC parameter outperformed individual MFCCs and acoustical measures in capturing STN-DBS conditions changes. The findings may assist in interpreting outcomes from clinical trials and improve the monitoring of disease progression.
Vojtech Illner, Petr Krýze, Jan Svihlík, Mário Sousa, Paul Krack, Elina Tripoliti, Robert Jech, Jan Rusz
INTERSPEECH8
2023 Glottal source analysis of voice deficits in basal ganglia dysfunction: evidence from de novo Parkinson's disease and Huntington's disease
abstract
Dysphonia is a common speech disruption in people with Parkinson's (PD) and Huntington's (HD). Though the glottal source analysis (GSS) yielded promising results in PD, no study analyzed utility of the GSS in HD. In addition, the potential GSS sex-dependency remains unknown. This study examines sustained vowel phonations provided by 40 PD, 40 HD and 40 age- and sex-matched healthy participants using six GSS features including normalized amplitude quotient, quasi-open quotient, magnitude difference of first two spectral peaks, harmonic richness factor, maximum dispersion quotient (MDQ), and peak slope. Our results showed significant differences in HD men and women compared to the healthy counterpart, suggesting breathiness (p < 0.01), tension (p < 0.001), and decreased timbre (p < 0.01) in HD. Reported sex-related differences highlighted the sensitivity of the GSS towards the speaker's sex. The correlation analysis revealed significant relationship between disease severity and MDQ in HD men.
Michal Novotny, Tereza Tykalová, Michal Simek, Tomás Kouba, Jan Rusz
INTERSPEECH5
2023 Automatic Classification of Hypokinetic and Hyperkinetic Dysarthria based on GMM-Supervectors
abstract
Hypokinetic and hyperkinetic dysarthria are motor speech disorders that appear in patients with Parkinson's and Huntington's disease, respectively. They are caused due to progressive lesions or alterations in the basal ganglia. In particular, Huntington's disease (HD) is known to be more invasive and difficult to treat than Parkinson's disease (PD), producing more aggressive motor and cognitive alterations. Since speech production requires the movement and control of many different muscles and limbs, it constitutes a highly complex motor activity that may reflect relevant aspects of the patient's health state. This paper proposes the discrimination between patients with PD, HD, and healthy controls (HC) based on different speech dimensions. Speaker models based on Gaussian-mixture model supervectors are created with the features extracted from each speech dimension. The results suggest that it is possible to distinguish between PD and HD patients using the supervectors-based approach.
Cristian D. Ríos-Urrego, Jan Rusz, Elmar Nöth, Juan Rafael Orozco-Arroyave
INTERSPEECH2
2023 Comparison of acoustic measures of dysphonia in Parkinson's disease and Huntington's disease: Effect of sex and speaking task
abstract
This study investigated whether voice quality is differentially affected in two distinct basal ganglia disorders causing hypokinetic and hyperkinetic dysarthria, including effects of gender and speaking task. The sustained vowel phonations and monologues of 40 de novo Parkinson's disease (PD) patients, 40 Huntington's disease (HD) patients, and 40 healthy control participants were evaluated. Using cepstral peak prominence extracted from sustained phonation, differences from controls were found for male and female HD patients (p < 0.05) but only male PD patients (p < 0.05). Using the glottal-to-noise excitation ratio obtained from monologue, differences from controls were detected for male and female PD groups (p < 0.05) but only male HD group (p < 0.05). In general, female patients show better voice quality. Our findings highlight that selecting suitable acoustic measures and speaking material is essential for adequate evaluation of dysphonia severity across differing etiologies.
Michal Simek, Tomás Kouba, Michal Novotny, Tereza Tykalová, Jan Rusz
INTERSPEECH5
2023 Relationship between LTAS-based spectral moments and acoustic parameters of hypokinetic dysarthria in Parkinson's disease
abstract
Although long-term averaged spectrum (LTAS) descriptors can detect the change in dysarthria of patients with Parkinson's disease (PD) due to subthalamic nucleus deep brain stimulation (STN-DBS), the relationship between LTAS variables with measures that relate to laryngeal physiology remain unknown. We aimed to find connections between LTAS-based moments and the main acoustic characteristics of hypokinetic dysarthria in PD as the response to STN-DBS stimulation changes. We analyzed reading passages of 23 PD patients in ON and OFF STN-DBS states compared to 23 healthy controls. We found a relation between the stimulation-induced change in several spectral moments and acoustic parameters representing voice quality, articulatory decay, articulation rate, and mean fundamental frequency. While the difference between PD and controls was significant across most acoustic descriptors, only the spectral mean and fundamental frequency variability could differentiate between ON and OFF conditions.
Jan Svihlík, Vojtech Illner, Petr Krýze, Mário Sousa, Paul Krack, Elina Tripoliti, Robert Jech, Jan Rusz
INTERSPEECH8
2021 Transfer learning helps to improve the accuracy to classify patients with different speech disorders in different languages
Juan Camilo Vásquez-Correa, Cristian D. Ríos-Urrego, Tomás Arias-Vergara, Maria Schuster, Jan Rusz, Elmar Nöth, Juan Rafael Orozco-Arroyave
Pattern Recognit. Lett.5
2019 Convolutional Neural Networks and a Transfer Learning Strategy to Classify Parkinson's Disease from Speech in Three Different Languages
Juan Camilo Vásquez-Correa, Tomás Arias-Vergara, Cristian D. Ríos-Urrego, Maria Schuster, Jan Rusz, Juan Rafael Orozco-Arroyave, Elmar Nöth
CIARP5
2019 Towards Disease-specific Speech Markers for Differential Diagnosis in Parkinsonism
abstract
Parkinsonism refers to Parkinson's Disease (PD) and Atypical Parkinsonian Syndromes (APS), such as Progressive Supranuclear Palsy (PSP) and Multiple System Atrophy (MSA). Discrimination between PD and APS and within APS groups in early disease stages is a very challenging task. Interestingly, speech disorder is frequently an early and prominent clinical feature of both PD and APS. This renders speech/voice analysis a promising tool for the development of an objective marker to assist neurologists in their diagnosis. This paper is a continuation of a recent work on speech-based differential diagnosis within APS. We address the difficult problem of defining disease-specific speech features which is crucial in the perspective of early differential diagnosis. We investigate this problem by considering the constraint that only a small amount of training data can be available in this setting. To do so, we perform univariate statistical analysis followed by a supervised learning that forces the designed new features to be 1-dimensional. We carry out experiments using speech recordings of MSA and PSP patients. We show that linear classification models allow the definition of new scalar variables which can be considered as speech features which are specific to each disease, MSA and PSP.
Biswajit Das, Khalid Daoudi, Jirí Klempír, Jan Rusz
ICASSP4
2018 Linear Classification in Speech-Based Objective Differential Diagnosis of Parkinsonism
abstract
Parkinsonism refers to Parkinsons disease (PD) and Atypical parkinsonian syndromes (APS). Speech disorder is a common and early symptom in Parkinsonism which makes speech analysis a very important research area for the purpose of early diagnosis. Most of research have however focused on discrimination between PD and healthy controls. Such research does not take into account the fact that PD and APS syndromes are very similar in early disease stages. The main problem that has to be addressed first is then differential diagnosis: discrimination between PD and APS and within APS. This paper is a continuation of an earlier pioneer work in differential diagnosis where we mostly address the machine learning problem due to the small amount of training data. We show that classical linear and generalized linear models can provide interpretable and robust classifiers in term of accuracy and generalization ability.
Gongfeng Li, Khalid Daoudi, Jirí Klempír, Jan Rusz
ICASSP4
2017 Dysprosody Differentiate Between Parkinson's Disease, Progressive Supranuclear Palsy, and Multiple System Atrophy
Jan Hlavnicka, Tereza Tykalová, Roman Cmejla, Jirí Klempír, Evzen Ruzicka, Jan Rusz
INTERSPEECH6
2017 Acoustic Evaluation of Nasality in Cerebellar Syndromes
Michal Novotny, Jan Rusz, K. Spálenka, Jirí Klempír, Dana Horáková, Evzen Ruzicka
INTERSPEECH2
2016 Towards an automatic monitoring of the neurological state of Parkinson's patients from speech
abstract
The suitability of articulation measures and speech intelligibility is evaluated to estimate the neurological state of patients with Parkinson's disease (PD). A set of measures recently introduced to model the articulatory capability of PD patients is considered. Additionally, the speech intelligibility in terms of the word accuracy obtained from the Google® speech recognizer is included. Recordings of patients in three different languages are considered: Spanish, German, and Czech. Additionally, the proposed approach is tested on data recently used in the INTERSPEECH 2015 Computational Paralinguistics Challenge. According to the results, it is possible to estimate the neurological state of PD patients from speech with a Spearman's correlation of up to 0.72 with respect to the evaluations performed by neurologist experts.
Juan Rafael Orozco-Arroyave, Juan Camilo Vásquez-Correa, Florian Hönig, Julián D. Arias-Londoño, Jesús Francisco Vargas-Bonilla, Sabine Skodda, Jan Rusz, Elmar Nöth
ICASSP7
2015 Automatic detection of voice onset time in dysarthric speech
abstract
Although a number of speech disorders reflect varying involvement of brain areas, recently published automatic speech analyses have primarily been limited to hypokinetic dysarthria in Parkinson's disease (PD). Therefore, the aim of the present study was to provide an automatic algorithm suitable for the assessment of voice onset time (VOT) in various dysarthria types. Twenty-four PD participants with hypokinetic dysarthria and 40 Huntington's disease (HD) subjects with hyperkinetic dysarthria were included. These two types of dysarthria were selected in the design of a robust algorithm as they contain most of the dysarthric patterns found among all dysarthria subtypes. For a 10 ms threshold, the proposed algorithm reached approximately 90% accuracy in PD speakers and 80% accuracy in HD speakers. The accuracy of 80% obtained in HD was superior to the performance of 55% achieved by a previous algorithm designed particularly for hypokinetic dysarthria in PD.
Michal Novotny, Jakub Pospisil, Roman Cmejla, Jan Rusz
ICASSP4
2015 Voiced/unvoiced transitions in speech as a potential bio-marker to detect parkinson's disease
abstract
Several studies have addressed the automatic classification of speakers with Parkinson’s disease (PD) and healthy controls (HC). Most of the studies are based on speech recordings of sustained vowels, isolated words, and single sentences. Only few investigations have considered read texts and/or sponta-neous speech. This paper addresses two main questions still open regarding the automatic analysis speech in patients with PD, (a) “Is it possible to classify PD patients and HC through running speech signals in multiple languages?”, and (b) “where is the information to discriminate between speech recordings of PD patients and HC? ” In this paper speech recordings of read texts and monologues spoken in three different languages are considered. The energy content of the borders between voiced and unvoiced sounds is modeled. According to the results with read texts it is possible to achieve accuracies ranging from 91% to 98 % depending on the language. With respect to the re-sults on monologues, the accuracies are above 98 % in all of the three languages. The presence of discriminant information in the voiced/unvoiced and unvoiced/voiced transitions is vali-dated here, evidencing the problems of PD patients to stop/start the vocal folds movement during the production of running speech. Index Terms: Parkinson’s disease, dysarthria, hesitation in speech, language and motor planning, energy content, voiced/unvoiced transitions. 1.
Juan Rafael Orozco-Arroyave, Florian Hönig, Julián D. Arias-Londoño, Jesús Francisco Vargas-Bonilla, Sabine Skodda, Jan Rusz, Elmar Nöth
INTERSPEECH6
2015 Characterization Methods for the Detection of Multiple Voice Disorders: Neurological, Functional, and Laryngeal Diseases
abstract
This paper evaluates the accuracy of different characterization methods for the automatic detection of multiple speech disorders. The speech impairments considered include dysphonia in people with Parkinson's disease (PD), dysphonia diagnosed in patients with different laryngeal pathologies (LP), and hypernasality in children with cleft lip and palate (CLP). Four different methods are applied to analyze the voice signals including noise content measures, spectral-cepstral modeling, nonlinear features, and measurements to quantify the stability of the fundamental frequency. These measures are tested in six databases: three with recordings of PD patients, two with patients with LP, and one with children with CLP. The abnormal vibration of the vocal folds observed in PD patients and in people with LP is modeled using the stability measures with accuracies ranging from 81% to 99% depending on the pathology. The spectral-cepstral features are used in this paper to model the voice spectrum with special emphasis around the first two formants. These measures exhibit accuracies ranging from 95% to 99% in the automatic detection of hypernasal voices, which confirms the presence of changes in the speech spectrum due to hypernasality. Noise measures suitably discriminate between dysphonic and healthy voices in both databases with speakers suffering from LP. The results obtained in this study suggest that it is not suitable to use every kind of features to model all of the voice pathologies; conversely, it is necessary to study the physiology of each impairment to choose the most appropriate set of features.
Juan Rafael Orozco-Arroyave, Elkyn Alexander Belalcázar-Bolaños, Julián D. Arias-Londoño, Jesús Francisco Vargas-Bonilla, Sabine Skodda, Jan Rusz, Khaled Daqrouq, Florian Hönig, Elmar Nöth
IEEE J. Biomed. Health Informatics6
2014 Automatic detection of parkinson's disease from words uttered in three different languages
abstract
About 90% of the people with Parkinson’s disease (PD) develop speech impairments such as monopitch, monoloudness, imprecise articulation, and other symptoms. There are several studies addressing the problem of the automatic detection of PD from speech signals in order to develop computer aided tools for the assessment and monitoring of the patients. Recent works have shown that it is possible to detect PD from speech with accuracies above 90%; however, it is still unclear whether it is possible to make the detection independent of the spoken language. This paper addresses the automatic detection of PD considering speech recordings of three languages: German, Spanish and Czech. According to the results it is possible to classify between speech of people with PD and healthy controls (HC) with accuracies ranging from 84% to 99%, depending on the utterance.
Juan Rafael Orozco-Arroyave, Florian Hönig, Julián D. Arias-Londoño, Jesús Francisco Vargas-Bonilla, Sabine Skodda, Jan Rusz, Elmar Nöth
INTERSPEECH6
2014 Automatic Evaluation of Articulatory Disorders in Parkinson's Disease
abstract
Although articulatory deficits represent an important manifestation of dysarthria in Parkinson’s disease (PD), the most widely used methods currently available for the automatic evaluation of speech performance are focused on the assessment of dysphonia. The aim of the present study was to design a reliable automatic approach for the precise estimation of articulatory deficits in PD. Twenty-four individuals diagnosed with de novo PD and twenty-two age-matched healthy controls were recruited. Each participant performed diadochokinetic tasks based upon the fast repetition of /pa/-/ta/-/ka/ syllables. All phonemes were manually labeled and an algorithm for their automatic detection was designed. Subsequently, 13 features describing six different articulatory aspects of speech including vowel quality, coordination of laryngeal and supralaryngeal activity, precision of consonant articulation, tongue movement, occlusion weakening, and speech timing were analyzed. In addition, a classification experiment using a support vector machine based on articulatory features was proposed to differentiate between PD patients and healthy controls. The proposed detection algorithm reached approximately 80% accuracy for a 5 ms threshold of absolute difference between manually labeled references and automatically detected positions. When compared to controls, PD patients showed impaired articulatory performance in all investigated speech dimensions ($p < 0.05$). Moreover, using the six features representing different aspects of articulation, the best overall classification result attained a success rate of 88% in separating PD from controls. Imprecise consonant articulation was found to be the most powerful indicator of PD-related dysarthria. We envisage our approach as the first step towards development of acoustic methods allowing the automated assessment of articulatory features in dysarthrias.
Michal Novotny, Jan Rusz, Roman Cmejla, Evzen Ruzicka
IEEE ACM Trans. Audio Speech Lang. Process.2
2013 Bayesian changepoint detection for the automatic assessment of fluency and articulatory disorders
Roman Cmejla, Jan Rusz, Petr Bergl, Jan Vokral
Speech Commun.2
2011 Detection of persons with Parkinson's disease by acoustic, vocal, and prosodic analysis
abstract
70% to 90% of patients with Parkinson's disease (PD) show an affected voice. Various studies revealed, that voice and prosody is one of the earliest indicators of PD. The issue of this study is to automatically detect whether the speech/voice of a person is affected by PD. We employ acoustic features, prosodic features and features derived from a two-mass model of the vocal folds on different kinds of speech tests: sustained phonations, syllable repetitions, read texts and monologues. Classification is performed in either case by SVMs. A correlation-based feature selection was performed, in order to identify the most important features for each of these systems. We report recognition results of 91% when trying to differentiate between normal speaking persons and speakers with PD in early stages with prosodic modeling. With acoustic modeling we achieved a recognition rate of 88% and with vocal modeling we achieved 79%. After feature selection these results could greatly be improved. But we expect those results to be too optimistic. We show that read texts and monologues are the most meaningful texts when it comes to the automatic detection of PD based on articulation, voice, and prosodic evaluations. The most important prosodic features were based on energy, pauses and F0. The masses and the compliances of spring were found to be the most important parameters of the two-mass vocal fold model.
Tobias Bocklet, Elmar Nöth, Georg Stemmer, Hana Ruzickova, Jan Rusz
ASRU5