Elmar Nöth

dblp:30/1288 · DBLP profile ↗
← Back
205ranked-venue papers
7as first author
57since 2021 · last 2026
0000-0002-3396-555XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 175 · 6 first-author · 46 since 2021Artificial intelligence and machine learning · 153 · 5 first-author · 48 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Computer networks · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 The PARLO Dementia Corpus: A German Multi-Center Resource for Alzheimer's Disease
abstract
Early and accessible detection of Alzheimer's disease (AD) remains a major challenge, as current diagnostic methods often rely on costly and invasive biomarkers. Speech and language analysis has emerged as a promising non-invasive and scalable approach to detecting cognitive impairment, but research in this area is hindered by the lack of publicly available datasets, especially for languages other than English. This paper introduces the PARLO Dementia Corpus (PDC), a new multi-center, clinically validated German resource for AD collected across nine academic memory clinics in Germany. The dataset comprises speech recordings from individuals with AD-related mild cognitive impairment and mild to moderate dementia, as well as cognitively healthy controls. Speech was elicited using a standardized test battery of eight neuropsychological tasks, including confrontation naming, verbal fluency, word repetition, picture description, story reading, and recall tasks. In addition to audio recordings, the dataset includes manually verified transcriptions and detailed demographic, clinical, and biomarker metadata. Baseline experiments on ASR benchmarking, automated test evaluation, and LLM-based classification illustrate the feasibility of automatic, speech-based cognitive assessment and highlight the diagnostic value of recall-driven speech production. The PDC thus establishes the first publicly available German benchmark for multi-modal and cross-lingual research on neurodegenerative diseases.
Franziska Braun, Christopher Witzl, Florian Hönig, Elmar Nöth, Tobias Bocklet, Korbinian Riedhammer
LREC4
2026 A speech-to-video synthesis approach using spatio-temporal diffusion for vocal tract MRI
Paula Andrea Pérez-Toro, Tomás Arias-Vergara, Fangxu Xing, Xiaofeng Liu 0001, Maureen Stone 0001, Jiachen Zhuo, Juan Rafael Orozco-Arroyave, Elmar Nöth, Jana Hutter, Jerry L. Prince, Andreas K. Maier, Jonghye Woo
Medical Image Anal.8
2025 Joint ASR and Speech Attribute Prediction for Conversational Dysarthric Speech Analysis with Multimodal Language Models
abstract
Dysarthric speech recognition systems often focus solely on transcription, limiting their applicability in clinical and assistive settings where assessments of perceptual attributes like intelligibility and naturalness are essential. We propose a multimodal conversational framework based on Phi-4-Multimodal that combines ASR with attribute rating prediction, enabling users to query both transcriptions and perceptual characteristics (e.g. “How intelligible is this utterance?”). Our multitask model adds auxiliary prediction heads for five clinically relevant attributes and is trained on the English Speech Accessibility Project dataset. The system achieves competitive ASR performance while delivering attribute-level feedback comparable to specialized classifiers. Additional experiments show improved ASR performance for German Parkinson’s speech, indicating preserved multilingual capabilities and partial cross-lingual transfer of dysarthric speech patterns.
Dominik Wagner 0002, Ilja Baumann, Natalie Engert, Elmar Nöth, Korbinian Riedhammer, Tobias Bocklet
ASRU4
2025 Same Semantics of the Signal - What Do We Cluster with what Representation
abstract
Semantic clustering of bioacoustic signals is crucial for a deeper understanding of intra-class differences. This is particularly important for understanding killer whale signals, as their vocalizations are learned behaviors and determination of matrilineal-specific dialects is reliant upon subtle differences within instances which may be characterized into a single larger category. Aspects of data collection may have an effect on how these calls are grouped, and it is therefore necessary to understand what the focus of the feature generation algorithm is. This study addresses the impact of factors such as recording conditions and environment, together with its respective relative noise levels, by first analyzing two different deep learning and data-driven feature representations, either derived by an undercomplete autoencoder or a supervised call type classifier. These are then compared with representations generated by two state-of-the-art transformer-based tools, namely HuBERT and Wav2Vec2.
Alexander Barnhill, Oliver Traub, Andreas K. Maier, Elmar Nöth, Christian Bergler
ICASSP4
2025 Personalized Fine-Tuning with Controllable Synthetic Speech from LLM-Generated Transcripts for Dysarthric Speech Recognition
abstract
In this work, we present our submission to the Speech Accessibility Project challenge for dysarthric speech recognition. We integrate parameter-efficient fine-tuning with latent audio representations to improve an encoder-decoder ASR system. Synthetic training data is generated by fine-tuning Parler-TTS to mimic dysarthric speech, using LLM-generated prompts for corpus-consistent target transcripts. Personalization with x-vectors consistently reduces word error rates (WERs) over non-personalized fine-tuning. AdaLoRA adapters outperform full fine-tuning and standard low-rank adaptation, achieving relative WER reductions of ∼23% and ∼22%, respectively. Further improvements (∼5% WER reduction) come from incorporating wav2vec 2.0-based audio representations. Training with synthetic dysarthric speech yields up to ∼7% relative WER improvement over personalized fine-tuning alone.
Dominik Wagner 0002, Ilja Baumann, Natalie Engert, Seanie Lee, Elmar Nöth, Korbinian Riedhammer, Tobias Bocklet
INTERSPEECH5
2025 Synchronous analysis of abnormal acoustic and linguistic production in Parkinson's speech
abstract
Parkinson's disease is a neurodegenerative disorder involving speech and language deficits. Often, these are separately studied as proxies of motor and non-motor (e.g., cognitive) symptoms, respectively. Conversely, links between both dimensions remain virtually uncharted. This paper introduces a methodology that enables the synchronous study of acoustic and linguistic patterns in Parkinson's speech. Our findings show that verbs and nouns provided relevant acoustic and linguistic information not only to model motor impairments but also to understand non-motor symptoms like those that appear when Parkinson's disease patients develop mild cognitive impairment.
Daniel Escobar-Grisales, Cristian D. Ríos-Urrego, Sabato Marco Siniscalchi, Adolfo M. García, Yamile Bocanegra, Leonardo Moreno, Elmar Nöth, Juan Rafael Orozco-Arroyave
INTERSPEECH7
2025 Towards Inclusive ASR: Investigating Voice Conversion for Dysarthric Speech Recognition in Low-Resource Languages
abstract
Automatic speech recognition (ASR) for dysarthric speech remains challenging due to data scarcity, particularly in non-English languages. To address this, we fine-tune a voice conversion model on English dysarthric speech (UASpeech) to encode both speaker characteristics and prosodic distortions, then apply it to convert healthy non-English speech (FLEURS) into non-English dysarthric-like speech. The generated data is then used to fine-tune a multilingual ASR model, Massively Multilingual Speech (MMS), for improved dysarthric speech recognition. Evaluation on PC-GITA (Spanish), EasyCall (Italian), and SSNCE (Tamil) demonstrates that VC with both speaker and prosody conversion significantly outperforms the off-the-shelf MMS performance and conventional augmentation techniques such as speed and tempo perturbation. Objective and subjective analyses of the generated data further confirm that the generated speech simulates dysarthric characteristics.
Chin-Jou Li, Eunjung Yeo, Kwanghee Choi, Paula Andrea Pérez-Toro, Masao Someki, Rohan Kumar Das, Zhengjun Yue, Juan Rafael Orozco-Arroyave, Elmar Nöth, David R. Mortensen
INTERSPEECH9
2024 Towards Interpretability of Automatic Phoneme Analysis in Cleft Lip and Palate Speech
abstract
Cleft Lip and Palate ranks among the most common congenital abnormalities and significantly influences speech articulation, resulting in varying phonemic impacts. In a clinical context, a detailed diagnosis is carried out by time-consuming perceptual evaluations. We use perceptual ratings of different articulatory modifications on phoneme-level as ground-truth and propose a system based on wav2vec 2.0, trained to the downstream task of classifying phonemic criteria as a multi-class and multi-label problem. The system is trained for detection on utterance level, without the usage of phoneme labels. To gain a clearer understanding of which areas of the speech signal have the greatest impact on classification, we assess the extent to which our system aligns with expert ratings at the phoneme level. Additionally, we examine which specific phonemes play a decisive role in determining the final classification of the labeled criteria. The results show that salient phonemes marked by experts contribute remarkably greater to the classification of the correct class using feature relevance explanation methods. To the best of our knowledge, this is the first study incorporating various utterance-level articulatory modifications classification and phoneme-level interpretation, offering a more comprehensive understanding for potential clinical applications.
Ilja Baumann, Dominik Wagner 0002, Maria Schuster, Elmar Nöth, Tobias Bocklet
ICASSP4
2024 Longitudinal Modeling of Depression Shifts Using Speech and Language
abstract
Speech analysis can provide a potential non-invasive and objective means of assessing and monitoring an individual’s mental health. Most studies to date have focused on cross-sectional analysis and have not explored the benefits of speech analysis as a longitudinal monitoring tool that can assist in the management of chronic conditions such as major depressive disorder (MDD). Objectively monitoring for shifts in depression symptom severity levels over time presents a notable challenge, which we address through an automated approach using longitudinal English and Spanish speech samples collected from a clinical population. We employ time–frequency representations and linguistic embeddings to enhance the early recognition of alterations in depression levels in individuals with MDD. We investigate the suitability of using siamese-based training for modeling these changes, intending to enable personalized and adaptive interventions.
Paula Andrea Pérez-Toro, Judith Dineley, Agnieszka Kaczkowska, Pauline Conde, Yuezhou Zhang 0001, Faith Matcham, Sara Siddi, Josep Maria Haro, Stuart Bruce, Til Wykes, Raquel Bailón, Srinivasan Vairavan, Richard J. B. Dobson, Andreas K. Maier, Elmar Nöth, Juan Rafael Orozco-Arroyave, Vaibhav A. Narayan, Nicholas Cummins
ICASSP15
2024 Utilizing Deep Incomplete Classifiers to Implement Semantic Clustering for Killer Whale Photo Identification Data
Alexander Barnhill, Jared R. Towers, Elmar Nöth, Andreas K. Maier, Christian Bergler
ICPR (1)3
2024 Large Language Models for Dysfluency Detection in Stuttered Speech
Dominik Wagner 0002, Sebastian P. Bayerl, Ilja Baumann, Elmar Nöth, Korbinian Riedhammer, Tobias Bocklet
INTERSPEECH4
2024 Contrastive Learning Approach for Assessment of Phonological Precision in Patients with Tongue Cancer Using MRI Data
abstract
Magnetic Resonance Imaging (MRI) allows analyzing speech production by capturing high-resolution images of the dynamic processes in the vocal tract. In clinical applications, combining MRI with synchronized speech recordings leads to improved patient outcomes, especially if a phonological-based approach is used for assessment. However, when audio signals are unavailable, the recognition accuracy of sounds is decreased when using only MRI data. We propose a contrastive learning approach to improve the detection of phonological classes from MRI data when acoustic signals are not available at inference time. We demonstrate that frame-wise recognition of phonological classes improves from an f1 of 0.74 to 0.85 when the contrastive loss approach is implemented. Furthermore, we show the utility of our approach in the clinical application of using such phonological classes to assess speech disorders in patients with tongue cancer, yielding promising results in the recognition task.
Tomás Arias-Vergara, Paula Andrea Pérez-Toro, Xiaofeng Liu 0001, Fangxu Xing, Maureen Stone 0001, Jiachen Zhuo, Jerry L. Prince, Maria Schuster, Elmar Nöth, Jonghye Woo, Andreas K. Maier
INTERSPEECH9
2024 ANIMAL-CLEAN - A Deep Denoising Toolkit for Animal-Independent Signal Enhancement
Alexander Barnhill, Elmar Nöth, Andreas K. Maier, Christian Bergler
INTERSPEECH2
2024 Towards Self-Attention Understanding for Automatic Articulatory Processes Analysis in Cleft Lip and Palate Speech
abstract
Cleft lip and palate (CLP) speech presents unique challenges for automatic phoneme analysis due to its distinct acoustic characteristics and articulatory anomalies. We perform phoneme analysis in CLP speech using a pre-trained wav2vec 2.0 model with a multi-head self-attention classification module to capture long-range dependencies within the speech signal, thereby enabling better contextual understanding of phoneme sequences. We demonstrate the effectiveness of our approach in the classification of various articulatory processes in CLP speech. Furthermore, we investigate the interpretability of self-attention to gain insights into the model’s understanding of CLP speech characteristics. Our findings highlight the potential of the selfattention mechanisms for improving automatic phoneme analysis in CLP speech, paving the way for enhanced diagnostics, adding interpretability for therapists and affected patients.
Ilja Baumann, Dominik Wagner 0002, Maria Schuster, Korbinian Riedhammer, Elmar Nöth, Tobias Bocklet
INTERSPEECH5
2024 It's Time to Take Action: Acoustic Modeling of Motor Verbs to Detect Parkinson's Disease
abstract
Pre-trained models generate speech representations that are used in different tasks, including the automatic detection of Parkinson’s disease (PD). Although these models can yield high accuracy, their interpretation is still challenging. This paper used a pre-trained Wav2vec 2.0 model to represent speech frames of 25ms length and perform a frame-by-frame discrimination between PD patients and healthy control (HC) subjects. This fine granularity prediction enabled us to identify specific linguistic segments with high discrimination capability. Speech representations of all produced verbs were compared w.r.t. nouns and the first ones yielded higher accuracies. To gaina deeper understanding of this pattern, representations of motor and non-motor verbs were compared and the first ones yielded better results, with accuracies of around 83% in an independent test set. These findings support well-established neurocognitive models about action-related language highlighted as key drivers of PD. Index Terms: computational paralinguistics, interpretability of pre-trained models, action verbs, Parkinson’s disease
Daniel Escobar-Grisales, Cristian D. Ríos-Urrego, Ilja Baumann, Korbinian Riedhammer, Elmar Nöth, Tobias Bocklet, Adolfo M. García, Juan Rafael Orozco-Arroyave
INTERSPEECH5
2024 Analysis of Pathological Speech - Pitfalls along the Way
Elmar Nöth
INTERSPEECH1
2024 Multilingual Speech and Language Analysis for the Assessment of Mild Cognitive Impairment: Outcomes from the Taukadial Challenge
Paula Andrea Pérez-Toro, Tomás Arias-Vergara, Philipp Klumpp, Tobias Weise, Maria Schuster, Elmar Nöth, Juan Rafael Orozco-Arroyave, Andreas K. Maier
INTERSPEECH6
2024 Speaker- and Text-Independent Estimation of Articulatory Movements and Phoneme Alignments from Speech
Tobias Weise, Philipp Klumpp, Kubilay Can Demir, Paula Andrea Pérez-Toro, Maria Schuster, Elmar Nöth, Björn Heismann, Andreas K. Maier, Seung-Hee Yang
INTERSPEECH6
2024 Representation learning strategies to model pathological speech: Effect of multiple spectral resolutions
abstract
This paper considers a representation learning strategy to model speech signals from patients with Parkinson’s disease, with the goal of predicting the presence of the disease, and evaluating the level of degradation of a patient’s speech. In particular, we propose a novel fusion strategy that combines wideband and narrowband spectral resolutions using a representation learning strategy based on autoencoders , called the multi-spectral autoencoder. The proposed model is able to classify the speech from Parkinson’s disease patients with accuracy up to 97%. The proposed model is also able to assess the dysarthria severity of Parkinson’s disease patients with a Spearman correlation up to 0.79. These results outperform those observed in literature where the same problem was addressed with the same corpus.
Gabriel F. Miller, Juan Camilo Vásquez-Correa, Juan Rafael Orozco-Arroyave, Elmar Nöth
Comput. Speech Lang.4
2023 Detection of Vowel Errors in Children's Speech using Synthetic Phonetic Transcripts
abstract
The analysis of phonological processes is crucial in evaluating speech development disorders in children, but encounters challenges due to limited children audio data. This work focuses on automatic vowel error detection using a two-stage pipeline. The first stage uses a fine-tuned cross-lingual phone recognizer (wav2vec 2.0) to extract phone sequences from audio. The second stage employs a language model (BERT) for classification from a phone sequence, entirely trained on synthetic transcripts, to counteract the very broad range of potential mistakes. We evaluate the system on nonword audio recordings recited by preschool children from a speech development test. The results show that the classifier trained on synthetic data performs well, but its efficacy relies on the quality of the phone recognizer. The best classifier achieves an 94.7% F1 score when evaluated against phonetic ground truths, whereas the F1 score is 76.2% when using automatically recognized phone sequences.
Ilja Baumann, Dominik Wagner 0002, Korbinian Riedhammer, Elmar Nöth, Tobias Bocklet
ASRU4
2023 Transferring Quantified Emotion Knowledge for the Detection of Depression in Alzheimer's Disease Using Forestnets
abstract
Progressive loss of memory is the most known symptom of Alzheimer’s Disease (AD); however, it also affects other cognitive skills and leads to depression symptoms. This paper presents a transfer learning strategy for automatically detecting AD and depression in AD patients using acoustic information and ForestNet, an artificial neural network that allows computing the contribution of a set of features to a model’s decision. The methodology consists of training ForestNet with a dataset commonly used for emotion recognition; then, we fine-tune the pre-trained model to detect AD and depression in AD. We trained the models with several acoustic features commonly used for emotion and AD applications. Unweighted average recalls of up to 0.87 were achieved to classify the disease and up to 0.82 to detect depression in AD. Our results indicate that the information obtained from the Arousal Valence plane may be suitable for discriminating and analyzing depression in AD.
Paula Andrea Pérez-Toro, Dalia Rodríguez-Salas, Tomás Arias-Vergara, Sebastian P. Bayerl, Philipp Klumpp, Korbinian Riedhammer, Maria Schuster, Elmar Nöth, Andreas K. Maier, Juan Rafael Orozco-Arroyave
ICASSP8
2023 Federated Learning for Secure Development of AI Models for Parkinson's Disease Detection Using Speech from Different Languages
abstract
Parkinson's disease (PD) is a neurological disorder impacting a person's speech. Among automatic PD assessment methods, deep learning models have gained particular interest. Recently, the community has explored cross-pathology and cross-language models which can improve diagnostic accuracy even further. However, strict patient data privacy regulations largely prevent institutions from sharing patient speech data with each other. In this paper, we employ federated learning (FL) for PD detection using speech signals from 3 real-world language corpora of German, Spanish, and Czech, each from a separate institution. Our results indicate that the FL model outperforms all the local models in terms of diagnostic accuracy, while not performing very differently from the model based on centrally combined training sets, with the advantage of not requiring any data sharing among collaborators. This will simplify inter-institutional collaborations, resulting in enhancement of patient outcomes.
Soroosh Tayebi Arasteh, Cristian D. Ríos-Urrego, Elmar Nöth, Andreas K. Maier, Seung-Hee Yang, Jan Rusz, Juan Rafael Orozco-Arroyave
INTERSPEECH3
2023 Measuring Phonological Precision in Children with Cleft Lip and Palate
Tomás Arias-Vergara, Elizabeth Londoño-Mora, Paula Andrea Pérez-Toro, Maria Schuster, Elmar Nöth, Juan Rafael Orozco-Arroyave, Andreas K. Maier
INTERSPEECH5
2023 Influence of Utterance and Speaker Characteristics on the Classification of Children with Cleft Lip and Palate
Ilja Baumann, Dominik Wagner 0002, Franziska Braun, Sebastian P. Bayerl, Elmar Nöth, Korbinian Riedhammer, Tobias Bocklet
INTERSPEECH5
2023 A Stutter Seldom Comes Alone - Cross-Corpus Stuttering Detection as a Multi-label Problem
Sebastian P. Bayerl, Dominik Wagner 0002, Ilja Baumann, Florian Hönig, Tobias Bocklet, Elmar Nöth, Korbinian Riedhammer
INTERSPEECH6
2023 Classifying Dementia in the Presence of Depression: A Cross-Corpus Study
abstract
Automated dementia screening enables early detection and intervention, reducing costs to healthcare systems and increasing quality of life for those affected. Depression has shared symptoms with dementia, adding complexity to diagnoses. The research focus so far has been on binary classification of dementia (DEM) and healthy controls (HC) using speech from picture description tests from a single dataset. In this work, we apply established baseline systems to discriminate cognitive impairment in speech from the semantic Verbal Fluency Test and the Boston Naming Test using text, audio and emotion embeddings in a 3-class classification problem (HC vs. MCI vs. DEM). We perform cross-corpus and mixed-corpus experiments on two independently recorded German datasets to investigate generalization to larger populations and different recording conditions. In a detailed error analysis, we look at depression as a secondary diagnosis to understand what our classifiers actually learn.
Franziska Braun, Sebastian P. Bayerl, Paula Andrea Pérez-Toro, Florian Hönig, Hartmut Lehfeld, Thomas Hillemacher, Elmar Nöth, Tobias Bocklet, Korbinian Riedhammer
INTERSPEECH7
2023 An Automatic Multimodal Approach to Analyze Linguistic and Acoustic Cues on Parkinson's Disease Patients
Daniel Escobar-Grisales, Tomás Arias-Vergara, Cristian D. Ríos-Urrego, Elmar Nöth, Adolfo M. García, Juan Rafael Orozco-Arroyave
INTERSPEECH4
2023 Speaking Clearly, Understanding Better: Predicting the L2 Narrative Comprehension of Chinese Bilingual Kindergarten Children Based on Speech Intelligibility Using a Machine Learning Approach
Hiuching Hung, Paula Andrea Pérez-Toro, Tomás Arias-Vergara, Andreas K. Maier, Elmar Nöth
INTERSPEECH5
2023 Automatic Assessment of Alzheimer's across Three Languages Using Speech and Language Features
Paula Andrea Pérez-Toro, Tomás Arias-Vergara, Franziska Braun, Florian Hönig, Carlos Tobon 0001, David Aguillón, Francisco Lopera, Liliana Hincapié-Henao, Maria Schuster, Korbinian Riedhammer, Andreas K. Maier, Elmar Nöth, Juan Rafael Orozco-Arroyave
INTERSPEECH12
2023 Automatic Classification of Hypokinetic and Hyperkinetic Dysarthria based on GMM-Supervectors
abstract
Hypokinetic and hyperkinetic dysarthria are motor speech disorders that appear in patients with Parkinson's and Huntington's disease, respectively. They are caused due to progressive lesions or alterations in the basal ganglia. In particular, Huntington's disease (HD) is known to be more invasive and difficult to treat than Parkinson's disease (PD), producing more aggressive motor and cognitive alterations. Since speech production requires the movement and control of many different muscles and limbs, it constitutes a highly complex motor activity that may reflect relevant aspects of the patient's health state. This paper proposes the discrimination between patients with PD, HD, and healthy controls (HC) based on different speech dimensions. Speaker models based on Gaussian-mixture model supervectors are created with the features extracted from each speech dimension. The results suggest that it is possible to distinguish between PD and HD patients using the supervectors-based approach.
Cristian D. Ríos-Urrego, Jan Rusz, Elmar Nöth, Juan Rafael Orozco-Arroyave
INTERSPEECH3
2023 Multi-class Detection of Pathological Speech with Latent Features: How does it perform on unseen data?
abstract
The detection of pathologies from speech features is usually defined as a binary classification task with one class representing a specific pathology and the other class representing healthy speech. In this work, we train neural networks, large margin classifiers, and tree boosting machines to distinguish between four pathologies: Parkinson's disease, laryngeal cancer, cleft lip and palate, and oral squamous cell carcinoma. We show that latent representations extracted at different layers of a pre-trained wav2vec 2.0 system can be effectively used to classify these types of pathological voices. We evaluate the robustness of our classifiers by adding room impulse responses to the test data and by applying them to unseen speech corpora. Our approach achieves unweighted average F1-Scores between 74.1% and 97.0%, depending on the model and the noise conditions used. The systems generalize and perform well on unseen data of healthy speakers sampled from a variety of different sources.
Dominik Wagner 0002, Ilja Baumann, Franziska Braun, Sebastian P. Bayerl, Elmar Nöth, Korbinian Riedhammer, Tobias Bocklet
INTERSPEECH5
2023 Classification of stuttering - The ComParE challenge and beyond
Sebastian P. Bayerl, Maurice Gerczuk, Anton Batliner, Christian Bergler, Shahin Amiriparian, Björn W. Schuller, Elmar Nöth, Korbinian Riedhammer
Comput. Speech Lang.7
2023 User State Modeling Based on the Arousal-Valence Plane: Applications in Customer Satisfaction and Health-Care
abstract
The acoustic analysis helps to discriminate emotions according to non-verbal information, while linguistics aims to capture verbal information from written sources. Acoustic and linguistic analyses can be addressed for different applications, where information related to emotions, mood, or affect are involved. The Arousal-Valence plane is commonly used to model emotional states in a multidimensional space. This study proposes a methodology focused on modeling the user’s state based on the Arousal-Valence plane in different scenarios. Acoustic and linguistic information are used as input to feed different deep learning architectures mainly based on convolutional and recurrent neural networks, which are trained to model the Arousal-Valence plane. The proposed approach is used for the evaluation of customer satisfaction in call-centers and for health-care applications in the assessment of depression in Parkinson’s disease and the discrimination of Alzheimer’s disease. F-scores of up to 0.89 are obtained for customer satisfaction, of up to 0.82 for depression in Parkinson’s patients, and of up to 0.80 for Alzheimer’s patients. The proposed approach confirms that there is information embedded in the Arousal-Valence plane that can be used for different purposes.
Paula Andrea Pérez-Toro, Juan Camilo Vásquez-Correa, Tobias Bocklet, Elmar Nöth, Juan Rafael Orozco-Arroyave
IEEE Trans. Affect. Comput.4
2022 ORCA-PARTY: An Automatic Killer Whale Sound Type Separation Toolkit Using Deep Learning
abstract
Data-driven and machine-based analysis of massive bioacoustic data collections, in particular acoustic regions containing a substantial number of vocalizations events, is essential and extremely valuable to identify recurring vocal paradigms. However, these acoustic sections are usually characterized by a strong incidence of overlapping vocalization events, a major problem severely affecting subsequent human-/machine-based analysis and interpretation. Robust machine-driven signal separation of species-specific call types is extremely challenging due to missing ground truth data, speaker/source-relevant information, limited knowledge about inter- and intra-call type variations, next to diverse recording conditions. The current study is the first introducing a fully-automated deep signal separation approach for overlapping orca vocalizations, addressing all of the previously mentioned challenges, together with one of the largest bioacoustic data archives recorded on killer whales (Orcinus Orca). Incorporating ORCA-PARTY as additional data enhancement step for downstream call type classification demonstrated to be extremely valuable. Besides the proof of cross-domain applicability and consistently promising results on non-overlapping signals, significant improvements were achieved when processing acoustic orca segments comprising a multitude of vocal activities. Apart from auspicious visual inspections, a final numerical evaluation on an unseen dataset proved that about 30 % more known sound patterns could be identified.
Christian Bergler, Manuel Schmitt, Andreas K. Maier, Rachael Cheng, Volker Barth, Elmar Nöth
ICASSP6
2022 Detecting Dysfluencies in Stuttering Therapy Using wav2vec 2.0
abstract
Stuttering is a varied speech disorder that harms an individual's communication ability. Persons who stutter (PWS) often use speech therapy to cope with their condition. Improving speech recognition systems for people with such non-typical speech or tracking the effectiveness of speech therapy would require systems that can detect dysfluencies while at the same time being able to detect speech techniques acquired in therapy. This paper shows that fine-tuning wav2vec 2.0 [1] for the classification of stuttering on a sizeable English corpus containing stuttered speech, in conjunction with multi-task learning, boosts the effectiveness of the general-purpose wav2vec 2.0 features for detecting stuttering in speech; both within and across languages. We evaluate our method on FluencyBank , [2] and the German therapy-centric Kassel State of Fluency (KSoF) [3] dataset by training Support Vector Machine classifiers using features extracted from the finetuned models for six different stuttering-related event types: blocks, prolongations, sound repetitions, word repetitions, interjections, and - specific to therapy - speech modifications. Using embeddings from the fine-tuned models leads to relative classification performance gains up to 27% w.r.t. F1-score.
Sebastian P. Bayerl, Dominik Wagner 0002, Elmar Nöth, Korbinian Riedhammer
INTERSPEECH3
2022 ORCA-WHISPER: An Automatic Killer Whale Sound Type Generation Toolkit Using Deep Learning
Christian Bergler, Alexander Barnhill, Dominik Perrin, Manuel Schmitt, Andreas K. Maier, Elmar Nöth
INTERSPEECH6
2022 Wav2vec behind the Scenes: How end2end Models learn Phonetics
Teena tom Dieck, Paula Andrea Pérez-Toro, Tomas Arias, Elmar Nöth, Philipp Klumpp
INTERSPEECH4
2022 Cross-lingual Self-Supervised Speech Representations for Improved Dysarthric Speech Recognition
abstract
State-of-the-art automatic speech recognition (ASR) systems perform well on healthy speech.However, the performance on impaired speech still remains an issue.The current study explores the usefulness of using Wav2Vec self-supervised speech representations as features for training an ASR system for dysarthric speech.Dysarthric speech recognition is particularly difficult as several aspects of speech such as articulation, prosody and phonation can be impaired.Specifically, we train an acoustic model with features extracted from Wav2Vec, Hubert, and the cross-lingual XLSR model.Results suggest that speech representations pretrained on large unlabelled data can improve word error rate (WER) performance.In particular, features from the multilingual model led to lower WERs than filterbanks (Fbank) or models trained on a single language.Improvements were observed in English speakers with cerebral palsy caused dysarthria (UASpeech corpus), Spanish speakers with Parkinsonian dysarthria (PC-GITA corpus) and Italian speakers with paralysis-based dysarthria (EasyCall corpus).Compared to using Fbank features, XLSR-based features reduced WERs by 6.8%, 22.0%, and 7.0% for the UASpeech, PC-GITA, and EasyCall corpus, respectively.
Abner Hernandez, Paula Andrea Pérez-Toro, Elmar Nöth, Juan Rafael Orozco-Arroyave, Andreas K. Maier, Seung-Hee Yang
INTERSPEECH3
2022 Alzheimer's Detection from English to Spanish Using Acoustic and Linguistic Embeddings
Paula Andrea Pérez-Toro, Philipp Klumpp, Abner Hernandez, Tomas Arias, Patricia Lillo, Andrea Slachevsky, Adolfo M. García, Maria Schuster, Andreas K. Maier, Elmar Nöth, Juan Rafael Orozco-Arroyave
INTERSPEECH10
2022 CoachLea: an Android Application to Evaluate the Speech Production and Perception of Children with Hearing Loss
P. Schäfer, Paula Andrea Pérez-Toro, Philipp Klumpp, Juan Rafael Orozco-Arroyave, Elmar Nöth, Andreas K. Maier, A. Abad, Maria Schuster, Tomás Arias-Vergara
INTERSPEECH5
2022 Disentangled Latent Speech Representation for Automatic Pathological Intelligibility Assessment
abstract
Speech intelligibility assessment plays an important role in the therapy of patients suffering from pathological speech disorders.Automatic and objective measures are desirable to assist therapists in their traditionally subjective and labor-intensive assessments.In this work, we investigate a novel approach for obtaining such a measure using the divergence in disentangled latent speech representations of a parallel utterance pair, obtained from a healthy reference and a pathological speaker.Experiments on an English database of Cerebral Palsy patients, using all available utterances per speaker, show high and significant correlation values (R = -0.9)with subjective intelligibility measures, while having only minimal deviation (±0.01) across four different reference speaker pairs.We also demonstrate the robustness of the proposed method (R = -0.89deviating ±0.02 over 1000 iterations) by considering a significantly smaller amount of utterances per speaker.Our results are among the first to show that disentangled speech representations can be used for automatic pathological speech intelligibility assessment, resulting in a reference speaker pair invariant method, applicable in scenarios with only few utterances available.
Tobias Weise, Philipp Klumpp, Andreas K. Maier, Elmar Nöth, Björn Heismann, Maria Schuster, Seung-Hee Yang
INTERSPEECH4
2022 KSoF: The Kassel State of Fluency Dataset - A Therapy Centered Dataset of Stuttering
abstract
Stuttering is a complex speech disorder that negatively affects an individual’s ability to communicate effectively. Persons who stutter (PWS) often suffer considerably under the condition and seek help through therapy. Fluency shaping is a therapy approach where PWSs learn to modify their speech to help them to overcome their stutter. Mastering such speech techniques takes time and practice, even after therapy. Shortly after therapy, success is evaluated highly, but relapse rates are high. To be able to monitor speech behavior over a long time, the ability to detect stuttering events and modifications in speech could help PWSs and speech pathologists to track the level of fluency. Monitoring could create the ability to intervene early by detecting lapses in fluency. To the best of our knowledge, no public dataset is available that contains speech from people who underwent stuttering therapy that changed the style of speaking. This work introduces the Kassel State of Fluency (KSoF), a therapy-based dataset containing over 5500 clips of PWSs. The clips were labeled with six stuttering-related event types: blocks, prolongations, sound repetitions, word repetitions, interjections, and – specific to therapy – speech modifications. The audio was recorded during therapy sessions at the Institut der Kasseler Stottertherapie. The data will be made available for research purposes upon request.
Sebastian P. Bayerl, Alexander W. von Gudenberg, Florian Hönig, Elmar Nöth, Korbinian Riedhammer
LREC4
2022 Common Phone: A Multilingual Dataset for Robust Acoustic Modelling
abstract
Current state of the art acoustic models can easily comprise more than 100 million parameters. This growing complexity demands larger training datasets to maintain a decent generalization of the final decision function. An ideal dataset is not necessarily large in size, but large with respect to the amount of unique speakers, utilized hardware and varying recording conditions. This enables a machine learning model to explore as much of the domain-specific input space as possible during parameter estimation. This work introduces Common Phone, a gender-balanced, multilingual corpus recorded from more than 76.000 contributors via Mozilla’s Common Voice project. It comprises around 116 hours of speech enriched with automatically generated phonetic segmentation. A Wav2Vec 2.0 acoustic model was trained with the Common Phone to perform phonetic symbol recognition and validate the quality of the generated phonetic annotation. The architecture achieved a PER of 18.1 % on the entire test set, computed with all 101 unique phonetic symbols, showing slight differences between the individual languages. We conclude that Common Phone provides sufficient variability and reliable phonetic annotation to help bridging the gap between research and application of acoustic models.
Philipp Klumpp, Tomás Arias-Vergara, Paula Andrea Pérez-Toro, Elmar Nöth, Juan Rafael Orozco-Arroyave
LREC4
2022 The phonetic footprint of Parkinson's disease
Philipp Klumpp, Tomás Arias-Vergara, Juan Camilo Vásquez-Correa, Paula Andrea Pérez-Toro, Juan Rafael Orozco-Arroyave, Anton Batliner, Elmar Nöth
Comput. Speech Lang.7
2022 Empirical Mode Decomposition articulation feature extraction on Parkinson's Diadochokinesia
Alice Rueda, Juan Camilo Vásquez-Correa, Juan Rafael Orozco-Arroyave, Elmar Nöth, Sridhar Krishnan 0001
Comput. Speech Lang.4
2022 Depression assessment in people with Parkinson's disease: The combination of acoustic features and natural language processing
Paula Andrea Pérez-Toro, Tomás Arias-Vergara, Philipp Klumpp, Juan Camilo Vásquez-Correa, Maria Schuster, Elmar Nöth, Juan Rafael Orozco-Arroyave
Speech Commun.6
2021 Colombian Dialect Recognition Based on Information Extracted from Speech and Text Signals
abstract
Dialect recognition is useful in many industrial sectors, par-ticularly with the aim of allowing a better interaction between customers and providers. The core idea is to improve or customize marketing and customer service strategies, de-pending on the geographic location, birthplace and culture. This study proposes different models to automatically dis-criminate between two Colombian dialects: “Antioqueño” and “Bogotano”, to the best of our knowledge this is the first work of Colombian dialect recognition based on real conver-sations from customer service centers. The proposed strategy consists of independent analyses, using information from speech recordings and their corresponding transliterations. On the one hand, classical approaches are used to model speech including prosody features, Mel frequency cepstral coefficients and the mean Hilbert envelope coefficients. For text models, Word2Vec and bidirectional encoding represen-tations from transformer embeddings are considered. On the other hand, a deep learning approach is applied by considering convolutional neural networks, which are trained using spectrograms and embedding matrices for speech and text, respectively. The implemented deep learning models seem to be more promising than the classical ones for the addressed problem. Further experiments will be considered to validate this claim in a wider spectrum of methods.
Daniel Escobar-Grisales, Cristian D. Ríos-Urrego, Diego Alexander Lopez-Santander, Jeferson David Gallo-Aristizábal, Juan Camilo Vásquez-Correa, Elmar Nöth, Juan Rafael Orozco-Arroyave
ASRU6
2021 Applying X-Vectors on Pathological Speech After Larynx Removal
abstract
Speaker embeddings extracted from time delayed neural networks (TDNNs) contributed to major recent advancements in speaker recognition and verification. We use an X-Vector system trained on augmented VoxCeleb1 and VoxCeleb2 data to obtain embeddings for pathological speech after total or partial larynx removal. We show that our model is able to effectively distinguish and visualize patient groups when generating embeddings. We further compare various regression models on the task of automatically predicting different perceptual ratings by speech therapists (intelligibility, vocal effort, and overall quality) based on the extracted speaker embeddings. For both patient groups we show Pearson correlations in the range of +0.8; we find that Random Forest and Support Vector Regression produce scores that best resemble the experts' assessments.
Ralph Scheuerer, Tino Haderlein, Elmar Nöth, Tobias Bocklet
ASRU3
2021 Acoustic and Linguistic Analyses to Assess Early-Onset and Genetic Alzheimer's Disease
abstract
The PSEN1-E280A or Paisa mutation is responsible for most of Early-Onset Alzheimer’s (EOA) disease cases in Colombia. It affects a large kindred of over 5000 members that present the same phenotype. The most common symptoms are related to language disorders, where speech fluency is also affected due to the difficulty to access semantic information intentionally. This study proposes the use of acoustic and linguistic methods to extract features from speech recordings and their transcriptions to discriminate people with conditions related to the Paisa mutation. We consider state-of-the-art word-embedding methods like Word2Vec and Bidirectional Encoder Representations from Transformer to process the transcripts. The speech signals are modeled by using traditional acoustic features and speaker embeddings. To the best of our knowledge, this is the first study focused on evaluating genetic Alzheimer’s and EOA using acoustics and linguistics.
Paula Andrea Pérez-Toro, Juan Camilo Vásquez-Correa, Tomás Arias-Vergara, Philipp Klumpp, M. Sierra-Castrillón, M. E. Roldán-López, David Aguillón, Liliana Hincapié-Henao, Carlos Tobon 0001, Tobias Bocklet, Maria Schuster, Juan Rafael Orozco-Arroyave, Elmar Nöth
ICASSP13
2021 End-2-End Modeling of Speech and Gait from Patients with Parkinson's Disease: Comparison Between High Quality Vs. Smartphone Data
abstract
Parkinson’s disease is a neurodegenerative disorder characterized by the presence of different motor impairments. Speech and gait signals have been analyzed to detect the presence of the disease and the severity in patients. However, most studies have been performed in controlled conditions using high quality data, which make those studies not suitable for a continuous at-home evaluation of the state of the patients. The developed technology should be evaluated in more realistic scenarios, for instance using smartphone data. We propose the use of state-of-the-art deep learning techniques to evaluate the speech and gait symptoms of patients. The proposed methods are evaluated in two scenarios to cover both high quality and smartphone data. The results indicate that it is possible to classify patients and healthy subjects with accuracies over 92% in both scenarios. The proposed methods are also promising to evaluate the severity of the speech symptoms and the global motor state of the patients.
Juan Camilo Vásquez-Correa, Tomás Arias-Vergara, Philipp Klumpp, Paula Andrea Pérez-Toro, Juan Rafael Orozco-Arroyave, Elmar Nöth
ICASSP6
2021 ORCA-SLANG: An Automatic Multi-Stage Semi-Supervised Deep Learning Framework for Large-Scale Killer Whale Call Type Identification
Christian Bergler, Manuel Schmitt, Andreas K. Maier, Helena Symonds, Paul Spong, Steven R. Ness, George Tzanetakis, Elmar Nöth
Interspeech8
2021 Modeling Dysphonia Severity as a Function of Roughness and Breathiness Ratings in the GRBAS Scale
Carlos A. Ferrer, Efren Aragón, María E. Hdez-Díaz, Marc De Bodt, Roman Cmejla, Marina Englert, Mara Behlau, Elmar Nöth
Interspeech8
2021 The Phonetic Footprint of Covid-19?
abstract
Against the background of the ongoing pandemic, this year’s Computational Paralinguistics Challenge featured a classification problem to detect Covid-19 from speech recordings. The presented approach is based on a phonetic analysis of speech samples, thus it enabled us not only to discriminate between Covid and non-Covid samples, but also to better understand how the condition influenced an individual’s speech signal. Our deep acoustic model was trained with datasets collected exclusively from healthy speakers. It served as a tool for segmentation and feature extraction on the samples from the challenge dataset. Distinct patterns were found in the embeddings of phonetic classes that have their place of articulation deep inside the vocal tract. We observed profound differences in classification results for development and test splits, similar to the baseline method. We concluded that, based on our phonetic findings, it was safe to assume that our classifier was able to reliably detect a pathological condition located in the respiratory tract. However, we found no evidence to claim that the system was able to discriminate between Covid-19 and other respiratory diseases.
Philipp Klumpp, Tobias Bocklet, Tomás Arias-Vergara, Juan Camilo Vásquez-Correa, Paula Andrea Pérez-Toro, Sebastian P. Bayerl, Juan Rafael Orozco-Arroyave, Elmar Nöth
Interspeech8
2021 Influence of the Interviewer on the Automatic Assessment of Alzheimer's Disease in the Context of the ADReSSo Challenge
abstract
Alzheimer’s Disease (AD) results from the progressive loss of neurons in the hippocampus, which affects the capability to produce coherent language. It affects lexical, grammatical, and semantic processes as well as speech fluency. This paper considers the analyses of speech and language for the assessment of AD in the context of the Alzheimer’s Dementia Recognition through Spontaneous Speech (ADReSSo) 2021 challenge. We propose to extract acoustic features such as X-vectors, prosody, and emotional embeddings as well as linguistic features such as perplexity, and word-embeddings. The data consist of speech recordings from AD patients and healthy controls. The transcriptions are obtained using a commercial automatic speech recognition system. We outperform baseline results on the test set, both for the classification and the Mini-Mental State Examination (MMSE) prediction. We achieved a classification accuracy of 80% and an RMSE of 4.56 in the regression. Additionally, we found strong evidence for the influence of the interviewer on classification results. In cross-validation on the training set, we get classification results of 85% accuracy using the combined speech of the interviewer and the participant. Using interviewer speech only we still get an accuracy of 78%. Thus, we provide strong evidence for interviewer influence on classification results.
Paula Andrea Pérez-Toro, Sebastian P. Bayerl, Tomás Arias-Vergara, Juan Camilo Vásquez-Correa, Philipp Klumpp, Maria Schuster, Elmar Nöth, Juan Rafael Orozco-Arroyave, Korbinian Riedhammer
Interspeech7
2021 On Modeling Glottal Source Information for Phonation Assessment in Parkinson's Disease
abstract
Parkinson's disease produces several motor symptoms, including different speech impairments that are known as hypokinetic dysarthria. Symptoms associated to dysarthria affect different dimensions of speech such as phonation, articulation, prosody, and intelligibility. Studies in the literature have mainly focused on the analysis of articulation and prosody because they seem to be the most prominent symptoms associated to dysarthria severity. However, phonation impairments also play a significant role to evaluate the global speech severity of Parkinson's patients. This paper proposes an extensive comparison of different methods to automatically evaluate the severity of specific phonation impairments in Parkinson's patients. The considered models include the computation of perturbation and glottal-based features, in addition to features extracted from a zero frequency filtered signals. We consider as well end-to-end models based on 1D CNNs, which are trained to learn features from the raw speech waveform, reconstructed glottal signals, and zero-frequency filtered signals. The results indicate that it is possible to automatically classify between speakers with low versus high phonation severity due to the presence of dysarthria and at the same time to evaluate the severity of the phonation impairments on a continuous scale, posed as a regression problem.
Juan Camilo Vásquez-Correa, Julian Fritsch, Juan Rafael Orozco-Arroyave, Elmar Nöth, Mathew Magimai-Doss
Interspeech4
2021 Multi-channel spectrograms for speech processing applications using deep learning methods
abstract
Abstract Time–frequency representations of the speech signals provide dynamic information about how the frequency component changes with time. In order to process this information, deep learning models with convolution layers can be used to obtain feature maps. In many speech processing applications, the time–frequency representations are obtained by applying the short-time Fourier transform and using single-channel input tensors to feed the models. However, this may limit the potential of convolutional networks to learn different representations of the audio signal. In this paper, we propose a methodology to combine three different time–frequency representations of the signals by computing continuous wavelet transform, Mel-spectrograms, and Gammatone spectrograms and combining then into 3D-channel spectrograms to analyze speech in two different applications: (1) automatic detection of speech deficits in cochlear implant users and (2) phoneme class recognition to extract phone-attribute features. For this, two different deep learning-based models are considered: convolutional neural networks and recurrent neural networks with convolution layers.
Tomás Arias-Vergara, Philipp Klumpp, Juan Camilo Vásquez-Correa, Elmar Nöth, Juan Rafael Orozco-Arroyave, Maria Schuster
Pattern Anal. Appl.4
2021 Transfer learning helps to improve the accuracy to classify patients with different speech disorders in different languages
Juan Camilo Vásquez-Correa, Cristian D. Ríos-Urrego, Tomás Arias-Vergara, Maria Schuster, Jan Rusz, Elmar Nöth, Juan Rafael Orozco-Arroyave
Pattern Recognit. Lett.6
2020 Comparison of User Models Based on GMM-UBM and I-Vectors for Speech, Handwriting, and Gait Assessment of Parkinson's Disease Patients
abstract
Parkinson's disease is a neurodegenerative disorder characterized by the presence of different motor impairments. Information from speech, handwriting, and gait signals have been considered to evaluate the neurological state of the patients. On the other hand, user models based on Gaussian mixture models - universal background models (GMMUBM) and i-vectors are considered the state-of-the-art in biometric applications like speaker verification because they are able to model specific speaker traits. This study introduces the use of GMM-UBM and i-vectors to evaluate the neurological state of Parkinson's patients using information from speech, handwriting, and gait. The results show the importance of different feature sets from each type of signal in the assessment of the neurological state of the patients.
Juan Camilo Vásquez-Correa, Tobias Bocklet, Juan Rafael Orozco-Arroyave, Elmar Nöth
ICASSP4
2020 ORCA-CLEAN: A Deep Denoising Toolkit for Killer Whale Communication
abstract
In bioacoustics, passive acoustic monitoring of animals living in the wild, both on land and underwater, leads to large data archives characterized by a strong imbalance between recorded animal sounds and ambient noises. Bioacoustic datasets suffer extremely from such large noise-variety, caused by a multitude of external influences and changing environmental conditions over years. This leads to significant deficiencies/problems concerning the analysis and interpretation of animal vocalizations by biologists and machine-learning algorithms. To counteract such huge noise diversity, it is essential to develop a denoising procedure enabling automated, efficient, and robust data enhancement. However, a fundamental problem is the lack of clean/denoised ground-truth samples. The current work is the first presenting a fully-automated deep denoising approach for bioacoustics, not requiring any clean ground-truth, together with one of the largest data archives recorded on killer whales (Orcinus Orca) – the Orchive. Therefor, an approach, originally developed for image restoration, known as Noise2Noise (N2N), was transferred to the field of bioacoustics, and extended by using automatic machine-generated binary masks as additional network attention mechanism. Besides a significant cross-domain signal enhancement, our previous results regarding supervised orca/noise segmentation and orca call type identification were outperformed by applying ORCACLEAN as additional data preprocessing/enhancement step
Christian Bergler, Manuel Schmitt, Andreas K. Maier, Simeon Smeele, Volker Barth, Elmar Nöth
INTERSPEECH6
2020 Surgical Mask Detection with Deep Recurrent Phonetic Models
Philipp Klumpp, Tomás Arias-Vergara, Juan Camilo Vásquez-Correa, Paula Andrea Pérez-Toro, Florian Hönig, Elmar Nöth, Juan Rafael Orozco-Arroyave
INTERSPEECH6
2020 Parallel Representation Learning for the Classification of Pathological Speech: Studies on Parkinson's Disease and Cleft Lip and Palate
Juan Camilo Vásquez-Correa, Tomás Arias-Vergara, Maria Schuster, Juan Rafael Orozco-Arroyave, Elmar Nöth
Speech Commun.5
2019 Multi-channel Convolutional Neural Networks for Automatic Detection of Speech Deficits in Cochlear Implant Users
Tomás Arias-Vergara, Juan Camilo Vásquez-Correa, Sandra Gollwitzer, Juan Rafael Orozco-Arroyave, Maria Schuster, Elmar Nöth
CIARP6
2019 Bidirectional Alignment of Glottal Pulse Length Sequences for the Evaluation of Pitch Detection Algorithms
Carlos A. Ferrer-Riesgo, Reinier Rodríguez Guillén, Elmar Nöth
CIARP3
2019 Analytical Solution for the Optimal Addition of an Item to a Composite of Scores for Maximum Reliability
Carlos A. Ferrer-Riesgo, Idileisy Torres-Rodríguez, Alberto Taboada-Crispí, Elmar Nöth
CIARP4
2019 Convolutional Neural Networks and a Transfer Learning Strategy to Classify Parkinson's Disease from Speech in Three Different Languages
Juan Camilo Vásquez-Correa, Tomás Arias-Vergara, Cristian D. Ríos-Urrego, Maria Schuster, Jan Rusz, Juan Rafael Orozco-Arroyave, Elmar Nöth
CIARP7
2019 Articulation and Empirical Mode Decomposition Features in Diadochokinetic Exercises for the Speech Assessment of Parkinson's Disease Patients
Juan Camilo Vásquez-Correa, Cristian D. Ríos-Urrego, Alice Rueda, Juan Rafael Orozco-Arroyave, Sri Krishnan, Elmar Nöth
CIARP6
2019 Automatic Diagnosis of Alzheimer's Disease Using Neural Network Language Models
abstract
In today's aging society, the number of neurodegenerative diseases such as Alzheimer's disease (AD) increases. Reliable tools for automatic early screening as well as monitoring of AD patients are necessary. For that, semantic deficits have been shown to be useful indicators. We present a way to significantly improve the method introduced by Wankerl et al. [1]. The purely statistical approach of n-gram language models (LMs) is enhanced by using the rwthlm toolkit to create neural network language models (NNLMs) with Long Short Term-Memory (LSTM) cells. The prediction is solely based on evaluating the perplexity of transliterations of descriptions of the Cookie Theft picture from DementiaBank's Pitt Corpus. Each transliteration is evaluated on LMs of both control and Alzheimer speakers in a leave-one-speaker-out cross-validation scheme. The resulting perplexity values reveal enough discrepancy to classify patients on just those two values with an accuracy of 85.6% at equal-error-rate.
Julian Fritsch, Sebastian Wankerl, Elmar Nöth
ICASSP3
2019 Segmentation, Classification, and Visualization of Orca Calls Using Deep Learning
abstract
Audiovisual media are increasingly used to study the communication and behavior of animal groups, e.g. by placing microphones in the animals habitat resulting in huge datasets with only a small amount of animal interactions. The Orcalab has recorded orca whales since 1973 using stationary underwater hydrophones and made it publicly available on the Orchive. There exist over 15 000 manually extracted orca/noise annotations and about 20 000 h unseen audio data. To analyze the behavior and communication of killer whales we need to interpret the different call types. In this work, we present a two-stage classification approach using the labeled call/noise files and a few labeled call-type files. Results indicate a reliable accuracy of 95.0 % for call segmentation and 87 % for classification of 12 call classes. We further visualize the learned orca call representations in the convolutional neural network (CNN) activations to explain the potential of CNN based recognition for bioaccousitc signals.
Hendrik Schröter, Elmar Nöth, Andreas K. Maier, Rachael Cheng, Volker Barth, Christian Bergler
ICASSP2
2019 Phone-Attribute Posteriors to Evaluate the Speech of Cochlear Implant Users
abstract
People with pre- and postlingual onset of deafness, i.e, age of occurrence of hearing loss, often present speech production\nproblems even after hearing rehabilitation by cochlear implantation. In this paper, the speech of 20 prelinguals (aged between 18 to 71 years old), 20 postlinguals (aged between 33 to 78 years old) and 20 healthy control (aged between 31 to 62 years old) German native speakers are analyzed considering phone-attribute features extracted with pre-trained Deep Neural Networks. Speech signals are analyzed with reference to the manner of articulation of consonants according to 5 groups: nasals, sibilants, fricatives, voiced-stops, and voiceless-stops. According to the results, it is possible to detect alterations in the consonant production of CI users when compared with healthy speakers. A comprehensive evaluation of speech changes of CI users will help in the rehabilitation after deafening.
Tomás Arias-Vergara, Juan Rafael Orozco-Arroyave, Milos Cernak, Sandra Gollwitzer, Maria Schuster, Elmar Nöth
INTERSPEECH6
2019 Deep Learning for Orca Call Type Identification - A Fully Unsupervised Approach
Christian Bergler, Manuel Schmitt, Rachael Cheng, Andreas K. Maier, Volker Barth, Elmar Nöth
INTERSPEECH6
2019 Feature Space Visualization with Spatial Similarity Maps for Pathological Speech Data
Philipp Klumpp, Juan Camilo Vásquez-Correa, Tino Haderlein, Elmar Nöth
INTERSPEECH4
2019 Feature Representation of Pathophysiology of Parkinsonian Dysarthria
Alice Rueda, Juan Camilo Vásquez-Correa, Cristian D. Ríos-Urrego, Juan Rafael Orozco-Arroyave, Sridhar Krishnan 0001, Elmar Nöth
INTERSPEECH6
2019 The INTERSPEECH 2019 Computational Paralinguistics Challenge: Styrian Dialects, Continuous Sleepiness, Baby Sounds & Orca Activity
abstract
The INTERSPEECH 2019 Computational Paralinguistics Challenge addresses four different problems for the first time in a research competition under well-defined conditions: In the Styrian Dialects Sub-Challenge, three types of Austrian-German dialects have to be classified; in the Continuous Sleepiness Sub-Challenge, the sleepiness of a speaker has to be assessed as regression problem; in the Baby Sound Sub-Challenge, five types of infant sounds have to be classified; and in the Orca Activity Sub-Challenge, orca sounds have to be detected.We describe the Sub-Challenges and baseline feature extraction and classifiers, which include data-learnt (supervised) feature representations by the 'usual' ComParE and BoAW features, and deep unsupervised representation learning using the AUDEEP toolkit.
Björn W. Schuller, Anton Batliner, Christian Bergler, Florian B. Pokorny, Jarek Krajewski, Margaret Cychosz, Ralf Vollmann, Sonja-Dana Roelen, Sebastian Schnieder, Elika Bergelson, Alejandrina Cristià, Amanda Seidl, Anne S. Warlaumont, Lisa Yankowitz, Elmar Nöth, Shahin Amiriparian, Simone Hantke, Maximilian Schmitt
INTERSPEECH15
2019 Apkinson: A Mobile Solution for Multimodal Assessment of Patients with Parkinson's Disease
Juan Camilo Vásquez-Correa, Tomás Arias-Vergara, Philipp Klumpp, M. Strauss, Arne Küderle, Nils Roth, Sebastian P. Bayerl, Nicanor García, Paula Andrea Pérez-Toro, L. Felipe Parra-Gallego, Cristian D. Ríos-Urrego, Daniel Escobar-Grisales, Juan Rafael Orozco-Arroyave, Björn M. Eskofier, Elmar Nöth
INTERSPEECH15
2019 Phonet: A Tool Based on Gated Recurrent Neural Networks to Extract Phonological Posteriors from Speech
Juan Camilo Vásquez-Correa, Philipp Klumpp, Juan Rafael Orozco-Arroyave, Elmar Nöth
INTERSPEECH4
2019 Multimodal Assessment of Parkinson's Disease: A Deep Learning Approach
abstract
Parkinson's disease is a neurodegenerative disorder characterized by a variety of motor symptoms. Particularly, difficulties to start/stop movements have been observed in patients. From a technical/diagnostic point of view, these movement changes can be assessed by modeling the transitions between voiced and unvoiced segments in speech, the movement when the patient starts or stops a new stroke in handwriting, or the movement when the patient starts or stops the walking process. This study proposes a methodology to model such difficulties to start or to stop movements considering information from speech, handwriting, and gait. We used those transitions to train convolutional neural networks to classify patients and healthy subjects. The neurological state of the patients was also evaluated according to different stages of the disease (initial, intermediate, and advanced). In addition, we evaluated the robustness of the proposed approach when considering speech signals in three different languages: Spanish, German, and Czech. According to the results, the fusion of information from the three modalities is highly accurate to classify patients and healthy subjects, and it shows to be suitable to assess the neurological state of the patients in several stages of the disease. We also aimed to interpret the feature maps obtained from the deep learning architectures with respect to the presence or absence of the disease and the neurological state of the patients. As far as we know, this is one of the first works that considers multimodal information to assess Parkinson's disease following a deep learning approach.
Juan Camilo Vásquez-Correa, Tomás Arias-Vergara, Juan Rafael Orozco-Arroyave, Björn M. Eskofier, Jochen Klucken, Elmar Nöth
IEEE J. Biomed. Health Informatics6
2018 Unobtrusive Monitoring of Speech Impairments of Parkinson'S Disease Patients Through Mobile Devices
abstract
Parkinson's disease (PD) produces several speech impairments in the patients. Automatic classification of PD patients is performed considering speech recordings collected in noncontrolled acoustic conditions during normal phone calls in a unobtrusive way. A speech enhancement algorithm is applied to improve the quality of the signals. Two different classification approaches are considered: the classification of PD patients and healthy speakers and a multi-class experiment to classify patients in several stages of the disease. According to the results it is possible to classify PD patients and healthy controls with a AUe of up to 0.87. This work is a step forward to the development of telemonitoring systems to assess the speech of the patients.
Tomás Arias-Vergara, Juan Camilo Vásquez-Correa, Juan Rafael Orozco-Arroyave, Philipp Klumpp, Elmar Nöth
ICASSP5
2018 Multimodal I-vectors to Detect and Evaluate Parkinson's Disease
Nicanor García, Juan Camilo Vásquez-Correa, Juan Rafael Orozco-Arroyave, Elmar Nöth
INTERSPEECH4
2018 A Multitask Learning Approach to Assess the Dysarthria Severity in Patients with Parkinson's Disease
Juan Camilo Vásquez-Correa, Tomás Arias-Vergara, Juan Rafael Orozco-Arroyave, Elmar Nöth
INTERSPEECH4
2018 Speaker models for monitoring Parkinson's disease progression considering different communication channels and acoustic conditions
abstract
Symptoms of Parkinson's disease vary from patient to patient. Additionally, the progression of those symptoms also differs among patients. Most of the studies on the analysis of speech of people with Parkinson's disease do not consider such an individual variation. This paper presents a methodology for the automatic and individual monitoring of speech disorders developed by PD patients. The neurological state and dysarthria level of the patients are evaluated. The proposed system is based on individual speaker models which are created for each patient. Two different models are evaluated, the classical GMM–UBM and the i–vectors approach. These two methods are compared with respect to a baseline found with a traditional Support Vector Regressor. Different speech aspects (phonation, articulation, and prosody) are considered to model recordings of spontaneous speech and a read text. A multi-aspect coefficient is proposed with the aim of incorporating information from all of these speech aspects into a single measure. Two different scenarios are considered to assess a set with seven PD patients: (1) the longitudinal test set which consists of speech recordings captured in five recording sessions distributed from 2012 to 2016, and (2) the at-home test set which consists of speech recordings captured in the home of the same seven patients during 4 months (one day per month, four times per day). The UBM is trained with the recordings of 100 speakers (50 with Parkinson's disease and 50 healthy speakers) captured with controlled acoustic conditions and a professional audio-setting. With the aim of evaluating the suitability of the proposed approaches and the possibility of extending this kind of systems to remotely assess the speech of the patients, a total of five different communication channels (sound-proof booth, Skype®, Hangouts®, mobile phone, and land-line) are considered to train and test the system. Due to the reduced number of recording sessions in the longitudinal test set, the experiments that involved this set are evaluated with the Pearson's correlation. The experiments with the at-home test set are evaluated with the Spearman's correlation. The results estimating the dysarthria level of the patients in the at-home test set indicate a correlation of 0.55 with a modified version of the Frenchay Dysarthria Assessment scale when the GMM-UBM model is applied upon the Skype® recordings. The results in the longitudinal test set indicate a correlation of 0.77 using a model based on i-vectors with recordings captured in the sound-proof-booth. The evaluation of the neurological state of the patients in the longitudinal test set shows correlations of up to 0.55 with the Movement Disorder Society - Unified Parkinson's Disease Rating Scale also using models based on i-vectors created with Skype® recordings. These results suggest that the i–vector approach is suitable when the acoustic conditions among recording sessions differ (longitudinal test set). The GMM-UBM approach seems to be more suitable when the acoustic conditions do not change a lot among recording sessions (at-home test set). Particularly, the best results were obtained with the Skype® calls, which can be explained due to several preprocessing stages that this codec applies to the audio signals. In general, the results suggest that the proposed approaches are suitable for tele-monitoring the dysarthria level and the neurological state of PD patients.
Tomás Arias-Vergara, Juan Camilo Vásquez-Correa, Juan Rafael Orozco-Arroyave, Elmar Nöth
Speech Commun.4
2017 On the impact of non-modal phonation on phonological features
abstract
Different modes of vibration of the vocal folds contribute significantly to the voice quality. The neutral mode phonation, often used in a modal voice, is one against which the other modes can be contrastively described, also called non-modal phonations. This paper investigates the impact of non-modal phonation on phonological posteriors, the probabilities of phonological features inferred from the speech signal using a deep learning approach. Five different non-modal phonations are considered: falsetto, creaky, harshness, tense and breathiness. The impact of such non-modal phonation on phonological features, the Sound Patterns of English (SPE), is investigated in both speech analysis and synthesis tasks. We found that breathy and tense phonation impact the SPE features less, creaky phonation impacts the features moderately, and harsh and falsetto phonation impact the phonological features the most. We also report invariant and the most different SPE features impacted by non-modal phonation.
Milos Cernak, Elmar Nöth, Frank Rudzicz, Heidi Christensen, Juan Rafael Orozco-Arroyave, Raman Arora, Tobias Bocklet, Hamid R. Chinaei, Julius Hannink, Phani S. Nidadavolu, Juan Camilo Vásquez-Correa, Maria Yancheva, Alyssa Vann, Nikolai Vogler
ICASSP2
2017 Multi-view representation learning via gcca for multimodal analysis of Parkinson's disease
abstract
Information from different bio-signals such as speech, handwriting, and gait have been used to monitor the state of Parkinson's disease (PD) patients, however, all the multimodal bio-signals may not always be available. We propose a method based on multi-view representation learning via generalized canonical correlation analysis (GCCA) for learning a representation of features extracted from handwriting and gait that can be used as a complement to speech-based features. Three different problems are addressed: classification of PD patients vs. healthy controls, prediction of the neurological state of PD patients according to the UPDRS score, and the prediction of a modified version of the Frenchay dysarthria assessment (m-FDA). According to the results, the proposed approach is suitable to improve the results in the addressed problems, specially in the prediction of the UPDRS, and m-FDA scores.
Juan Camilo Vásquez-Correa, Juan Rafael Orozco-Arroyave, Raman Arora, Elmar Nöth, Najim Dehak, Heidi Christensen, Frank Rudzicz, Tobias Bocklet, Milos Cernak, Hamid R. Chinaei, Julius Hannink, Phani S. Nidadavolu, Maria Yancheva, Alyssa Vann, Nikolai Vogler
ICASSP4
2017 Effect of acoustic conditions on algorithms to detect Parkinson's disease from speech
abstract
Automatic detection of Parkinson's disease (PD) from speech is a basic step towards computer-aided tools supporting the diagnosis and monitoring of the disease. Although several methods have been proposed, their applicability to real-world situations is still unclear. In particular, the effect of acoustic conditions is not well understood. In this paper, the effects on the accuracy of five different methods to detect PD from speech are evaluated. Among the considered conditions, background noise produces the worst effect, while dynamic compression or some speech codecs can even have a marginal positive impact. We also consider, for the first time in this context, the problem of mismatches, i.e., when train/test acoustic conditions are different, and observe a high negative impact on all considered methods. Overall, this study is a step forward in performing a continuous monitoring of the neurological state of the patients in non-controlled acoustic conditions.
Juan Camilo Vásquez-Correa, Joan Serrà, Juan Rafael Orozco-Arroyave, Jesús Francisco Vargas-Bonilla, Elmar Nöth
ICASSP5
2017 Evaluation of the Neurological State of People with Parkinson's Disease Using i-Vectors
Nicanor García, Juan Rafael Orozco-Arroyave, Luis Fernando D'Haro, Najim Dehak, Elmar Nöth
INTERSPEECH5
2017 Apkinson - A Mobile Monitoring Solution for Parkinson's Disease
Philipp Klumpp, Thomas Janu, Tomás Arias-Vergara, Juan Camilo Vásquez-Correa, Juan Rafael Orozco-Arroyave, Elmar Nöth
INTERSPEECH6
2017 Convolutional Neural Network to Model Articulation Impairments in Patients with Parkinson's Disease
Juan Camilo Vásquez-Correa, Juan Rafael Orozco-Arroyave, Elmar Nöth
INTERSPEECH3
2017 An N-Gram Based Approach to the Automatic Diagnosis of Alzheimer's Disease from Spoken Language
Sebastian Wankerl, Elmar Nöth, Stefan Evert
INTERSPEECH2
2017 Characterisation of voice quality of Parkinson's disease using differential phonological posterior features
Milos Cernak, Juan Rafael Orozco-Arroyave, Frank Rudzicz, Heidi Christensen, Juan Camilo Vásquez-Correa, Elmar Nöth
Comput. Speech Lang.6
2017 Detection of different voice diseases based on the nonlinear characterization of speech signals
Carlos Manuel Travieso-González, Jesús B. Alonso, Juan Rafael Orozco-Arroyave, Jesús Francisco Vargas-Bonilla, Elmar Nöth, Antonio G. Ravelo-García
Expert Syst. Appl.5
2016 Towards an automatic monitoring of the neurological state of Parkinson's patients from speech
abstract
The suitability of articulation measures and speech intelligibility is evaluated to estimate the neurological state of patients with Parkinson's disease (PD). A set of measures recently introduced to model the articulatory capability of PD patients is considered. Additionally, the speech intelligibility in terms of the word accuracy obtained from the Google® speech recognizer is included. Recordings of patients in three different languages are considered: Spanish, German, and Czech. Additionally, the proposed approach is tested on data recently used in the INTERSPEECH 2015 Computational Paralinguistics Challenge. According to the results, it is possible to estimate the neurological state of PD patients from speech with a Spearman's correlation of up to 0.72 with respect to the evaluations performed by neurologist experts.
Juan Rafael Orozco-Arroyave, Juan Camilo Vásquez-Correa, Florian Hönig, Julián D. Arias-Londoño, Jesús Francisco Vargas-Bonilla, Sabine Skodda, Jan Rusz, Elmar Nöth
ICASSP8
2016 Parkinson's Disease Progression Assessment from Speech Using GMM-UBM
Tomás Arias-Vergara, Juan Camilo Vásquez-Correa, Juan Rafael Orozco-Arroyave, Jesús Francisco Vargas-Bonilla, Elmar Nöth
INTERSPEECH5
2016 Automatic Detection of Parkinson's Disease Based on Modulated Vowels
Daria Hemmerling, Juan Rafael Orozco-Arroyave, Andrzej Skalski, Janusz Gajda, Elmar Nöth
INTERSPEECH5
2016 Combining Semantic Word Classes and Sub-Word Unit Speech Recognition for Robust OOV Detection
abstract
Out-of-vocabulary words (OOVs) are often the main reason for the failure of tasks like automated voice searches or humanmachine dialogs.This is especially true if rare but task-relevant content words, e.g.person or location names, are not in the recognizer's vocabulary.Since applications like spoken dialog systems use the result of the speech recognizer to extract a semantic representation of a user utterance, the detection of OOVs as well as their (semantic) word class can support to manage a dialog successfully.In this paper we suggest to combine two wellknown approaches in the context of OOV detection: semantic word classes and OOV models based on sub-word units.With our system, which builds upon the widely used Kaldi speech recognition toolkit, we show on two different data sets that -compared to other methods -such a combination improves OOV detection performance for open word classes at a given false alarm rate.Another result of our approach is a reduction of the word error rate (WER).
Axel Horndasch, Anton Batliner, Caroline Kaufhold, Elmar Nöth
INTERSPEECH4
2016 Assessing the Prosody of Non-Native Speakers of English: Measures and Feature Sets
Eduardo Coutinho, Florian Hönig, Yue Zhang 0014, Simone Hantke, Anton Batliner, Elmar Nöth, Björn W. Schuller
LREC6
2015 PATSY - it's all about pronunciation!
Caroline Kaufhold, Vadim Gamidov, Andreas Kießling 0001, Klaus Reinhard, Elmar Nöth
INTERSPEECH5
2015 Language-independent method for analysis of German stuttering recordings
abstract
The paper describes experiments where automatic acoustic al-gorithms initially intended to be used on Czech stuttering speak-ers were applied on recordings of German stuttering speak-ers. Four algorithms based on voice activity and abrupt spectral changes detection are introduced. The database consists of 34 speakers. The measure, the number of abrupt spectral changes in speech segments, reached a correlation with fluency rating of 0.85. The other measures have also good agreement with subjective evaluation. Results indicate that it could be basically possible to do language–independent analysis of stuttering, here demonstrated on read recordings of German speakers. Index Terms: automatic algorithms, stuttering, disfluency, Czech, German, language–independent
Tomas Lustyk, Petr Bergl, Tino Haderlein, Elmar Nöth, Roman Cmejla
INTERSPEECH4
2015 Voiced/unvoiced transitions in speech as a potential bio-marker to detect parkinson's disease
abstract
Several studies have addressed the automatic classification of speakers with Parkinson’s disease (PD) and healthy controls (HC). Most of the studies are based on speech recordings of sustained vowels, isolated words, and single sentences. Only few investigations have considered read texts and/or sponta-neous speech. This paper addresses two main questions still open regarding the automatic analysis speech in patients with PD, (a) “Is it possible to classify PD patients and HC through running speech signals in multiple languages?”, and (b) “where is the information to discriminate between speech recordings of PD patients and HC? ” In this paper speech recordings of read texts and monologues spoken in three different languages are considered. The energy content of the borders between voiced and unvoiced sounds is modeled. According to the results with read texts it is possible to achieve accuracies ranging from 91% to 98 % depending on the language. With respect to the re-sults on monologues, the accuracies are above 98 % in all of the three languages. The presence of discriminant information in the voiced/unvoiced and unvoiced/voiced transitions is vali-dated here, evidencing the problems of PD patients to stop/start the vocal folds movement during the production of running speech. Index Terms: Parkinson’s disease, dysarthria, hesitation in speech, language and motor planning, energy content, voiced/unvoiced transitions. 1.
Juan Rafael Orozco-Arroyave, Florian Hönig, Julián D. Arias-Londoño, Jesús Francisco Vargas-Bonilla, Sabine Skodda, Jan Rusz, Elmar Nöth
INTERSPEECH7
2015 The INTERSPEECH 2015 computational paralinguistics challenge: nativeness, parkinson's & eating condition
abstract
The INTERSPEECH 2015 Computational Paralinguistics Challenge addresses three different problems for the first time in research competition under well-defined conditions: the estimation of the degree of nativeness, the neurological state of patients with Parkinson’s condition, and the eating conditions of speakers, i. e., whether and which food type they are eating in a seven-class problem. In this paper, we describe these sub-challenges, their conditions, and the baseline feature extraction and classifiers, as provided to the participants. Index Terms: Computational Paralinguistics, Challenge, Degree of Nativeness, Parkinson’s Condition, Eating Condition
Björn W. Schuller, Stefan Steidl, Anton Batliner, Simone Hantke, Florian Hönig, Juan Rafael Orozco-Arroyave, Elmar Nöth, Yue Zhang 0014, Felix Weninger
INTERSPEECH7
2015 Automatic detection of parkinson's disease from continuous speech recorded in non-controlled noise conditions
Juan Camilo Vásquez-Correa, Tomás Arias-Vergara, Juan Rafael Orozco-Arroyave, Jesús Francisco Vargas-Bonilla, Julián D. Arias-Londoño, Elmar Nöth
INTERSPEECH6
2015 Low-frequency components analysis in running speech for the automatic detection of parkinson's disease
Tatiana Villa-Cañas, Julián D. Arias-Londoño, Juan Rafael Orozco-Arroyave, Jesús Francisco Vargas-Bonilla, Elmar Nöth
INTERSPEECH5
2015 Visual comparison of speaker groups
Sebastian Wankerl, Florian Hönig, Anton Batliner, Juan Rafael Orozco-Arroyave, Elmar Nöth
INTERSPEECH5
2015 A Survey on perceived speaker traits: Personality, likability, pathology, and the first challenge
Björn W. Schuller, Stefan Steidl, Anton Batliner, Elmar Nöth, Alessandro Vinciarelli, Felix Burkhardt, R. J. J. H. van Son, Felix Weninger, Florian Eyben, Tobias Bocklet, Gelareh Mohammadi, Benjamin Weiss 0001
Comput. Speech Lang.4
2015 Spectral and cepstral analyses for Parkinson's disease detection in Spanish vowels and words
abstract
Abstract About 1% of people older than 65 years suffer from Parkinson's disease (PD) and 90% of them develop several speech impairments, affecting phonation, articulation, prosody and fluency. Computer‐aided tools for the automatic evaluation of speech can provide useful information to the medical experts to perform a more accurate and objective diagnosis and monitoring of PD patients and can help also to evaluate the correctness and progress of their therapy. Although there are several studies that consider spectral and cepstral information to perform automatic classification of speech of people with PD, so far it is not known which is the most discriminative, spectral or cepstral analysis. In this paper, the discriminant capability of six sets of spectral and cepstral coefficients is evaluated, considering speech recordings of the five Spanish vowels and a total of 24 isolated words. According to the results, linear predictive cepstral coefficients are the most robust and exhibit values of the area under the receiver operating characteristic curve above 0.85 in 6 of the 24 words.
Juan Rafael Orozco-Arroyave, Florian Hönig, Julián D. Arias-Londoño, Jesús Francisco Vargas-Bonilla, Elmar Nöth
Expert Syst. J. Knowl. Eng.5
2015 Characterization Methods for the Detection of Multiple Voice Disorders: Neurological, Functional, and Laryngeal Diseases
abstract
This paper evaluates the accuracy of different characterization methods for the automatic detection of multiple speech disorders. The speech impairments considered include dysphonia in people with Parkinson's disease (PD), dysphonia diagnosed in patients with different laryngeal pathologies (LP), and hypernasality in children with cleft lip and palate (CLP). Four different methods are applied to analyze the voice signals including noise content measures, spectral-cepstral modeling, nonlinear features, and measurements to quantify the stability of the fundamental frequency. These measures are tested in six databases: three with recordings of PD patients, two with patients with LP, and one with children with CLP. The abnormal vibration of the vocal folds observed in PD patients and in people with LP is modeled using the stability measures with accuracies ranging from 81% to 99% depending on the pathology. The spectral-cepstral features are used in this paper to model the voice spectrum with special emphasis around the first two formants. These measures exhibit accuracies ranging from 95% to 99% in the automatic detection of hypernasal voices, which confirms the presence of changes in the speech spectrum due to hypernasality. Noise measures suitably discriminate between dysphonic and healthy voices in both databases with speakers suffering from LP. The results obtained in this study suggest that it is not suitable to use every kind of features to model all of the voice pathologies; conversely, it is necessary to study the physiology of each impairment to choose the most appropriate set of features.
Juan Rafael Orozco-Arroyave, Elkyn Alexander Belalcázar-Bolaños, Julián D. Arias-Londoño, Jesús Francisco Vargas-Bonilla, Sabine Skodda, Jan Rusz, Khaled Daqrouq, Florian Hönig, Elmar Nöth
IEEE J. Biomed. Health Informatics9
2014 A phonetic similarity based noisy channel approach to ASR hypothesis re-ranking and error detection
abstract
We present a new method to augment the correct transcript from automatic speech recognition (ASR) output containing multiple hypotheses. The error-prone ASR process is taken as black box and modeled as a noisy channel on phoneme level. The probabilities of the individual phoneme errors are assigned according to phonetic confusability. We score potential candidate hypotheses by their posterior probability of being the channel input given the competing ASR hypotheses as observed output. The resulting scores provide useful information not included in traditional confidence measures. We investigated the usefulness of the method for rescoring, re-ranking and word error detection. The method alone is not powerful enough to improve the recognition results, but by employing a decision tree classifier it is possible to isolate cases where the method works very well. Our results show that the combination with other knowledge sources and postprocessing techniques can lead to promising improvements.
Martin Hacker, Elmar Nöth
ICASSP2
2014 Are men more sleepy than women or does it only look like - Automatic analysis of sleepy speech
abstract
The degree of sleepiness in the Sleepy Language Corpus from the Interspeech 2011 Speaker State Challenge is predicted with regression and a very large feature vector. Most notable is the great gender difference which can mainly be attributed to females showing their sleepiness less than males do.
Florian Hönig, Anton Batliner, Tobias Bocklet, Georg Stemmer, Elmar Nöth, Sebastian Schnieder, Jarek Krajewski
ICASSP5
2014 Automatic modelling of depressed speech: relevant features and relevance of gender
abstract
Depression is an affective disorder characterised by psychomotor retardation; in speech, this shows up in reduction of pitch (variation, range), loudness, and tempo, and in voice qualities different from those of typical modal speech.A similar reduction can be observed in sleepy speech (relaxation).In this paper, we employ a small group of acoustic features modelling prosody and spectrum that have been proven successful in the modelling of sleepy speech, enriched with voice quality features, for the modelling of depressed speech within a regression approach.This knowledge-based approach is complemented by and compared with brute-forcing and automatic feature selection.We further discuss gender differences and the contributions of (groups of) features both for the modelling of depression and across depression and sleepiness.
Florian Hönig, Anton Batliner, Elmar Nöth, Sebastian Schnieder, Jarek Krajewski
INTERSPEECH3
2014 Automatic detection of parkinson's disease from words uttered in three different languages
abstract
About 90% of the people with Parkinson’s disease (PD) develop speech impairments such as monopitch, monoloudness, imprecise articulation, and other symptoms. There are several studies addressing the problem of the automatic detection of PD from speech signals in order to develop computer aided tools for the assessment and monitoring of the patients. Recent works have shown that it is possible to detect PD from speech with accuracies above 90%; however, it is still unclear whether it is possible to make the detection independent of the spoken language. This paper addresses the automatic detection of PD considering speech recordings of three languages: German, Spanish and Czech. According to the results it is possible to classify between speech of people with PD and healthy controls (HC) with accuracies ranging from 84% to 99%, depending on the utterance.
Juan Rafael Orozco-Arroyave, Florian Hönig, Julián D. Arias-Londoño, Jesús Francisco Vargas-Bonilla, Sabine Skodda, Jan Rusz, Elmar Nöth
INTERSPEECH7
2014 Erlangen-CLP: A Large Annotated Corpus of Speech from Children with Cleft Lip and Palate
Tobias Bocklet, Andreas K. Maier, Korbinian Riedhammer, Ulrich Eysholdt, Elmar Nöth
LREC5
2014 New Spanish speech corpus database for the analysis of people suffering from Parkinson's disease
Juan Rafael Orozco-Arroyave, Julián D. Arias-Londoño, Jesús Francisco Vargas-Bonilla, María Claudia Gonzalez-Rátiva, Elmar Nöth
LREC5
2013 Automatic phoneme analysis in children with Cleft Lip and Palate
abstract
Cleft Lip and Palate (CLP) is among the most frequent congenital abnormalities. The impaired facial development affects the articulation, with different phonemes being impacted inhomogeneously among different patients. This work focuses on automatic phoneme analysis of children with CLP for a detailed diagnosis and therapy control. In clinical routine, the state-of-the-art evaluation is based on perceptual evaluations. Perceptual ratings act as ground-truth throughout this work, with the goal to build an automatic system that is as reliable as humans. We propose two different automatic systems focusing on modeling the articulatory space of a speaker: one system models a speaker by a GMM, the other system employs a speech recognition system and estimates fMLLR matrices for each speaker. SVR is then used to predict the perceptual ratings. We show that the fMLLR-based system is able to achieve automatic phoneme evaluation results that are in the same range as perceptual inter-rater-agreements.
Tobias Bocklet, Korbinian Riedhammer, Ulrich Eysholdt, Elmar Nöth
ICASSP4
2013 Automatic evaluation of parkinson's speech - acoustic, prosodic and voice related cues
abstract
Articulation and phonation is affected in 70 % to 90 % of patients with Parkinson’s disease (PD). This study focuses on the question whether speech carries information about 1. PD being present at a speaker or not, and 2. estimating the sever-ity of PD (if present). We first perform classification experi-ments focusing on the automatic detection of PD as a 2-class problem (PD vs. healthy speakers). The detection of severity is described as a 3-class task based on the Unified Parkinson’s Disease Rating Scale (UPDRS) ratings. We employ acous-tic, prosodic and glottal features on different kinds of speech tests: various syllable repetition tasks, read sentences and texts, and monologues. Classification is performed in either case by SVMs. We report recognition results of 81.9 % when trying to differentiate between normally speaking persons and speakers with PD. With system fusion we achieved a recognition results of 59.1 % on the task of UPDRS classification. Index Terms: Parkinson’s Disease, pathologic speech, speech analysis
Tobias Bocklet, Stefan Steidl, Elmar Nöth, Sabine Skodda
INTERSPEECH3
2012 A software kit for automatic voice descrambling
abstract
Voice scrambling is widely used to add privacy to the radio communication of various authorities - but is also used by criminals to evade prosecution. In this article, we consider various analog voice scrambling techniques such as fixed frequency inversion, splitband inversion and rolling code scramblers. We explain how to break them using automatically extracted measures and scoring algorithms, and evaluate the proposed system using simulated data. While the simple inversion can be easily broken, the more advanced techniques require additional work prior to unsupervised automatization; the presented user interface allows the user to refine the automatic results to obtain a high quality solution.
Korbinian Riedhammer, Matthias Ring, Elmar Nöth, Dirk Kolb
ICC3
2012 The Automatic Assessment of Non-native Prosody: Combining Classical Prosodic Analysis with Acoustic Modelling
abstract
In earlier studies, we employed a large prosodic feature vector to assess the quality of L2 learner's utterances with respect to sentence melody and rhythm.In this paper, we combine these features with two standard approaches in paralinguistic analysis: (1) features derived from a Gaussian Mixture Model used as Universal Background Model (GMM-UBM), and (2) openSMILE, an open-source toolkit for extracting acoustic features.We evaluate our approach with English speech from 94 non-native speakers perceptually scored by 62 native labellers.GMM-UBM or openSMILE modelling alone yields lower performance than our prosodic feature vector; however, adding information from the GMM-UBM modelling or openSMILE by late fusion improves results.
Florian Hönig, Tobias Bocklet, Korbinian Riedhammer, Anton Batliner, Elmar Nöth
INTERSPEECH5
2012 Automatic detection of hypernasal speech signals using nonlinear and entropy measurements
abstract
Automatic hypernasality detection in children with Cleft Lip and Palate is classically performed by means of acoustic analysis; however, recent findings indicate that nonlinear dynamics features could be useful for this task. In order to continue deepening in this issue, in this paper the discriminant capability of 4 different nonlinear dynamics features along with a set of 6 entropy measurements is studied. The whole set of features is optimized using an automatic feature selection technique based on principal component analysis. The decision about the presence or absence of hypernasality is made by employing a support vector machine. The system is tested over two databases, one considers the five Spanish vowels and the words /coco/ and /gato/, and the other one considers different German words. The performance of the system is presented in terms of accuracy, sensitivity, specificity and receiver operating curves. According to the results, the accuracy of system increases when nonlinear and entropy measures are combined.
Juan Rafael Orozco-Arroyave, Julián D. Arias-Londoño, Jesús Francisco Vargas-Bonilla, Elmar Nöth
INTERSPEECH4
2012 The INTERSPEECH 2012 Speaker Trait Challenge
abstract
LIDIAP
Björn W. Schuller, Stefan Steidl, Anton Batliner, Elmar Nöth, Alessandro Vinciarelli, Felix Burkhardt, R. J. J. H. van Son, Felix Weninger, Florian Eyben, Tobias Bocklet, Gelareh Mohammadi, Benjamin Weiss 0001
INTERSPEECH4
2012 The FAU Video Lecture Browser system
abstract
A growing number of universities and other educational institutions provide recordings of lectures and seminars as an additional resource to the students. In contrast to educational films that are scripted, directed and often shot by film professionals, these plain recordings are typically not post-processed in an editorial sense. Thus, the videos often contain longer periods of inactivity or silence, unnecessary repetitions, or corrections of prior mistakes. This paper describes the FAU Video Lecture Browser system, a web-based platform for the interactive assessment of video lectures, that helps to close the gap between a plain recording and a useful e-learning resource by displaying automatically extracted and ranked key phrases on an augmented time line based on stream graphs. In a pilot study, users of the interface were able to complete a topic localization task about 29 % faster than users provided with the video only while achieving about the same accuracy. The user interactions can be logged on the server to collect data to evaluate the quality of the phrases and rankings, and to train systems that produce customized phrase rankings.
Korbinian Riedhammer, Martin Gropp, Elmar Nöth
SLT3
2011 Detection of persons with Parkinson's disease by acoustic, vocal, and prosodic analysis
abstract
70% to 90% of patients with Parkinson's disease (PD) show an affected voice. Various studies revealed, that voice and prosody is one of the earliest indicators of PD. The issue of this study is to automatically detect whether the speech/voice of a person is affected by PD. We employ acoustic features, prosodic features and features derived from a two-mass model of the vocal folds on different kinds of speech tests: sustained phonations, syllable repetitions, read texts and monologues. Classification is performed in either case by SVMs. A correlation-based feature selection was performed, in order to identify the most important features for each of these systems. We report recognition results of 91% when trying to differentiate between normal speaking persons and speakers with PD in early stages with prosodic modeling. With acoustic modeling we achieved a recognition rate of 88% and with vocal modeling we achieved 79%. After feature selection these results could greatly be improved. But we expect those results to be too optimistic. We show that read texts and monologues are the most meaningful texts when it comes to the automatic detection of PD based on articulation, voice, and prosodic evaluations. The most important prosodic features were based on energy, pauses and F0. The masses and the compliances of spring were found to be the most important parameters of the two-mass vocal fold model.
Tobias Bocklet, Elmar Nöth, Georg Stemmer, Hana Ruzickova, Jan Rusz
ASRU2
2011 Associating children's non-verbal and verbal behaviour: Body movements, emotions, and laughter in a human-robot interaction
abstract
In this article, we associate different types of vocal behaviour denoting emotional user states and laughter with different types of body movements such as gestures, forward bends, or liveliness. Our subjects are German children giving commands to Sony's Aibo robot; the data are fully realistic. The analysis reveals characteristic and significant co-occurrences of body movements and vocal events.
Anton Batliner, Stefan Steidl, Elmar Nöth
ICASSP3
2011 Compensation of extrinsic variability in speaker verification systems on simulated Skype and HF channel data
abstract
In this work we focus on speaker verification on channels of varying quality, namely Skype and high frequency (HF) radio. In our setup, we assume to have telephone recordings of speakers for training, but recordings of different channels for testing with varying (lower) signal quality. Starting from a Gaussian mixture / support vector machine (GMM/SVM) baseline, we evaluate multi-condition training (MCT), an ideal channel classification approach (ICC), and nuisance attribute projection (NAP) to compensate for the loss of information due to the transmission. In an evaluation on Switchboard-2 data using Skype and HF channel simulators, we show that, for good signal quality, NAP improves the baseline system performance from 5% EER to 3.33% EER (for both Skype and HF). For strongly distorted data, MCT or, if adequate, ICC turn out to be the method of choice.
Korbinian Riedhammer, Tobias Bocklet, Elmar Nöth
ICASSP3
2011 Drink and Speak: On the Automatic Classification of Alcohol Intoxication by Acoustic, Prosodic and Text-Based Features
abstract
This paper focuses on the automatic detection of a person’s blood level alcohol based on automatic speech processing ap-proaches. We compare 5 different feature types with different ways of modeling. Experiments are based on the ALC corpus of IS2011 Speaker State Challenge. The classification task is restricted to the detection of a blood alcohol level above 0.5‰. Three feature sets are based on spectral observations: MFCCs, PLPs, TRAPS. These are modeled by GMMs. Classification is either done by a Gaussian classifier or by SVMs. In the later case classification is based on GMM-based supervectors, i.e. concatenation of GMM mean vectors. A prosodic system extracts a 292-dimensional feature vector based on a voiced-unvoiced decision. A transcription-based system makes use of text transcriptions related to phoneme durations and textual structure. We compare the stand-alone performances of these systems and combine them on score level by logistic regres-sion. The best stand-alone performance is the transcription-based system which outperforms the baseline by 4.8 % on the development set. A Combination on score level gave a huge boost when the spectral-based systems were added (73.6 %). This is a relative improvement of 12.7 % to the baseline. On the test-set we achieved an UA of 68.6 % which is a significant improvement of 4.1 % to the baseline system. Index Terms: GMM, alcohol intoxication, system fusion 1.
Tobias Bocklet, Korbinian Riedhammer, Elmar Nöth
INTERSPEECH3
2011 Does it Groove or does it Stumble - Automatic Classification of Alcoholic Intoxication using Prosodic Features
abstract
This paper studies how prosodic features can help in the automatic detection of alcoholic intoxication.We compute features that have recently been proposed to model speech rhythm such as the pair-wise variability index for consonantal and vocalic segments (PVI) and study their aptness for the task.Further, we use a large prosodic feature vector modelling the usual candidates -pitch, intensity, and duration -and apply it onto different units such as words, syllables and stressed syllables to create generalizations of the rhythm features mentioned.The results show that the prosodic features computed are suitable for detecting alcoholic intoxication and add complementary information to state-of-the-art features.The database is the intoxication database provided by the organizers of the 2011 Interspeech Speaker State Challenge.
Florian Hönig, Anton Batliner, Elmar Nöth
INTERSPEECH3
2011 Combining Phonological and Acoustic ASR-Free Features for Pathological Speech Intelligibility Assessment
abstract
Intelligibility is widely used to measure the severity of articulatory problems in pathological speech. Recently, a number of automatic intelligibility assessment tools have been developed. Most of them use automatic speech recognizers (ASR) to compare the patient's utterance with the target text. These methods are bound to one language and tend to be less accurate when speakers hesitate or make reading errors. To circumvent these problems, two different ASR-free methods were developed over the last few years, only making use of the acoustic or phonological properties of the utterance. In this paper, we demonstrate that these ASR-free techniques are also able to predict intelligibility in other languages. Moreover, they show to be complementary, resulting in even better intelligibility predictions when both methods are combined.
Catherine Middag, Tobias Bocklet, Jean-Pierre Martens, Elmar Nöth
INTERSPEECH4
2011 Java Visual Speech Components for Rapid Application Development of GUI Based Speech Processing Applications
abstract
In this paper, we describe a new Java framework for an easy and efficient way of developing new GUI based speech processing applications. Standard components are provided to display the speech signal, the power plot, and the spectrogram. Furthermore, a component to create a new transcription and to display and manipulate an existing transcription is provided, as well as a component to display and manually correct external pitch values. These Swing components can be easily embedded into own Java programs. They can be synchronized to display the same region of the speech file. The object-oriented design provides base classes for rapid development of own components.
Stefan Steidl, Korbinian Riedhammer, Tobias Bocklet, Florian Hönig, Elmar Nöth
INTERSPEECH5
2011 A scalable architecture for multilingual speech recognition on embedded devices
Martin Raab, Rainer Gruhn, Elmar Nöth
Speech Commun.3
2010 Clap your hands! Calibrating spectral subtraction for dereverberation
abstract
Reverberation effects as observed by room microphones severely degrade the performance of automatic speech recognition systems. We investigate the use of dereverberation by spectral subtraction as proposed by Lebart and Boucher and introduce a simple approach to estimate the required decay parameter by clapping hands. Experiments on small vocabulary continuous speech recognition task on read speech show that using the calibrated dereverberation improves WER from 73.2 to 54.7 for the best microphone. In combination with system adaptation, the WER could be reduced to 28.2, which is only a 16% relative loss of performance comparison to using a headset instead of a room microphone.
Uwe Zah, Korbinian Riedhammer, Tobias Bocklet, Elmar Nöth
ICASSP4
2010 Age and gender recognition based on multiple systems - early vs. late fusion
abstract
This paper focuses on the automatic recognition of a per-son’s age and gender based only on his or her voice. Up to five different systems are compared and combined in dif-ferent configurations: three systems model the speaker’s characteristics in different feature spaces, i.e., MFCC, PLP, TRAPS, by Gaussian mixture models. The features of these systems are the concatenated mean vectors. Sys-tem number 4 uses a physical two-mass vocal model and estimates in a data-driven optimization procedure 9 glot-tal features from voiced speech sections. For each ut-terance the minimum, maximum and mean vectors form a 27-dimensional feature vector. The last system calcu-lates a 219-dimensional prosodic feature set for each ut-terance based on voice and unvoiced speech segments. We compare two different ways to fuse the different sys-tems: First, we concatenate the system on feature level. The second way of combination is performed on score level by multi-class logistic regression. Despite there are just minor differences between the two approaches, late fusion is slightly superior. On the development set of the Interspeech Agender challenge we achieved an un-weighted recall of 46.1 % with early fusion and 47.8% with late fusion.
Tobias Bocklet, Georg Stemmer, Viktor Zeißler, Elmar Nöth
INTERSPEECH4
2010 FAU IISAH Corpus -- A German Speech Database Consisting of Human-Machine and Human-Human Interaction Acquired by Close-Talking and Far-Distance Microphones
Werner Spiegl, Korbinian Riedhammer, Stefan Steidl, Elmar Nöth
LREC4
2010 Improvement of a speech recognizer for standardized medical assessment of children's speech by integration of prior knowledge
abstract
Speech recognition of children is a more difficult task than speech recognition of adults. This problem is amplified for children with articulation disorders like cleft lip and palate (CLP). In this work we improved our automatic speech recognition system by integrating prior knowledge. Prior knowledge focuses on two different aspects: A test-dependent language modeling and an age-dependent acoustic modeling. These two approaches are merged at the end to different test- and age-dependent recognizers. We evaluated our system on a dataset of 35 children with CLP. Significant improvements could be found on this dataset. With our baseline system we achieved a negative word accuarcy (WA) of -11.0%. By an extended language modeling we achieved 27.5%. The age-dependent recognition system gains a huge improvement and achieves aWA of 42.6%. With the significant improvements in WA it is possible to perform an automatic detection and identification of specific words. Thus, we took the first step towards a speech assessment on word and subword level.
Tobias Bocklet, Andreas K. Maier, Ulrich Eysholdt, Elmar Nöth
SLT4
2009 A language-independent feature set for the automatic evaluation of prosody
abstract
In second language learning, the correct use of prosody plays a vital role.Therefore, an automatic method to evaluate the naturalness of the prosody of a speaker is desirable.We present a novel method to model prosody independently of the text and thus independently of the language as well.For this purpose, the voiced and unvoiced speech segments are extracted and a 187-dimensional feature vector is computed for each voiced segment.This approach is compared to word based prosodic features on a German text passage.Both are confronted with the perceptive evaluation of two native speakers of German.The word-based feature set yielded correlations of up to 0.92 while the text-independent feature set yielded 0.88.This is in the same range as the inter-rater correlation with 0.88.Furthermore, the text-independent features were computed for a Japanese translation of the passage which was also rated by two native speakers of Japanese.Again, the correlation between the automatic system and the human perception of the naturalness was high with 0.83 and not significantly lower than the inter-rater correlation of 0.92.
Andreas K. Maier, Florian Hönig, Viktor Zeißler, Anton Batliner, Erik Körner, Nobuyuki Yamanaka, Peter Ackermann, Elmar Nöth
INTERSPEECH8
2009 A microphone-independent visualization technique for speech disorders
abstract
In this paper we introduce a novel method for the visualization of speech disorders. We demonstrate the method with disordered speech and a control group. However, both groups were recorded using two different microphones. The projection of the patient data using a single microphone yields significant correlations between the coordinates on the map and certain criteria of the disorder which were perceptually rated. However, projection of data from multiple microphones reduces this correlation. Usually, the acoustical mismatch between the microphones is greater than the mismatch between the speakers, i.e., not the disorders but the microphones form clusters in the visualization. Based on an extension of the Sammon mapping, we are able to create a map which projects the same speakers onto the same position even if multiple microphones are used. Furthermore, our method also restores the correlation between the map coordinates and the perceptual assessment. Index Terms: visualization, robustness, speech processing.
Andreas K. Maier, Stefan Wenhardt, Tino Haderlein, Maria Schuster, Elmar Nöth
INTERSPEECH5
2009 Online generation of acoustic models for multilingual speech recognition
abstract
Our goal is to provide a multilingual speech based Human Machine Interface for in-car infotainment and navigation systems. The multilinguality is for example needed for music player control via speech as artist and song names in the globalized music market come from many languages. Another frequent use case is the input of foreign navigation destinations via speech. In this paper we propose approximated projections between mixtures of Gaussians that allow the generation of the multilingual system from monolingual systems. This makes the creation of the multilingual systems on an embedded system possible with the benefit that training and maintenance effort remain unchanged compared to the provision of monolingual systems. We also sketch how this algorithm can help together with our previous work to have an efficient architecture for multilingual speech recognition on embedded devices.
Martin Raab, Guillermo Aradilla, Rainer Gruhn, Elmar Nöth
INTERSPEECH4
2009 Intelligibility assessment in children with cleft lip and palate in Italian and German
abstract
Current research has shown that the speech intelligibility in children with cleft lip and palate (CLP) can be estimated automatically using speech recognition methods. On German CLP data high and significant correlations between human ratings and the recognition accuracy of a speech recognition system were already reported. In this paper we investigate whether the approach is also suitable for other languages. Therefore, we compare the correlations obtained on German data with the correlations on Italian data. A high and significant correlation (r=0.76; p 0.05). Index Terms: speech recognition, speech intelligibility, cleft lip and palate.
Marcello Scipioni, Matteo Gerosa, Diego Giuliani, Elmar Nöth, Andreas K. Maier
INTERSPEECH4
2009 Analyzing features for automatic age estimation on cross-sectional data
abstract
We develop an acoustic feature set for the estimation of a per-son’s age from a recorded speech signal. The baseline features are Mel-frequency cepstral coefficients (MFCCs) which are ex-tended by various prosodic features, pitch and formant frequen-cies. From experiments on the University of Florida Vocal Ag-ing Database we can draw different conclusions. On the one hand, adding prosodic, pitch and formant features to the MFCC baseline leads to relative reductions of the mean absolute error between 4-20%. Improvements are even larger when percep-tual age labels are taken as a reference. On the other hand, reasonable results with a mean absolute error in age estimation of about 12 years are already achieved using a simple gender-independent setup and MFCCs only. Future experiments will evaluate the robustness of the prosodic features against channel variability on other databases and investigate the differences be-tween perceptual and chronological age labels.
Werner Spiegl, Georg Stemmer, Eva Lasarcyk, Varada Kolhatkar, Andrew S. Cassidy, Blaise Potard, Stephen H. Shum, Young Chol Song, Puyang Xu, Peter Beyerlein, James D. Harnsberger, Elmar Nöth
INTERSPEECH12
2009 Automatic pronunciation scoring of words and sentences independent from the non-native's first language
Tobias Cincarek, Rainer Gruhn, Christian Hacker, Elmar Nöth, Satoshi Nakamura 0001
Comput. Speech Lang.4
2009 PEAKS - A system for the automatic evaluation of voice and speech disorders
Andreas K. Maier, Tino Haderlein, Ulrich Eysholdt, Frank Rosanowski, Anton Batliner, Maria Schuster, Elmar Nöth
Speech Commun.7
2008 Age and gender recognition for telephone applications based on GMM supervectors and support vector machines
abstract
This paper compares two approaches of automatic age and gender classification with 7 classes. The first approach are Gaussian mixture models (GMMs) with universal background models (UBMs), which is well known for the task of speaker identification/verification. The training is performed by the EM algorithm or MAP adaptation respectively. For the second approach for each speaker of the test and training set a GMM model is trained. The means of each model are extracted and concatenated, which results in a GMM supervector for each speaker. These supervectors are then used in a support vector machine (SVM). Three different kernels were employed for the SVM approach: a polynomial kernel (with different polynomials), an RBF kernel and a linear GMM distance kernel, based on the KL divergence. With the SVM approach we improved the recognition rate to 74% (p < 0.001) and are in the same range as humans.
Tobias Bocklet, Andreas K. Maier, Josef G. Bauer, Felix Burkhardt, Elmar Nöth
ICASSP5
2008 Multilingual weighted codebooks
abstract
In this paper we present an approach for speech recognition of multiple languages with constrained resources on embedded devices. Examples of such systems are navigation systems, mobile phones and MP3 players. Speech recognizers on such systems are typically to-date semi-continuous speech recognizers, which are based on vector quantization. Typical vector quantization algorithms can only generate vector quantization prototypes that are optimal for one language. We hypothesize and provide evidence that a certain fixed vector quantization is responsible for a significant drop of recognition performance when a recognizer is extended to recognize multiple languages at the same time. This paper proposes an algorithm for the construction of multilingual weighted codebooks (MWCs). These MWCs have the advantage that they offer significantly improved performance for the recognition of multiple languages.
Martin Raab, Rainer Gruhn, Elmar Nöth
ICASSP3
2008 Automatic evaluation of characteristic speech disorders in children with cleft lip and palate
abstract
Abstract This paper discusses the automatic evaluation of speech of chil-dren with cleft lip and palate (CLP). CLP speech shows specialcharacteristics such as hypernasality, backing, and weakeningof plosives. In total ve criteria were subjectively assessed byan experienced speech expert on the phone level. This subjec-tive evaluation was used as a gold standard to train a classi-cation system. The automatic system achieves recognition re-sults on frame, phone, and word level of up to 75.8% CL. Onspeaker level signicant and high correlations between the sub-jective evaluation and the automatic system of up to 0.89 areobtained. Index Terms : pathologic speech, speech assessment, pronun-ciation scoring, children’s speech 1. Introduction Cleft Lip and Palate (CLP) is the most common malformationof the head. It constitutes almost two-thirds of the major facialdefects and almost 80% of all orofacial clefts [1]. Its prevalencediffers in different populations from 1 in 400 to 500 newborns inAsians to 1 in 1500 to 2000 in African Americans. The preva-lence in Caucasians is 1 in 750 to 900 births [2, 3].In clinical practice, articulation disorders are mainly eval-uated by subjective tools. The simplest method is the audi-tive perception, mostly performed by a speech therapist. Pre-vious studies have shown that experience is an important fac-tor that inuences the subjective estimation of speech disorderswhich leads to inaccurate evaluation by persons with only fewyears of experience as speech therapist [4]. Until now, objectivemeans exist only for quantitative measurements of nasal emis-sions [5, 6, 7] and for the detection of secondary voice disorders[8]. But other specic articulation disorders in CLP cannot besufciently quantied.In this paper, we present a new technical procedure for themeasurement and evaluation of specic speech disorders andcompare the results obtained with subjective ratings of an expe-rienced speech therapist.
Andreas K. Maier, Florian Hönig, Christian Hacker, Maria Schuster, Elmar Nöth
INTERSPEECH5
2008 Private emotions versus social interaction: a data-driven approach towards analysing emotion in speech
Anton Batliner, Stefan Steidl, Christian Hacker, Elmar Nöth
User Model. User Adapt. Interact.4
2007 Non-native speech databases
abstract
This paper presents a review of already collected non-native speech databases. Although the number of non-native speech databases is significantly less than the one of common speech databases, there were already a lot of data collection efforts taken at different institutes and companies. Because of the comparably small size of the databases, many of them are not available through the common distributors of speech corpora like ELDA or LDC. This leads to the fact that it is hard to keep an overview of what kind of databases have already been collected, and for what purposes there are still no collections. With this paper we hope to provide a useful resource regarding this issue.
Martin Raab, Rainer Gruhn, Elmar Nöth
ASRU3
2007 Towards robust automatic evaluation of pathologic telephone speech
abstract
For many aspects of speech therapy an objective evaluation of the intelligibility of a patient's speech is needed. We investigate the evaluation of the intelligibility of speech by means of automatic speech recognition. Previous studies have shown that measures like word accuracy are consistent with human experts' ratings. To ease the patient's burden, it is highly desirable to conduct the assessment via phone. However, the telephone channel influences the quality of the speech signal which negatively affects the results. To reduce inaccuracies, we propose a combination of two speech recognizers. Experiments on two sets of pathological speech show that the combination results in consistent improvements in the correlation between the automatic evaluation and the ratings by human experts. Furthermore, the approach leads to reductions of 10% and 25% of the maximum error of the intelligibility measure.
Korbinian Riedhammer, Georg Stemmer, Tino Haderlein, Maria Schuster, Frank Rosanowski, Elmar Nöth, Andreas K. Maier
ASRU6
2007 Boosting of Prosodic and Pronunciation Features to Detect Mispronunciations of Non-Native Children
abstract
Commercial products that support L2-learners with computer assisted pronunciation training usually focus per exercise only on one possible pronunciation mistake that is typical for speakers of the respective L1 group. Acoustic models for words with wrong pronunciation are added to the system. In the present paper a more general approach with features that have proved to be widely independent of the learners' mother tongue is proposed. It is able to take various possible mistakes into consideration all at once. High dimensional feature vectors that encode prosodic varieties and differences of reference and recognized sentences are analyzed. With the ADABOOST algorithm those features are found, which contain the most important information to assess German children learning English. With 35 features 89 % of the agreement of experts is achieved.
Christian Hacker, Tobias Cincarek, Andreas K. Maier, Andre Heßler, Elmar Nöth
ICASSP (4)5
2007 Automatic scoring of the intelligibility in patients with cancer of the oral cavity
abstract
After surgical treatment of cancer of the oral cavity patients often suffer from functional restrictions such as speech disorders.In this paper we present a novel approach to assess the outcome of the treatment w.r.t. the intelligibility of the patient using the result of an automatic speech recognition system.The word recognition rate was taken as intelligibility score.Compared to four speech experts this method yields results that are as good as the best speech expert compared to the other experts.The correlation between our system and the mean opinion of the experts is .92.Furthermore we show that our system has better performance than the average expert and is more reliable.
Andreas K. Maier, Maria Schuster, Anton Batliner, Elmar Nöth, Emeka Nkenke
INTERSPEECH4
2006 Phoneme-to-grapheme mapping for spoken inquiries to the semantic web
abstract
Automatic methods for grapheme-to-phoneme (G2P) and phoneme-to-grapheme (P2G) conversion have become very popular in recent years.Their performance has improved considerably, while at the same time these developments required less input from expert lexicographers.Continuing in this tradition we will present in this paper a data-driven, language-independent approach called MASSIVE 1 with which it is possible to create efficient online modules for automatic symbol mapping.Our framework is solely based on statistical methods for training and run-time and has been optimized for P2G conversion in the context of spoken inquiries to the Semantic Web, an issue researched in the SmartWeb project 2 .MASSIVE systems can be trained using a pronunciation lexicon, the output of a phone recognizer or any other suitable set of corresponding symbol strings.Successful tests have been performed on German and English data sets.
Axel Horndasch, Elmar Nöth, Anton Batliner, Volker Warnke
INTERSPEECH2
2005 Can you Understand him? Let's Look at his Word Accuracy - Automatic Evaluation of Tracheoesophageal Speech
abstract
Tracheoesophageal (TE) speech is a possibility to restore the ability to speak after laryngectomy. TE speech often shows low intelligibility. An objective means to determine and quantify the intelligibility does not exist until now and an automation of this procedure is desirable. We used a speech recognizer trained on normal, non-pathologic voices. We compared intelligibility scores for TE speech from five experienced raters with the word accuracy (WA) of our speech recognizer. A correlation coefficient of -0.84 shows that WA can be a good indicator of intelligibility for pathologic voices. An outlook for future work is presented.
Maria Schuster, Elmar Nöth, Tino Haderlein, Stefan Steidl, Anton Batliner, Frank Rosanowski
ICASSP (1)2
2005 "Of All Things the Measure Is Man" : Automatic Classification of Emotions and Inter-Labeler Consistency
abstract
In traditional classification problems, the reference needed for training a classifier is given and considered to be absolutely correct. However, this does not apply to all tasks. In emotion recognition in non-acted speech, for instance, one often does not know which emotion was really intended by the speaker. Hence, the data is annotated by a group of human labelers who do not agree on one common class in most cases. Often, similar classes are confused systematically. We propose a new entropy-based method to evaluate classification results taking into account these systematic confusions. We can show that a classifier which achieves a recognition rate of "only" about 60 % on a four-class-problem performs as well as our five human labelers on average.
Stefan Steidl, Michael Levit, Anton Batliner, Elmar Nöth, Heinrich Niemann
ICASSP (1)4
2005 Tales of tuning - prototyping for automatic classification of emotional user states
abstract
Classification performance for emotional user states found in the few realistic, spontaneous databases available is as yet not very high.We present a database with emotional children's speech in a human-robot scenario.Baseline classification performance for seven classes is 44.5%, for four classes 59.2%.We discuss possible strategies for tuning, e.g., using only prototypes (based on annotation correspondence or classification scores), or taking into account requirements and feasibility in possible applications (weighting of false alarms or speakerspecific overall frequencies).
Anton Batliner, Stefan Steidl, Christian Hacker, Elmar Nöth, Heinrich Niemann
INTERSPEECH4
2004 Aspects of named entity processing
abstract
In this paper we investigate the utility of three aspects of named entity processing: detection, localization and value extraction. We corroborate this task categorization by providing examples of practical applications for each of these subtasks. We also suggest methods for tackling these subtasks, giving particular attention to working with speech data. We employ Support Vector Machines to solve the detection task and show how localization and value extraction can successfully be dealt with using a combination of grammar-based and statistical methods. 1.
Michael Levit, Allen L. Gorin, Patrick Haffner, Hiyan Alshawi, Elmar Nöth
INTERSPEECH5
2004 Adaptation in the pronunciation space for non-native speech recognition
abstract
We introduce a new technique to improve the recognition of non-native speech. The underlying assumption is that for each non-native pronunciation of a speech sound, there is at least one sound in the target language that has a similar native pronunciation. The adaptation is performed by HMM interpolation between adequate native acoustic models. The interpolation partners are determined automatically in a data-driven manner. Our experiments show that this technique is suitable for both the offline adaptation to a whole group of speakers as well as for the unsupervised online adaptation to a single speaker. Results are given both for spontaneous non-native English speech as well as for a set of read non-native German utterances.
Georg Stemmer, Stefan Steidl, Christian Hacker, Elmar Nöth
INTERSPEECH4
2004 "You Stupid Tin Box" - Children Interacting with the AIBO Robot: A Cross-linguistic Emotional Speech Corpus
Anton Batliner, Christian Hacker, Stefan Steidl, Elmar Nöth, Shona D'Arcy, Martin J. Russell
LREC4
2003 Optimizing Eigenfaces by Face Masks for Facial Expression Recognition
Carmen Frank, Elmar Nöth
CAIP2
2003 A phone recognizer helps to recognize words better
abstract
For most speech recognition systems dynamic features are the only way to incorporate temporal context into the output distributions of the HMM. In this paper we propose an efficient method to utilize a large context in the recognition process. State scores of a phone recognizer which runs in parallel to the word recognizer are computed. Integrating these scores in the HMM of the word recognizer makes their output densities context-dependent. The approach is evaluated on a set of spontaneous utterances which have been recorded with our spoken dialogue system. A significant reduction of the word error rate has been achieved.
Georg Stemmer, Viktor Zeißler, Christian Hacker, Elmar Nöth, Heinrich Niemann
ICASSP (1)4
2003 We are not amused - but how do you know? user states in a multi-modal dialogue system
abstract
For the multi-modal dialogue system SmartKom, emotional user states in a Wizard-of-Oz experiment as, e.g., joyful, angry, helpless, are annotated holistically and based purely on facial expressions; other phenomena (prosodic peculiarities, offtalk, i.e., speaking aside, etc.) are labelled as well.We present the correlations between these different annotations and report classification results using a large prosodic feature vector.The performance of the user state classification is not yet satisfactory; possible reasons and remedies are discussed.
Anton Batliner, Viktor Zeißler, Carmen Frank, Johann Adelhardt, Rui Ping Shi, Elmar Nöth
INTERSPEECH6
2003 Context-sensitive evaluation and correction of phone recognition output
abstract
In speech and language processing, information about the errors made by a learning system is commonly used to assess and improve its performance. Because of high computational complexity, the context of the errors is usually either ignored, or exploited in a simplistic form. The complexity becomes tractable, however, for phone recognition because of the small lexicon. For phonebased systems, an exhaustive modeling of local context is possible. Furthermore, recent research studies have shown phone recognition to be useful for several spoken language processing tasks. In this paper, we present a mechanism which learns patterns of context-sensitive errors from ASR-output aligned with the “true” phone transcriptions. We also show how this information, encoded as a context-sensitive weighted transducer, can provide a modest improvement to phone recognition accuracy even when no transcriptions are available for the domain of interest.
Michael Levit, Hiyan Alshawi, Allen L. Gorin, Elmar Nöth
INTERSPEECH4
2003 Acoustic normalization of children's speech
abstract
Young speakers are not represented adequately in current speech recognizers. In this paper we focus on the problem to adapt the acoustic frontend of a speech recognizer which has been trained on adults ’ speech to achieve a better performance on speech from children. We introduce and evaluate a method to perform non-linear VTLN by an unconstrained data-driven op-timization of the filterbank. A second approach normalizes the speaking rate of the young speakers with the PSOLA algorithm. Significant reductions in word error rate have been achieved.
Georg Stemmer, Christian Hacker, Stefan Steidl, Elmar Nöth
INTERSPEECH4
2003 Context-dependent output densities for hidden Markov models in speech recognition
abstract
In this paper we propose an efficient method to utilize context in the output densities of HMMs. State scores of a phone recognizer are integrated into the HMMs of a word recognizer which makes their output densities context-dependent. A significant reduction of the word error rate has been achieved when the approach is evaluated on a set of spontaneous speech utterances. As we can expect that context is more important for some phone models than for others, we further extend the approach by statedependent weighting factors which are used to control the influence of the different information sources. A small additional improvement has been achieved.
Georg Stemmer, Viktor Zeißler, Christian Hacker, Elmar Nöth, Heinrich Niemann
INTERSPEECH4
2003 MOBSY: Integration of vision and dialogue in service robots
Matthias Zobel, Joachim Denzler, Benno Heigl, Elmar Nöth, Dietrich Paulus, Jochen Schmidt, Georg Stemmer
Mach. Vis. Appl.4
2003 How to find trouble in communication
Anton Batliner, K. Fischer, Richard Huber, Jörg Spilker, Elmar Nöth
Speech Commun.5
2002 Using EM-trained string-edit distances for approximate matching of acoustic morphemes
abstract
Our research concerns spoken language understanding within the domain of automated telecommunication services. In the recent papers we presented a new methodology for training of statistical language models for recognition and understanding of utterances from large corpora of phone sequences obtained as the output of a task-independent ASR-system. The advantage of this strategy compared to the traditional word-based strategy is that we don't have to manually transcribe large amounts of data in order to extract acoustic morphemes to train the classifier. Since the baseline strategy suffered high False Rejection Rates caused by finding no acoustic morphemes in the test data, we describe in this paper how approximate matching can be incorporated in the Bayes-classifier to reduce FRR. The experiments are evaluated for "How May I Help You?"-task.
Michael Levit, Elmar Nöth, Allen L. Gorin
INTERSPEECH2
2002 Integrated recognition of words and prosodic phrase boundaries
Florian Gallwitz, Heinrich Niemann, Elmar Nöth, Volker Warnke
Speech Commun.3
2002 On the use of prosody in automatic dialogue understanding
Elmar Nöth, Anton Batliner, Volker Warnke, Manuela Boros, Jan Buckow, Richard Huber, Florian Gallwitz, M. Nutt, Heinrich Niemann
Speech Commun.1
2001 Boiling down prosody for the classification of boundaries and accents in German and English
abstract
In the focus of this paper is a comparison of the most relevant prosodic features/feature classes for the classification of boundaries and accents in German and in English.Principal components were computed based on a large prosodic feature vector; these principal components were used as predictor variables in a Linear Discriminant analysis as well as in a Classification and Regression Tree.The number of the most relevant principal components was between three and five; for both languages and for boundary and accent classification alike, most important were principal components modelling duration, in combination with energy, followed by pauses and F0.
Anton Batliner, Jan Buckow, Richard Huber, Volker Warnke, Elmar Nöth, Heinrich Niemann
INTERSPEECH5
2001 Prosodic models, automatic speech understanding, and speech synthesis: towards the common ground
abstract
Automatic speech understanding and speech synthesis, two of the major speech processing applications, impose strikingly different constraints and requirements on prosodic models.The prevalent models of prosody and intonation fail to offer a unified solution to these conflicting constraints.As a consequence, prosodic models have been applied only occasionally in end-toend automatic speech understanding systems; in contrast, they have been applied extensively in speech synthesis systems.In this paper we want to discuss the reasons for this state of affairs as well as possible strategies to overcome the shortcomings of the use of prosodic modelling in automatic speech processing.
Anton Batliner, Bernd Möbius, Gregor Möhler, Antje Schweitzer, Elmar Nöth
INTERSPEECH5
2001 Acoustic modeling of foreign words in a German speech recognition system
abstract
... foreign words for a German speech recognizer. The recognition quality of foreign words is crucial for the overall performance of a system in application fields like spoken dialogue systems, when foreign words occur as proper names. One of the main problems in the modeling of foreign words is the limitation of training data, which must contain samples of the non-native pronunciation of the foreign sounds. In order to obtain robust acoustic models, which are still precise enough, we compare several methods to map or to merge the models of phonemes, which are pronounced in a similar way by German speakers. We utilize an entropy-based distance measure between sets of phoneme models. The best approach yields a reduction of 16.5% word error rate, when compared to a baseline system.
Georg Stemmer, Elmar Nöth, Heinrich Niemann
INTERSPEECH2
2000 Recognition of emotion in a realistic dialogue scenario
abstract
Nowadays modern automatic dialogue systems are able to understand complex sentences instead of only a few commands like Stop or No.In a call-center, such a system should be able to determine in a critical phase of the dialogue if the call should be passed over to a human operator.Such a critical phase can be indicated by the customer's vocal expression.Other studies prooved that it is possible to distinguish between anger and neutral speech w i t h prosodic features alone.Subjects in these studies were mostly people acting or simulating emotions like anger.In this paper we use data from a so-called Wizard of O z (WoZ) scenario to get more realistic data instead of simulated anger.As shown below, the classi cation rate for the two classes "emotion" (class E) and "neutral" (class :E) is signi cantly worse for these more realistic data.Furthermore the classi cation results are heavily speaker dependent.Prosody alone might t h us not be sucient and has to be supplemented by the use of other knowledge sources such as the detection of repetitions, reformulations, swear words, and dialogue acts.
Richard Huber, Anton Batliner, Jan Buckow, Elmar Nöth, Volker Warnke, Heinrich Niemann
INTERSPEECH4
2000 Automatic stuttering recognition using hidden Markov models
Elmar Nöth, Heinrich Niemann, Tino Haderlein, Michael Decher, Ulrich Eysholdt, Frank Rosanowski, Thomas Wittenberg
INTERSPEECH1
2000 Labeling of Prosodic Events in Slovenian Speech Database GOPOLIS
France Mihelic, Jerneja Zganec-Gros, Elmar Nöth, Volker Warnke
LREC3
2000 VERBMOBIL: the use of prosody in the linguistic components of a speech understanding system
abstract
We show how prosody can be used in speech understanding systems. This is demonstrated with the VERBMOBIL speech to-speech translation system which, to our knowledge, is the first complete system which successfully uses prosodic information in the linguistic analysis. Prosody is used by computing probabilities for clause boundaries, accentuation, and different types, of sentence mood for each of the word hypotheses computed by the word recognizer. These probabilities guide the search of the linguistic analysis. Disambiguation is already achieved during the analysis and not by a prosodic verification of different linguistic hypotheses. So far, the most useful prosodic information is provided by clause boundaries. These are detected with a recognition rate of 94%. For the parsing of word hypotheses graphs, the use of clause boundary probabilities yields a speed-up of 92% and a 96% reduction of alternative readings.
Elmar Nöth, Anton Batliner, Andreas Kießling 0001, Ralf Kompe, Heinrich Niemann
IEEE Trans. Speech Audio Process.1
1999 Discriminative estimation of interpolation parameters for language model classifiers
abstract
In this paper we present a new approach for estimating the interpolation parameters of language models (LM) which are used as classifiers. With the classical maximum likelihood (ML) estimation theoretically one needs to have a huge amount of data and the fundamental density assumption has to be correct. Usually one of these conditions is violated, so different optimization techniques like maximum mutual information (MMI) and minimum classification error (MCE) can be used instead, where the interpolation parameters are not optimized on their own but in consideration of all models together. In this paper we present how MCE and MMI techniques can be applied to two different kind of interpolation strategies: the linear interpolation, which is the standard interpolation method and the rational interpolation. We compare ML, MCE and MMI on the German part of the Verbmobil corpus, where we get a reduction of 3% of classification error when discriminating between 18 dialog act classes.
Volker Warnke, Stefan Harbeck, Elmar Nöth, Heinrich Niemann, Michael Levit
ICASSP3
1999 Automatic annotation and classification of phrase accents in spontaneous speech
abstract
During the last years, we have been working on the automatic classification of boundaries and accents in the German VERBMOBIL (VM) project (human-human communication, appointment scheduling dialogues). A sub-corpus was annotated manually with prosodic boundary and accent labels, and neural networks (NN) trained with a large set of prosodic features were used for automatic classification. The classification of boundaries could be improved markedly with a combination of the NN with a language model (LM) that was trained with manually annotated syntactic-prosodic boundary labels in a much larger sub-corpus. Here we show how a combination of NN with LM along similar lines can be used for an improvement of accent classification as well. For the training of the LM, accents are annotated automatically in the transliteration with the help of a rule--based system that uses part--of--speech (POS) as well as other linguistic /phonological information. 1. INTRODUCTION This research has been condu...
Anton Batliner, M. Nutt, Volker Warnke, Elmar Nöth, Jan Buckow, Richard Huber, Heinrich Niemann
EUROSPEECH4
1999 Learning of domain dependent knowledge in semantic networks
abstract
In speech technology more and more databases of spoken language are becoming available. For research the availability of these data offers the possibility to study huge corpora. Apart from the fact that these corpora may be represented in different formats, it is sometimes difficult to relate annotations of one corpus to those of another corpus. This contribution argues for a representation of information in speech corpora that allows for the integrated representation of information on various levels of description in XML. Secondly, the study of huge amounts of speech data requires adequate retrieval mechanisms. A query architecture is described that allows for the retrieval of encoded entities by specifying their properties or various relations to other entities. The output of the query processor is represented in XML and thus can be used for further queries or a new level of description. The work presented here is part of the results of the MATE project (http://mate.mip.ou.dk).
Frank Deinzer, Julia Fischer, U. Ahlrichs, Elmar Nöth
EUROSPEECH4
1999 A hybrid approach to spoken dialogue understanding: prosody, statistics and partial parsing
Elmar Nöth, Volker Warnke, Florian Gallwitz, Manuela Boros
EUROSPEECH1
1999 Integrating multiple knowledge sources for word hypotheses graph interpretation
abstract
We present a n i n tegrated approach for the interpretation of word hypotheses graphs (WHGs) using multiple knowledge sources.Commonly, dierent knowledge sources in speech understanding are applied sequentially.Typically, speech understanding systems, such as the Verbmobil speech-to-speech translation system, rst use a word recognizer to determine word hypotheses, only based on acoustic and language model (LM) information.The resulting word sequences or WHGs are then segmented according to syntactic and/or prosodic information.Finally, these segments are interpreted by a parser or a stochastic process.Thus, it is impossible to use the knowledge of the syntactic-prosodic process, the parser or any other subsequent process to nd the best word sequence.In our new approach w e use acoustic, prosodic and LM information to determine the best word chain, to detect syntactic/prosodic/pragmatic phrase boundaries and to classify dialog acts in one integrated search procedure, based on a WHG or a word lattice.
Volker Warnke, Florian Gallwitz, Anton Batliner, Jan Buckow, Richard Huber, Elmar Nöth, A. Höthker
EUROSPEECH6
1999 Interpolated markov chains for eukaryotic promoter recognition
abstract
MOTIVATION: We describe a new content-based approach for the detection of promoter regions of eukaryotic protein encoding genes. Our system is based on three interpolated Markov chains (IMCs) of different order which are trained on coding, non-coding and promoter sequences. It was recently shown that the interpolation of Markov chains leads to stable parameters and improves on the results in microbial gene finding (Salzberg et al., Nucleic Acids Res., 26, 544-548, 1998). Here, we present new methods for an automated estimation of optimal interpolation parameters and show how the IMCs can be applied to detect promoters in contiguous DNA sequences. Our interpolation approach can also be employed to obtain a reliable scoring function for human coding DNA regions, and the trained models can easily be incorporated in the general framework for gene recognition systems. RESULTS: A 5-fold cross-validation evaluation of our IMC approach on a representative sequence set yielded a mean correlation coefficient of 0.84 (promoter versus coding sequences) and 0.53 (promoter versus non-coding sequences). Applied to the task of eukaryotic promoter region identification in genomic DNA sequences, our classifier identifies 50% of the promoter regions in the sequences used in the most recent review and comparison by Fickett and Hatzigeorgiou ( Genome Res., 7, 861-878, 1997), while having a false-positive rate of 1/849 bp.
Uwe Ohler, Stefan Harbeck, Heinrich Niemann, Elmar Nöth, Martin G. Reese
Bioinform.4
1998 SQEL: a multilingual and multifunctional dialogue system
abstract
Within the EC-funded project SQEL, the German EVAR spoken dialogue system has been extended with respect to multilinguality and multifunctionality. The current demonstrator can handle four different languages and domains: German, Slovak, and Czech (and their national train connections), and Slovenian (European flights). The SQEL demonstrator can also access databases on the WWW, which enables users without an internet connection to meet their information needs by just using the phone. The system starts up with a German opening phrase and the user is free to use any of the implemented languages. Amultilingual word recognizer implicitly identifies the language, which is then associated with the appropriate domain and database. For the remainder of the dialogue, the corresponding monolingual recognizer is used instead. Experiments to date have shown that the multilingual and the (respective) monolingual recognizers attain comparable word accuracy rates, although the former is less efficient. The existence of language-independent task parameters, such as goal and source location, has meant that porting the system to a new language involves mainly the development of lexica and grammars (apart from the word recognizers) and not an extensive restructuring of the interpretation process within the Dialogue Manager. The latter is sufficiently flexible to switchbetween the different domains and languages.
Maria Aretoulaki, Stefan Harbeck, Florian Gallwitz, Elmar Nöth, Heinrich Niemann, Jozef Ivanecký, Ivo Ipsic, Nikola Pavesic, Václav Matousek
ICSLP4
1998 Dovetailing of acoustics and prosody in spontaneous speech recognition
abstract
Prosody can be applied to improve the performance of spontaneous speech translation systems like VERBMOBIL.In VERB-MOBIL we previously augmented the output of a word recognizer with prosodic information.Here we present a new approach of interleaving word recognition and prosodic processing.While we still use the output of a word recognizer to determine phrase boundaries, we do not wait until the end of the utterance before we start processing.Instead we intercept chunks of word hypotheses during the forward search of the recognizer.Neural networks and language models are used to predict phrase boundaries.Those boundary hypotheses, in turn, are used by the recognizer to cut the stream of incoming speech into syntactic-prosodic phrases.Thus, incremental processing is possible.We investigate which features are suited for incremental prosodic processing and compare them w.r.t.classification performance and efficiency.We show that with a set of features that can be computed efficiently classification results are achieved which are almost as good as those with the previously used computationally more expensive features.
Jan Buckow, Anton Batliner, Richard Huber, Elmar Nöth, Volker Warnke, Heinrich Niemann
ICSLP4
1998 Empowering knowledge based speech understanding through statistics
abstract
In this paper we present an innovative approach to speech understanding which is based on a fine-grained knowledge representation automatically compiled from a semantic network and on iterative optimization. Besides allowing an efficient exploitation of parallelism, any-time capability is provided since after each iteration step a (sub-)optimal solution is always available. We apply this approach to a real--world task, which is a dialog system able to answer queries about the German train timetable. In order to speed up the search for the best interpretation of an utterance we make use of statistical methods, e.g. neural networks, n-grams, and classification trees, which are trained on application relevant utterances collected over the public telephone network. At the moment the real--time factor for interpreting the initial user's utterance is 0.7.
Julia Fischer, Elmar Nöth, Heinrich Niemann, Frank Deinzer
ICSLP3
1998 Integrated recognition of words and phrase boundaries
abstract
In this paper we present an integrated approach for recognizing both the word sequence and the syntactic-prosodic structure of a spontaneous utterance. We take into account the fact that a spontaneous utterance is not merely an unstructured sequence of words by incorporating phrase boundary information into the language model and by providing HMMs to model boundaries. This allows for a distinction between word transitions across phrase boundaries and transitions within a phrase. During recognition, the syntactic-prosodic structure of the utterance is determined implicitly. Without any increase in computational effort, this leads to a 4% reduction of word error rate, and, at the same time, syntactic-prosodic boundary labels are provided for subsequent processing. The boundaries are recognized with a precision and recall rate of about 75% each. They can be used to reduce drastically the computational effort for parsing spontaneous utterances. We also present a system architecture to inco...
Florian Gallwitz, Anton Batliner, Jan Buckow, Richard Huber, Heinrich Niemann, Elmar Nöth
ICSLP6
1998 A bootstrap training approach for language model classifiers
abstract
In this paper, we present a bootstrap training approach for lan-guage model (LM) classifiers. Training class dependent LM and running them in parallel, LM can serve as classifiers with any kind of symbol sequence, e.g., word or phoneme sequences for tasks like topic spotting or language identification (LID). Irre-spective of the special symbol sequence used for a LM classifier, the training of a LM is done with a manually labeled training set for each class obtained from not necessarily cooperative speak-ers. Therefore, we have to face some erroneous labels and devia-tions from the originally intended class specification. Both facts can worsen classification. It might therefore be better not to use all utterances for training but to automatically select those utter-ances that improve recognition accuracy; this can be done by a bootstrap procedure. We present the results achieved with our best approach on the VERBMOBIL corpus for the tasks dialog act classification and LID. 1.
Volker Warnke, Elmar Nöth, Jan Buckow, Stefan Harbeck, Heinrich Niemann
ICSLP2
1998 M = Syntax + Prosody: A syntactic-prosodic labelling scheme for large spontaneous speech databases
Anton Batliner, Ralf Kompe, Andreas Kießling 0001, Marion Mast, Heinrich Niemann, Elmar Nöth
Speech Commun.6
1997 Improving parsing of spontaneous speech with the help of prosodic boundaries
abstract
Parsing can be improved in automatic speech understanding if prosodic boundary marking is taken into account, because syntactic boundaries are often marked by prosodic means. Because large databases are needed for the training of statistical models for prosodic boundaries, we developed a labeling scheme for syntactic-prosodic boundaries within the German Verbmobil project (automatic speech-to-speech translation). We compare the results of classifiers (multi-layer perceptrons and language models) trained on these syntactic-prosodic boundary labels with classifiers trained on perceptual-prosodic and purely syntactic labels. Recognition rates of up to 96% were achieved. The turns that we need to parse consist of 20 words on the average and frequently contain sequences of partial sentence equivalents due to restarts, ellipsis, etc. For this material, the boundary scores computed by our classifiers can successfully be integrated into the syntactic parsing of word graphs; currently, they improve the parse time by 92% and reduce the number of parse trees by 96%. This is achieved by introducing a special prosodic syntactic clause boundary (PSCB) symbol into our grammar and guiding the search for the best word chain with the prosodic boundary scores.
Ralf Kompe, Andreas Kießling 0001, Heinrich Niemann, Elmar Nöth, Anton Batliner, Stefanie Schachtl, Tobias Ruland, Hans Ulrich Block
ICASSP4
1997 Prosodic processing and its use in VERBMOBIL
abstract
We present the prosody module of the VERBMOBIL speech-to-speech translation system, the world wide first complete system, which successfully uses prosodic information in linguistic analysis. This is achieved by computing probabilities for clause boundaries, accentuation, and different types of sentence mood for each of the word hypotheses computed by the word recognizer. These probabilities guide the search of the linguistic analysis. Disambiguation is already achieved during the analysis and not by a prosodic verification of different linguistic hypotheses. So far, the most useful prosodic information is provided by clause boundaries. These are detected with a recognition rate of 94%. For the parsing of word hypotheses graphs, the use of clause boundary probabilities yields a speed-up of 92% and a 96% reduction of alternative readings.
Heinrich Niemann, Elmar Nöth, Andreas Kießling 0001, Ralf Kompe, Anton Batliner
ICASSP2
1997 Tempo and its change in spontaneous speech
abstract
In this paper, we give a first account of speech tempo and its change in spontaneous speech in a very large data base (Verbmobil, i.e., human-human appointment dialogs). As features representing speech tempo, we computed mean normalized speech duration (speaking rate) and normalized phone duration in different ways. The importance of these features is evaluated with an automatic classification of boundaries and accents where different sets of prosodic features (including also information about F0, energy, pause, etc.) were used. The best results (83% for accents, 88% for boundaries, two classes each) could be achieved when all features were used. For the 2nd issue change of tempo was labelled manually. We present the characterizing feature values for changes from slow to fast and from fast to slow, as well as the results of an automatic classification of change of tempo (72% for three classes). Finally, we discuss the possible function of change of tempo and its use in automatic speech processing. (orig.)
Anton Batliner, Andreas Kießling 0001, Ralf Kompe, Heinrich Niemann, Elmar Nöth
EUROSPEECH5
1997 Semantic processing of out-of-vocabulary words in a spoken dialogue system
abstract
One of the most important causes of failure in spoken dialogue systems is usually neglected: the problem of words that are not covered by the system's vocabulary (out-of-vocabulary or OOV words). In this paper a methodology is described for the detection, classification and processing of OOV words in an automatic train timetable information system [2]. The various extensions that had to be effected on the different modules of the system are reported, resulting in the design of appropriate dialogue strategies, as are encouraging evaluation results on the new versions of the word recogniser and the linguistic processor. 1. INTRODUCTION The majority of speech understanding systems have to face the problem of words that are not covered by their current lexicon, i.e. OOV words. In such a case the word recogniser usually recognises one or more different words with a similar acoustic profile to the unknown. These misrecognitions often result in possibly irreparable misunderstandings between ...
Manuela Boros, Maria Aretoulaki, Florian Gallwitz, Elmar Nöth, Heinrich Niemann
EUROSPEECH4
1997 A frame and segment based approach for topic spotting
abstract
In this paper we present a new approach for topic spotting based on subword units (phonemes and feature vectors) instead of words. Classification of topics is done by running topic dependent polygram language models over these symbol sequences and deciding for the one with the best score. We trained and tested the two methods on three different corpora. The first is a part of a media corpus which contains data from TV shows for three different topics (IDS), the second is part of the Switchboard corpus, the third is a collection of human machine dialogs about train timetable information (EVAR corpus) . The results on Switchboard are compared with phoneme based approaches which were made at CRIM (Montr'eal) and DRA (Malvern) and are presented as ROC curves; the results on IDS and EVAR are compared with a word based approach and presented as confusion tables. We show that a surprisingly little amount of recognition accuracy is lost when going from word to subword based topic spotting.
Elmar Nöth, Stefan Harbeck, Heinrich Niemann, Volker Warnke
EUROSPEECH1
1997 Integrated dialog act segmentation and classification using prosodic features and language models
abstract
This paper presents an integrated approach for the segmentation and classification of dialog acts (DA) in the Verbmobil project. In Verbmobil it is often sufficient to recognize the sequence of DAs occurring during a dialog between the two partners. In our previous work we segmented and classified a dialog in two steps: first we calculated hypotheses for the segment boundaries and decided for a boundary if the probabilities exceeded a predefined threshold level. Second we classified the segments into DAs using semantic classification trees or stochastic language models. In our new approach we integrate the segmentation and classification in the A*-algorithm to search for the optimal segmentation and classification of DAs on the basis of word hypotheses graphs (WHGs). The hypotheses for the segment boundaries are calculated with the help of a stochastic language model operating on the word chain and a multi-layer perceptron (MLP) classifying prosodic feature. The DA classification is done using a category based language model for each DA. For our experiments we used data from the Verbobil-corpus. (orig.)
Volker Warnke, Ralf Kompe, Heinrich Niemann, Elmar Nöth
EUROSPEECH4
1996 Integrating Syntactic and Prosodic Information for the Efficient Detection of Empty Categories
Anton Batliner, Anke Feldhaus, Stefan Geißler, Andreas Kießling 0001, Tibor Kiss, Ralf Kompe, Elmar Nöth
COLING7
1996 An integrated model of acoustics and language using semantic classification trees
abstract
We propose multilevel semantic classification trees to combine different information sources for predicting speech events (e.g. word chains, phrases, etc.). Traditionally in speech recognition systems these information sources (acoustic evidence, language model) are calculated independently and combined via Bayes rule. The proposed approach allows one to combine sources of different types it is no longer necessary for each source to yield a probability. Moreover the tree can look at several information sources simultaneously. The approach is demonstrated for the prediction of prosodically marked phrase boundaries, combining information about the spoken word chain, word category information, prosodic parameters, and the result of a neural network predicting the boundary on the basis of acoustic-prosodic features. The recognition rates of up to 90% for the two class problem boundary vs. no boundary are already comparable to results achieved with the above mentioned Bayes rule approach that combines the acoustic classifier with a 5-gram categorical language model. This is remarkable, since so far only a small set of questions combining information from different sources have been implemented.
Elmar Nöth, Renato De Mori, Julia Fischer, Arnd Gebhard, Stefan Harbeck, Ralf Kompe, Roland Kuhn 0001, Heinrich Niemann, Marion Mast
ICASSP1
1996 Comparison of two tree-structured approaches for grapheme-to-phoneme conversion
abstract
Recently, we described a two-step self-learning approach for grapheme-to-phoneme (G2P) conversion [1].In the first step, grapheme and phoneme strings in the training data are aligned via an iterative Viterbi procedure that may insert graphemic and phonemic nulls where required.In the second step, a Trie structure encoding pronunciation rules is generated.In this paper we describe the alignment module, and give alignment accuracies on the NETtalk database.We also compare transcription accuracies for two approaches to the second step on three databases: the NETtalk database, the CMU dictionary and the French part of the ONOMASTICA lexicon.The two transcription approaches applied in this research are a Trie approach [1] and an approach based on binary decision trees grown by means of the Gelfand-Ravishankar-Delp algorithm [2,3,4].We discuss the choice of questions for these decision trees -it may be possible to formulate questions about groups of characters (e.g., "is the next letter a vowel?") that yield better trees than those that only use questions about individual characters (e.g., "is the next letter an 'A' ?").Finally, we discuss the implications of our work for G2P conversion.
Ove Andersen, Roland Kuhn 0001, Ariane Lazaridès, Paul Dalsgaard, Elmar Nöth
ICSLP6
1996 Prosody, empty categories and parsing - a success story
abstract
We describe a number of experiments that demonstrate the usefulness of prosodic information for a processing module which parses spoken utterances with a feature-based grammar employing empty categories.We show that by requiring certain prosodic properties from those positions in the input, where the presence of an empty category has to be hypothesized, a derivation can be accomplished more eciently.The approach has been implemented in the machine translation project Verbmobil and results in a signicant reduction of the work-load for the parser.
Anton Batliner, Anke Feldhaus, Stefan Geißler, Tibor Kiss, Ralf Kompe, Elmar Nöth
ICSLP6
1996 Syntactic-prosodic labeling of large spontaneous speech data-bases
Anton Batliner, Ralf Kompe, Andreas Kießling 0001, Heinrich Niemann, Elmar Nöth
ICSLP5
1996 A category based approach for recognition of out-of-vocabulary words
Florian Gallwitz, Elmar Nöth, Heinrich Niemann
ICSLP2
1996 Dialog act classification with the help of prosody
Marion Mast, Ralf Kompe, Stefan Harbeck, Andreas Kießling 0001, Heinrich Niemann, Elmar Nöth, Ernst Günter Schukat-Talamazzini, Volker Warnke
ICSLP6
1995 Robust pitch period detection using dynamic programming with an ANN cost function
Stefan Harbeck, Andreas Kießling 0001, Ralf Kompe, Heinrich Niemann, Elmar Nöth
EUROSPEECH5
1995 Prosodic scoring of word hypotheses graphs
Ralf Kompe, Andreas Kießling 0001, Heinrich Niemann, Elmar Nöth, Ernst Günter Schukat-Talamazzini, A. Zottmann, Anton Batliner
EUROSPEECH4
1994 Automatic classification of prosodically marked phrase boundaries in German
abstract
A large corpus has been created automatically and read by 100 speakers. Phrase boundaries were labeled in the sentences automatically during sentence generation. Perception experiments on a subset of 500 utterances showed a high agreement between the automatically generated boundary markers and the ones perceived by listeners. Gaussian distribution and polynomial classifiers were trained on a set of prosodic features computed from the speech signal using the automatically generated boundary markers. Comparing the classification results with the judgments of the listeners yielded in a recognition rate of 87%. A combination with stochastic language models improved the recognition rate to 90%. We found that the pause and the durational features are most important for the classification, but that the influence of F0 is not neglectable.>
Ralf Kompe, Anton Batliner, Andreas Kießling 0001, Ute Kilian, Heinrich Niemann, Elmar Nöth, Peter Regel-Brietzmann
ICASSP (2)6
1994 Improving parsing by incorporating 'prosodic clause boundaries into a grammar
abstract
In written language, punctuation is used to separate main and subordinate clause.In spoken language, ambiguities arise due to missing punctuation, but clause boundaries are often marked prosodically and can be used instead.We detect PCBs (Prosodically marked Clause Boundaries) by using prosodic features (duration, intonation, energy, and pause information) with a neural network, achieving a recognition rate of 82%.PCBs are integrated into our grammar using a special syntactic category 'break' that can be used in the phrase-structure rules of the grammar in a similar way as punctuation is used in grammars for written language.Whereas punctuation in most cases is obligatory, PCBs are sometimes optional.Moreover, they can in principle occur everywhere in the sentence due e.g. to hesitations or misrecognition.To cope with these problems we tested two different approaches: A slightly modified parser for word chains containing PCBs and a word graph parser that takes the probabilities of PCBs into account.Tests were conducted on a subset of infinitive subordinate clauses from a large speech database containing sentences from the domain of train table inquiries.The average number of syntactic derivations could be reduced by about 70 % even when working on recognized word graphs.
Gabriele Bakenecker, Hans Ulrich Block, Anton Batliner, Ralf Kompe, Elmar Nöth, Peter Regel-Brietzmann
ICSLP5
1994 Automatic labeling of phrase accents in German
abstract
In this paper a method for the automatic labeling of phrase accents is described, based on a large text corpus that has been generated automatically and read by 100 speakers.Perception experiments on a subset of 500 utterances show a high agreement between the automatically generated accent labels and the judgment scores obtained.We computed different prosodic feature vectors from the speech signal for each syllable and trained different Gaussian distribution classifiers and artificial neural networks using the automatically generated accent labels.Recognition rates of up to 83% could be achieved for the distinction of accentuated vs. unaccentuated syllables.Similar results could be obtained for the comparison of the listeners judgments with the automatic classification.
Andreas Kießling 0001, Ralf Kompe, Anton Batliner, Heinrich Niemann, Elmar Nöth
ICSLP5
1994 Prosody takes over: Towards a prosodically guided dialog system
Ralf Kompe, Elmar Nöth, Andreas Kießling 0001, Thomas Kuhn 0002, Marion Mast, Heinrich Niemann, K. Ott, Anton Batliner
Speech Communication2
1993 Going back to the source: inverse filtering of the speech signal with ANNs
abstract
In this paper we present a new method transforming speech signals to voice source signals (VSS) using articial neural networks (ANN).We will point out that the ANN mapping of speech signals into source signals is quite accurate, and most of the irregularities in the speech signal will lead to an irregularity in the source signal, produced by the ANN (ANN-VSS).We will show that the mapping of the ANN is robust with respect to untrained speakers, di erent recording conditions and facilities, and di erent v ocabularies.We will also present preliminary results which show that from the ANN source signal pitch periods can be determined accurately.
Joachim Denzler, Ralf Kompe, Andreas Kießling 0001, Heinrich Niemann, Elmar Nöth
EUROSPEECH5
1993 Prosody takes over: a prosodically guided dialog system
abstract
In this paper rst experiments with naive persons using the speech understanding and dialog system EVAR are discussed.The domain of EVAR is train table inquiry.We observed that in real human-human dialogs when the o cer transmits the information the customer very often interrupts.Many of these interruptions are just repetitions of the time of day given by the o cer.The functional role of these interruptions is determined b y prosodic cues only.An important result of the experiments with EVAR is that it is hard to follow the system giving the train connection via speech synthesis.In this case it is even more important than in human-human dialogs that the user has the opportunity to interact during the answer phase.Therefore we extended the dialog m o dule to allow the user to repeat the time of day and we added a p r osody module guiding the continuation of the dialog.
Ralf Kompe, Andreas Kießling 0001, Thomas Kuhn 0002, Marion Mast, Heinrich Niemann, Elmar Nöth, K. Ott, Anton Batliner
EUROSPEECH6
1992 DP-based determination of F0 contours from speech signals
abstract
A new algorithm for the determination of fundamental frequency (F/sub 0/) contours is presented. For each voiced frame appropriate divisors of the frequency with the maximum energy in the spectrum are taken as F/sub 0/ candidates. An F/sub 0/ contour is computed using a dynamic programming (DP) method by minimizing a weighted sum of the difference between consecutive candidates and the distance of the candidates to a predetermined local target value. With this algorithm a coarse error rate of 0.6% on the frame level and of 6.4% on the sentence level is achieved on a German speech database. On the average the difference to the reference is 1.9 Hz. The algorithm outperforms two conventional algorithms tested on the same data.>
Andreas Kießling 0001, Ralf Kompe, Heinrich Niemann, Elmar Nöth, Anton Batliner
ICASSP4
1992 The dialog module of the speech recognition and dialog system EVAR
abstract
This article describes the dialog module of EVAR. The dialog is seen as a sequence of dialog acts uttered by the system and the user. From a corpus of real humanhuman dialogs a model was extracted. The model covers all sequences of dialog acts observed in the corpus. Each dialog act is modelled as a set of pragmatic, semantic and syntactic concepts. The properties of the concepts and the current dialog state are used to identify the actual dialog act. Each user utterance is interpreted by the system and the system reacts with the appropriate dialog act. After the user has defined his request completly, the system starts a database query and finally generates a synthesized natural language answer. The answer generation is realised with sentence masks, where the information from the database enquiry is filled in at the appropriate slots. A dialog memory which allows the interpretation of elliptical sentences and proforms is updated after each utterance.
Marion Mast, Ralf Kompe, Franz Kummert, Heinrich Niemann, Elmar Nöth
ICSLP5
1989 The prediction of focus
abstract
We present results on how focus is marked intonationally in German.Several speakers produced a large corpus of sentences.The corpus was constructed in a way that sentence modality and place of focus could only be differentiated by intona tional means.Acoustic features representing the intonational para meters pitch, duration, and intensity, were extracted manually or automatically.The relevance of these features and the effect of several transformations were tested with statistical methods.Percep tual experiments where the listeners had to judge the naturalness and categories of the utterances were performed as well.By calcula ting average values for the (appropriately transformed) relevant features we found "normal", prototypical cases.We will show that by looking at utterances where all listeners agreed on the naturalness and (intended) categories we arrived at coinciding results.At the same time we found ''unusual" but regular productions. MATERIAL AND PROCEDURESThis paper is concerned with the prediction of focus; focus is the part of an utterance which is semantically most important.On the phonetic surface focus is marked by the focal accent (FA).To be more exact, we will try to predict the phrase that carries the FA.
Anton Batliner, Elmar Nöth
EUROSPEECH2