VLDB 2026 Research / reviewers in the wild / expert
Tobias Bocklet
dblp:92/1549
· DBLP profile ↗
68ranked-venue papers
9as first author
39since 2021 · last 2026
0009-0008-7780-8821ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 60 · 8 first-author · 33 since 2021Artificial intelligence and machine learning · 48 · 5 first-author · 31 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Evaluation Pitfalls and Challenges in Multimedia Event ExtractionabstractMultimedia event extraction aims to jointly identify events and their arguments across multiple modalities, such as text and images, to support more comprehensive event understanding.While recent work reports steady and substantial progress, the reliability and comparability of these results critically depend on consistent and rigorous evaluation.In this work, we present the first systematic analysis of evaluation pitfalls in multimedia event extraction and identify three major sources of issues: inconsistent data processing, inconsistent task assumptions, and overly relaxed evaluation settings.We demonstrate, through a series of controlled experiments under a strict evaluation framework, that minor evaluation choices can cause large performance variations and lead to overestimation of a model's ability to ground real-world events across modalities.Our findings highlight the need for comparable evaluation standards and encourage a shift toward more rigorous evaluation in multimedia event extraction. Philipp Seeberger, Steffen Freisinger, Tobias Bocklet, Korbinian Riedhammer |
ACL (1) | 3 |
| 2026 | The PARLO Dementia Corpus: A German Multi-Center Resource for Alzheimer's DiseaseabstractEarly and accessible detection of Alzheimer's disease (AD) remains a major challenge, as current diagnostic methods often rely on costly and invasive biomarkers. Speech and language analysis has emerged as a promising non-invasive and scalable approach to detecting cognitive impairment, but research in this area is hindered by the lack of publicly available datasets, especially for languages other than English. This paper introduces the PARLO Dementia Corpus (PDC), a new multi-center, clinically validated German resource for AD collected across nine academic memory clinics in Germany. The dataset comprises speech recordings from individuals with AD-related mild cognitive impairment and mild to moderate dementia, as well as cognitively healthy controls. Speech was elicited using a standardized test battery of eight neuropsychological tasks, including confrontation naming, verbal fluency, word repetition, picture description, story reading, and recall tasks. In addition to audio recordings, the dataset includes manually verified transcriptions and detailed demographic, clinical, and biomarker metadata. Baseline experiments on ASR benchmarking, automated test evaluation, and LLM-based classification illustrate the feasibility of automatic, speech-based cognitive assessment and highlight the diagnostic value of recall-driven speech production. The PDC thus establishes the first publicly available German benchmark for multi-modal and cross-lingual research on neurodegenerative diseases. Franziska Braun, Christopher Witzl, Florian Hönig, Elmar Nöth, Tobias Bocklet, Korbinian Riedhammer |
LREC | 5 |
| 2025 | Text-Guided Speech Representations for Language Acquisition AssessmentabstractAssessing spoken language abilities in children is critical for early detection of developmental delays. We propose a two-stage framework for automatic pronunciation assessment, targeting classification of items in a language development screening. In the first stage, we fine-tune a sentence transformer using phonetic and label-informed losses to create a structured latent text space. In the second stage, we fine-tune an audio encoder to project spoken utterances into the text space. Our approach enables inference from audio alone while preserving alignment with phonetically meaningful structure. Compared to baseline systems, our method achieves the best classification performance and provides a latent space suitable for downstream analysis. This facilitates the identification of phonological or grammatical error patterns from cluster structure, paving the way for potentially interpretable assessment of early language acquisition. Ilja Baumann, Dominik Wagner 0002, Philipp Seeberger, Korbinian Riedhammer, Tobias Bocklet |
ASRU | 5 |
| 2025 | On the Difficulty of Token-Level Modeling of Dysfluency and Fluency Shaping ArtifactsabstractAutomatic transcription of stuttered speech remains a challenge, even for modern end-to-end (E2E) automatic speech recognition (ASR) frameworks. Dysfluencies and fluency-shaping artifacts are often overlooked, resulting in non-verbatim transcriptions with limited clinical and research value. We propose a parameter-efficient adaptation method to decode dysfluencies and fluency modifications as special tokens within transcriptions, evaluated on simulated (LibriStutter, English) and natural (KSoF, German) stuttered speech datasets. To mitigate ASR performance disparities and bias towards English, we introduce a multi-step fine-tuning strategy with language-adaptive pretraining. Tokenization analysis further highlights the tokenizer’s English-centric bias, which poses challenges for improving performance on German data. Our findings demonstrate the effectiveness of lightweight adaptation techniques for dysfluency-aware ASR while exposing key limitations in multilingual E2E systems. Kashaf Gulzar, Dominik Wagner 0002, Sebastian P. Bayerl, Florian Hönig, Tobias Bocklet, Korbinian Riedhammer |
ASRU | 5 |
| 2025 | Improving Multimodal Speech-To-Slide Alignment for Academic Lectures with Vision LLMsabstractWe enhance the MaViLS multimodal algorithm to improve speech-to-slide alignment for lecture podcasts. Our approach integrates vision large language models for optical character recognition and automatic generation of lecture transcripts from slide content, coupled with a multilingual multimodal embedding model for text and image alignment. By combining slide-extracted text with automatically generated lecture transcripts and captioned slide images, we generate enhanced audio features that better capture speech-slide mapping. Our method improves the average F1 score for audio feature alignment on the English MaViLS dataset from 0.51 to 0.71 and on a newly created German podcast lectures dataset from 0.65 to 0.84. Thomas Ranzenberger, Dominik Wagner 0002, Steffen Freisinger, Tobias Bocklet, Korbinian Riedhammer |
ASRU | 4 |
| 2025 | Joint ASR and Speech Attribute Prediction for Conversational Dysarthric Speech Analysis with Multimodal Language ModelsabstractDysarthric speech recognition systems often focus solely on transcription, limiting their applicability in clinical and assistive settings where assessments of perceptual attributes like intelligibility and naturalness are essential. We propose a multimodal conversational framework based on Phi-4-Multimodal that combines ASR with attribute rating prediction, enabling users to query both transcriptions and perceptual characteristics (e.g. “How intelligible is this utterance?”). Our multitask model adds auxiliary prediction heads for five clinically relevant attributes and is trained on the English Speech Accessibility Project dataset. The system achieves competitive ASR performance while delivering attribute-level feedback comparable to specialized classifiers. Additional experiments show improved ASR performance for German Parkinson’s speech, indicating preserved multilingual capabilities and partial cross-lingual transfer of dysarthric speech patterns. Dominik Wagner 0002, Ilja Baumann, Natalie Engert, Elmar Nöth, Korbinian Riedhammer, Tobias Bocklet |
ASRU | 6 |
| 2025 | Optimized Self-supervised Training with BEST-RQ for Speech RecognitionabstractSelf-supervised learning has been successfully used for various speech related tasks, including automatic speech recognition. BERT-based Speech pre-Training with Random-projection Quantizer (BEST-RQ) has achieved state-of-the-art results in speech recognition. In this work, we further optimize the BEST-RQ approach using Kullback-Leibler divergence as an additional regularizing loss and multi-codebook extension per cluster derived from low-level feature clustering. Preliminary experiments on train-100 split of LibriSpeech result in a relative improvement of 11.2% on test-clean by using multiple codebooks, utilizing a combination of cross-entropy and Kullback-Leibler divergence further reduces the word error rate by 4.5%. The proposed optimizations on full LibriSpeech pre-training and fine-tuning result in relative word error rate improvements of up to 23.8% on test-clean and 30.6% on test-other using 6 codebooks. Furthermore, the proposed setup leads to faster convergence in pre-training and fine-tuning and additionally stabilizes the pre-training. Ilja Baumann, Dominik Wagner 0002, Korbinian Riedhammer, Tobias Bocklet |
ICASSP | 4 |
| 2025 | Digital Operating Mode Classification of Real-World Amateur Radio TransmissionsabstractThis study presents an ML approach for classifying digital radio operating modes evaluated on real-world transmissions. We generated 98 different parameterized radio signals from 17 digital operating modes, transmitted each of them on the 70 cm (UHF) amateur radio band, and recorded our transmissions with two different architectures of SDR receivers. Three lightweight ML models were trained exclusively on spectrograms of limited non-transmitted signals with random characters as payloads. This training involved an online data augmentation pipeline to simulate various radio channel impairments. Our best model, EfficientNetB0, achieved an accuracy of 93.80% across the 17 operating modes and 85.47% across all 98 parameterized radio signals, evaluated on our real-world transmissions with Wikipedia articles as payloads. Furthermore, we analyzed the impact of varying signal durations & the number of FFT bins on classification, assessed the effectiveness of our simulated channel impairments, and tested our models across multiple simulated SNRs. Maximilian Bundscherer, Thomas H. Schmitt, Ilja Baumann, Tobias Bocklet |
ICASSP | 4 |
| 2025 | Adapter-Based Multi-Agent AVSR Extension for Pre-Trained ASR ModelsabstractWe present an approach to Audio-Visual Speech Recognition that builds on a pre-trained Whisper model. To infuse visual information into this audio-only model, we extend it with an AV fusion module and LoRa adapters, one of the most up-to-date adapter approaches. One advantage of adapter-based approaches, is that only a relatively small number of parameters are trained, while the basic model remains unchanged. Common AVSR approaches train single models to handle several noise categories and noise levels simultaneously. Taking advantage of the lightweight nature of adapter approaches, we train noise-scenario-specific adapter-sets, each covering individual noise-categories or a specific noise-level range. The most suitable adapter-set is selected by previously classifying the noise-scenario. This enables our models to achieve an optimum coverage across different noise-categories and noise-levels, while training only a minimum number of parameters.Compared to a full fine-tuning approach with SOTA performance our models achieve almost comparable results over the majority of the tested noise-categories and noise-levels, with up to 88.5% less trainable parameters. Our approach can be extended by further noise-specific adapter-sets to cover additional noise scenarios. It is also possible to utilize the underlying powerful ASR model when no visual information is available, as it remains unchanged. Christopher Simic, Korbinian Riedhammer, Tobias Bocklet |
ICASSP | 3 |
| 2025 | Personalized Fine-Tuning with Controllable Synthetic Speech from LLM-Generated Transcripts for Dysarthric Speech RecognitionabstractIn this work, we present our submission to the Speech Accessibility Project challenge for dysarthric speech recognition. We integrate parameter-efficient fine-tuning with latent audio representations to improve an encoder-decoder ASR system. Synthetic training data is generated by fine-tuning Parler-TTS to mimic dysarthric speech, using LLM-generated prompts for corpus-consistent target transcripts. Personalization with x-vectors consistently reduces word error rates (WERs) over non-personalized fine-tuning. AdaLoRA adapters outperform full fine-tuning and standard low-rank adaptation, achieving relative WER reductions of ∼23% and ∼22%, respectively. Further improvements (∼5% WER reduction) come from incorporating wav2vec 2.0-based audio representations. Training with synthetic dysarthric speech yields up to ∼7% relative WER improvement over personalized fine-tuning alone. Dominik Wagner 0002, Ilja Baumann, Natalie Engert, Seanie Lee, Elmar Nöth, Korbinian Riedhammer, Tobias Bocklet |
INTERSPEECH | 7 |
| 2025 | Pathology-Aware Speech Encoding and Data Augmentation for Dysarthric Speech RecognitionabstractAutomatic speech recognition (ASR) for pathologic speech remains a major challenge due to high variability in articulation, phonation, and prosody distortions. In this work, we propose a pathology-aware speech encoder based on BEST-RQ pre-training, which incorporates 46k hours of speech, including pathologic and atypical speech. We continue pre-training for domain adaptation and experiment with etiology-specific codebooks. We achieve a 13.2% relative word error rate (WER) improvement using the pathology-aware speech encoder with etiology-specific continued pre-training. Additionally, we examine the impact of incorporating synthetic and out-of-domain (OOD) data to further enhance ASR performance. Synthetic data reduces WER by up to 8.7%, while OOD data improves WER by 12.2%. Finally, we introduce a semantic similaritybased data augmentation technique to optimize data selection, achieving a WER improvement of up to 9.7% while minimizing the need for additional training data. Ilja Baumann, Dominik Wagner 0002, Korbinian Riedhammer, Tobias Bocklet |
INTERSPEECH | 4 |
| 2025 | Pitfalls and Limits in Automatic Dementia AssessmentabstractCurrent work on speech-based dementia assessment focuses on either feature extraction to predict assessment scales, or on the automation of existing test procedures. Most research uses public data unquestioningly and rarely performs a detailed error analysis, focusing primarily on numerical performance. We perform an in-depth analysis of an automated standardized dementia assessment, the Syndrom-Kurz-Test. We find that while there is a high overall correlation with human annotators, due to certain artifacts, we observe high correlations for the severely impaired individuals, which is less true for the healthy or mildly impaired ones. Speech production decreases with cognitive decline, leading to overoptimistic correlations when test scoring relies on word naming. Depending on the test design, fallback handling introduces further biases that favor certain groups. These pitfalls remain independent of group distributions in datasets and require differentiated analysis of target groups. Franziska Braun, Christopher Witzl, Andreas Erzigkeit, Hartmut Lehfeld, Thomas Hillemacher, Tobias Bocklet, Korbinian Riedhammer |
INTERSPEECH | 6 |
| 2025 | Towards Multi-Level Transcript Segmentation: LoRA Fine-Tuning for Table-of-Contents GenerationabstractSegmenting speech transcripts into thematic sections benefits both downstream processing and users who depend on written text for accessibility. We introduce a novel approach to hierarchical topic segmentation in transcripts, generating multi-level tables of contents that capture both topic and subtopic boundaries. We compare zero-shot prompting and LoRA fine-tuning on large language models, while also exploring the integration of high-level speech pause features. Evaluations on English meeting recordings and multilingual lecture transcripts (Portuguese, German) show significant improvements over established topic segmentation baselines. Additionally, we adapt a common evaluation measure for multi-level segmentation, taking into account all hierarchical levels within one metric. Steffen Freisinger, Philipp Seeberger, Thomas Ranzenberger, Tobias Bocklet, Korbinian Riedhammer |
INTERSPEECH | 4 |
| 2025 | FedSVD: Adaptive Orthogonalization for Private Federated Learning with LoRAabstractLow-Rank Adaptation (LoRA), which introduces a product of two trainable low-rank matrices into frozen pre-trained weights, is widely used for efficient fine-tuning of language models in federated learning (FL).
However, when combined with differentially private stochastic gradient descent (DP-SGD), LoRA faces substantial noise amplification: DP-SGD perturbs per-sample gradients, and the matrix multiplication of the LoRA update ($BA$) intensifies this effect. Freezing one matrix (*e.g.*, $A$) reduces the noise but restricts model expressiveness, often resulting in suboptimal adaptation.
To address this, we propose $\texttt{FedSVD}$, a simple yet effective method that introduces a global reparameterization based on singular value decomposition (SVD).
In our approach, each client optimizes only the $B$ matrix and transmits it to the server.
The server aggregates the $B$ matrices, computes the product $BA$ using the previous $A$, and refactorizes the result via SVD.
This yields a new adaptive $A$ composed of the orthonormal right singular vectors of $BA$, and an updated $B$ containing the remaining SVD components.
This reparameterization avoids quadratic noise amplification, while allowing $A$ to better capture the principal directions of the aggregate updates.
Moreover, the orthonormal structure of $A$ bounds the gradient norms of $B$ and preserves more signal under DP-SGD, as confirmed by our theoretical analysis.
As a result, $\texttt{FedSVD}$ consistently improves stability and performance across a variety of privacy settings and benchmarks, outperforming relevant baselines under both private and non-private regimes. Seanie Lee, Dong Bok Lee, Dominik Wagner 0002, Haebin Seong, Tobias Bocklet, Juho Lee 0001, Sung Ju Hwang |
NeurIPS | 6 |
| 2024 | Optimized Speculative Sampling for GPU Hardware AcceleratorsabstractIn this work, we optimize speculative sampling for parallel hardware accelerators to improve sampling speed.We notice that substantial portions of the intermediate matrices necessary for speculative sampling can be computed concurrently.This allows us to distribute the workload across multiple GPU threads, enabling simultaneous operations on matrix segments within thread blocks.This results in profiling time improvements ranging from 6% to 13% relative to the baseline implementation, without compromising accuracy.To further accelerate speculative sampling, probability distributions parameterized by softmax are approximated by sigmoid.This approximation approach results in significantly greater relative improvements in profiling time, ranging from 37% to 94%, with a minor decline in accuracy.We conduct extensive experiments on both automatic speech recognition and summarization tasks to validate the effectiveness of our optimization methods. Dominik Wagner 0002, Seanie Lee, Ilja Baumann, Philipp Seeberger, Korbinian Riedhammer, Tobias Bocklet |
EMNLP | 6 |
| 2024 | Towards Interpretability of Automatic Phoneme Analysis in Cleft Lip and Palate SpeechabstractCleft Lip and Palate ranks among the most common congenital abnormalities and significantly influences speech articulation, resulting in varying phonemic impacts. In a clinical context, a detailed diagnosis is carried out by time-consuming perceptual evaluations. We use perceptual ratings of different articulatory modifications on phoneme-level as ground-truth and propose a system based on wav2vec 2.0, trained to the downstream task of classifying phonemic criteria as a multi-class and multi-label problem. The system is trained for detection on utterance level, without the usage of phoneme labels. To gain a clearer understanding of which areas of the speech signal have the greatest impact on classification, we assess the extent to which our system aligns with expert ratings at the phoneme level. Additionally, we examine which specific phonemes play a decisive role in determining the final classification of the labeled criteria. The results show that salient phonemes marked by experts contribute remarkably greater to the classification of the correct class using feature relevance explanation methods. To the best of our knowledge, this is the first study incorporating various utterance-level articulatory modifications classification and phoneme-level interpretation, offering a more comprehensive understanding for potential clinical applications. Ilja Baumann, Dominik Wagner 0002, Maria Schuster, Elmar Nöth, Tobias Bocklet |
ICASSP | 5 |
| 2024 | Self-Supervised Adaptive AV Fusion Module for Pre-Trained ASR ModelsabstractAutomatic speech recognition (ASR) has reached a level of accuracy in recent years, that even outperforms humans in transcribing speech to text. Nevertheless, all current ASR approaches show a certain weakness against ambient noise. To reduce this weakness, audio-visual speech recognition (AVSR) approaches additionally consider visual information from lip movements for transcription. This additional modality increases the computational cost for training models from scratch. We propose an approach, that builds on a pre-trained ASR model and extends it with an adaptive upstream module, that fuses audio and visual information. Since we do not need to train the transformer structure from scratch, our approach requires a fraction of the computational resources compared to traditional AVSR models. Compared to current SOTA systems like AV-HuBERT, our approach achieves an average improvement of 8.3 % in word error rate across different model sizes, noise categories and broad SNR range. The approach allows up to 21 % smaller models and requires only a fraction of the computational resources for training and inference compared to common AVSR approaches. Christopher Simic, Tobias Bocklet |
ICASSP | 2 |
| 2024 | Large Language Models for Dysfluency Detection in Stuttered Speech
Dominik Wagner 0002, Sebastian P. Bayerl, Ilja Baumann, Elmar Nöth, Korbinian Riedhammer, Tobias Bocklet |
INTERSPEECH | 6 |
| 2024 | Outlier Reduction with Gated Attention for Improved Post-training Quantization in Large Sequence-to-sequence Speech Foundation ModelsabstractThis paper explores the improvement of post-training quantization (PTQ) after knowledge distillation in the Whisper speech foundation model family. We address the challenge of outliers in weights and activation tensors, known to impede quantization quality in transformer-based language and vision models. Extending this observation to Whisper, we demonstrate that these outliers are also present when transformer-based models are trained to perform automatic speech recognition, necessitating mitigation strategies for PTQ. We show that outliers can be reduced by a recently proposed gating mechanism in the attention blocks of the student model, enabling effective 8-bit quantization, and lower word error rates compared to student models without the gating mechanism in place. Dominik Wagner 0002, Ilja Baumann, Korbinian Riedhammer, Tobias Bocklet |
INTERSPEECH | 4 |
| 2024 | Towards Self-Attention Understanding for Automatic Articulatory Processes Analysis in Cleft Lip and Palate SpeechabstractCleft lip and palate (CLP) speech presents unique challenges for automatic phoneme analysis due to its distinct acoustic characteristics and articulatory anomalies. We perform phoneme analysis in CLP speech using a pre-trained wav2vec 2.0 model with a multi-head self-attention classification module to capture long-range dependencies within the speech signal, thereby enabling better contextual understanding of phoneme sequences. We demonstrate the effectiveness of our approach in the classification of various articulatory processes in CLP speech. Furthermore, we investigate the interpretability of self-attention to gain insights into the model’s understanding of CLP speech characteristics. Our findings highlight the potential of the selfattention mechanisms for improving automatic phoneme analysis in CLP speech, paving the way for enhanced diagnostics, adding interpretability for therapists and affected patients. Ilja Baumann, Dominik Wagner 0002, Maria Schuster, Korbinian Riedhammer, Elmar Nöth, Tobias Bocklet |
INTERSPEECH | 6 |
| 2024 | Automatic Evaluation of a Sentence Memory Test for Preschool ChildrenabstractAssessment of memory capabilities in preschool-aged children is crucial for early detection of potential speech development impairments or delays. We present an approach for the automatic evaluation of a standardized sentence memory test specifically for preschool children. Our methodology leverages automatic transcription of recited sentences and evaluation based on natural language processing techniques. We demonstrate the effectiveness of our approach on a dataset comprised of recited sentences from preschool-aged children, incorporating ratings of semantic and syntactic correctness. The best performing systems achieve an F1 score of 91.7% for semantic correctness and 86.1% for syntactic correctness using automatic transcripts. Our results showcase the potential of automated evaluation systems in providing reliable and efficient assessments of memory capabilities in early childhood, facilitating timely interventions and support for children with language development needs. Ilja Baumann, Nicole Unger, Dominik Wagner 0002, Korbinian Riedhammer, Tobias Bocklet |
INTERSPEECH | 5 |
| 2024 | Infusing Acoustic Pause Context into Text-Based Dementia AssessmentabstractSpeech pauses, alongside content and structure, offer a valuable and non-invasive biomarker for detecting dementia. This work investigates the use of pause-enriched transcripts in transformer-based language models to differentiate the cognitive states of subjects with no cognitive impairment, mild cognitive impairment, and Alzheimer's dementia based on their speech from a clinical assessment. We address three binary classification tasks: Onset, monitoring, and dementia exclusion. The performance is evaluated through experiments on a German Verbal Fluency Test and a Picture Description Test, comparing the model's effectiveness across different speech production contexts. Starting from a textual baseline, we investigate the effect of incorporation of pause information and acoustic context. We show the test should be chosen depending on the task, and similarly, lexical pause information and acoustic cross-attention contribute differently. Franziska Braun, Sebastian P. Bayerl, Florian Hönig, Hartmut Lehfeld, Thomas Hillemacher, Tobias Bocklet, Korbinian Riedhammer |
INTERSPEECH | 6 |
| 2024 | It's Time to Take Action: Acoustic Modeling of Motor Verbs to Detect Parkinson's DiseaseabstractPre-trained models generate speech representations that are used in different tasks, including the automatic detection of Parkinson’s disease (PD). Although these models can yield high accuracy, their interpretation is still challenging. This paper used a pre-trained Wav2vec 2.0 model to represent speech frames of 25ms length and perform a frame-by-frame discrimination between PD patients and healthy control (HC) subjects. This fine granularity prediction enabled us to identify specific linguistic segments with high discrimination capability. Speech representations of all produced verbs were compared w.r.t. nouns and the first ones yielded higher accuracies. To gaina deeper understanding of this pattern, representations of motor and non-motor verbs were compared and the first ones yielded better results, with accuracies of around 83% in an independent test set. These findings support well-established neurocognitive models about action-related language highlighted as key drivers of PD. Index Terms: computational paralinguistics, interpretability of pre-trained models, action verbs, Parkinson’s disease Daniel Escobar-Grisales, Cristian D. Ríos-Urrego, Ilja Baumann, Korbinian Riedhammer, Elmar Nöth, Tobias Bocklet, Adolfo M. García, Juan Rafael Orozco-Arroyave |
INTERSPEECH | 6 |
| 2024 | Personalizing Large Sequence-to-Sequence Speech Foundation Models With Speaker RepresentationsabstractWe present a method to personalize large transformer-based encoder-decoder speech foundation models without the need for changes in the underlying model structure or training from scratch. This is achieved by projecting speaker-specific information into the latent space of the transformer decoder via a small neural network and learning to process the speaker information along with domain-specific information via parameter-efficient finetuning. We use this method to improve the automatic speech recognition results of spoken academic German and English. Our approach yields average relative word error rate (WER) improvements of approximately 29% on German academic speech and 25% on English academic speech. It also translates well to conversational speech, achieving relative WER improvements of up to 36%, and demonstrates modest gains of up to 5% on read speech. Moreover, we observe that incorporating utterances from the recent past as personalization context yields the most significant overall improvements and that changes in voice characteristics resulting from prolonged speaking have a minimal effect on the personalization quality of academic lectures. Dominik Wagner 0002, Ilja Baumann, Thomas Ranzenberger, Korbinian Riedhammer, Tobias Bocklet |
SLT | 5 |
| 2023 | Detection of Vowel Errors in Children's Speech using Synthetic Phonetic TranscriptsabstractThe analysis of phonological processes is crucial in evaluating speech development disorders in children, but encounters challenges due to limited children audio data. This work focuses on automatic vowel error detection using a two-stage pipeline. The first stage uses a fine-tuned cross-lingual phone recognizer (wav2vec 2.0) to extract phone sequences from audio. The second stage employs a language model (BERT) for classification from a phone sequence, entirely trained on synthetic transcripts, to counteract the very broad range of potential mistakes. We evaluate the system on nonword audio recordings recited by preschool children from a speech development test. The results show that the classifier trained on synthetic data performs well, but its efficacy relies on the quality of the phone recognizer. The best classifier achieves an 94.7% F1 score when evaluated against phonetic ground truths, whereas the F1 score is 76.2% when using automatically recognized phone sequences. Ilja Baumann, Dominik Wagner 0002, Korbinian Riedhammer, Elmar Nöth, Tobias Bocklet |
ASRU | 5 |
| 2023 | Speaker Adaptation for End-to-End Speech Recognition Systems in Noisy EnvironmentsabstractWe analyze the impact of speaker adaptation in end-to-end automatic speech recognition models based on transformers and wav2vec 2.0 under different noise conditions. By including speaker embeddings obtained from x-vector and ECAPA-TDNN systems, as well as i-vectors, we achieve relative word error rate improvements of up to 16.3% on LibriSpeech and up to 14.5% on Switchboard. We show that the proven method of concatenating speaker vectors to the acoustic features and supplying them as auxiliary model inputs remains a viable option to increase the robustness of end-to-end architectures. The effect on transformer models is stronger, when more noise is added to the input speech. The most substantial benefits for systems based on wav2vec 2.0 are achieved under moderate or no noise conditions. Both x-vectors and ECAPA-TDNN embeddings outperform i-vectors as speaker representations. The optimal embedding size depends on the dataset and also varies with the noise condition. Dominik Wagner 0002, Ilja Baumann, Sebastian P. Bayerl, Korbinian Riedhammer, Tobias Bocklet |
ASRU | 5 |
| 2023 | Semmeldetector: Application of Machine Learning in Commercial BakeriesabstractThe Semmeldetector, is a machine learning application that utilizes object detection models to detect, classify and count baked goods in images. Our application allows commercial bakers to track unsold baked goods, which allows them to op-timize production and increase resource efficiency. We compiled a dataset comprising 1151 images that distinguishes between 18 different types of baked goods to train our detection models. To facilitate model training, we used a Copy-Paste augmentation pipeline to expand our dataset. We trained the state-of-the-art object detection model YOLOv8 on our detection task. We tested the impact of different training data, model scale, and online image augmentation pipelines on model performance. Our overall best performing model, achieved an$AP_{{0.5}}$of 89.1 % on our test set. Based on our results, we conclude that machine learning can be a valuable tool even for unforeseen industries like bakeries, even with very limited datasets. Thomas H. Schmitt, Maximilian Bundscherer, Tobias Bocklet |
ICMLA | 3 |
| 2023 | Influence of Utterance and Speaker Characteristics on the Classification of Children with Cleft Lip and Palate
Ilja Baumann, Dominik Wagner 0002, Franziska Braun, Sebastian P. Bayerl, Elmar Nöth, Korbinian Riedhammer, Tobias Bocklet |
INTERSPEECH | 7 |
| 2023 | A Stutter Seldom Comes Alone - Cross-Corpus Stuttering Detection as a Multi-label Problem
Sebastian P. Bayerl, Dominik Wagner 0002, Ilja Baumann, Florian Hönig, Tobias Bocklet, Elmar Nöth, Korbinian Riedhammer |
INTERSPEECH | 5 |
| 2023 | Classifying Dementia in the Presence of Depression: A Cross-Corpus StudyabstractAutomated dementia screening enables early detection and intervention, reducing costs to healthcare systems and increasing quality of life for those affected. Depression has shared symptoms with dementia, adding complexity to diagnoses. The research focus so far has been on binary classification of dementia (DEM) and healthy controls (HC) using speech from picture description tests from a single dataset. In this work, we apply established baseline systems to discriminate cognitive impairment in speech from the semantic Verbal Fluency Test and the Boston Naming Test using text, audio and emotion embeddings in a 3-class classification problem (HC vs. MCI vs. DEM). We perform cross-corpus and mixed-corpus experiments on two independently recorded German datasets to investigate generalization to larger populations and different recording conditions. In a detailed error analysis, we look at depression as a secondary diagnosis to understand what our classifiers actually learn. Franziska Braun, Sebastian P. Bayerl, Paula Andrea Pérez-Toro, Florian Hönig, Hartmut Lehfeld, Thomas Hillemacher, Elmar Nöth, Tobias Bocklet, Korbinian Riedhammer |
INTERSPEECH | 8 |
| 2023 | Real Time Detection of Soft Voice for Speech Enhancement
Héctor A. Cordourier, Georg Stemmer, Sinem Aslan, Tobias Bocklet, Himanshu Bhalla |
INTERSPEECH | 4 |
| 2023 | Detection of Emotional Hotspots in Meetings Using a Cross-Corpus Approach
Georg Stemmer, Paulo Lopez-Meyer, Juan A. del Hoyo Ontiveros, Jose A. Lopez, Héctor A. Cordourier, Tobias Bocklet |
INTERSPEECH | 6 |
| 2023 | Multi-class Detection of Pathological Speech with Latent Features: How does it perform on unseen data?abstractThe detection of pathologies from speech features is usually defined as a binary classification task with one class representing a specific pathology and the other class representing healthy speech. In this work, we train neural networks, large margin classifiers, and tree boosting machines to distinguish between four pathologies: Parkinson's disease, laryngeal cancer, cleft lip and palate, and oral squamous cell carcinoma. We show that latent representations extracted at different layers of a pre-trained wav2vec 2.0 system can be effectively used to classify these types of pathological voices. We evaluate the robustness of our classifiers by adding room impulse responses to the test data and by applying them to unseen speech corpora. Our approach achieves unweighted average F1-Scores between 74.1% and 97.0%, depending on the model and the noise conditions used. The systems generalize and perform well on unseen data of healthy speakers sampled from a variety of different sources. Dominik Wagner 0002, Ilja Baumann, Franziska Braun, Sebastian P. Bayerl, Elmar Nöth, Korbinian Riedhammer, Tobias Bocklet |
INTERSPEECH | 7 |
| 2023 | User State Modeling Based on the Arousal-Valence Plane: Applications in Customer Satisfaction and Health-CareabstractThe acoustic analysis helps to discriminate emotions according to non-verbal information, while linguistics aims to capture verbal information from written sources. Acoustic and linguistic analyses can be addressed for different applications, where information related to emotions, mood, or affect are involved. The Arousal-Valence plane is commonly used to model emotional states in a multidimensional space. This study proposes a methodology focused on modeling the user’s state based on the Arousal-Valence plane in different scenarios. Acoustic and linguistic information are used as input to feed different deep learning architectures mainly based on convolutional and recurrent neural networks, which are trained to model the Arousal-Valence plane. The proposed approach is used for the evaluation of customer satisfaction in call-centers and for health-care applications in the assessment of depression in Parkinson’s disease and the discrimination of Alzheimer’s disease. F-scores of up to 0.89 are obtained for customer satisfaction, of up to 0.82 for depression in Parkinson’s patients, and of up to 0.80 for Alzheimer’s patients. The proposed approach confirms that there is information embedded in the Arousal-Valence plane that can be used for different purposes. Paula Andrea Pérez-Toro, Juan Camilo Vásquez-Correa, Tobias Bocklet, Elmar Nöth, Juan Rafael Orozco-Arroyave |
IEEE Trans. Affect. Comput. | 3 |
| 2022 | Nonwords Pronunciation Classification in Language Development Tests for Preschool ChildrenabstractThis work aims to automatically evaluate whether the language development of children is age-appropriate. Validated speech and language tests are used for this purpose to test the auditory memory. In this work, the task is to determine whether spoken nonwords have been uttered correctly. We compare different approaches that are motivated to model specific language structures: Low-level features (FFT), speaker embeddings (ECAPA-TDNN), grapheme-motivated embeddings (wav2vec 2.0), and phonetic embeddings in form of senones (ASR acoustic model). Each of the approaches provides input for VGG-like 5-layer CNN classifiers. We also examine the adaptation per nonword. The evaluation of the proposed systems was performed using recordings from different kindergartens of spoken nonwords. ECAPA-TDNN and low-level FFT features do not explicitly model phonetic information; wav2vec2.0 is trained on grapheme labels, our ASR acoustic model features contain (sub-)phonetic information. We found that the more granular the phonetic modeling is, the higher are the achieved recognition rates. The best system trained on ASR acoustic model features with VTLN achieved an accuracy of 89.4% and an area under the ROC (Receiver Operating Characteristic) curve (AUC) of 0.923. This corresponds to an improvement in accuracy of 20.2% and AUC of 0.309 relative compared to the FFT-baseline. Ilja Baumann, Dominik Wagner 0002, Sebastian P. Bayerl, Tobias Bocklet |
INTERSPEECH | 4 |
| 2022 | Generative Models for Improved Naturalness, Intelligibility, and Voicing of Whispered SpeechabstractThis work adapts two recent architectures of generative models and evaluates their effectiveness for the conversion of whispered speech to normal speech. We incorporate the normal target speech into the training criterion of vector-quantized variational autoencoders (VQ-VAEs) and Mel-GANs, thereby conditioning the systems to recover voiced speech from whispered inputs. Objective and subjective quality measures indicate that both VQ-VAEs and MelGANs can be modified to perform the conversion task. We find that the proposed approaches significantly improve the Mel cepstral distortion (MCD) metric by at least 25% relative to a Disco-GAN baseline. Subjective listening tests suggest that the MelGAN-based system significantly improves naturalness, intelligibility, and voicing compared to the whispered input speech. A novel evaluation measure based on differences between latent speech representations also indicates that our MelGAN-based approach yields improvements relative to the baseline. Dominik Wagner 0002, Sebastian P. Bayerl, Héctor A. Cordourier, Tobias Bocklet |
SLT | 4 |
| 2021 | Applying X-Vectors on Pathological Speech After Larynx RemovalabstractSpeaker embeddings extracted from time delayed neural networks (TDNNs) contributed to major recent advancements in speaker recognition and verification. We use an X-Vector system trained on augmented VoxCeleb1 and VoxCeleb2 data to obtain embeddings for pathological speech after total or partial larynx removal. We show that our model is able to effectively distinguish and visualize patient groups when generating embeddings. We further compare various regression models on the task of automatically predicting different perceptual ratings by speech therapists (intelligibility, vocal effort, and overall quality) based on the extracted speaker embeddings. For both patient groups we show Pearson correlations in the range of +0.8; we find that Random Forest and Support Vector Regression produce scores that best resemble the experts' assessments. Ralph Scheuerer, Tino Haderlein, Elmar Nöth, Tobias Bocklet |
ASRU | 4 |
| 2021 | Acoustic and Linguistic Analyses to Assess Early-Onset and Genetic Alzheimer's DiseaseabstractThe PSEN1-E280A or Paisa mutation is responsible for most of Early-Onset Alzheimer’s (EOA) disease cases in Colombia. It affects a large kindred of over 5000 members that present the same phenotype. The most common symptoms are related to language disorders, where speech fluency is also affected due to the difficulty to access semantic information intentionally. This study proposes the use of acoustic and linguistic methods to extract features from speech recordings and their transcriptions to discriminate people with conditions related to the Paisa mutation. We consider state-of-the-art word-embedding methods like Word2Vec and Bidirectional Encoder Representations from Transformer to process the transcripts. The speech signals are modeled by using traditional acoustic features and speaker embeddings. To the best of our knowledge, this is the first study focused on evaluating genetic Alzheimer’s and EOA using acoustics and linguistics. Paula Andrea Pérez-Toro, Juan Camilo Vásquez-Correa, Tomás Arias-Vergara, Philipp Klumpp, M. Sierra-Castrillón, M. E. Roldán-López, David Aguillón, Liliana Hincapié-Henao, Carlos Tobon 0001, Tobias Bocklet, Maria Schuster, Juan Rafael Orozco-Arroyave, Elmar Nöth |
ICASSP | 10 |
| 2021 | The Phonetic Footprint of Covid-19?abstractAgainst the background of the ongoing pandemic, this year’s Computational Paralinguistics Challenge featured a classification problem to detect Covid-19 from speech recordings. The presented approach is based on a phonetic analysis of speech samples, thus it enabled us not only to discriminate between Covid and non-Covid samples, but also to better understand how the condition influenced an individual’s speech signal. Our deep acoustic model was trained with datasets collected exclusively from healthy speakers. It served as a tool for segmentation and feature extraction on the samples from the challenge dataset. Distinct patterns were found in the embeddings of phonetic classes that have their place of articulation deep inside the vocal tract. We observed profound differences in classification results for development and test splits, similar to the baseline method. We concluded that, based on our phonetic findings, it was safe to assume that our classifier was able to reliably detect a pathological condition located in the respiratory tract. However, we found no evidence to claim that the system was able to discriminate between Covid-19 and other respiratory diseases. Philipp Klumpp, Tobias Bocklet, Tomás Arias-Vergara, Juan Camilo Vásquez-Correa, Paula Andrea Pérez-Toro, Sebastian P. Bayerl, Juan Rafael Orozco-Arroyave, Elmar Nöth |
Interspeech | 2 |
| 2020 | Comparison of User Models Based on GMM-UBM and I-Vectors for Speech, Handwriting, and Gait Assessment of Parkinson's Disease PatientsabstractParkinson's disease is a neurodegenerative disorder characterized by the presence of different motor impairments. Information from speech, handwriting, and gait signals have been considered to evaluate the neurological state of the patients. On the other hand, user models based on Gaussian mixture models - universal background models (GMMUBM) and i-vectors are considered the state-of-the-art in biometric applications like speaker verification because they are able to model specific speaker traits. This study introduces the use of GMM-UBM and i-vectors to evaluate the neurological state of Parkinson's patients using information from speech, handwriting, and gait. The results show the importance of different feature sets from each type of signal in the assessment of the neurological state of the patients. Juan Camilo Vásquez-Correa, Tobias Bocklet, Juan Rafael Orozco-Arroyave, Elmar Nöth |
ICASSP | 2 |
| 2020 | Length- and Noise-Aware Training Techniques for Short-Utterance Speaker RecognitionabstractSpeaker recognition performance has been greatly improved with the emergence of deep learning. Deep neural networks show the capacity to effectively deal with impacts of noise and reverberation, making them attractive to far-field speaker recognition systems. The x-vector framework is a popular choice for generating speaker embeddings in recent literature due to its robust training mechanism and excellent performance in various test sets. In this paper, we start with early work on including invariant representation learning (IRL) to the loss function and modify the approach with centroid alignment (CA) and length variability cost (LVC) techniques to further improve robustness in noisy, far-field applications. This work mainly focuses on improvements for short-duration test utterances (1-8s). We also present improved results on long-duration tasks. In addition, this work discusses a novel self-attention mechanism. On the VOiCES far-field corpus, the combination of the proposed techniques achieves relative improvements of 7.0% for extremely short and 8.2% for full-duration test utterances on equal error rate (EER) over our baseline system. Wenda Chen, Jonathan Huang, Tobias Bocklet |
INTERSPEECH | 3 |
| 2020 | Compact Speaker Embedding: lrx-VectorabstractDeep neural networks (DNN) have recently been widely used in speaker recognition systems, achieving state-of-the-art performance on various benchmarks. The x-vector architecture is especially popular in this research community, due to its excellent performance and manageable computational complexity. In this paper, we present the lrx-vector system, which is the low-rank factorized version of the x-vector embedding network. The primary objective of this topology is to further reduce the memory requirement of the speaker recognition system. We discuss the deployment of knowledge distillation for training the lrx-vector system and compare against low-rank factorization with SVD. On the VOiCES 2019 far-field corpus we were able to reduce the weights by 28% compared to the full-rank x-vector system while keeping the recognition rate constant (1.83% EER). Munir Georges, Jonathan Huang, Tobias Bocklet |
INTERSPEECH | 3 |
| 2020 | State Sequence Pooling Training of Acoustic Models for Keyword SpottingabstractWe propose a new training method to improve HMM-based keyword spotting. The loss function is based on a score computed with the keyword/filler model from the entire input sequence. It is equivalent to max/attention pooling but is based on prior acoustic knowledge. We also employ a multi-task learning setup by predicting both LVCSR and keyword posteriors. We compare our model to a baseline trained on frame-wise cross entropy, with and without per-class weighting. We employ a low-footprint TDNN for acoustic modeling. The proposed training yields significant and consistent improvement over the baseline in adverse noise conditions. The FRR on cafeteria noise is reduced from 13.07% to 5.28% at 9 dB SNR and from 37.44% to 6.78% at 5 dB SNR. We obtain these results with only 600 unique training keyword samples. The training method is independent of the frontend and acoustic model topology. Kuba Lopatka, Tobias Bocklet |
INTERSPEECH | 2 |
| 2019 | Ultra-Compact NLU: Neuronal Network Binarization as Regularization
Munir Georges, Krzysztof Czarnowski, Tobias Bocklet |
INTERSPEECH | 3 |
| 2019 | Intel Far-Field Speaker Recognition System for VOiCES Challenge 2019
Jonathan Huang, Tobias Bocklet |
INTERSPEECH | 2 |
| 2017 | On the impact of non-modal phonation on phonological featuresabstractDifferent modes of vibration of the vocal folds contribute significantly to the voice quality. The neutral mode phonation, often used in a modal voice, is one against which the other modes can be contrastively described, also called non-modal phonations. This paper investigates the impact of non-modal phonation on phonological posteriors, the probabilities of phonological features inferred from the speech signal using a deep learning approach. Five different non-modal phonations are considered: falsetto, creaky, harshness, tense and breathiness. The impact of such non-modal phonation on phonological features, the Sound Patterns of English (SPE), is investigated in both speech analysis and synthesis tasks. We found that breathy and tense phonation impact the SPE features less, creaky phonation impacts the features moderately, and harsh and falsetto phonation impact the phonological features the most. We also report invariant and the most different SPE features impacted by non-modal phonation. Milos Cernak, Elmar Nöth, Frank Rudzicz, Heidi Christensen, Juan Rafael Orozco-Arroyave, Raman Arora, Tobias Bocklet, Hamid R. Chinaei, Julius Hannink, Phani S. Nidadavolu, Juan Camilo Vásquez-Correa, Maria Yancheva, Alyssa Vann, Nikolai Vogler |
ICASSP | 7 |
| 2017 | Multi-view representation learning via gcca for multimodal analysis of Parkinson's diseaseabstractInformation from different bio-signals such as speech, handwriting, and gait have been used to monitor the state of Parkinson's disease (PD) patients, however, all the multimodal bio-signals may not always be available. We propose a method based on multi-view representation learning via generalized canonical correlation analysis (GCCA) for learning a representation of features extracted from handwriting and gait that can be used as a complement to speech-based features. Three different problems are addressed: classification of PD patients vs. healthy controls, prediction of the neurological state of PD patients according to the UPDRS score, and the prediction of a modified version of the Frenchay dysarthria assessment (m-FDA). According to the results, the proposed approach is suitable to improve the results in the addressed problems, specially in the prediction of the UPDRS, and m-FDA scores. Juan Camilo Vásquez-Correa, Juan Rafael Orozco-Arroyave, Raman Arora, Elmar Nöth, Najim Dehak, Heidi Christensen, Frank Rudzicz, Tobias Bocklet, Milos Cernak, Hamid R. Chinaei, Julius Hannink, Phani S. Nidadavolu, Maria Yancheva, Alyssa Vann, Nikolai Vogler |
ICASSP | 8 |
| 2017 | Speech Recognition and Understanding on Hardware-Accelerated DSP
Georg Stemmer, Munir Georges, Joachim Hofer, Piotr Rozen, Josef G. Bauer, Jakub Nowicki, Tobias Bocklet, Hannah R. Colett, Ohad Falik, Michael Deisher, Sylvia J. Downing |
INTERSPEECH | 7 |
| 2015 | A Survey on perceived speaker traits: Personality, likability, pathology, and the first challenge
Björn W. Schuller, Stefan Steidl, Anton Batliner, Elmar Nöth, Alessandro Vinciarelli, Felix Burkhardt, R. J. J. H. van Son, Felix Weninger, Florian Eyben, Tobias Bocklet, Gelareh Mohammadi, Benjamin Weiss 0001 |
Comput. Speech Lang. | 10 |
| 2014 | Are men more sleepy than women or does it only look like - Automatic analysis of sleepy speechabstractThe degree of sleepiness in the Sleepy Language Corpus from the Interspeech 2011 Speaker State Challenge is predicted with regression and a very large feature vector. Most notable is the great gender difference which can mainly be attributed to females showing their sleepiness less than males do. Florian Hönig, Anton Batliner, Tobias Bocklet, Georg Stemmer, Elmar Nöth, Sebastian Schnieder, Jarek Krajewski |
ICASSP | 3 |
| 2014 | Erlangen-CLP: A Large Annotated Corpus of Speech from Children with Cleft Lip and Palate
Tobias Bocklet, Andreas K. Maier, Korbinian Riedhammer, Ulrich Eysholdt, Elmar Nöth |
LREC | 1 |
| 2013 | Automatic phoneme analysis in children with Cleft Lip and PalateabstractCleft Lip and Palate (CLP) is among the most frequent congenital abnormalities. The impaired facial development affects the articulation, with different phonemes being impacted inhomogeneously among different patients. This work focuses on automatic phoneme analysis of children with CLP for a detailed diagnosis and therapy control. In clinical routine, the state-of-the-art evaluation is based on perceptual evaluations. Perceptual ratings act as ground-truth throughout this work, with the goal to build an automatic system that is as reliable as humans. We propose two different automatic systems focusing on modeling the articulatory space of a speaker: one system models a speaker by a GMM, the other system employs a speech recognition system and estimates fMLLR matrices for each speaker. SVR is then used to predict the perceptual ratings. We show that the fMLLR-based system is able to achieve automatic phoneme evaluation results that are in the same range as perceptual inter-rater-agreements. Tobias Bocklet, Korbinian Riedhammer, Ulrich Eysholdt, Elmar Nöth |
ICASSP | 1 |
| 2013 | Automatic evaluation of parkinson's speech - acoustic, prosodic and voice related cuesabstractArticulation and phonation is affected in 70 % to 90 % of patients with Parkinson’s disease (PD). This study focuses on the question whether speech carries information about 1. PD being present at a speaker or not, and 2. estimating the sever-ity of PD (if present). We first perform classification experi-ments focusing on the automatic detection of PD as a 2-class problem (PD vs. healthy speakers). The detection of severity is described as a 3-class task based on the Unified Parkinson’s Disease Rating Scale (UPDRS) ratings. We employ acous-tic, prosodic and glottal features on different kinds of speech tests: various syllable repetition tasks, read sentences and texts, and monologues. Classification is performed in either case by SVMs. We report recognition results of 81.9 % when trying to differentiate between normally speaking persons and speakers with PD. With system fusion we achieved a recognition results of 59.1 % on the task of UPDRS classification. Index Terms: Parkinson’s Disease, pathologic speech, speech analysis Tobias Bocklet, Stefan Steidl, Elmar Nöth, Sabine Skodda |
INTERSPEECH | 1 |
| 2012 | Revisiting semi-continuous hidden Markov modelsabstractIn the past decade, semi-continuous hidden Markov models (SCHMMs) have not attracted much attention in the speech recognition community. Growing amounts of training data and increasing sophistication of model estimation led to the impression that continuous HMMs are the best choice of acoustic model. However, recent work on recognition of under-resourced languages faces the same old problem of estimating a large number of parameters from limited amounts of transcribed speech. This has led to a renewed interest in methods of reducing the number of parameters while maintaining or extending the modeling capabilities of continuous models. In this work, we compare classic and multiple-codebook semi-continuous models using diagonal and full covariance matrices with continuous HMMs and subspace Gaussian mixture models. Experiments on the RM and WSJ corpora show that while a classical semicontinuous system does not perform as well as a continuous one, multiple-codebook semi-continuous systems can perform better, particular when using full-covariance Gaussians. Korbinian Riedhammer, Tobias Bocklet, Arnab Ghoshal, Daniel Povey |
ICASSP | 2 |
| 2012 | The Automatic Assessment of Non-native Prosody: Combining Classical Prosodic Analysis with Acoustic ModellingabstractIn earlier studies, we employed a large prosodic feature vector to assess the quality of L2 learner's utterances with respect to sentence melody and rhythm.In this paper, we combine these features with two standard approaches in paralinguistic analysis: (1) features derived from a Gaussian Mixture Model used as Universal Background Model (GMM-UBM), and (2) openSMILE, an open-source toolkit for extracting acoustic features.We evaluate our approach with English speech from 94 non-native speakers perceptually scored by 62 native labellers.GMM-UBM or openSMILE modelling alone yields lower performance than our prosodic feature vector; however, adding information from the GMM-UBM modelling or openSMILE by late fusion improves results. Florian Hönig, Tobias Bocklet, Korbinian Riedhammer, Anton Batliner, Elmar Nöth |
INTERSPEECH | 2 |
| 2012 | The INTERSPEECH 2012 Speaker Trait ChallengeabstractLIDIAP Björn W. Schuller, Stefan Steidl, Anton Batliner, Elmar Nöth, Alessandro Vinciarelli, Felix Burkhardt, R. J. J. H. van Son, Felix Weninger, Florian Eyben, Tobias Bocklet, Gelareh Mohammadi, Benjamin Weiss 0001 |
INTERSPEECH | 10 |
| 2011 | Detection of persons with Parkinson's disease by acoustic, vocal, and prosodic analysisabstract70% to 90% of patients with Parkinson's disease (PD) show an affected voice. Various studies revealed, that voice and prosody is one of the earliest indicators of PD. The issue of this study is to automatically detect whether the speech/voice of a person is affected by PD. We employ acoustic features, prosodic features and features derived from a two-mass model of the vocal folds on different kinds of speech tests: sustained phonations, syllable repetitions, read texts and monologues. Classification is performed in either case by SVMs. A correlation-based feature selection was performed, in order to identify the most important features for each of these systems. We report recognition results of 91% when trying to differentiate between normal speaking persons and speakers with PD in early stages with prosodic modeling. With acoustic modeling we achieved a recognition rate of 88% and with vocal modeling we achieved 79%. After feature selection these results could greatly be improved. But we expect those results to be too optimistic. We show that read texts and monologues are the most meaningful texts when it comes to the automatic detection of PD based on articulation, voice, and prosodic evaluations. The most important prosodic features were based on energy, pauses and F0. The masses and the compliances of spring were found to be the most important parameters of the two-mass vocal fold model. Tobias Bocklet, Elmar Nöth, Georg Stemmer, Hana Ruzickova, Jan Rusz |
ASRU | 1 |
| 2011 | Compensation of extrinsic variability in speaker verification systems on simulated Skype and HF channel dataabstractIn this work we focus on speaker verification on channels of varying quality, namely Skype and high frequency (HF) radio. In our setup, we assume to have telephone recordings of speakers for training, but recordings of different channels for testing with varying (lower) signal quality. Starting from a Gaussian mixture / support vector machine (GMM/SVM) baseline, we evaluate multi-condition training (MCT), an ideal channel classification approach (ICC), and nuisance attribute projection (NAP) to compensate for the loss of information due to the transmission. In an evaluation on Switchboard-2 data using Skype and HF channel simulators, we show that, for good signal quality, NAP improves the baseline system performance from 5% EER to 3.33% EER (for both Skype and HF). For strongly distorted data, MCT or, if adequate, ICC turn out to be the method of choice. Korbinian Riedhammer, Tobias Bocklet, Elmar Nöth |
ICASSP | 2 |
| 2011 | Drink and Speak: On the Automatic Classification of Alcohol Intoxication by Acoustic, Prosodic and Text-Based FeaturesabstractThis paper focuses on the automatic detection of a person’s blood level alcohol based on automatic speech processing ap-proaches. We compare 5 different feature types with different ways of modeling. Experiments are based on the ALC corpus of IS2011 Speaker State Challenge. The classification task is restricted to the detection of a blood alcohol level above 0.5‰. Three feature sets are based on spectral observations: MFCCs, PLPs, TRAPS. These are modeled by GMMs. Classification is either done by a Gaussian classifier or by SVMs. In the later case classification is based on GMM-based supervectors, i.e. concatenation of GMM mean vectors. A prosodic system extracts a 292-dimensional feature vector based on a voiced-unvoiced decision. A transcription-based system makes use of text transcriptions related to phoneme durations and textual structure. We compare the stand-alone performances of these systems and combine them on score level by logistic regres-sion. The best stand-alone performance is the transcription-based system which outperforms the baseline by 4.8 % on the development set. A Combination on score level gave a huge boost when the spectral-based systems were added (73.6 %). This is a relative improvement of 12.7 % to the baseline. On the test-set we achieved an UA of 68.6 % which is a significant improvement of 4.1 % to the baseline system. Index Terms: GMM, alcohol intoxication, system fusion 1. Tobias Bocklet, Korbinian Riedhammer, Elmar Nöth |
INTERSPEECH | 1 |
| 2011 | Combining Phonological and Acoustic ASR-Free Features for Pathological Speech Intelligibility AssessmentabstractIntelligibility is widely used to measure the severity of articulatory problems in pathological speech. Recently, a number of automatic intelligibility assessment tools have been developed. Most of them use automatic speech recognizers (ASR) to compare the patient's utterance with the target text. These methods are bound to one language and tend to be less accurate when speakers hesitate or make reading errors. To circumvent these problems, two different ASR-free methods were developed over the last few years, only making use of the acoustic or phonological properties of the utterance. In this paper, we demonstrate that these ASR-free techniques are also able to predict intelligibility in other languages. Moreover, they show to be complementary, resulting in even better intelligibility predictions when both methods are combined. Catherine Middag, Tobias Bocklet, Jean-Pierre Martens, Elmar Nöth |
INTERSPEECH | 2 |
| 2011 | Java Visual Speech Components for Rapid Application Development of GUI Based Speech Processing ApplicationsabstractIn this paper, we describe a new Java framework for an easy and efficient way of developing new GUI based speech processing applications. Standard components are provided to display the speech signal, the power plot, and the spectrogram. Furthermore, a component to create a new transcription and to display and manipulate an existing transcription is provided, as well as a component to display and manually correct external pitch values. These Swing components can be easily embedded into own Java programs. They can be synchronized to display the same region of the speech file. The object-oriented design provides base classes for rapid development of own components. Stefan Steidl, Korbinian Riedhammer, Tobias Bocklet, Florian Hönig, Elmar Nöth |
INTERSPEECH | 3 |
| 2010 | Clap your hands! Calibrating spectral subtraction for dereverberationabstractReverberation effects as observed by room microphones severely degrade the performance of automatic speech recognition systems. We investigate the use of dereverberation by spectral subtraction as proposed by Lebart and Boucher and introduce a simple approach to estimate the required decay parameter by clapping hands. Experiments on small vocabulary continuous speech recognition task on read speech show that using the calibrated dereverberation improves WER from 73.2 to 54.7 for the best microphone. In combination with system adaptation, the WER could be reduced to 28.2, which is only a 16% relative loss of performance comparison to using a headset instead of a room microphone. Uwe Zah, Korbinian Riedhammer, Tobias Bocklet, Elmar Nöth |
ICASSP | 3 |
| 2010 | Age and gender recognition based on multiple systems - early vs. late fusionabstractThis paper focuses on the automatic recognition of a per-son’s age and gender based only on his or her voice. Up to five different systems are compared and combined in dif-ferent configurations: three systems model the speaker’s characteristics in different feature spaces, i.e., MFCC, PLP, TRAPS, by Gaussian mixture models. The features of these systems are the concatenated mean vectors. Sys-tem number 4 uses a physical two-mass vocal model and estimates in a data-driven optimization procedure 9 glot-tal features from voiced speech sections. For each ut-terance the minimum, maximum and mean vectors form a 27-dimensional feature vector. The last system calcu-lates a 219-dimensional prosodic feature set for each ut-terance based on voice and unvoiced speech segments. We compare two different ways to fuse the different sys-tems: First, we concatenate the system on feature level. The second way of combination is performed on score level by multi-class logistic regression. Despite there are just minor differences between the two approaches, late fusion is slightly superior. On the development set of the Interspeech Agender challenge we achieved an un-weighted recall of 46.1 % with early fusion and 47.8% with late fusion. Tobias Bocklet, Georg Stemmer, Viktor Zeißler, Elmar Nöth |
INTERSPEECH | 1 |
| 2010 | Improvement of a speech recognizer for standardized medical assessment of children's speech by integration of prior knowledgeabstractSpeech recognition of children is a more difficult task than speech recognition of adults. This problem is amplified for children with articulation disorders like cleft lip and palate (CLP). In this work we improved our automatic speech recognition system by integrating prior knowledge. Prior knowledge focuses on two different aspects: A test-dependent language modeling and an age-dependent acoustic modeling. These two approaches are merged at the end to different test- and age-dependent recognizers. We evaluated our system on a dataset of 35 children with CLP. Significant improvements could be found on this dataset. With our baseline system we achieved a negative word accuarcy (WA) of -11.0%. By an extended language modeling we achieved 27.5%. The age-dependent recognition system gains a huge improvement and achieves aWA of 42.6%. With the significant improvements in WA it is possible to perform an automatic detection and identification of specific words. Thus, we took the first step towards a speech assessment on word and subword level. Tobias Bocklet, Andreas K. Maier, Ulrich Eysholdt, Elmar Nöth |
SLT | 1 |
| 2009 | Speaker recognition using syllable-based constraints for cepstral frame selectionabstractWe describe a new GMM-UBM speaker recognition system that uses standard cepstral features, but selects different frames of speech for different subsystems. Subsystems, or ldquoconstraintsrdquo, are based on syllable-level information and combined at the score level. Results on both the NIST 2006 and 2008 test data sets for the English telephone train and test condition reveal that a set of eight constraints performs extremely well, resulting in better performance than other commonly-used cepstral models. Given the still largely-unexplored world of possible constraints and combinations, it is likely that the approach can be even further improved. Tobias Bocklet, Elizabeth Shriberg |
ICASSP | 1 |
| 2009 | THE SRI NIST 2008 speaker recognition evaluation systemabstractThe SRI speaker recognition system for the 2008 NIST speaker recognition evaluation (SRE) incorporates a variety of models and features, both cepstral and stylistic. We highlight the improvements made to specific subsystems and analyze the performance of various subsystem combinations in different data conditions. We show the importance of language and nativeness conditioning, as well as the role of ASR for speaker verification. Sachin S. Kajarekar, Nicolas Scheffer, Martin Graciarena, Elizabeth Shriberg, Andreas Stolcke, Luciana Ferrer, Tobias Bocklet |
ICASSP | 7 |
| 2009 | Feature-based and channel-based analyses of intrinsic variability in speaker verificationabstractWe explore how intrinsic variations (those associated with the speaker rather than the recording environment) affect textindependent speaker verification performance. In a previous paper we introduced the SRI-FRTIV corpus and provided speaker verification results using a Gaussian mixture model (GMM) system on telephone-channel speech. In this paper we explore the use of other speaker verification systems on the telephone channel data and compare against the GMM baseline. We found the GMM system to be one of the more robust across all conditions. Systems relying on recognition hypotheses had a significant degradation in low vocal effort conditions. We also explore the use of the GMM system on several other channels. We found improved performance on table-top microphones compared to the telephone channel in furtive conditions and gradual degradations as a function of the distance from the microphone to the speaker. Therefore distant microphones further degrade the speaker verification performance due to intrinsic variability. Index Terms: speaker recognition, vocal effort, speaking style, intrinsic variation, furtive speech, interview speech, read Martin Graciarena, Tobias Bocklet, Elizabeth Shriberg, Andreas Stolcke, Sachin S. Kajarekar |
INTERSPEECH | 2 |
| 2008 | Age and gender recognition for telephone applications based on GMM supervectors and support vector machinesabstractThis paper compares two approaches of automatic age and gender classification with 7 classes. The first approach are Gaussian mixture models (GMMs) with universal background models (UBMs), which is well known for the task of speaker identification/verification. The training is performed by the EM algorithm or MAP adaptation respectively. For the second approach for each speaker of the test and training set a GMM model is trained. The means of each model are extracted and concatenated, which results in a GMM supervector for each speaker. These supervectors are then used in a support vector machine (SVM). Three different kernels were employed for the SVM approach: a polynomial kernel (with different polynomials), an RBF kernel and a linear GMM distance kernel, based on the KL divergence. With the SVM approach we improved the recognition rate to 74% (p < 0.001) and are in the same range as humans. Tobias Bocklet, Andreas K. Maier, Josef G. Bauer, Felix Burkhardt, Elmar Nöth |
ICASSP | 1 |