Dominik Wagner 0002

dblp:180/8364-2 · DBLP profile ↗
← Back
26ranked-venue papers
11as first author
26since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 23 · 10 first-author · 23 since 2021Artificial intelligence and machine learning · 20 · 7 first-author · 20 since 2021
YearPublicationVenuePosition
2025 Text-Guided Speech Representations for Language Acquisition Assessment
abstract
Assessing spoken language abilities in children is critical for early detection of developmental delays. We propose a two-stage framework for automatic pronunciation assessment, targeting classification of items in a language development screening. In the first stage, we fine-tune a sentence transformer using phonetic and label-informed losses to create a structured latent text space. In the second stage, we fine-tune an audio encoder to project spoken utterances into the text space. Our approach enables inference from audio alone while preserving alignment with phonetically meaningful structure. Compared to baseline systems, our method achieves the best classification performance and provides a latent space suitable for downstream analysis. This facilitates the identification of phonological or grammatical error patterns from cluster structure, paving the way for potentially interpretable assessment of early language acquisition.
Ilja Baumann, Dominik Wagner 0002, Philipp Seeberger, Korbinian Riedhammer, Tobias Bocklet
ASRU2
2025 On the Difficulty of Token-Level Modeling of Dysfluency and Fluency Shaping Artifacts
abstract
Automatic transcription of stuttered speech remains a challenge, even for modern end-to-end (E2E) automatic speech recognition (ASR) frameworks. Dysfluencies and fluency-shaping artifacts are often overlooked, resulting in non-verbatim transcriptions with limited clinical and research value. We propose a parameter-efficient adaptation method to decode dysfluencies and fluency modifications as special tokens within transcriptions, evaluated on simulated (LibriStutter, English) and natural (KSoF, German) stuttered speech datasets. To mitigate ASR performance disparities and bias towards English, we introduce a multi-step fine-tuning strategy with language-adaptive pretraining. Tokenization analysis further highlights the tokenizer’s English-centric bias, which poses challenges for improving performance on German data. Our findings demonstrate the effectiveness of lightweight adaptation techniques for dysfluency-aware ASR while exposing key limitations in multilingual E2E systems.
Kashaf Gulzar, Dominik Wagner 0002, Sebastian P. Bayerl, Florian Hönig, Tobias Bocklet, Korbinian Riedhammer
ASRU2
2025 Improving Multimodal Speech-To-Slide Alignment for Academic Lectures with Vision LLMs
abstract
We enhance the MaViLS multimodal algorithm to improve speech-to-slide alignment for lecture podcasts. Our approach integrates vision large language models for optical character recognition and automatic generation of lecture transcripts from slide content, coupled with a multilingual multimodal embedding model for text and image alignment. By combining slide-extracted text with automatically generated lecture transcripts and captioned slide images, we generate enhanced audio features that better capture speech-slide mapping. Our method improves the average F1 score for audio feature alignment on the English MaViLS dataset from 0.51 to 0.71 and on a newly created German podcast lectures dataset from 0.65 to 0.84.
Thomas Ranzenberger, Dominik Wagner 0002, Steffen Freisinger, Tobias Bocklet, Korbinian Riedhammer
ASRU2
2025 Joint ASR and Speech Attribute Prediction for Conversational Dysarthric Speech Analysis with Multimodal Language Models
abstract
Dysarthric speech recognition systems often focus solely on transcription, limiting their applicability in clinical and assistive settings where assessments of perceptual attributes like intelligibility and naturalness are essential. We propose a multimodal conversational framework based on Phi-4-Multimodal that combines ASR with attribute rating prediction, enabling users to query both transcriptions and perceptual characteristics (e.g. “How intelligible is this utterance?”). Our multitask model adds auxiliary prediction heads for five clinically relevant attributes and is trained on the English Speech Accessibility Project dataset. The system achieves competitive ASR performance while delivering attribute-level feedback comparable to specialized classifiers. Additional experiments show improved ASR performance for German Parkinson’s speech, indicating preserved multilingual capabilities and partial cross-lingual transfer of dysarthric speech patterns.
Dominik Wagner 0002, Ilja Baumann, Natalie Engert, Elmar Nöth, Korbinian Riedhammer, Tobias Bocklet
ASRU1
2025 SELMA: A Speech-Enabled Language Model for Virtual Assistant Interactions
abstract
In this work, we present and evaluate SELMA, a Speech-Enabled Language Model for virtual Assistant interactions that integrates audio and text as inputs to a Large Language Model (LLM). SELMA is designed to handle three primary and two auxiliary tasks related to interactions with virtual assistants simultaneously within a single end-to-end model. We employ low-rank adaptation modules for parameter-efficient training of both the audio encoder and the LLM. Additionally, we implement a feature pooling strategy enabling the system to recognize global patterns and improve accuracy on tasks less reliant on individual sequence elements. Experimental results on Voice Trigger (VT) detection, Device-Directed Speech Detection (DDSD), and Automatic Speech Recognition (ASR), demonstrate that our approach both simplifies the typical input processing pipeline of virtual assistants significantly and also improves performance compared to dedicated models for each individual task. SELMA yields relative Equal-Error Rate improvements of 64% on the VT detection task, and 22% on DDSD, while also achieving word error rates close to the baseline.
Dominik Wagner 0002, Alexander W. Churchill, Siddharth Sigtia, Erik Marchi
ICASSP1
2025 Optimized Self-supervised Training with BEST-RQ for Speech Recognition
abstract
Self-supervised learning has been successfully used for various speech related tasks, including automatic speech recognition. BERT-based Speech pre-Training with Random-projection Quantizer (BEST-RQ) has achieved state-of-the-art results in speech recognition. In this work, we further optimize the BEST-RQ approach using Kullback-Leibler divergence as an additional regularizing loss and multi-codebook extension per cluster derived from low-level feature clustering. Preliminary experiments on train-100 split of LibriSpeech result in a relative improvement of 11.2% on test-clean by using multiple codebooks, utilizing a combination of cross-entropy and Kullback-Leibler divergence further reduces the word error rate by 4.5%. The proposed optimizations on full LibriSpeech pre-training and fine-tuning result in relative word error rate improvements of up to 23.8% on test-clean and 30.6% on test-other using 6 codebooks. Furthermore, the proposed setup leads to faster convergence in pre-training and fine-tuning and additionally stabilizes the pre-training.
Ilja Baumann, Dominik Wagner 0002, Korbinian Riedhammer, Tobias Bocklet
ICASSP2
2025 HarmAug: Effective Data Augmentation for Knowledge Distillation of Safety Guard Models
abstract
Safety guard models that detect malicious queries aimed at large language models (LLMs) are essential for ensuring the secure and responsible deployment of LLMs in real-world applications. However, deploying existing safety guard models with billions of parameters alongside LLMs on mobile devices is impractical due to substantial memory requirements and latency. To reduce this cost, we distill a large teacher safety guard model into a smaller one using a labeled dataset of instruction-response pairs with binary harmfulness labels. Due to the limited diversity of harmful instructions in the existing labeled dataset, naively distilled models tend to underperform compared to larger models. To bridge the gap between small and large models, we propose **HarmAug**, a simple yet effective data augmentation method that involves jailbreaking an LLM and prompting it to generate harmful instructions. Given a prompt such as, "Make a single harmful instruction prompt that would elicit offensive content", we add an affirmative prefix (e.g., "I have an idea for a prompt:") to the LLM's response. This encourages the LLM to continue generating the rest of the response, leading to sampling harmful instructions. Another LLM generates a response to the harmful instruction, and the teacher model labels the instruction-response pair. We empirically show that our HarmAug outperforms other relevant baselines. Moreover, a 435-million-parameter safety guard model trained with HarmAug achieves an F1 score comparable to larger models with over 7 billion parameters, and even outperforms them in AUPRC, while operating at less than 25\% of their computational cost. Our [code](https://anonymous.4open.science/r/HarmAug/), [safety guard model](https://huggingface.co/AnonHB/HarmAug_Guard_Model_deberta_v3_large_finetuned), and [synthetic dataset](https://huggingface.co/datasets/AnonHB/HarmAug_generated_dataset) are publicly available.
Seanie Lee, Haebin Seong, Dong Bok Lee, Minki Kang, Xiaoyin Chen, Dominik Wagner 0002, Yoshua Bengio, Juho Lee 0001, Sung Ju Hwang
ICLR6
2025 Personalized Fine-Tuning with Controllable Synthetic Speech from LLM-Generated Transcripts for Dysarthric Speech Recognition
abstract
In this work, we present our submission to the Speech Accessibility Project challenge for dysarthric speech recognition. We integrate parameter-efficient fine-tuning with latent audio representations to improve an encoder-decoder ASR system. Synthetic training data is generated by fine-tuning Parler-TTS to mimic dysarthric speech, using LLM-generated prompts for corpus-consistent target transcripts. Personalization with x-vectors consistently reduces word error rates (WERs) over non-personalized fine-tuning. AdaLoRA adapters outperform full fine-tuning and standard low-rank adaptation, achieving relative WER reductions of ∼23% and ∼22%, respectively. Further improvements (∼5% WER reduction) come from incorporating wav2vec 2.0-based audio representations. Training with synthetic dysarthric speech yields up to ∼7% relative WER improvement over personalized fine-tuning alone.
Dominik Wagner 0002, Ilja Baumann, Natalie Engert, Seanie Lee, Elmar Nöth, Korbinian Riedhammer, Tobias Bocklet
INTERSPEECH1
2025 Pathology-Aware Speech Encoding and Data Augmentation for Dysarthric Speech Recognition
abstract
Automatic speech recognition (ASR) for pathologic speech remains a major challenge due to high variability in articulation, phonation, and prosody distortions. In this work, we propose a pathology-aware speech encoder based on BEST-RQ pre-training, which incorporates 46k hours of speech, including pathologic and atypical speech. We continue pre-training for domain adaptation and experiment with etiology-specific codebooks. We achieve a 13.2% relative word error rate (WER) improvement using the pathology-aware speech encoder with etiology-specific continued pre-training. Additionally, we examine the impact of incorporating synthetic and out-of-domain (OOD) data to further enhance ASR performance. Synthetic data reduces WER by up to 8.7%, while OOD data improves WER by 12.2%. Finally, we introduce a semantic similaritybased data augmentation technique to optimize data selection, achieving a WER improvement of up to 9.7% while minimizing the need for additional training data.
Ilja Baumann, Dominik Wagner 0002, Korbinian Riedhammer, Tobias Bocklet
INTERSPEECH2
2025 FedSVD: Adaptive Orthogonalization for Private Federated Learning with LoRA
abstract
Low-Rank Adaptation (LoRA), which introduces a product of two trainable low-rank matrices into frozen pre-trained weights, is widely used for efficient fine-tuning of language models in federated learning (FL). However, when combined with differentially private stochastic gradient descent (DP-SGD), LoRA faces substantial noise amplification: DP-SGD perturbs per-sample gradients, and the matrix multiplication of the LoRA update ($BA$) intensifies this effect. Freezing one matrix (*e.g.*, $A$) reduces the noise but restricts model expressiveness, often resulting in suboptimal adaptation. To address this, we propose $\texttt{FedSVD}$, a simple yet effective method that introduces a global reparameterization based on singular value decomposition (SVD). In our approach, each client optimizes only the $B$ matrix and transmits it to the server. The server aggregates the $B$ matrices, computes the product $BA$ using the previous $A$, and refactorizes the result via SVD. This yields a new adaptive $A$ composed of the orthonormal right singular vectors of $BA$, and an updated $B$ containing the remaining SVD components. This reparameterization avoids quadratic noise amplification, while allowing $A$ to better capture the principal directions of the aggregate updates. Moreover, the orthonormal structure of $A$ bounds the gradient norms of $B$ and preserves more signal under DP-SGD, as confirmed by our theoretical analysis. As a result, $\texttt{FedSVD}$ consistently improves stability and performance across a variety of privacy settings and benchmarks, outperforming relevant baselines under both private and non-private regimes.
Seanie Lee, Dong Bok Lee, Dominik Wagner 0002, Haebin Seong, Tobias Bocklet, Juho Lee 0001, Sung Ju Hwang
NeurIPS4
2024 Optimized Speculative Sampling for GPU Hardware Accelerators
abstract
In this work, we optimize speculative sampling for parallel hardware accelerators to improve sampling speed.We notice that substantial portions of the intermediate matrices necessary for speculative sampling can be computed concurrently.This allows us to distribute the workload across multiple GPU threads, enabling simultaneous operations on matrix segments within thread blocks.This results in profiling time improvements ranging from 6% to 13% relative to the baseline implementation, without compromising accuracy.To further accelerate speculative sampling, probability distributions parameterized by softmax are approximated by sigmoid.This approximation approach results in significantly greater relative improvements in profiling time, ranging from 37% to 94%, with a minor decline in accuracy.We conduct extensive experiments on both automatic speech recognition and summarization tasks to validate the effectiveness of our optimization methods.
Dominik Wagner 0002, Seanie Lee, Ilja Baumann, Philipp Seeberger, Korbinian Riedhammer, Tobias Bocklet
EMNLP1
2024 Towards Interpretability of Automatic Phoneme Analysis in Cleft Lip and Palate Speech
abstract
Cleft Lip and Palate ranks among the most common congenital abnormalities and significantly influences speech articulation, resulting in varying phonemic impacts. In a clinical context, a detailed diagnosis is carried out by time-consuming perceptual evaluations. We use perceptual ratings of different articulatory modifications on phoneme-level as ground-truth and propose a system based on wav2vec 2.0, trained to the downstream task of classifying phonemic criteria as a multi-class and multi-label problem. The system is trained for detection on utterance level, without the usage of phoneme labels. To gain a clearer understanding of which areas of the speech signal have the greatest impact on classification, we assess the extent to which our system aligns with expert ratings at the phoneme level. Additionally, we examine which specific phonemes play a decisive role in determining the final classification of the labeled criteria. The results show that salient phonemes marked by experts contribute remarkably greater to the classification of the correct class using feature relevance explanation methods. To the best of our knowledge, this is the first study incorporating various utterance-level articulatory modifications classification and phoneme-level interpretation, offering a more comprehensive understanding for potential clinical applications.
Ilja Baumann, Dominik Wagner 0002, Maria Schuster, Elmar Nöth, Tobias Bocklet
ICASSP2
2024 A Multimodal Approach to Device-Directed Speech Detection with Large Language Models
abstract
Interactions with virtual assistants typically start with a predefined trigger phrase followed by the user command. To make interactions with the assistant more intuitive, we explore whether it is feasible to drop the requirement that users must begin each command with a trigger phrase. We explore this task in three ways: First, we train classifiers using only acoustic information obtained from the audio waveform. Second, we take the decoder outputs of an automatic speech recognition (ASR) system, such as 1-best hypotheses, as input features to a large language model (LLM). Finally, we explore a multimodal system that combines acoustic and lexical features, as well as ASR decoder signals in an LLM. Using multimodal information yields relative equal-error-rate improvements over text-only and audio-only models of up to 39% and 61%. Increasing the size of the LLM and training with low-rank adaption leads to further relative EER reductions of up to 18% on our dataset.
Dominik Wagner 0002, Alexander W. Churchill, Siddharth Sigtia, Panayiotis G. Georgiou, Matt Mirsamadi, Aarshee Mishra, Erik Marchi
ICASSP1
2024 Large Language Models for Dysfluency Detection in Stuttered Speech
Dominik Wagner 0002, Sebastian P. Bayerl, Ilja Baumann, Elmar Nöth, Korbinian Riedhammer, Tobias Bocklet
INTERSPEECH1
2024 Outlier Reduction with Gated Attention for Improved Post-training Quantization in Large Sequence-to-sequence Speech Foundation Models
abstract
This paper explores the improvement of post-training quantization (PTQ) after knowledge distillation in the Whisper speech foundation model family. We address the challenge of outliers in weights and activation tensors, known to impede quantization quality in transformer-based language and vision models. Extending this observation to Whisper, we demonstrate that these outliers are also present when transformer-based models are trained to perform automatic speech recognition, necessitating mitigation strategies for PTQ. We show that outliers can be reduced by a recently proposed gating mechanism in the attention blocks of the student model, enabling effective 8-bit quantization, and lower word error rates compared to student models without the gating mechanism in place.
Dominik Wagner 0002, Ilja Baumann, Korbinian Riedhammer, Tobias Bocklet
INTERSPEECH1
2024 Towards Self-Attention Understanding for Automatic Articulatory Processes Analysis in Cleft Lip and Palate Speech
abstract
Cleft lip and palate (CLP) speech presents unique challenges for automatic phoneme analysis due to its distinct acoustic characteristics and articulatory anomalies. We perform phoneme analysis in CLP speech using a pre-trained wav2vec 2.0 model with a multi-head self-attention classification module to capture long-range dependencies within the speech signal, thereby enabling better contextual understanding of phoneme sequences. We demonstrate the effectiveness of our approach in the classification of various articulatory processes in CLP speech. Furthermore, we investigate the interpretability of self-attention to gain insights into the model’s understanding of CLP speech characteristics. Our findings highlight the potential of the selfattention mechanisms for improving automatic phoneme analysis in CLP speech, paving the way for enhanced diagnostics, adding interpretability for therapists and affected patients.
Ilja Baumann, Dominik Wagner 0002, Maria Schuster, Korbinian Riedhammer, Elmar Nöth, Tobias Bocklet
INTERSPEECH2
2024 Automatic Evaluation of a Sentence Memory Test for Preschool Children
abstract
Assessment of memory capabilities in preschool-aged children is crucial for early detection of potential speech development impairments or delays. We present an approach for the automatic evaluation of a standardized sentence memory test specifically for preschool children. Our methodology leverages automatic transcription of recited sentences and evaluation based on natural language processing techniques. We demonstrate the effectiveness of our approach on a dataset comprised of recited sentences from preschool-aged children, incorporating ratings of semantic and syntactic correctness. The best performing systems achieve an F1 score of 91.7% for semantic correctness and 86.1% for syntactic correctness using automatic transcripts. Our results showcase the potential of automated evaluation systems in providing reliable and efficient assessments of memory capabilities in early childhood, facilitating timely interventions and support for children with language development needs.
Ilja Baumann, Nicole Unger, Dominik Wagner 0002, Korbinian Riedhammer, Tobias Bocklet
INTERSPEECH3
2024 Personalizing Large Sequence-to-Sequence Speech Foundation Models With Speaker Representations
abstract
We present a method to personalize large transformer-based encoder-decoder speech foundation models without the need for changes in the underlying model structure or training from scratch. This is achieved by projecting speaker-specific information into the latent space of the transformer decoder via a small neural network and learning to process the speaker information along with domain-specific information via parameter-efficient finetuning. We use this method to improve the automatic speech recognition results of spoken academic German and English. Our approach yields average relative word error rate (WER) improvements of approximately 29% on German academic speech and 25% on English academic speech. It also translates well to conversational speech, achieving relative WER improvements of up to 36%, and demonstrates modest gains of up to 5% on read speech. Moreover, we observe that incorporating utterances from the recent past as personalization context yields the most significant overall improvements and that changes in voice characteristics resulting from prolonged speaking have a minimal effect on the personalization quality of academic lectures.
Dominik Wagner 0002, Ilja Baumann, Thomas Ranzenberger, Korbinian Riedhammer, Tobias Bocklet
SLT1
2023 Detection of Vowel Errors in Children's Speech using Synthetic Phonetic Transcripts
abstract
The analysis of phonological processes is crucial in evaluating speech development disorders in children, but encounters challenges due to limited children audio data. This work focuses on automatic vowel error detection using a two-stage pipeline. The first stage uses a fine-tuned cross-lingual phone recognizer (wav2vec 2.0) to extract phone sequences from audio. The second stage employs a language model (BERT) for classification from a phone sequence, entirely trained on synthetic transcripts, to counteract the very broad range of potential mistakes. We evaluate the system on nonword audio recordings recited by preschool children from a speech development test. The results show that the classifier trained on synthetic data performs well, but its efficacy relies on the quality of the phone recognizer. The best classifier achieves an 94.7% F1 score when evaluated against phonetic ground truths, whereas the F1 score is 76.2% when using automatically recognized phone sequences.
Ilja Baumann, Dominik Wagner 0002, Korbinian Riedhammer, Elmar Nöth, Tobias Bocklet
ASRU2
2023 Speaker Adaptation for End-to-End Speech Recognition Systems in Noisy Environments
abstract
We analyze the impact of speaker adaptation in end-to-end automatic speech recognition models based on transformers and wav2vec 2.0 under different noise conditions. By including speaker embeddings obtained from x-vector and ECAPA-TDNN systems, as well as i-vectors, we achieve relative word error rate improvements of up to 16.3% on LibriSpeech and up to 14.5% on Switchboard. We show that the proven method of concatenating speaker vectors to the acoustic features and supplying them as auxiliary model inputs remains a viable option to increase the robustness of end-to-end architectures. The effect on transformer models is stronger, when more noise is added to the input speech. The most substantial benefits for systems based on wav2vec 2.0 are achieved under moderate or no noise conditions. Both x-vectors and ECAPA-TDNN embeddings outperform i-vectors as speaker representations. The optimal embedding size depends on the dataset and also varies with the noise condition.
Dominik Wagner 0002, Ilja Baumann, Sebastian P. Bayerl, Korbinian Riedhammer, Tobias Bocklet
ASRU1
2023 Influence of Utterance and Speaker Characteristics on the Classification of Children with Cleft Lip and Palate
Ilja Baumann, Dominik Wagner 0002, Franziska Braun, Sebastian P. Bayerl, Elmar Nöth, Korbinian Riedhammer, Tobias Bocklet
INTERSPEECH2
2023 A Stutter Seldom Comes Alone - Cross-Corpus Stuttering Detection as a Multi-label Problem
Sebastian P. Bayerl, Dominik Wagner 0002, Ilja Baumann, Florian Hönig, Tobias Bocklet, Elmar Nöth, Korbinian Riedhammer
INTERSPEECH2
2023 Multi-class Detection of Pathological Speech with Latent Features: How does it perform on unseen data?
abstract
The detection of pathologies from speech features is usually defined as a binary classification task with one class representing a specific pathology and the other class representing healthy speech. In this work, we train neural networks, large margin classifiers, and tree boosting machines to distinguish between four pathologies: Parkinson's disease, laryngeal cancer, cleft lip and palate, and oral squamous cell carcinoma. We show that latent representations extracted at different layers of a pre-trained wav2vec 2.0 system can be effectively used to classify these types of pathological voices. We evaluate the robustness of our classifiers by adding room impulse responses to the test data and by applying them to unseen speech corpora. Our approach achieves unweighted average F1-Scores between 74.1% and 97.0%, depending on the model and the noise conditions used. The systems generalize and perform well on unseen data of healthy speakers sampled from a variety of different sources.
Dominik Wagner 0002, Ilja Baumann, Franziska Braun, Sebastian P. Bayerl, Elmar Nöth, Korbinian Riedhammer, Tobias Bocklet
INTERSPEECH1
2022 Nonwords Pronunciation Classification in Language Development Tests for Preschool Children
abstract
This work aims to automatically evaluate whether the language development of children is age-appropriate. Validated speech and language tests are used for this purpose to test the auditory memory. In this work, the task is to determine whether spoken nonwords have been uttered correctly. We compare different approaches that are motivated to model specific language structures: Low-level features (FFT), speaker embeddings (ECAPA-TDNN), grapheme-motivated embeddings (wav2vec 2.0), and phonetic embeddings in form of senones (ASR acoustic model). Each of the approaches provides input for VGG-like 5-layer CNN classifiers. We also examine the adaptation per nonword. The evaluation of the proposed systems was performed using recordings from different kindergartens of spoken nonwords. ECAPA-TDNN and low-level FFT features do not explicitly model phonetic information; wav2vec2.0 is trained on grapheme labels, our ASR acoustic model features contain (sub-)phonetic information. We found that the more granular the phonetic modeling is, the higher are the achieved recognition rates. The best system trained on ASR acoustic model features with VTLN achieved an accuracy of 89.4% and an area under the ROC (Receiver Operating Characteristic) curve (AUC) of 0.923. This corresponds to an improvement in accuracy of 20.2% and AUC of 0.309 relative compared to the FFT-baseline.
Ilja Baumann, Dominik Wagner 0002, Sebastian P. Bayerl, Tobias Bocklet
INTERSPEECH2
2022 Detecting Dysfluencies in Stuttering Therapy Using wav2vec 2.0
abstract
Stuttering is a varied speech disorder that harms an individual's communication ability. Persons who stutter (PWS) often use speech therapy to cope with their condition. Improving speech recognition systems for people with such non-typical speech or tracking the effectiveness of speech therapy would require systems that can detect dysfluencies while at the same time being able to detect speech techniques acquired in therapy. This paper shows that fine-tuning wav2vec 2.0 [1] for the classification of stuttering on a sizeable English corpus containing stuttered speech, in conjunction with multi-task learning, boosts the effectiveness of the general-purpose wav2vec 2.0 features for detecting stuttering in speech; both within and across languages. We evaluate our method on FluencyBank , [2] and the German therapy-centric Kassel State of Fluency (KSoF) [3] dataset by training Support Vector Machine classifiers using features extracted from the finetuned models for six different stuttering-related event types: blocks, prolongations, sound repetitions, word repetitions, interjections, and - specific to therapy - speech modifications. Using embeddings from the fine-tuned models leads to relative classification performance gains up to 27% w.r.t. F1-score.
Sebastian P. Bayerl, Dominik Wagner 0002, Elmar Nöth, Korbinian Riedhammer
INTERSPEECH2
2022 Generative Models for Improved Naturalness, Intelligibility, and Voicing of Whispered Speech
abstract
This work adapts two recent architectures of generative models and evaluates their effectiveness for the conversion of whispered speech to normal speech. We incorporate the normal target speech into the training criterion of vector-quantized variational autoencoders (VQ-VAEs) and Mel-GANs, thereby conditioning the systems to recover voiced speech from whispered inputs. Objective and subjective quality measures indicate that both VQ-VAEs and MelGANs can be modified to perform the conversion task. We find that the proposed approaches significantly improve the Mel cepstral distortion (MCD) metric by at least 25% relative to a Disco-GAN baseline. Subjective listening tests suggest that the MelGAN-based system significantly improves naturalness, intelligibility, and voicing compared to the whispered input speech. A novel evaluation measure based on differences between latent speech representations also indicates that our MelGAN-based approach yields improvements relative to the baseline.
Dominik Wagner 0002, Sebastian P. Bayerl, Héctor A. Cordourier, Tobias Bocklet
SLT1