EDBT 2026 Demo / reviewers in the wild / expert
Ilja Baumann
dblp:318/1385
· DBLP profile ↗
20ranked-venue papers
9as first author
20since 2021 · last 2025
0000-0002-1991-3710ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 19 · 9 first-author · 19 since 2021Artificial intelligence and machine learning · 16 · 7 first-author · 16 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Text-Guided Speech Representations for Language Acquisition AssessmentabstractAssessing spoken language abilities in children is critical for early detection of developmental delays. We propose a two-stage framework for automatic pronunciation assessment, targeting classification of items in a language development screening. In the first stage, we fine-tune a sentence transformer using phonetic and label-informed losses to create a structured latent text space. In the second stage, we fine-tune an audio encoder to project spoken utterances into the text space. Our approach enables inference from audio alone while preserving alignment with phonetically meaningful structure. Compared to baseline systems, our method achieves the best classification performance and provides a latent space suitable for downstream analysis. This facilitates the identification of phonological or grammatical error patterns from cluster structure, paving the way for potentially interpretable assessment of early language acquisition. Ilja Baumann, Dominik Wagner 0002, Philipp Seeberger, Korbinian Riedhammer, Tobias Bocklet |
ASRU | 1 |
| 2025 | Joint ASR and Speech Attribute Prediction for Conversational Dysarthric Speech Analysis with Multimodal Language ModelsabstractDysarthric speech recognition systems often focus solely on transcription, limiting their applicability in clinical and assistive settings where assessments of perceptual attributes like intelligibility and naturalness are essential. We propose a multimodal conversational framework based on Phi-4-Multimodal that combines ASR with attribute rating prediction, enabling users to query both transcriptions and perceptual characteristics (e.g. “How intelligible is this utterance?”). Our multitask model adds auxiliary prediction heads for five clinically relevant attributes and is trained on the English Speech Accessibility Project dataset. The system achieves competitive ASR performance while delivering attribute-level feedback comparable to specialized classifiers. Additional experiments show improved ASR performance for German Parkinson’s speech, indicating preserved multilingual capabilities and partial cross-lingual transfer of dysarthric speech patterns. Dominik Wagner 0002, Ilja Baumann, Natalie Engert, Elmar Nöth, Korbinian Riedhammer, Tobias Bocklet |
ASRU | 2 |
| 2025 | Optimized Self-supervised Training with BEST-RQ for Speech RecognitionabstractSelf-supervised learning has been successfully used for various speech related tasks, including automatic speech recognition. BERT-based Speech pre-Training with Random-projection Quantizer (BEST-RQ) has achieved state-of-the-art results in speech recognition. In this work, we further optimize the BEST-RQ approach using Kullback-Leibler divergence as an additional regularizing loss and multi-codebook extension per cluster derived from low-level feature clustering. Preliminary experiments on train-100 split of LibriSpeech result in a relative improvement of 11.2% on test-clean by using multiple codebooks, utilizing a combination of cross-entropy and Kullback-Leibler divergence further reduces the word error rate by 4.5%. The proposed optimizations on full LibriSpeech pre-training and fine-tuning result in relative word error rate improvements of up to 23.8% on test-clean and 30.6% on test-other using 6 codebooks. Furthermore, the proposed setup leads to faster convergence in pre-training and fine-tuning and additionally stabilizes the pre-training. Ilja Baumann, Dominik Wagner 0002, Korbinian Riedhammer, Tobias Bocklet |
ICASSP | 1 |
| 2025 | Digital Operating Mode Classification of Real-World Amateur Radio TransmissionsabstractThis study presents an ML approach for classifying digital radio operating modes evaluated on real-world transmissions. We generated 98 different parameterized radio signals from 17 digital operating modes, transmitted each of them on the 70 cm (UHF) amateur radio band, and recorded our transmissions with two different architectures of SDR receivers. Three lightweight ML models were trained exclusively on spectrograms of limited non-transmitted signals with random characters as payloads. This training involved an online data augmentation pipeline to simulate various radio channel impairments. Our best model, EfficientNetB0, achieved an accuracy of 93.80% across the 17 operating modes and 85.47% across all 98 parameterized radio signals, evaluated on our real-world transmissions with Wikipedia articles as payloads. Furthermore, we analyzed the impact of varying signal durations & the number of FFT bins on classification, assessed the effectiveness of our simulated channel impairments, and tested our models across multiple simulated SNRs. Maximilian Bundscherer, Thomas H. Schmitt, Ilja Baumann, Tobias Bocklet |
ICASSP | 3 |
| 2025 | Personalized Fine-Tuning with Controllable Synthetic Speech from LLM-Generated Transcripts for Dysarthric Speech RecognitionabstractIn this work, we present our submission to the Speech Accessibility Project challenge for dysarthric speech recognition. We integrate parameter-efficient fine-tuning with latent audio representations to improve an encoder-decoder ASR system. Synthetic training data is generated by fine-tuning Parler-TTS to mimic dysarthric speech, using LLM-generated prompts for corpus-consistent target transcripts. Personalization with x-vectors consistently reduces word error rates (WERs) over non-personalized fine-tuning. AdaLoRA adapters outperform full fine-tuning and standard low-rank adaptation, achieving relative WER reductions of ∼23% and ∼22%, respectively. Further improvements (∼5% WER reduction) come from incorporating wav2vec 2.0-based audio representations. Training with synthetic dysarthric speech yields up to ∼7% relative WER improvement over personalized fine-tuning alone. Dominik Wagner 0002, Ilja Baumann, Natalie Engert, Seanie Lee, Elmar Nöth, Korbinian Riedhammer, Tobias Bocklet |
INTERSPEECH | 2 |
| 2025 | Pathology-Aware Speech Encoding and Data Augmentation for Dysarthric Speech RecognitionabstractAutomatic speech recognition (ASR) for pathologic speech remains a major challenge due to high variability in articulation, phonation, and prosody distortions. In this work, we propose a pathology-aware speech encoder based on BEST-RQ pre-training, which incorporates 46k hours of speech, including pathologic and atypical speech. We continue pre-training for domain adaptation and experiment with etiology-specific codebooks. We achieve a 13.2% relative word error rate (WER) improvement using the pathology-aware speech encoder with etiology-specific continued pre-training. Additionally, we examine the impact of incorporating synthetic and out-of-domain (OOD) data to further enhance ASR performance. Synthetic data reduces WER by up to 8.7%, while OOD data improves WER by 12.2%. Finally, we introduce a semantic similaritybased data augmentation technique to optimize data selection, achieving a WER improvement of up to 9.7% while minimizing the need for additional training data. Ilja Baumann, Dominik Wagner 0002, Korbinian Riedhammer, Tobias Bocklet |
INTERSPEECH | 1 |
| 2024 | Optimized Speculative Sampling for GPU Hardware AcceleratorsabstractIn this work, we optimize speculative sampling for parallel hardware accelerators to improve sampling speed.We notice that substantial portions of the intermediate matrices necessary for speculative sampling can be computed concurrently.This allows us to distribute the workload across multiple GPU threads, enabling simultaneous operations on matrix segments within thread blocks.This results in profiling time improvements ranging from 6% to 13% relative to the baseline implementation, without compromising accuracy.To further accelerate speculative sampling, probability distributions parameterized by softmax are approximated by sigmoid.This approximation approach results in significantly greater relative improvements in profiling time, ranging from 37% to 94%, with a minor decline in accuracy.We conduct extensive experiments on both automatic speech recognition and summarization tasks to validate the effectiveness of our optimization methods. Dominik Wagner 0002, Seanie Lee, Ilja Baumann, Philipp Seeberger, Korbinian Riedhammer, Tobias Bocklet |
EMNLP | 3 |
| 2024 | Towards Interpretability of Automatic Phoneme Analysis in Cleft Lip and Palate SpeechabstractCleft Lip and Palate ranks among the most common congenital abnormalities and significantly influences speech articulation, resulting in varying phonemic impacts. In a clinical context, a detailed diagnosis is carried out by time-consuming perceptual evaluations. We use perceptual ratings of different articulatory modifications on phoneme-level as ground-truth and propose a system based on wav2vec 2.0, trained to the downstream task of classifying phonemic criteria as a multi-class and multi-label problem. The system is trained for detection on utterance level, without the usage of phoneme labels. To gain a clearer understanding of which areas of the speech signal have the greatest impact on classification, we assess the extent to which our system aligns with expert ratings at the phoneme level. Additionally, we examine which specific phonemes play a decisive role in determining the final classification of the labeled criteria. The results show that salient phonemes marked by experts contribute remarkably greater to the classification of the correct class using feature relevance explanation methods. To the best of our knowledge, this is the first study incorporating various utterance-level articulatory modifications classification and phoneme-level interpretation, offering a more comprehensive understanding for potential clinical applications. Ilja Baumann, Dominik Wagner 0002, Maria Schuster, Elmar Nöth, Tobias Bocklet |
ICASSP | 1 |
| 2024 | Large Language Models for Dysfluency Detection in Stuttered Speech
Dominik Wagner 0002, Sebastian P. Bayerl, Ilja Baumann, Elmar Nöth, Korbinian Riedhammer, Tobias Bocklet |
INTERSPEECH | 3 |
| 2024 | Outlier Reduction with Gated Attention for Improved Post-training Quantization in Large Sequence-to-sequence Speech Foundation ModelsabstractThis paper explores the improvement of post-training quantization (PTQ) after knowledge distillation in the Whisper speech foundation model family. We address the challenge of outliers in weights and activation tensors, known to impede quantization quality in transformer-based language and vision models. Extending this observation to Whisper, we demonstrate that these outliers are also present when transformer-based models are trained to perform automatic speech recognition, necessitating mitigation strategies for PTQ. We show that outliers can be reduced by a recently proposed gating mechanism in the attention blocks of the student model, enabling effective 8-bit quantization, and lower word error rates compared to student models without the gating mechanism in place. Dominik Wagner 0002, Ilja Baumann, Korbinian Riedhammer, Tobias Bocklet |
INTERSPEECH | 2 |
| 2024 | Towards Self-Attention Understanding for Automatic Articulatory Processes Analysis in Cleft Lip and Palate SpeechabstractCleft lip and palate (CLP) speech presents unique challenges for automatic phoneme analysis due to its distinct acoustic characteristics and articulatory anomalies. We perform phoneme analysis in CLP speech using a pre-trained wav2vec 2.0 model with a multi-head self-attention classification module to capture long-range dependencies within the speech signal, thereby enabling better contextual understanding of phoneme sequences. We demonstrate the effectiveness of our approach in the classification of various articulatory processes in CLP speech. Furthermore, we investigate the interpretability of self-attention to gain insights into the model’s understanding of CLP speech characteristics. Our findings highlight the potential of the selfattention mechanisms for improving automatic phoneme analysis in CLP speech, paving the way for enhanced diagnostics, adding interpretability for therapists and affected patients. Ilja Baumann, Dominik Wagner 0002, Maria Schuster, Korbinian Riedhammer, Elmar Nöth, Tobias Bocklet |
INTERSPEECH | 1 |
| 2024 | Automatic Evaluation of a Sentence Memory Test for Preschool ChildrenabstractAssessment of memory capabilities in preschool-aged children is crucial for early detection of potential speech development impairments or delays. We present an approach for the automatic evaluation of a standardized sentence memory test specifically for preschool children. Our methodology leverages automatic transcription of recited sentences and evaluation based on natural language processing techniques. We demonstrate the effectiveness of our approach on a dataset comprised of recited sentences from preschool-aged children, incorporating ratings of semantic and syntactic correctness. The best performing systems achieve an F1 score of 91.7% for semantic correctness and 86.1% for syntactic correctness using automatic transcripts. Our results showcase the potential of automated evaluation systems in providing reliable and efficient assessments of memory capabilities in early childhood, facilitating timely interventions and support for children with language development needs. Ilja Baumann, Nicole Unger, Dominik Wagner 0002, Korbinian Riedhammer, Tobias Bocklet |
INTERSPEECH | 1 |
| 2024 | It's Time to Take Action: Acoustic Modeling of Motor Verbs to Detect Parkinson's DiseaseabstractPre-trained models generate speech representations that are used in different tasks, including the automatic detection of Parkinson’s disease (PD). Although these models can yield high accuracy, their interpretation is still challenging. This paper used a pre-trained Wav2vec 2.0 model to represent speech frames of 25ms length and perform a frame-by-frame discrimination between PD patients and healthy control (HC) subjects. This fine granularity prediction enabled us to identify specific linguistic segments with high discrimination capability. Speech representations of all produced verbs were compared w.r.t. nouns and the first ones yielded higher accuracies. To gaina deeper understanding of this pattern, representations of motor and non-motor verbs were compared and the first ones yielded better results, with accuracies of around 83% in an independent test set. These findings support well-established neurocognitive models about action-related language highlighted as key drivers of PD. Index Terms: computational paralinguistics, interpretability of pre-trained models, action verbs, Parkinson’s disease Daniel Escobar-Grisales, Cristian D. Ríos-Urrego, Ilja Baumann, Korbinian Riedhammer, Elmar Nöth, Tobias Bocklet, Adolfo M. García, Juan Rafael Orozco-Arroyave |
INTERSPEECH | 3 |
| 2024 | Personalizing Large Sequence-to-Sequence Speech Foundation Models With Speaker RepresentationsabstractWe present a method to personalize large transformer-based encoder-decoder speech foundation models without the need for changes in the underlying model structure or training from scratch. This is achieved by projecting speaker-specific information into the latent space of the transformer decoder via a small neural network and learning to process the speaker information along with domain-specific information via parameter-efficient finetuning. We use this method to improve the automatic speech recognition results of spoken academic German and English. Our approach yields average relative word error rate (WER) improvements of approximately 29% on German academic speech and 25% on English academic speech. It also translates well to conversational speech, achieving relative WER improvements of up to 36%, and demonstrates modest gains of up to 5% on read speech. Moreover, we observe that incorporating utterances from the recent past as personalization context yields the most significant overall improvements and that changes in voice characteristics resulting from prolonged speaking have a minimal effect on the personalization quality of academic lectures. Dominik Wagner 0002, Ilja Baumann, Thomas Ranzenberger, Korbinian Riedhammer, Tobias Bocklet |
SLT | 2 |
| 2023 | Detection of Vowel Errors in Children's Speech using Synthetic Phonetic TranscriptsabstractThe analysis of phonological processes is crucial in evaluating speech development disorders in children, but encounters challenges due to limited children audio data. This work focuses on automatic vowel error detection using a two-stage pipeline. The first stage uses a fine-tuned cross-lingual phone recognizer (wav2vec 2.0) to extract phone sequences from audio. The second stage employs a language model (BERT) for classification from a phone sequence, entirely trained on synthetic transcripts, to counteract the very broad range of potential mistakes. We evaluate the system on nonword audio recordings recited by preschool children from a speech development test. The results show that the classifier trained on synthetic data performs well, but its efficacy relies on the quality of the phone recognizer. The best classifier achieves an 94.7% F1 score when evaluated against phonetic ground truths, whereas the F1 score is 76.2% when using automatically recognized phone sequences. Ilja Baumann, Dominik Wagner 0002, Korbinian Riedhammer, Elmar Nöth, Tobias Bocklet |
ASRU | 1 |
| 2023 | Speaker Adaptation for End-to-End Speech Recognition Systems in Noisy EnvironmentsabstractWe analyze the impact of speaker adaptation in end-to-end automatic speech recognition models based on transformers and wav2vec 2.0 under different noise conditions. By including speaker embeddings obtained from x-vector and ECAPA-TDNN systems, as well as i-vectors, we achieve relative word error rate improvements of up to 16.3% on LibriSpeech and up to 14.5% on Switchboard. We show that the proven method of concatenating speaker vectors to the acoustic features and supplying them as auxiliary model inputs remains a viable option to increase the robustness of end-to-end architectures. The effect on transformer models is stronger, when more noise is added to the input speech. The most substantial benefits for systems based on wav2vec 2.0 are achieved under moderate or no noise conditions. Both x-vectors and ECAPA-TDNN embeddings outperform i-vectors as speaker representations. The optimal embedding size depends on the dataset and also varies with the noise condition. Dominik Wagner 0002, Ilja Baumann, Sebastian P. Bayerl, Korbinian Riedhammer, Tobias Bocklet |
ASRU | 2 |
| 2023 | Influence of Utterance and Speaker Characteristics on the Classification of Children with Cleft Lip and Palate
Ilja Baumann, Dominik Wagner 0002, Franziska Braun, Sebastian P. Bayerl, Elmar Nöth, Korbinian Riedhammer, Tobias Bocklet |
INTERSPEECH | 1 |
| 2023 | A Stutter Seldom Comes Alone - Cross-Corpus Stuttering Detection as a Multi-label Problem
Sebastian P. Bayerl, Dominik Wagner 0002, Ilja Baumann, Florian Hönig, Tobias Bocklet, Elmar Nöth, Korbinian Riedhammer |
INTERSPEECH | 3 |
| 2023 | Multi-class Detection of Pathological Speech with Latent Features: How does it perform on unseen data?abstractThe detection of pathologies from speech features is usually defined as a binary classification task with one class representing a specific pathology and the other class representing healthy speech. In this work, we train neural networks, large margin classifiers, and tree boosting machines to distinguish between four pathologies: Parkinson's disease, laryngeal cancer, cleft lip and palate, and oral squamous cell carcinoma. We show that latent representations extracted at different layers of a pre-trained wav2vec 2.0 system can be effectively used to classify these types of pathological voices. We evaluate the robustness of our classifiers by adding room impulse responses to the test data and by applying them to unseen speech corpora. Our approach achieves unweighted average F1-Scores between 74.1% and 97.0%, depending on the model and the noise conditions used. The systems generalize and perform well on unseen data of healthy speakers sampled from a variety of different sources. Dominik Wagner 0002, Ilja Baumann, Franziska Braun, Sebastian P. Bayerl, Elmar Nöth, Korbinian Riedhammer, Tobias Bocklet |
INTERSPEECH | 2 |
| 2022 | Nonwords Pronunciation Classification in Language Development Tests for Preschool ChildrenabstractThis work aims to automatically evaluate whether the language development of children is age-appropriate. Validated speech and language tests are used for this purpose to test the auditory memory. In this work, the task is to determine whether spoken nonwords have been uttered correctly. We compare different approaches that are motivated to model specific language structures: Low-level features (FFT), speaker embeddings (ECAPA-TDNN), grapheme-motivated embeddings (wav2vec 2.0), and phonetic embeddings in form of senones (ASR acoustic model). Each of the approaches provides input for VGG-like 5-layer CNN classifiers. We also examine the adaptation per nonword. The evaluation of the proposed systems was performed using recordings from different kindergartens of spoken nonwords. ECAPA-TDNN and low-level FFT features do not explicitly model phonetic information; wav2vec2.0 is trained on grapheme labels, our ASR acoustic model features contain (sub-)phonetic information. We found that the more granular the phonetic modeling is, the higher are the achieved recognition rates. The best system trained on ASR acoustic model features with VTLN achieved an accuracy of 89.4% and an area under the ROC (Receiver Operating Characteristic) curve (AUC) of 0.923. This corresponds to an improvement in accuracy of 20.2% and AUC of 0.309 relative compared to the FFT-baseline. Ilja Baumann, Dominik Wagner 0002, Sebastian P. Bayerl, Tobias Bocklet |
INTERSPEECH | 1 |