VLDB 2026 Research / reviewers in the wild / expert
Nick Rossenbach
dblp:205/9028
· DBLP profile ↗
9ranked-venue papers
5as first author
8since 2021 · last 2026
0009-0007-0704-6625ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 4 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Supplementary Resources and Analysis for Automatic Speech Recognition Systems Trained on the Loquacious Dataset
Nick Rossenbach, Robin Schmitt, Tina Raissi, Simon Berger, Larissa Kleppel, Ralf Schlüter |
LREC | 1 |
| 2025 | Analysis of Domain Shift across ASR Architectures via TTS-Enabled Separation of Target Domain and Acoustic ConditionsabstractWe analyze automatic speech recognition (ASR) modeling choices under domain mismatch, comparing classic modular and novel sequence-to-sequence (seq2seq) architectures. Across the different ASR architectures, we examine a spectrum of modeling choices, including label units, context length, and topology. To isolate language domain effects from acoustic variation, we synthesize target domain audio using a text-to-speech system trained on LibriSpeech. We incorporate target domain ngram and neural language models for domain adaptation without retraining the acoustic model. To our knowledge, this is the first controlled comparison of optimized ASR systems across state-of-the-art architectures under domain shift, offering insights into their generalization. The results show that, under domain shift, rather than the decoder architecture choice or the distinction between classic modular and novel seq2seq models, it is specific modeling choices that influence performance. Tina Raissi, Nick Rossenbach, Ralf Schlüter |
ASRU | 2 |
| 2025 | Analyzing the Importance of Blank for CTC-Based Knowledge Distillation
Benedikt Hilmes, Nick Rossenbach, Ralf Schlüter |
INTERSPEECH | 2 |
| 2025 | Running Conventional Automatic Speech Recognition on Memristor Hardware: A Simulated Approach
Nick Rossenbach, Benedikt Hilmes, Leon Brackmann, Moritz Gunz, Ralf Schlüter |
INTERSPEECH | 1 |
| 2023 | On the Relevance of Phoneme Duration Variability of Synthesized Training Data for Automatic Speech RecognitionabstractSynthetic data generated by text-to-speech (TTS) systems can be used to improve automatic speech recognition (ASR) systems in low-resource or domain mismatch tasks. It has been shown that TTS-generated outputs still do not have the same qualities as real data. In this work we focus on the temporal structure of synthetic data and its relation to ASR training. By using a novel oracle setup we show how much the degradation of synthetic data quality is influenced by duration modeling in non-autoregressive (NAR) TTS. To get reference phoneme durations we use two common alignment methods, a hidden Markov Gaussian-mixture model (HMM-GMM) aligner and a neural connectionist temporal classification (CTC) aligner. Using a simple algorithm based on random walks we shift phoneme duration distributions of the TTS system closer to real durations, resulting in an improvement of an ASR system using synthetic data in a semi-supervised setting. Nick Rossenbach, Benedikt Hilmes, Ralf Schlüter |
ASRU | 1 |
| 2023 | Take the Hint: Improving Arabic Diacritization with Partially-Diacritized Text
Parnia Bahar, Mattia Antonino Di Gangi, Nick Rossenbach, Mohammad Zeineldeen |
INTERSPEECH | 3 |
| 2022 | Automatic Video Dubbing at AppTekabstractVideo dubbing is the activity of revoicing a video while offering a viewing experience equivalent to the original video. The revoicing usually comes with a changed script, mostly in a different language, and the revoicing should reproduce the original emotions, coherent with the body language, and lip synchronized. In this project, we aim to build an AD system in three phases: (1) voice-over; (2) emotional voice-over; (3) full dubbing, while enhancing the system with human-in-the-loop capabilities for a higher quality. Mattia Antonino Di Gangi, Nick Rossenbach, Parnia Bahar, Eugen Beck, Patrick Wilken, Evgeny Matusov |
EAMT | 2 |
| 2021 | Comparing the Benefit of Synthetic Training Data for Various Automatic Speech Recognition ArchitecturesabstractRecent publications on automatic-speech-recognition (ASR) have a strong focus on attention encoder-decoder (AED) architectures which tend to suffer from over-fitting in low resource scenarios. One solution to tackle this issue is to generate synthetic data with a trained text-to-speech system (TTS) if additional text is available. This was successfully applied in many publications with AED systems, but only very limited in the context of other ASR architectures. We investigate the effect of varying pre-processing, the speaker embedding and input encoding of the TTS system w.r.t. the effectiveness of the synthesized data for AED-ASR training. Additionally, we also consider internal language model subtraction for the first time, resulting in up to 38% relative improvement. We compare the AED results to a state-of-the-art hybrid ASR system, a monophone based system using connectionist-temporal-classification (CTC) and a monotonic transducer based system. We show that for the later systems the addition of synthetic data has no relevant effect, but they still outperform the AED systems on LibriSpeech-100h. We achieve a final word-error-rate of 3.3%/10.0% with a hybrid system on the clean/noisy test-sets, surpassing any previous state-of-the-art systems on Librispeech-100h that do not include unlabeled audio data. Nick Rossenbach, Mohammad Zeineldeen, Benedikt Hilmes, Ralf Schlüter, Hermann Ney |
ASRU | 1 |
| 2020 | Generating Synthetic Audio Data for Attention-Based Speech Recognition SystemsabstractRecent advances in text-to-speech (TTS) led to the development of flexible multi-speaker end-to-end TTS systems. We extend state-of-the-art attention-based automatic speech recognition (ASR) systems with synthetic audio generated by a TTS system trained only on the ASR corpora itself. ASR and TTS systems are built separately to show that text-only data can be used to enhance existing end-to-end ASR systems without the necessity of parameter or architecture changes. We compare our method with language model integration of the same text data and with simple data augmentation methods like SpecAugment and show that performance improvements are mostly independent. We achieve improvements of up to 33% relative in word-error-rate (WER) over a strong baseline with data-augmentation in a low-resource environment (LibriSpeech-100h), closing the gap to a comparable oracle experiment by more than 50%. We also show improvements of up to 5% relative WER over our most recent ASR baseline on LibriSpeech-960h. Nick Rossenbach, Albert Zeyer, Ralf Schlüter, Hermann Ney |
ICASSP | 1 |