Simon Berger

dblp:136/0623 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
7since 2021 · last 2026
0000-0003-1333-4029ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 4 since 2021
YearPublicationVenuePosition
2026 Supplementary Resources and Analysis for Automatic Speech Recognition Systems Trained on the Loquacious Dataset
Nick Rossenbach, Robin Schmitt, Tina Raissi, Simon Berger, Larissa Kleppel, Ralf Schlüter
LREC4
2024 Investigating the Effect of Label Topology and Training Criterion on ASR Performance and Alignment Quality
Tina Raissi, Christoph Lüscher, Simon Berger, Ralf Schlüter, Hermann Ney
INTERSPEECH3
2024 Combining TF-GridNet And Mixture Encoder For Continuous Speech Separation For Meeting Transcription
abstract
Many real-life applications of automatic speech recognition (ASR) require processing of overlapped speech. A common method involves first separating the speech into overlap-free streams on which ASR is performed. Recently, TF-GridNet has shown impressive performance in speech separation in real reverberant conditions. Furthermore, a mixture encoder was proposed that leverages the mixed speech to mitigate the effect of separation artifacts. In this work, we extended the mixture encoder from a static two-speaker scenario to a natural meeting context featuring an arbitrary number of speakers and varying degrees of overlap. We further demonstrate its limits by the integration with separators of varying strength including TF-GridNet. Our experiments result in a new state-of-the-art performance on LibriCSS using a single microphone. They show that TF-GridNet largely closes the gap between previous methods and oracle separation independent of mixture encoding. We further investigate the remaining potential for improvement.
Peter Vieting, Simon Berger, Thilo von Neumann, Christoph Böddeker, Ralf Schlüter, Reinhold Häb-Umbach
SLT2
2023 RASR2: The RWTH ASR Toolkit for Generic Sequence-to-sequence Speech Recognition
abstract
Modern public ASR tools usually provide rich support for training various sequence-to-sequence (S2S) models, but rather simple support for decoding open-vocabulary scenarios only.For closed-vocabulary scenarios, public tools supporting lexicalconstrained decoding are usually only for classical ASR, or do not support all S2S models.To eliminate this restriction on research possibilities such as modeling unit choice, we present RASR2 in this work, a research-oriented generic S2S decoder implemented in C++.It offers a strong flexibility/compatibility for various S2S models, language models, label units/topologies and neural network architectures.It provides efficient decoding for both open-and closed-vocabulary scenarios based on a generalized search framework with rich support for different search modes and settings.We evaluate RASR2 with a wide range of experiments on both switchboard and Librispeech corpora.Our source code is public online.
Wei Zhou 0043, Eugen Beck, Simon Berger, Ralf Schlüter, Hermann Ney
INTERSPEECH3
2023 Mixture Encoder for Joint Speech Separation and Recognition
Simon Berger, Peter Vieting, Christoph Böddeker, Ralf Schlüter, Reinhold Häb-Umbach
INTERSPEECH1
2022 HMM vs. CTC for Automatic Speech Recognition: Comparison Based on Full-Sum Training from Scratch
abstract
In this work, we compare from-scratch sequence-level cross-entropy (full-sum) training of Hidden Markov Model (HMM) and Connectionist Temporal Classification (CTC) topologies for automatic speech recognition (ASR). Besides accuracy, we further analyze their capability for generating high-quality time alignment between the speech signal and the transcription, which can be crucial for many subsequent applications. Moreover, we propose several methods to improve convergence of from-scratch full-sum training by addressing the alignment modeling issue. Systematic comparison is conducted on both Switchboard and LibriSpeech corpora across CTC, posterior HMM with and w/o transition probabilities, and standard hybrid HMM. We also provide a detailed analysis of both Viterbi forced-alignment and Baum-Welch full-sum occupation probabilities.
Tina Raissi, Wei Zhou 0043, Simon Berger, Ralf Schlüter, Hermann Ney
SLT3
2021 Phoneme Based Neural Transducer for Large Vocabulary Speech Recognition
abstract
To join the advantages of classical and end-to-end approaches for speech recognition, we present a simple, novel and competitive approach for phoneme-based neural transducer modeling. Different alignment label topologies are compared and word-end-based phoneme label augmentation is proposed to improve performance. Utilizing the local dependency of phonemes, we adopt a simplified neural network structure and a straightforward integration with the external word-level language model to preserve the consistency of seq-to-seq modeling. We also present a simple, stable and efficient training procedure using frame-wise cross-entropy loss. A phonetic context size of one is shown to be sufficient for the best performance. A simplified scheduled sampling approach is applied for further improvement and different decoding approaches are briefly compared. The overall performance of our best model is comparable to state-of-the-art (SOTA) results for the TED-LIUM Release 2 and Switchboard corpora.
Wei Zhou 0043, Simon Berger, Ralf Schlüter, Hermann Ney
ICASSP2
2013 Cognitive Parameter Adaption in Regular Control Structures - Using Process Knowledge for Parameter Adaption
Simon Berger, Gunther Reinhart
ICINCO (1)2