VLDB 2026 Research / reviewers in the wild / expert
Eugen Beck
dblp:140/2765
· DBLP profile ↗
17ranked-venue papers
3as first author
6since 2021 · last 2025
0000-0003-1641-6628ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 13 · 3 first-author · 4 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Efficient Supernet Training with Orthogonal Softmax for Scalable ASR Model CompressionabstractASR systems are deployed across diverse environments, each with specific hardware constraints. We use supernet training to jointly train multiple encoders of varying sizes, enabling dynamic model size adjustment to fit hardware constraints without redundant training. Moreover, we introduce a novel method called OrthoSoftmax, which applies multiple orthogonal softmax functions to efficiently identify optimal subnets within the supernet, avoiding resource-intensive search. This approach also enables more flexible and precise subnet selection by allowing selection based on various criteria and levels of granularity. Our results with CTC on Librispeech and TED-LIUM-v2 show that FLOPs-aware component-wise selection achieves the best overall performance. With the same number of training updates from one single job, WERs for all model sizes are comparable to or slightly better than those of individually trained models. Furthermore, we analyze patterns in the selected components and reveal interesting insights. Jingjing Xu 0002, Eugen Beck, Ralf Schlüter |
ICASSP | 2 |
| 2025 | Dynamic Acoustic Model Architecture Optimization in Training for ASR
Jingjing Xu 0002, Albert Zeyer, Eugen Beck, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 4 |
| 2024 | Dynamic Encoder Size Based on Data-Driven Layer-wise Pruning for Speech Recognition
Jingjing Xu 0002, Wei Zhou 0043, Eugen Beck, Ralf Schlüter |
INTERSPEECH | 4 |
| 2023 | RASR2: The RWTH ASR Toolkit for Generic Sequence-to-sequence Speech RecognitionabstractModern public ASR tools usually provide rich support for training various sequence-to-sequence (S2S) models, but rather simple support for decoding open-vocabulary scenarios only.For closed-vocabulary scenarios, public tools supporting lexicalconstrained decoding are usually only for classical ASR, or do not support all S2S models.To eliminate this restriction on research possibilities such as modeling unit choice, we present RASR2 in this work, a research-oriented generic S2S decoder implemented in C++.It offers a strong flexibility/compatibility for various S2S models, language models, label units/topologies and neural network architectures.It provides efficient decoding for both open-and closed-vocabulary scenarios based on a generalized search framework with rich support for different search modes and settings.We evaluate RASR2 with a wide range of experiments on both switchboard and Librispeech corpora.Our source code is public online. Wei Zhou 0043, Eugen Beck, Simon Berger, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 2 |
| 2022 | Automatic Video Dubbing at AppTekabstractVideo dubbing is the activity of revoicing a video while offering a viewing experience equivalent to the original video. The revoicing usually comes with a changed script, mostly in a different language, and the revoicing should reproduce the original emotions, coherent with the body language, and lip synchronized. In this project, we aim to build an AD system in three phases: (1) voice-over; (2) emotional voice-over; (3) full dubbing, while enhancing the system with human-in-the-loop capabilities for a higher quality. Mattia Antonino Di Gangi, Nick Rossenbach, Parnia Bahar, Eugen Beck, Patrick Wilken, Evgeny Matusov |
EAMT | 5 |
| 2022 | Improving Factored Hybrid HMM Acoustic Modeling without State TyingabstractIn this work, we show that a factored hybrid hidden Markov model (FH-HMM) which is defined without any phonetic state-tying outperforms a state-of-the-art hybrid HMM. The factored hybrid HMM provides a link to transducer models in the way it models phonetic (label) context while preserving the strict separation of acoustic and language model of the hybrid HMM approach. Furthermore, we show that the factored hybrid model can be trained from scratch without using phonetic state-tying in any of the training steps. Our modeling approach enables triphone context while avoiding phonetic state-tying by a decomposition into locally normalized factored posteriors for monophones/HMM states in phoneme context. Experimental results are provided for Switchboard 300h and LibriSpeech. On the former task we also show that by avoiding the phonetic state-tying step, the factored hybrid can take better advantage of regularization techniques during training, compared to the standard hybrid HMM with phonetic state-tying based on classification and regression trees (CART). Tina Raissi, Eugen Beck, Ralf Schlüter, Hermann Ney |
ICASSP | 2 |
| 2020 | Domain Robust, Fast, and Compact Neural Language ModelsabstractDespite advances in neural language modeling, obtaining a good model on a large scale multi-domain dataset still remains a difficult task. We propose training methods for building neural language models for such a task, which are not only domain robust, but reasonable in model size and fast for evaluation. We combine knowledge distillation from pretrained domain expert language models with the noise contrastive estimation (NCE) loss. Knowledge distillation allows to train a single student model which is both compact and domain robust, while the use of NCE loss makes the model self-normalized, which enables fast evaluation. We conduct experiments on a large English multi-domain speech recognition dataset provided by AppTek. The resulting student model is of the size of one domain expert, while it gives similar perplexities as various teacher models on their expert domain; the model is self-normalized, allowing for 30% faster first pass decoding than the naive models which require the full soft- max computation, and finally it gives improvements of more than 8% relative in terms of word error rate over a large multidomain 4-gram count model trained on more than 10 B words. Alexander Gerstenberger, Kazuki Irie, Pavel Golik, Eugen Beck, Hermann Ney |
ICASSP | 4 |
| 2020 | LVCSR with Transformer Language Models
Eugen Beck, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 1 |
| 2020 | Context-Dependent Acoustic Modeling Without Explicit Phone ClusteringabstractPhoneme-based acoustic modeling of large vocabulary automatic speech recognition takes advantage of phoneme context. The large number of context-dependent (CD) phonemes and their highly varying statistics require tying or smoothing to enable robust training. Usually, classification and regression trees are used for phonetic clustering, which is standard in hidden Markov model (HMM)-based systems. However, this solution introduces a secondary training objective and does not allow for end-to-end training. In this work, we address a direct phonetic context modeling for the hybrid deep neural network (DNN)/HMM, that does not build on any phone clustering algorithm for the determination of the HMM state inventory. By performing different decompositions of the joint probability of the center phoneme state and its left and right contexts, we obtain a factorized network consisting of different components, trained jointly. Moreover, the representation of the phonetic context for the network relies on phoneme embeddings. The recognition accuracy of our proposed models on the Switchboard task is comparable and outperforms slightly the hybrid model using the standard state-tying decision trees. Tina Raissi, Eugen Beck, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 2 |
| 2019 | RWTH ASR Systems for LibriSpeech: Hybrid vs AttentionabstractWe present state-of-the-art automatic speech recognition (ASR) systems employing a standard hybrid DNN/HMM architecture compared to an attention-based encoder-decoder design for the LibriSpeech task. Detailed descriptions of the system development, including model design, pretraining schemes, training schedules, and optimization approaches are provided for both system architectures. Both hybrid DNN/HMM and attention-based systems employ bi-directional LSTMs for acoustic modeling/encoding. For language modeling, we employ both LSTM and Transformer based architectures. All our systems are built using RWTHs open-source toolkits RASR and RETURNN. To the best knowledge of the authors, the results obtained when training on the full LibriSpeech training set, are the best published currently, both for the hybrid DNN/HMM and the attention-based systems. Our single hybrid system even outperforms previous results obtained from combining eight single systems. Our comparison shows that on the LibriSpeech 960h task, the hybrid DNN/HMM system outperforms the attention-based system by 15% relative on the clean and 40% relative on the other test sets in terms of word error rate. Moreover, experiments on a reduced 100h-subset of the LibriSpeech training corpus even show a more pronounced margin between the hybrid DNN/HMM and attention-based architectures. Christoph Lüscher, Eugen Beck, Kazuki Irie, Markus Kitza, Wilfried Michel, Albert Zeyer, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 2 |
| 2019 | Rescoring Keyword Search Confidence Estimates with Graph-Based Re-Ranking Using Acoustic Word Embeddings
Anna Piunova, Eugen Beck, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 2 |
| 2019 | Upper and Lower Tight Error Bounds for Feature Omission with an Extension to Context ReductionabstractIn this work, fundamental analytic results in the form of error bounds are presented that quantify the effect of feature omission and selection for pattern classification in general, as well as the effect of context reduction in string classification, like automatic speech recognition, printed/handwritten character recognition, or statistical machine translation. A general simulation framework is introduced that supports discovery and proof of error bounds, which lead to the error bounds presented here. Initially derived tight lower and upper bounds for feature omission are generalized to feature selection, followed by another extension to context reduction of string class priors (aka language models) in string classification. For string classification, the quantitative effect of string class prior context reduction on symbol-level Bayes error is presented. The tightness of the original feature omission bounds seems lost in this case, as further simulations indicate. However, combining both feature omission andcontext reduction, the tightness of the bounds is retained. A central result of this work is the proof of the existence, and the amount of a statistical threshold w.r.t. the introduction of additional features in general pattern classification, or the increase of context in string classification beyond which a decrease in Bayes error is guaranteed. Ralf Schlüter, Eugen Beck, Hermann Ney |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2018 | Segmental Encoder-Decoder Models for Large Vocabulary Automatic Speech RecognitionabstractIt has been known for a long time that the classic Hidden-Markov-Model (HMM) derivation for speech recognition contains assumptions such as independence of observation vectors and weak duration modeling that are practical but unrealistic.When using the hybrid approach this is amplified by trying to fit a discriminative model into a generative one.Hidden Conditional Random Fields (CRFs) and segmental models (e.g.Semi-Markov CRFs / Segmental CRFs) have been proposed as an alternative, but for a long time have failed to get traction until recently.In this paper we explore different length modeling approaches for segmental models, their relation to attention-based systems.Furthermore we show experimental results on a handwriting recognition task and to the best of our knowledge the first reported results on the Switchboard 300h speech recognition corpus using this approach. Eugen Beck, Mirko Hannemann, Patrick Doetsch, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 1 |
| 2017 | CTC in the Context of Generalized Full-Sum HMM TrainingabstractWe formulate a generalized hybrid HMM-NN training procedure using the full-sum over the hidden state-sequence and identify CTC as a special case of it.We present an analysis of the alignment behavior of such a training procedure and explain the strong localization of label output behavior of full-sum training (also referred to as peaky or spiky behavior).We show how to avoid that behavior by using a state prior.We discuss the temporal decoupling between output label position/time-frame, and the corresponding evidence in the input observations when this is trained with BLSTM models.We also show a way how to overcome this by jointly training a FFNN.We implemented the Baum-Welch alignment algorithm in CUDA to be able to do fast soft realignments on GPU.We have published this code along with some of our experiments as part of RETURNN, RWTH's extensible training framework for universal recurrent neural networks.We finish with experimental validation of our study on WSJ and Switchboard. Albert Zeyer, Eugen Beck, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 2 |
| 2015 | Error bounds for context reduction and feature omission
Eugen Beck, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 1 |
| 2013 | Relative error bounds for statistical classifiers based on the f-divergenceabstractIn language classification, measures like perplexity and Kullback-Leibler divergence are used to compare language models. While this bears the advantage of isolating the effect of the language model in speech and language processing problems, the measures have no clear relation to the corresponding classification error. In practice, an improvement in terms of perplexity does not necessarily correspond to an improvement in the error rate. It is well-known that Bayes decision rule is optimal if the true distribution is used for classification. Since the true distribution is unknown in practice, a model distribution is used instead, introducing suboptimality. We focus on the degradation introduced by a model distribution, and provide an upper bound on the error difference between Bayes decision and a modelbased decision rule in terms of the f-Divergence between the true and model distributions. Simulations are first presented to reveal a special case of the bound, followed by an analytic proof of the generalized bound and its tightness. In addition, the conditions that result in the boundary cases will be discussed. Several instances of the bound will be verified using simulations, and the bound will be used to study the effect of the language model on the classification error. Index Terms: generalization bounds, language modeling, perplexity, confidence measures, f-Divergence, error mismatch. Markus Nußbaum-Thom, Eugen Beck, Tamer Alkhouli, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 2 |
| 2013 | Novel tight classification error bounds under mismatch conditions based on f-DivergenceabstractBy default, statistical classification/multiple hypothesis testing is faced with the model mismatch introduced by replacing the true distributions in Bayes decision rule by model distributions estimated on training samples. Although a large number of statistical measures exist w.r.t. to the mismatch introduced, these works rarely relate to the mismatch in accuracy, i.e. the difference between model error and Bayes error. In this work, the accuracy mismatch between the ideal Bayes decision rule/Bayes test and a mismatched decision rule in statistical classification/multiple hypothesis testing is investigated explicitly. A proof of a novel generalized tight statistical bound on the accuracy mismatch is presented. This result is compared to existing statistical bounds related to the total variational distance that can be extended to bounds of the accuracy mismatch. The analytic results are supported by distribution simulations. Ralf Schlüter, Markus Nußbaum-Thom, Eugen Beck, Tamer Alkhouli, Hermann Ney |
ITW | 3 |