EDBT 2026 Demo / reviewers in the wild / expert
Wilfried Michel
dblp:212/6375
· DBLP profile ↗
14ranked-venue papers
3as first author
8since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 10 · 2 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | End-To-End Training of a Neural HMM with Label and Transition ProbabilitiesabstractWe investigate a novel modeling approach for end-to-end neural network training using hidden Markov models (HMM) where the transition probabilities between hidden states are modeled and learned explicitly. Most contemporary sequence-to-sequence models allow for from-scratch training by summing over all possible label segmentations in a given topology. In our approach there are explicit, learnable probabilities for transitions between segments as opposed to a blank label that implicitly encodes duration statistics.We implement a GPU-based forward-backward algorithm that enables the simultaneous training of label and transition probabilities.We investigate recognition results and additionally Viterbi alignments of our models. We find that while the transition model training does not improve recognition performance, it has a positive impact on the alignment quality. The generated alignments are shown to be viable targets in state-of-the-art Viterbi trainings. Daniel Mann, Tina Raissi, Wilfried Michel, Ralf Schlüter, Hermann Ney |
ASRU | 3 |
| 2022 | Efficient Sequence Training of Attention Models Using Approximative RecombinationabstractSequence discriminative training is a great tool to improve the performance of an automatic speech recognition system. It does, however, necessitate a sum over all possible word sequences, which is intractable to compute in practice. Current state-of-the-art systems with unlimited label context circumvent this problem by limiting the summation to an n-best list of relevant competing hypotheses obtained from beam search.This work proposes to perform (approximative) recombinations of hypotheses during beam search, if they share a common local history. The error that is incurred by the approximation is analyzed and it is shown that using this technique the effective beam size can be increased by several orders of magnitude without significantly increasing the computational requirements. Lastly, it is shown that this technique can be used to effectively perform sequence discriminative training for attention-based encoder-decoder acoustic models on the LibriSpeech task. Nils-Philipp Wynands, Wilfried Michel, Jan Rosendahl, Ralf Schlüter, Hermann Ney |
ICASSP | 2 |
| 2022 | Conformer-Based Hybrid ASR System For Switchboard DatasetabstractThe recently proposed conformer architecture has been successfully used for end-to-end automatic speech recognition (ASR) architectures achieving state-of-the-art performance on different datasets. To our best knowledge, the impact of using conformer acoustic model for hybrid ASR is not investigated. In this paper, we present and evaluate a competitive conformer-based hybrid model training recipe. We study different training aspects and methods to improve worderror-rate as well as to increase training speed. We apply time downsampling methods for efficient training and use transposed convolutions to upsample the output sequence again. We conduct experiments on Switchboard 300h dataset and our conformer-based hybrid model achieves competitive results compared to other architectures. It generalizes very well on Hub5’01 test set and outperforms the BLSTM-based hybrid model significantly. Mohammad Zeineldeen, Jingjing Xu 0002, Christoph Lüscher, Wilfried Michel, Alexander Gerstenberger, Ralf Schlüter, Hermann Ney |
ICASSP | 4 |
| 2022 | Automatic Learning of Subword Dependent Model ScalesabstractTo improve the performance of state-of-the-art automatic speech recognition systems it is common practice to include external knowledge sources such as language models or prior corrections.This is usually done via log-linear model combination using separate scaling parameters for each model.Typically these parameters are manually optimized on some held-out data.In this work we propose to optimize these scaling parameters via automatic differentiation and stochastic gradient decent similar to the neural network model parameters.We show on the LibriSpeech (LBS) and Switchboard (SWB) corpora that the model scales for a combination of attentionbased encoder-decoder acoustic model and language model can be learned as effectively as with manual tuning.We further extend this approach to subword dependent model scales which could not be tuned manually which leads to 7% improvement on LBS and 3% on SWB.We also show that joint training of scales and model parameters is possible and gives additional 6% improvement on LBS. Felix Meyer, Wilfried Michel, Mohammad Zeineldeen, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 2 |
| 2022 | Efficient Training of Neural Transducer for Speech RecognitionabstractAs one of the most popular sequence-to-sequence modeling approaches for speech recognition, the RNN-Transducer has achieved evolving performance with more and more sophisticated neural network models of growing size and increasing training epochs. While strong computation resources seem to be the prerequisite of training superior models, we try to overcome it by carefully designing a more efficient training pipeline. In this work, we propose an efficient 3-stage progressive training pipeline to build highly-performing neural transducer models from scratch with very limited computation resources in a reasonable short time period. The effectiveness of each stage is experimentally verified on both Librispeech and Switchboard corpora. The proposed pipeline is able to train transducer models approaching state-of-the-art performance with a single GPU in just 2-3 weeks. Our best conformer transducer achieves 4.1% WER on Librispeech test-other with only 35 epochs of training. Wei Zhou 0043, Wilfried Michel, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 2 |
| 2021 | On Architectures and Training for Raw Waveform Feature Extraction in ASRabstractWith the success of neural network based modeling in auto-matic speech recognition (ASR), many studies investigated acoustic modeling and learning of feature extractors directly based on the raw waveform. Recently, one line of research has focused on unsupervised pre-training of feature extractors on audio-only data to improve downstream ASR performance. In this work, we investigate the usefulness of one of these front-end frameworks, namely wav2vec, in a setting without additional untranscribed data for hybrid ASR systems. We compare this framework both to the manually defined stan-dard Gammatone feature set, as well as to features extracted as part of the acoustic model of an ASR system trained su-pervised. We study the benefits of using the pre-trained feature extractor and explore how to additionally exploit an ex-isting acoustic model trained with different features. Finally, we systematically examine combinations of the described features in order to further advance the performance. Peter Vieting, Christoph Lüscher, Wilfried Michel, Ralf Schlüter, Hermann Ney |
ASRU | 3 |
| 2021 | Investigating Methods to Improve Language Model Integration for Attention-Based Encoder-Decoder ASR ModelsabstractAttention-based encoder-decoder (AED) models learn an implicit internal language model (ILM) from the training transcriptions. The integration with an external LM trained on much more unpaired text usually leads to better performance. A Bayesian interpretation as in the hybrid autoregressive transducer (HAT) suggests dividing by the prior of the discriminative acoustic model, which corresponds to this implicit LM, similarly as in the hybrid hidden Markov model approach. The implicit LM cannot be calculated efficiently in general and it is yet unclear what are the best methods to estimate it. In this work, we compare different approaches from the literature and propose several novel methods to estimate the ILM directly from the AED model. Our proposed methods outperform all previous approaches. We also investigate other methods to suppress the ILM mainly by decreasing the capacity of the AED model, limiting the label context, and also by training the AED model together with a pre-existing LM. Mohammad Zeineldeen, Aleksandr Glushko, Wilfried Michel, Albert Zeyer, Ralf Schlüter, Hermann Ney |
Interspeech | 3 |
| 2021 | Librispeech Transducer Model with Internal Language Model Prior CorrectionabstractWe present our transducer model on Librispeech. We study variants to include an external language model (LM) with shallow fusion and subtract an estimated internal LM. This is justified by a Bayesian interpretation where the transducer model prior is given by the estimated internal LM. The subtraction of the internal LM gives us over 14% relative improvement over normal shallow fusion. Our transducer has a separate probability distribution for the non-blank labels which allows for easier combination with the external LM, and easier estimation of the internal LM. We additionally take care of including the end-of-sentence (EOS) probability of the external LM in the last blank probability which further improves the performance. All our code and setups are published. Albert Zeyer, André Merboldt, Wilfried Michel, Ralf Schlüter, Hermann Ney |
Interspeech | 3 |
| 2020 | Frame-Level MMI as A Sequence Discriminative Training Criterion for LVCSRabstractIn this work we present frame-level maximum mutual information (MMI) as a novel sequence discriminative training criterion for hybrid HMM-DNN acoustic models. Compared to the standard, sequence-level MMI criterion we show that frame-level MMI has increased robustness towards missing cross-entropy (CE) smoothing and can converge even without interpolation. Using model free optimization, we show that in the asymptotic case of an infinite amount of training data models trained using this criterion are equal to the true class posterior distribution, whereas training using the state-level minimum Bayes risk (sMBR) criterion leads to a distorted function of the true class posterior distribution. This analytical result is backed by experimental evidence. We further propose a generalized class of training criteria, that continuously interpolates between frame-level MMI and sMBR criterion. Wilfried Michel, Ralf Schlüter, Hermann Ney |
ICASSP | 1 |
| 2020 | The Rwth Asr System for Ted-Lium Release 2: Improving Hybrid Hmm With SpecaugmentabstractWe present a complete training pipeline to build a state-of-the-art hybrid HMM-based ASR system on the 2nd release of the TED-LIUM corpus. Data augmentation using SpecAugment is successfully applied to improve performance on top of our best SAT model using i-vectors. By investigating the effect of different maskings, we achieve improvements from SpecAugment on hybrid HMM models without increasing model size and training time. A subsequent sMBR training is applied to fine-tune the final acoustic model, and both LSTM and Transformer language models are trained and evaluated. Our best system achieves a 5.6% WER on the test set, which outperforms the previous state-of-the-art by 27% relative. Wei Zhou 0043, Wilfried Michel, Kazuki Irie, Markus Kitza, Ralf Schlüter, Hermann Ney |
ICASSP | 2 |
| 2020 | Early Stage LM Integration Using Local and Global Log-Linear CombinationabstractSequence-to-sequence models with an implicit alignment mechanism (e.g. attention) are closing the performance gap towards traditional hybrid hidden Markov models (HMM) for the task of automatic speech recognition. One important factor to improve word error rate in both cases is the use of an external language model (LM) trained on large text-only corpora. Language model integration is straightforward with the clear separation of acoustic model and language model in classical HMM-based modeling. In contrast, multiple integration schemes have been proposed for attention models. In this work, we present a novel method for language model integration into implicit-alignment based sequence-to-sequence models. Log-linear model combination of acoustic and language model is performed with a per-token renormalization. This allows us to compute the full normalization term efficiently both in training and in testing. This is compared to a global renormalization scheme which is equivalent to applying shallow fusion in training. The proposed methods show good improvements over standard model combination (shallow fusion) on our state-of-the-art Librispeech system. Furthermore, the improvements are persistent even if the LM is exchanged for a more powerful one after training. Wilfried Michel, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 1 |
| 2019 | RWTH ASR Systems for LibriSpeech: Hybrid vs AttentionabstractWe present state-of-the-art automatic speech recognition (ASR) systems employing a standard hybrid DNN/HMM architecture compared to an attention-based encoder-decoder design for the LibriSpeech task. Detailed descriptions of the system development, including model design, pretraining schemes, training schedules, and optimization approaches are provided for both system architectures. Both hybrid DNN/HMM and attention-based systems employ bi-directional LSTMs for acoustic modeling/encoding. For language modeling, we employ both LSTM and Transformer based architectures. All our systems are built using RWTHs open-source toolkits RASR and RETURNN. To the best knowledge of the authors, the results obtained when training on the full LibriSpeech training set, are the best published currently, both for the hybrid DNN/HMM and the attention-based systems. Our single hybrid system even outperforms previous results obtained from combining eight single systems. Our comparison shows that on the LibriSpeech 960h task, the hybrid DNN/HMM system outperforms the attention-based system by 15% relative on the clean and 40% relative on the other test sets in terms of word error rate. Moreover, experiments on a reduced 100h-subset of the LibriSpeech training corpus even show a more pronounced margin between the hybrid DNN/HMM and attention-based architectures. Christoph Lüscher, Eugen Beck, Kazuki Irie, Markus Kitza, Wilfried Michel, Albert Zeyer, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 5 |
| 2019 | Comparison of Lattice-Free and Lattice-Based Sequence Discriminative Training Criteria for LVCSRabstractSequence discriminative training criteria have long been a standard tool in automatic speech recognition for improving the performance of acoustic models over their maximum likelihood / cross entropy trained counterparts. While previously a lattice approximation of the search space has been necessary to reduce computational complexity, recently proposed methods use other approximations to dispense of the need for the computationally expensive step of separate lattice creation. In this work we present a memory efficient implementation of the forward-backward computation that allows us to use uni-gram word-level language models in the denominator calculation while still doing a full summation on GPU. This allows for a direct comparison of lattice-based and lattice-free sequence discriminative training criteria such as MMI and sMBR, both using the same language model during training. We compared performance, speed of convergence, and stability on large vocabulary continuous speech recognition tasks like Switchboard and Quaero. We found that silence modeling seriously impacts the performance in the lattice-free case and needs special treatment. In our experiments lattice-free MMI comes on par with its lattice-based counterpart. Lattice-based sMBR still outperforms all lattice-free training criteria. Wilfried Michel, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 1 |
| 2017 | Parallel Neural Network Features for Improved Tandem Acoustic ModelingabstractThe combination of acoustic models or features is a standard approach to exploit various knowledge sources.This paper investigates the concatenation of different bottleneck (BN) neural network (NN) outputs for tandem acoustic modeling.Thus, combination of NN features is performed via Gaussian mixture models (GMM).Complementarity between the NN feature representations is attained by using various network topologies: LSTM recurrent, feed-forward, and hierarchical, as well as different non-linearities: hyperbolic tangent, sigmoid, and rectified linear units.Speech recognition experiments are carried out on various tasks: telephone conversations, Skype calls, as well as broadcast news and conversations.Results indicate that LSTM based tandem approach is still competitive, and such tandem model can challenge comparable hybrid systems.The traditional steps of tandem modeling, speaker adaptive and sequence discriminative GMM training, improve the tandem results further.Furthermore, these "old-fashioned" steps remain applicable after the concatenation of multiple neural network feature streams.Exploiting the parallel processing of input feature streams, it is shown that 2-5% relative improvement could be achieved over the single best BN feature set.Finally, we also report results after neural network based language model rescoring and examine the system combination possibilities using such complex tandem models. Zoltán Tüske, Wilfried Michel, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 2 |