Ralf Schlüter

dblp:43/2928 · DBLP profile ↗
← Back
238ranked-venue papers
16as first author
46since 2021 · last 2026
0000-0003-2839-9247ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 223 · 12 first-author · 43 since 2021Artificial intelligence and machine learning · 145 · 10 first-author · 28 since 2021Theory of computation · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Supplementary Resources and Analysis for Automatic Speech Recognition Systems Trained on the Loquacious Dataset
Nick Rossenbach, Robin Schmitt, Tina Raissi, Simon Berger, Larissa Kleppel, Ralf Schlüter
LREC6
2025 Analysis of Domain Shift across ASR Architectures via TTS-Enabled Separation of Target Domain and Acoustic Conditions
abstract
We analyze automatic speech recognition (ASR) modeling choices under domain mismatch, comparing classic modular and novel sequence-to-sequence (seq2seq) architectures. Across the different ASR architectures, we examine a spectrum of modeling choices, including label units, context length, and topology. To isolate language domain effects from acoustic variation, we synthesize target domain audio using a text-to-speech system trained on LibriSpeech. We incorporate target domain ngram and neural language models for domain adaptation without retraining the acoustic model. To our knowledge, this is the first controlled comparison of optimized ASR systems across state-of-the-art architectures under domain shift, offering insights into their generalization. The results show that, under domain shift, rather than the decoder architecture choice or the distinction between classic modular and novel seq2seq models, it is specific modeling choices that influence performance.
Tina Raissi, Nick Rossenbach, Ralf Schlüter
ASRU3
2025 Efficient Supernet Training with Orthogonal Softmax for Scalable ASR Model Compression
abstract
ASR systems are deployed across diverse environments, each with specific hardware constraints. We use supernet training to jointly train multiple encoders of varying sizes, enabling dynamic model size adjustment to fit hardware constraints without redundant training. Moreover, we introduce a novel method called OrthoSoftmax, which applies multiple orthogonal softmax functions to efficiently identify optimal subnets within the supernet, avoiding resource-intensive search. This approach also enables more flexible and precise subnet selection by allowing selection based on various criteria and levels of granularity. Our results with CTC on Librispeech and TED-LIUM-v2 show that FLOPs-aware component-wise selection achieves the best overall performance. With the same number of training updates from one single job, WERs for all model sizes are comparable to or slightly better than those of individually trained models. Furthermore, we analyze patterns in the selected components and reveal interesting insights.
Jingjing Xu 0002, Eugen Beck, Ralf Schlüter
ICASSP4
2025 Right Label Context in End-to-End Training of Time-Synchronous ASR Models
abstract
Current time-synchronous sequence-to-sequence automatic speech recognition (ASR) models are trained by using sequence level cross-entropy that sums over all alignments. Due to the discriminative formulation, incorporating the right label context into the training criterion’s gradient causes normalization problems and is not mathematically well-defined. The classic hybrid neural network hidden Markov model (NN-HMM) with its inherent generative formulation enables conditioning on the right label context. However, due to the HMM state-tying the identity of the right label context is never modeled explicitly. In this work, we propose a factored loss with auxiliary left and right label contexts that sums over all alignments. We show that the inclusion of the right label context is particularly beneficial when training data resources are limited. Moreover, we also show that it is possible to build a factored hybrid HMM system by relying exclusively on the full-sum criterion. Experiments were conducted on Switchboard 300h and LibriSpeech 960h.
Tina Raissi, Ralf Schlüter, Hermann Ney
ICASSP2
2025 The Conformer Encoder May Reverse the Time Dimension
abstract
We sometimes observe monotonically decreasing cross-attention weights in our Conformer-based global attention-based encoder-decoder (AED) models, negatively affecting performance compared to monotonically increasing attention weights. Further investigation shows that the Conformer encoder reverses the sequence in the time dimension. We analyze the initial behavior of the decoder cross-attention mechanism and find that it encourages the Conformer encoder self-attention to build a connection between the initial frames and all other informative frames. Furthermore, we show that, at some point in training, the self-attention module of the Conformer starts dominating the output over the preceding feed-forward module, which then only allows the reversed information to pass through. We propose methods and ideas of how this flipping can be avoided and investigate a novel method to obtain label-frame-position alignments by using the gradients of the label log probabilities w.r.t. the encoder input frames.
Robin Schmitt, Albert Zeyer, Mohammad Zeineldeen, Ralf Schlüter, Hermann Ney
ICASSP4
2025 Classification Error Bound for Low Bayes Error Conditions in Machine Learning
abstract
In statistical classification and machine learning, classification error is an important performance measure, which is minimized by the Bayes decision rule. In practice, the unknown true distribution is usually replaced with a model distribution estimated from the training data in the Bayes decision rule. This substitution introduces a mismatch between the Bayes error and the model-based classification error. In this work, we apply classification error bounds to study the relationship between the error mismatch and the Kullback-Leibler divergence in machine learning. Motivated by recent observations of low model-based classification errors in many machine learning tasks, bounding the Bayes error to be lower, we propose a linear approximation of the classification error bound for low Bayes error conditions. Then, the bound for class priors are discussed. Moreover, we extend the classification error bound for sequences. Using automatic speech recognition as a representative example of machine learning applications, this work analytically discusses the correlations among different performance measures with extended bounds, including cross-entropy loss, language model perplexity, and word error rate.
Vahe Eminyan, Ralf Schlüter, Hermann Ney
ICASSP3
2025 Dynamic Acoustic Model Architecture Optimization in Training for ASR
Jingjing Xu 0002, Albert Zeyer, Eugen Beck, Ralf Schlüter, Hermann Ney
INTERSPEECH5
2025 Analyzing the Importance of Blank for CTC-Based Knowledge Distillation
Benedikt Hilmes, Nick Rossenbach, Ralf Schlüter
INTERSPEECH3
2025 Running Conventional Automatic Speech Recognition on Memristor Hardware: A Simulated Approach
Nick Rossenbach, Benedikt Hilmes, Leon Brackmann, Moritz Gunz, Ralf Schlüter
INTERSPEECH5
2025 Regularizing Learnable Feature Extraction for Automatic Speech Recognition
abstract
Neural front-ends are an appealing alternative to traditional, fixed feature extraction pipelines for automatic speech recognition (ASR) systems since they can be directly trained to fit the acoustic model. However, their performance often falls short compared to classical methods, which we show is largely due to their increased susceptibility to overfitting. This work therefore investigates regularization methods for training ASR models with learnable feature extraction front-ends. First, we examine audio perturbation methods and show that larger relative improvements can be obtained for learnable features. Additionally, we identify two limitations in the standard use of SpecAugment for these front-ends and propose masking in the short time Fourier transform (STFT)-domain as a simple but effective modification to address these challenges. Finally, integrating both regularization approaches effectively closes the performance gap between traditional and learnable features.
Peter Vieting, Maximilian Kannen, Benedikt Hilmes, Ralf Schlüter, Hermann Ney
INTERSPEECH4
2025 Label-Context-Dependent Internal Language Model Estimation for CTC
Minh-Nghia Phan, Ralf Schlüter, Hermann Ney
INTERSPEECH3
2024 On the Relation Between Internal Language Model and Sequence Discriminative Training for Neural Transducers
abstract
Internal language model (ILM) subtraction has been widely applied to improve the performance of the RNN-Transducer with external language model (LM) fusion for speech recognition. In this work, we show that sequence discriminative training has a strong correlation with ILM subtraction from both theoretical and empirical points of view. Theoretically, we derive that the global optimum of maximum mutual information (MMI) training shares a similar formula as ILM subtraction. Empirically, we show that ILM subtraction and sequence discriminative training achieve similar effects across a wide range of experiments on Librispeech, including both MMI and minimum Bayes risk (MBR) criteria, as well as neural transducers and LMs of both full and limited context. The benefit of ILM subtraction also becomes much smaller after sequence discriminative training. We also provide an indepth study to show that sequence discriminative training has a minimal effect on the commonly used zero-encoder ILM estimation, but a joint effect on both encoder and prediction + joint network for posterior probability reshaping including both ILM and blank suppression.
Wei Zhou 0043, Ralf Schlüter, Hermann Ney
ICASSP3
2024 Chunked Attention-Based Encoder-Decoder Model for Streaming Speech Recognition
abstract
We study a streamable attention-based encoder-decoder model in which either the decoder, or both the encoder and decoder, operate on pre-defined, fixed-size windows called chunks. A special end-of-chunk (EOC) symbol advances from one chunk to the next chunk, effectively replacing the conventional end-of-sequence symbol. This modification, while minor, situates our model as equivalent to a transducer model that operates on chunks instead of frames, where EOC corresponds to the blank symbol. We further explore the remaining differences between a standard transducer and our model. Additionally, we examine relevant aspects such as long-form speech generalization, beam size, and length normalization. Through experiments on Librispeech and TED-LIUM-v2, and by concatenating consecutive sequences for long-form trials, we find that our streamable model maintains competitive performance compared to the non-streamable variant and generalizes very well to long-form speech.
Mohammad Zeineldeen, Albert Zeyer, Ralf Schlüter, Hermann Ney
ICASSP3
2024 Dynamic Encoder Size Based on Data-Driven Layer-wise Pruning for Speech Recognition
Jingjing Xu 0002, Wei Zhou 0043, Eugen Beck, Ralf Schlüter
INTERSPEECH5
2024 Investigating the Effect of Label Topology and Training Criterion on ASR Performance and Alignment Quality
Tina Raissi, Christoph Lüscher, Simon Berger, Ralf Schlüter, Hermann Ney
INTERSPEECH4
2024 Refined Statistical Bounds for Classification Error Mismatches with Constrained Bayes Error
abstract
In statistical classification/multiple hypothesis testing and machine learning, a model distribution estimated from the training data is usually applied to replace the unknown true distribution in the Bayes decision rule, which introduces a mismatch between the Bayes error and the model-based classification error. In this work, we derive the classification error bound to study the relationship between the Kullback-Leibler divergence and the classification error mismatch. We first reconsider the statistical bounds based on classification error mismatch derived in previous works, employing a different method of derivation. Then, motivated by the observation that the Bayes error is typically low in machine learning tasks like speech recognition and pattern recognition, we derive a refined Kullback-Leibler-divergence-based bound on the error mismatch with the constraint that the Bayes error is lower than a threshold.
Vahe Eminyan, Ralf Schlüter, Hermann Ney
ITW3
2024 Combining TF-GridNet And Mixture Encoder For Continuous Speech Separation For Meeting Transcription
abstract
Many real-life applications of automatic speech recognition (ASR) require processing of overlapped speech. A common method involves first separating the speech into overlap-free streams on which ASR is performed. Recently, TF-GridNet has shown impressive performance in speech separation in real reverberant conditions. Furthermore, a mixture encoder was proposed that leverages the mixed speech to mitigate the effect of separation artifacts. In this work, we extended the mixture encoder from a static two-speaker scenario to a natural meeting context featuring an arbitrary number of speakers and varying degrees of overlap. We further demonstrate its limits by the integration with separators of varying strength including TF-GridNet. Our experiments result in a new state-of-the-art performance on LibriCSS using a single microphone. They show that TF-GridNet largely closes the gap between previous methods and oracle separation independent of mixture encoding. We further investigate the remaining potential for improvement.
Peter Vieting, Simon Berger, Thilo von Neumann, Christoph Böddeker, Ralf Schlüter, Reinhold Häb-Umbach
SLT5
2024 End-to-End Speech Recognition: A Survey
abstract
In the last decade of automatic speech recognition (ASR) research, the introduction of deep learning has brought considerable reductions in word error rate of more than 50% relative, compared to modeling without deep learning. In the wake of this transition, a number of all-neural ASR architectures have been introduced. These so-calledend-to-end(E2E) models provide highly integrated, completely neural ASR models, which rely strongly on general machine learning knowledge, learn more consistently from data, with lower dependence on ASR domain-specific experience. The success and enthusiastic adoption of deep learning, accompanied by more generic model architectures has led to E2E models now becoming the prominent ASR approach. The goal of this survey is to provide a taxonomy of E2E ASR models and corresponding improvements, and to discuss their properties and their relationship to classical hidden Markov model (HMM) based ASR architectures. All relevant aspects of E2E ASR are covered in this work: modeling, training, decoding, and external language model integration, discussions of performance and deployment opportunities, as well as an outlook into potential future developments.
Rohit Prabhavalkar, Takaaki Hori, Tara N. Sainath, Ralf Schlüter, Shinji Watanabe 0001
IEEE ACM Trans. Audio Speech Lang. Process.4
2023 End-To-End Training of a Neural HMM with Label and Transition Probabilities
abstract
We investigate a novel modeling approach for end-to-end neural network training using hidden Markov models (HMM) where the transition probabilities between hidden states are modeled and learned explicitly. Most contemporary sequence-to-sequence models allow for from-scratch training by summing over all possible label segmentations in a given topology. In our approach there are explicit, learnable probabilities for transitions between segments as opposed to a blank label that implicitly encodes duration statistics.We implement a GPU-based forward-backward algorithm that enables the simultaneous training of label and transition probabilities.We investigate recognition results and additionally Viterbi alignments of our models. We find that while the transition model training does not improve recognition performance, it has a positive impact on the alignment quality. The generated alignments are shown to be viable targets in state-of-the-art Viterbi trainings.
Daniel Mann, Tina Raissi, Wilfried Michel, Ralf Schlüter, Hermann Ney
ASRU4
2023 On the Relevance of Phoneme Duration Variability of Synthesized Training Data for Automatic Speech Recognition
abstract
Synthetic data generated by text-to-speech (TTS) systems can be used to improve automatic speech recognition (ASR) systems in low-resource or domain mismatch tasks. It has been shown that TTS-generated outputs still do not have the same qualities as real data. In this work we focus on the temporal structure of synthetic data and its relation to ASR training. By using a novel oracle setup we show how much the degradation of synthetic data quality is influenced by duration modeling in non-autoregressive (NAR) TTS. To get reference phoneme durations we use two common alignment methods, a hidden Markov Gaussian-mixture model (HMM-GMM) aligner and a neural connectionist temporal classification (CTC) aligner. Using a simple algorithm based on random walks we shift phoneme duration distributions of the TTS system closer to real durations, resulting in an improvement of an ASR system using synthetic data in a semi-supervised setting.
Nick Rossenbach, Benedikt Hilmes, Ralf Schlüter
ASRU3
2023 Investigating The Effect of Language Models in Sequence Discriminative Training For Neural Transducers
abstract
In this work, we investigate the effect of language models (LMs) with different context lengths and label units (phoneme vs. word) used in sequence discriminative training for phoneme-based neural transducers. Both lattice-free and N-best-list approaches are examined. For lattice-free methods with phoneme-level LMs, we propose a method to approximate the context history to employ LMs with full-context dependency. This approximation can be extended to arbitrary context length and enables the usage of word-level LMs in lattice-free methods. Moreover, a systematic comparison is conducted across lattice-free and N-best-list-based methods. Experimental results on Librispeech show that using the word-level LM in training outperforms the phoneme-level LM. Besides, we find that the context size of the LM used for probability computation has a limited effect on performance. Moreover, our results reveal the pivotal importance of the hypothesis space quality in sequence discriminative training.
Wei Zhou 0043, Ralf Schlüter, Hermann Ney
ASRU3
2023 Lattice-Free Sequence Discriminative Training for Phoneme-Based Neural Transducers
abstract
Recently, RNN-Transducers have achieved remarkable results on various automatic speech recognition tasks. However, lattice-free sequence discriminative training methods, which obtain superior performance in hybrid models, are rarely investigated in RNN-Transducers. In this work, we propose three lattice-free training objectives, namely lattice-free maximum mutual information, lattice-free segment-level minimum Bayes risk, and lattice-free minimum Bayes risk, which are used for the final posterior output of the phoneme-based neural transducer with a limited context dependency. Compared to criteria using N-best lists, lattice-free methods eliminate the decoding step for hypotheses generation during training, which leads to more efficient training. Experimental results show that lattice-free methods gain up to 6.5% relative improvement in word error rate compared to a sequence-level cross-entropy trained model. Compared to the N-best-list based minimum Bayes risk objectives, lattice-free methods gain 40% - 70% relative training time speedup with a small degradation in performance.
Wei Zhou 0043, Ralf Schlüter, Hermann Ney
ICASSP3
2023 Enhancing and Adversarial: Improve ASR with Speaker Labels
abstract
ASR can be improved by multi-task learning (MTL) with domain enhancing or domain adversarial training, which are two opposite objectives with the aim to increase/decrease domain variance towards domain-aware/agnostic ASR, respectively. In this work, we study how to best apply these two opposite objectives with speaker labels to improve conformer-based ASR. We also propose a novel adaptive gradient reversal layer for stable and effective adversarial training without tuning effort. Detailed analysis and experimental verification are conducted to show the optimal positions in the ASR neural network (NN) to apply speaker enhancing and adversarial training. We also explore their combination for further improvement, achieving the same performance as i-vectors plus adversarial training. Our best speaker-based MTL achieves 7% relative improvement on the Switchboard Hub5’00 set. We also investigate the effect of such speaker-based MTL w.r.t. cleaner dataset and weaker ASR NN.
Wei Zhou 0043, Jingjing Xu 0002, Mohammad Zeineldeen, Christoph Lüscher, Ralf Schlüter, Hermann Ney
ICASSP6
2023 RASR2: The RWTH ASR Toolkit for Generic Sequence-to-sequence Speech Recognition
abstract
Modern public ASR tools usually provide rich support for training various sequence-to-sequence (S2S) models, but rather simple support for decoding open-vocabulary scenarios only.For closed-vocabulary scenarios, public tools supporting lexicalconstrained decoding are usually only for classical ASR, or do not support all S2S models.To eliminate this restriction on research possibilities such as modeling unit choice, we present RASR2 in this work, a research-oriented generic S2S decoder implemented in C++.It offers a strong flexibility/compatibility for various S2S models, language models, label units/topologies and neural network architectures.It provides efficient decoding for both open-and closed-vocabulary scenarios based on a generalized search framework with rich support for different search modes and settings.We evaluate RASR2 with a wide range of experiments on both switchboard and Librispeech corpora.Our source code is public online.
Wei Zhou 0043, Eugen Beck, Simon Berger, Ralf Schlüter, Hermann Ney
INTERSPEECH4
2023 Mixture Encoder for Joint Speech Separation and Recognition
Simon Berger, Peter Vieting, Christoph Böddeker, Ralf Schlüter, Reinhold Häb-Umbach
INTERSPEECH4
2023 Competitive and Resource Efficient Factored Hybrid HMM Systems are Simpler Than You Think
Tina Raissi, Christoph Lüscher, Moritz Gunz, Ralf Schlüter, Hermann Ney
INTERSPEECH4
2022 Improving Factored Hybrid HMM Acoustic Modeling without State Tying
abstract
In this work, we show that a factored hybrid hidden Markov model (FH-HMM) which is defined without any phonetic state-tying outperforms a state-of-the-art hybrid HMM. The factored hybrid HMM provides a link to transducer models in the way it models phonetic (label) context while preserving the strict separation of acoustic and language model of the hybrid HMM approach. Furthermore, we show that the factored hybrid model can be trained from scratch without using phonetic state-tying in any of the training steps. Our modeling approach enables triphone context while avoiding phonetic state-tying by a decomposition into locally normalized factored posteriors for monophones/HMM states in phoneme context. Experimental results are provided for Switchboard 300h and LibriSpeech. On the former task we also show that by avoiding the phonetic state-tying step, the factored hybrid can take better advantage of regularization techniques during training, compared to the standard hybrid HMM with phonetic state-tying based on classification and regression trees (CART).
Tina Raissi, Eugen Beck, Ralf Schlüter, Hermann Ney
ICASSP3
2022 Efficient Sequence Training of Attention Models Using Approximative Recombination
abstract
Sequence discriminative training is a great tool to improve the performance of an automatic speech recognition system. It does, however, necessitate a sum over all possible word sequences, which is intractable to compute in practice. Current state-of-the-art systems with unlimited label context circumvent this problem by limiting the summation to an n-best list of relevant competing hypotheses obtained from beam search.This work proposes to perform (approximative) recombinations of hypotheses during beam search, if they share a common local history. The error that is incurred by the approximation is analyzed and it is shown that using this technique the effective beam size can be increased by several orders of magnitude without significantly increasing the computational requirements. Lastly, it is shown that this technique can be used to effectively perform sequence discriminative training for attention-based encoder-decoder acoustic models on the LibriSpeech task.
Nils-Philipp Wynands, Wilfried Michel, Jan Rosendahl, Ralf Schlüter, Hermann Ney
ICASSP4
2022 Conformer-Based Hybrid ASR System For Switchboard Dataset
abstract
The recently proposed conformer architecture has been successfully used for end-to-end automatic speech recognition (ASR) architectures achieving state-of-the-art performance on different datasets. To our best knowledge, the impact of using conformer acoustic model for hybrid ASR is not investigated. In this paper, we present and evaluate a competitive conformer-based hybrid model training recipe. We study different training aspects and methods to improve worderror-rate as well as to increase training speed. We apply time downsampling methods for efficient training and use transposed convolutions to upsample the output sequence again. We conduct experiments on Switchboard 300h dataset and our conformer-based hybrid model achieves competitive results compared to other architectures. It generalizes very well on Hub5’01 test set and outperforms the BLSTM-based hybrid model significantly.
Mohammad Zeineldeen, Jingjing Xu 0002, Christoph Lüscher, Wilfried Michel, Alexander Gerstenberger, Ralf Schlüter, Hermann Ney
ICASSP6
2022 On Language Model Integration for RNN Transducer Based Speech Recognition
abstract
The mismatch between an external language model (LM) and the implicitly learned internal LM (ILM) of RNN-Transducer (RNN-T) can limit the performance of LM integration such as simple shallow fusion. A Bayesian interpretation suggests to remove this sequence prior as ILM correction. In this work, we study various ILM correction-based LM integration methods formulated in a common RNN-T framework. We provide a decoding interpretation on two major reasons for performance improvement with ILM correction, which is further experimentally verified with detailed analysis. We also propose an exact-ILM training framework by extending the proof given in the hybrid autoregressive transducer, which enables a theoretical justification for other ILM approaches. Systematic comparison is conducted for both in-domain and cross-domain evaluation on the Librispeech and TED-LIUM Release 2 corpora, respectively. Our proposed exact-ILM training can further improve the best ILM method.
Wei Zhou 0043, Zuoyun Zheng, Ralf Schlüter, Hermann Ney
ICASSP3
2022 Automatic Learning of Subword Dependent Model Scales
abstract
To improve the performance of state-of-the-art automatic speech recognition systems it is common practice to include external knowledge sources such as language models or prior corrections.This is usually done via log-linear model combination using separate scaling parameters for each model.Typically these parameters are manually optimized on some held-out data.In this work we propose to optimize these scaling parameters via automatic differentiation and stochastic gradient decent similar to the neural network model parameters.We show on the LibriSpeech (LBS) and Switchboard (SWB) corpora that the model scales for a combination of attentionbased encoder-decoder acoustic model and language model can be learned as effectively as with manual tuning.We further extend this approach to subword dependent model scales which could not be tuned manually which leads to 7% improvement on LBS and 3% on SWB.We also show that joint training of scales and model parameters is possible and gives additional 6% improvement on LBS.
Felix Meyer, Wilfried Michel, Mohammad Zeineldeen, Ralf Schlüter, Hermann Ney
INTERSPEECH4
2022 Self-Normalized Importance Sampling for Neural Language Modeling
abstract
To mitigate the problem of having to traverse over the full vocabulary in the softmax normalization of a neural language model, sampling-based training criteria are proposed and investigated in the context of large vocabulary word-based neural language models. These training criteria typically enjoy the benefit of faster training and testing, at a cost of slightly degraded performance in terms of perplexity and almost no visible drop in word error rate. While noise contrastive estimation is one of the most popular choices, recently we show that other sampling-based criteria can also perform well, as long as an extra correction step is done, where the intended class posterior probability is recovered from the raw model outputs. In this work, we propose self-normalized importance sampling. Compared to our previous work, the criteria considered in this work are self-normalized and there is no need to further conduct a correction step. Through self-normalized language model training as well as lattice rescoring experiments, we show that our proposed self-normalized importance sampling is competitive in both research-oriented and production-oriented automatic speech recognition tasks.
Yingbo Gao, Alexander Gerstenberger, Jintao Jiang, Ralf Schlüter, Hermann Ney
INTERSPEECH5
2022 Improving the Training Recipe for a Robust Conformer-based Hybrid Model
Mohammad Zeineldeen, Jingjing Xu 0002, Christoph Lüscher, Ralf Schlüter, Hermann Ney
INTERSPEECH4
2022 Efficient Training of Neural Transducer for Speech Recognition
abstract
As one of the most popular sequence-to-sequence modeling approaches for speech recognition, the RNN-Transducer has achieved evolving performance with more and more sophisticated neural network models of growing size and increasing training epochs. While strong computation resources seem to be the prerequisite of training superior models, we try to overcome it by carefully designing a more efficient training pipeline. In this work, we propose an efficient 3-stage progressive training pipeline to build highly-performing neural transducer models from scratch with very limited computation resources in a reasonable short time period. The effectiveness of each stage is experimentally verified on both Librispeech and Switchboard corpora. The proposed pipeline is able to train transducer models approaching state-of-the-art performance with a single GPU in just 2-3 weeks. Our best conformer transducer achieves 4.1% WER on Librispeech test-other with only 35 epochs of training.
Wei Zhou 0043, Wilfried Michel, Ralf Schlüter, Hermann Ney
INTERSPEECH3
2022 HMM vs. CTC for Automatic Speech Recognition: Comparison Based on Full-Sum Training from Scratch
abstract
In this work, we compare from-scratch sequence-level cross-entropy (full-sum) training of Hidden Markov Model (HMM) and Connectionist Temporal Classification (CTC) topologies for automatic speech recognition (ASR). Besides accuracy, we further analyze their capability for generating high-quality time alignment between the speech signal and the transcription, which can be crucial for many subsequent applications. Moreover, we propose several methods to improve convergence of from-scratch full-sum training by addressing the alignment modeling issue. Systematic comparison is conducted on both Switchboard and LibriSpeech corpora across CTC, posterior HMM with and w/o transition probabilities, and standard hybrid HMM. We also provide a detailed analysis of both Viterbi forced-alignment and Baum-Welch full-sum occupation probabilities.
Tina Raissi, Wei Zhou 0043, Simon Berger, Ralf Schlüter, Hermann Ney
SLT4
2022 Monotonic Segmental Attention for Automatic Speech Recognition
abstract
We introduce a novel segmental-attention model for automatic speech recognition. We restrict the decoder attention to segments to avoid quadratic runtime of global attention, better generalize to long sequences, and eventually enable streaming. We directly compare global-attention and different segmental-attention modeling variants. We develop and compare two separate time-synchronous decoders, one specifically taking the segmental nature into account, yielding further improvements. Using time-synchronous decoding for segmental models is novel and a step towards streaming applications. Our experiments show the importance of a length model to predict the segment boundaries. The final best segmental-attention model using segmental decoding performs better than global-attention, in contrast to other monotonic attention approaches in the literature. Further, we observe that the segmental model generalizes much better to long sequences of up to several minutes.
Albert Zeyer, Robin Schmitt, Wei Zhou 0043, Ralf Schlüter, Hermann Ney
SLT4
2021 Comparing the Benefit of Synthetic Training Data for Various Automatic Speech Recognition Architectures
abstract
Recent publications on automatic-speech-recognition (ASR) have a strong focus on attention encoder-decoder (AED) architectures which tend to suffer from over-fitting in low resource scenarios. One solution to tackle this issue is to generate synthetic data with a trained text-to-speech system (TTS) if additional text is available. This was successfully applied in many publications with AED systems, but only very limited in the context of other ASR architectures. We investigate the effect of varying pre-processing, the speaker embedding and input encoding of the TTS system w.r.t. the effectiveness of the synthesized data for AED-ASR training. Additionally, we also consider internal language model subtraction for the first time, resulting in up to 38% relative improvement. We compare the AED results to a state-of-the-art hybrid ASR system, a monophone based system using connectionist-temporal-classification (CTC) and a monotonic transducer based system. We show that for the later systems the addition of synthetic data has no relevant effect, but they still outperform the AED systems on LibriSpeech-100h. We achieve a final word-error-rate of 3.3%/10.0% with a hybrid system on the clean/noisy test-sets, surpassing any previous state-of-the-art systems on Librispeech-100h that do not include unlabeled audio data.
Nick Rossenbach, Mohammad Zeineldeen, Benedikt Hilmes, Ralf Schlüter, Hermann Ney
ASRU4
2021 On Architectures and Training for Raw Waveform Feature Extraction in ASR
abstract
With the success of neural network based modeling in auto-matic speech recognition (ASR), many studies investigated acoustic modeling and learning of feature extractors directly based on the raw waveform. Recently, one line of research has focused on unsupervised pre-training of feature extractors on audio-only data to improve downstream ASR performance. In this work, we investigate the usefulness of one of these front-end frameworks, namely wav2vec, in a setting without additional untranscribed data for hybrid ASR systems. We compare this framework both to the manually defined stan-dard Gammatone feature set, as well as to features extracted as part of the acoustic model of an ASR system trained su-pervised. We study the benefits of using the pre-trained feature extractor and explore how to additionally exploit an ex-isting acoustic model trained with different features. Finally, we systematically examine combinations of the described features in order to further advance the performance.
Peter Vieting, Christoph Lüscher, Wilfried Michel, Ralf Schlüter, Hermann Ney
ASRU4
2021 Phoneme Based Neural Transducer for Large Vocabulary Speech Recognition
abstract
To join the advantages of classical and end-to-end approaches for speech recognition, we present a simple, novel and competitive approach for phoneme-based neural transducer modeling. Different alignment label topologies are compared and word-end-based phoneme label augmentation is proposed to improve performance. Utilizing the local dependency of phonemes, we adopt a simplified neural network structure and a straightforward integration with the external word-level language model to preserve the consistency of seq-to-seq modeling. We also present a simple, stable and efficient training procedure using frame-wise cross-entropy loss. A phonetic context size of one is shown to be sufficient for the best performance. A simplified scheduled sampling approach is applied for further improvement and different decoding approaches are briefly compared. The overall performance of our best model is comparable to state-of-the-art (SOTA) results for the TED-LIUM Release 2 and Switchboard corpora.
Wei Zhou 0043, Simon Berger, Ralf Schlüter, Hermann Ney
ICASSP3
2021 On Sampling-Based Training Criteria for Neural Language Modeling
abstract
As the vocabulary size of modern word-based language models becomes ever larger, many sampling-based training criteria are proposed and investigated.The essence of these sampling methods is that the softmax-related traversal over the entire vocabulary can be simplified, giving speedups compared to the baseline.A problem we notice about the current landscape of such sampling methods is the lack of a systematic comparison and some myths about preferring one over another.In this work, we consider Monte Carlo sampling, importance sampling, a novel method we call compensated partial summation, and noise contrastive estimation.Linking back to the three traditional criteria, namely mean squared error, binary cross-entropy, and crossentropy, we derive the theoretical solutions to the training problems.Contrary to some common belief, we show that all these sampling methods can perform equally well, as long as we correct for the intended class posterior probabilities.Experimental results in language modeling and automatic speech recognition on Switchboard and LibriSpeech support our claim, with all sampling-based methods showing similar perplexities and word error rates while giving the expected speedups.
Yingbo Gao, David Thulke, Alexander Gerstenberger, Khoa Viet Tran, Ralf Schlüter, Hermann Ney
Interspeech5
2021 The Impact of ASR on the Automatic Analysis of Linguistic Complexity and Sophistication in Spontaneous L2 Speech
abstract
In recent years, automated approaches to assessing linguistic complexity in second language (L2) writing have made significant progress in gauging learner performance, predicting human ratings of the quality of learner productions, and benchmarking L2 development.In contrast, there is comparatively little work in the area of speaking, particularly with respect to fully automated approaches to assessing L2 spontaneous speech.While the importance of a well-performing ASR system is widely recognized, little research has been conducted to investigate the impact of its performance on subsequent automatic text analysis.In this paper, we focus on this issue and examine the impact of using a state-of-the-art ASR system for subsequent automatic analysis of linguistic complexity in spontaneously produced L2 speech.A set of 30 selected measures were considered, falling into four categories: syntactic, lexical, n-gram frequency, and information-theoretic measures.The agreement between the scores for these measures obtained on the basis of ASR-generated vs. manual transcriptions was determined through correlation analysis.A more differential effect of ASR performance on specific types of complexity measures when controlling for task type effects is also presented.
Yu Qiao 0005, Wei Zhou 0043, Elma Kerz, Ralf Schlüter
Interspeech4
2021 Investigating Methods to Improve Language Model Integration for Attention-Based Encoder-Decoder ASR Models
abstract
Attention-based encoder-decoder (AED) models learn an implicit internal language model (ILM) from the training transcriptions. The integration with an external LM trained on much more unpaired text usually leads to better performance. A Bayesian interpretation as in the hybrid autoregressive transducer (HAT) suggests dividing by the prior of the discriminative acoustic model, which corresponds to this implicit LM, similarly as in the hybrid hidden Markov model approach. The implicit LM cannot be calculated efficiently in general and it is yet unclear what are the best methods to estimate it. In this work, we compare different approaches from the literature and propose several novel methods to estimate the ILM directly from the AED model. Our proposed methods outperform all previous approaches. We also investigate other methods to suppress the ILM mainly by decreasing the capacity of the AED model, limiting the label context, and also by training the AED model together with a pre-existing LM.
Mohammad Zeineldeen, Aleksandr Glushko, Wilfried Michel, Albert Zeyer, Ralf Schlüter, Hermann Ney
Interspeech5
2021 Librispeech Transducer Model with Internal Language Model Prior Correction
abstract
We present our transducer model on Librispeech. We study variants to include an external language model (LM) with shallow fusion and subtract an estimated internal LM. This is justified by a Bayesian interpretation where the transducer model prior is given by the estimated internal LM. The subtraction of the internal LM gives us over 14% relative improvement over normal shallow fusion. Our transducer has a separate probability distribution for the non-blank labels which allows for easier combination with the external LM, and easier estimation of the internal LM. We additionally take care of including the end-of-sentence (EOS) probability of the external LM in the last blank probability which further improves the performance. All our code and setups are published.
Albert Zeyer, André Merboldt, Wilfried Michel, Ralf Schlüter, Hermann Ney
Interspeech4
2021 Equivalence of Segmental and Neural Transducer Modeling: A Proof of Concept
abstract
With the advent of direct models in automatic speech recognition (ASR), the formerly prevalent frame-wise acoustic modeling based on hidden Markov models (HMM) diversified into a number of modeling architectures like encoder-decoder attention models, transducer models and segmental models (direct HMM).While transducer models stay with a frame-level model definition, segmental models are defined on the level of label segments directly.While (soft-)attention-based models avoid explicit alignment, transducer and segmental approach internally do model alignment, either by segment hypotheses or, more implicitly, by emitting so-called blank symbols.In this work, we prove that the widely used class of RNN-Transducer models and segmental models (direct HMM) are equivalent and therefore show equal modeling power.It is shown that blank probabilities translate into segment length probabilities and vice versa.In addition, we provide initial experiments investigating decoding and beam-pruning, comparing time-synchronous and label-/segment-synchronous search strategies and their properties using the same underlying model.
Wei Zhou 0043, Albert Zeyer, André Merboldt, Ralf Schlüter, Hermann Ney
Interspeech4
2021 Acoustic Data-Driven Subword Modeling for End-to-End Speech Recognition
abstract
Subword units are commonly used for end-to-end automatic speech recognition (ASR), while a fully acoustic-oriented subword modeling approach is somewhat missing.We propose an acoustic data-driven subword modeling (ADSM) approach that adapts the advantages of several text-based and acousticbased subword methods into one pipeline.With a fully acousticoriented label design and learning process, ADSM produces acoustic-structured subword units and acoustic-matched target sequence for further ASR training.The obtained ADSM labels are evaluated with different end-to-end ASR approaches including CTC, RNN-Transducer and attention models.Experiments on the LibriSpeech corpus show that ADSM clearly outperforms both byte pair encoding (BPE) and pronunciationassisted subword modeling (PASM) in all cases.Detailed analysis shows that ADSM achieves acoustically more logical word segmentation and more balanced sequence length, and thus, is suitable for both time-synchronous and label-synchronous models.We also briefly describe how to apply acoustic-based subword regularization and unseen text segmentation using ADSM.
Wei Zhou 0043, Mohammad Zeineldeen, Zuoyun Zheng, Ralf Schlüter, Hermann Ney
Interspeech4
2021 Tight Integrated End-to-End Training for Cascaded Speech Translation
abstract
A cascaded speech translation model relies on discrete and non-differentiable transcription, which provides a supervision signal from the source side and helps the transformation between source speech and target text. Such modeling suffers from error propagation between ASR and MT models. Direct speech translation is an alternative method to avoid error propagation; however, its performance is often behind the cascade system. To use an intermediate representation and preserve the end-to-end trainability, previous studies have proposed using two-stage models by passing the hidden vectors of the recognizer into the decoder of the MT model and ignoring the MT encoder. This work explores the feasibility of collapsing the entire cascade components into a single end-to-end trainable model by optimizing all parameters of ASR and MT models jointly without ignoring any learned parameters. It is a tightly integrated method that passes renormalized source word posterior distributions as a soft decision instead of one-hot vectors and enables backpropagation. Therefore, it provides both transcriptions and translations and achieves strong consistency between them. Our experiments on four tasks with different data scenarios show that the model outperforms cascade models up to 1.8% in BLEU and 2.0% in TER and is superior compared to direct models.
Parnia Bahar, Tobias Bieschke, Ralf Schlüter, Hermann Ney
SLT3
2020 Exploring A Zero-Order Direct Hmm Based on Latent Attention for Automatic Speech Recognition
abstract
In this paper, we study a simple yet elegant latent variable attention model for automatic speech recognition (ASR) which enables an integration of attention sequence modeling into the direct hidden Markov model (HMM) concept. We use a sequence of hidden variables that establishes a mapping from output labels to input frames. Inspired by the direct HMM model, we assume a decomposition of the label sequence posterior into emission and transition probabilities using zero-order assumption and incorporate both Transformer and LSTM attention models into it. The method keeps the explicit alignment as part of the stochastic model and combines the ease of the end-to-end training of the attention model as well as an efficient and simple beam search. To study the effect of the latent model, we qualitatively analyze the alignment behavior of the different approaches. Our experiments on three ASR tasks show promising results in WER with more focused alignments in comparison to the attention models.
Parnia Bahar, Nikita Makarov, Albert Zeyer, Ralf Schlüter, Hermann Ney
ICASSP4
2020 A Comprehensive Study of Residual CNNS for Acoustic Modeling in ASR
abstract
Long short-term memory (LSTM) networks are the dominant architecture for large vocabulary continuous speech recognition (LVCSR) acoustic modeling due to their good performance. However, LSTMs are hard to tune and computationally expensive. To build a system with lower computational costs and which allows online streaming applications, we explore convolutional neural networks (CNN). To the best of our knowledge there is no overview on CNN hyper-parameter tuning for LVCSR in the literature, so we present our results explicitly. Apart from recognition performance, we focus on the training and evaluation speed and provide a time-efficient setup for CNNs. We faced an overfitting problem in training and solved it with data augmentation, namely SpecAugment. The system achieves results competitive with the top LSTM results. We significantly increased the speed of CNN in training and decoding approaching the speed of the offline LSTM.
Vitalii Bozheniuk, Albert Zeyer, Ralf Schlüter, Hermann Ney
ICASSP3
2020 How Much Self-Attention Do We Need? Trading Attention for Feed-Forward Layers
abstract
We propose simple architectural modifications in the standard Transformer with the goal to reduce its total state size (defined as the number of self-attention layers times the sum of the key and value dimensions, times position) without loss of performance. Large scale Transformer language models have been empirically proved to give very good performance. However, scaling up results in a model that needs to store large states at evaluation time. This can increase the memory requirement dramatically for search e.g., in speech recognition (first pass decoding, lattice rescoring, or shallow fusion). In order to efficiently increase the model capacity without increasing the state size, we replace the single-layer feed-forward module in the Transformer layer by a deeper network, and decrease the total number of layers. In addition, we also evaluate the effect of key-value tying which directly divides the state size in half. On TED-LIUM 2, we obtain a model of state size 4 times smaller than the standard Transformer, with only 2% relative loss in terms of perplexity, which makes the deployment of Transformer language models more convenient.
Kazuki Irie, Alexander Gerstenberger, Ralf Schlüter, Hermann Ney
ICASSP3
2020 Frame-Level MMI as A Sequence Discriminative Training Criterion for LVCSR
abstract
In this work we present frame-level maximum mutual information (MMI) as a novel sequence discriminative training criterion for hybrid HMM-DNN acoustic models. Compared to the standard, sequence-level MMI criterion we show that frame-level MMI has increased robustness towards missing cross-entropy (CE) smoothing and can converge even without interpolation. Using model free optimization, we show that in the asymptotic case of an infinite amount of training data models trained using this criterion are equal to the true class posterior distribution, whereas training using the state-level minimum Bayes risk (sMBR) criterion leads to a distorted function of the true class posterior distribution. This analytical result is backed by experimental evidence. We further propose a generalized class of training criteria, that continuously interpolates between frame-level MMI and sMBR criterion.
Wilfried Michel, Ralf Schlüter, Hermann Ney
ICASSP2
2020 Generating Synthetic Audio Data for Attention-Based Speech Recognition Systems
abstract
Recent advances in text-to-speech (TTS) led to the development of flexible multi-speaker end-to-end TTS systems. We extend state-of-the-art attention-based automatic speech recognition (ASR) systems with synthetic audio generated by a TTS system trained only on the ASR corpora itself. ASR and TTS systems are built separately to show that text-only data can be used to enhance existing end-to-end ASR systems without the necessity of parameter or architecture changes. We compare our method with language model integration of the same text data and with simple data augmentation methods like SpecAugment and show that performance improvements are mostly independent. We achieve improvements of up to 33% relative in word-error-rate (WER) over a strong baseline with data-augmentation in a low-resource environment (LibriSpeech-100h), closing the gap to a comparable oracle experiment by more than 50%. We also show improvements of up to 5% relative WER over our most recent ASR baseline on LibriSpeech-960h.
Nick Rossenbach, Albert Zeyer, Ralf Schlüter, Hermann Ney
ICASSP3
2020 Layer-Normalized LSTM for Hybrid-Hmm and End-To-End ASR
abstract
Training deep neural networks is often challenging in terms of training stability. It often requires careful hyperparameter tuning or a pretraining scheme to converge. Layer normalization (LN) has shown to be a crucial ingredient in training deep encoder-decoder models. We explore various LN long short-term memory (LSTM) recurrent neural networks (RNN) variants by applying LN to different parts of the internal recurrency of LSTMs. There is no previous work that investigates this. We carry out experiments on the Switchboard 300h task for both hybrid and end-to-end ASR models and we show that LN improves the final word error rate (WER), the stability during training, allows to train even deeper models, requires less hyperparameter tuning, and works well even without pre-training. We find that applying LN to both forward and recurrent inputs globally, which we denoted by Global Joined Norm variant, gives a 10% relative improvement in WER.
Mohammad Zeineldeen, Albert Zeyer, Ralf Schlüter, Hermann Ney
ICASSP3
2020 The Rwth Asr System for Ted-Lium Release 2: Improving Hybrid Hmm With Specaugment
abstract
We present a complete training pipeline to build a state-of-the-art hybrid HMM-based ASR system on the 2nd release of the TED-LIUM corpus. Data augmentation using SpecAugment is successfully applied to improve performance on top of our best SAT model using i-vectors. By investigating the effect of different maskings, we achieve improvements from SpecAugment on hybrid HMM models without increasing model size and training time. A subsequent sMBR training is applied to fine-tune the final acoustic model, and both LSTM and Transformer language models are trained and evaluated. Our best system achieves a 5.6% WER on the test set, which outperforms the previous state-of-the-art by 27% relative.
Wei Zhou 0043, Wilfried Michel, Kazuki Irie, Markus Kitza, Ralf Schlüter, Hermann Ney
ICASSP5
2020 Full-Sum Decoding for Hybrid Hmm Based Speech Recognition Using LSTM Language Model
abstract
In hybrid HMM based speech recognition, LSTM language models have been widely applied and achieved large improvements. The theoretical capability of modeling any unlimited context suggests that no recombination should be applied in decoding. This motivates to reconsider full summation over the HMM-state sequences instead of Viterbi approximation in decoding. We explore the potential gain from more accurate probabilities in terms of decision making and apply the full-sum decoding with a modified prefix-tree search framework. The proposed full-sum decoding is evaluated on both Switchboard and Librispeech corpora. Different models using CE and sMBR training criteria are used. Additionally, both MAP and confusion network decoding as approximated variants of general Bayes decision rule are evaluated. Consistent improvements over strong baselines are achieved in almost all cases without extra cost. We also discuss tuning effort, efficiency and some limitations of full-sum decoding.
Wei Zhou 0043, Ralf Schlüter, Hermann Ney
ICASSP2
2020 LVCSR with Transformer Language Models
Eugen Beck, Ralf Schlüter, Hermann Ney
INTERSPEECH2
2020 Investigation of Large-Margin Softmax in Neural Language Modeling
Jingjing Huo, Yingbo Gao, Weiyue Wang 0001, Ralf Schlüter, Hermann Ney
INTERSPEECH4
2020 Early Stage LM Integration Using Local and Global Log-Linear Combination
abstract
Sequence-to-sequence models with an implicit alignment mechanism (e.g. attention) are closing the performance gap towards traditional hybrid hidden Markov models (HMM) for the task of automatic speech recognition. One important factor to improve word error rate in both cases is the use of an external language model (LM) trained on large text-only corpora. Language model integration is straightforward with the clear separation of acoustic model and language model in classical HMM-based modeling. In contrast, multiple integration schemes have been proposed for attention models. In this work, we present a novel method for language model integration into implicit-alignment based sequence-to-sequence models. Log-linear model combination of acoustic and language model is performed with a per-token renormalization. This allows us to compute the full normalization term efficiently both in training and in testing. This is compared to a global renormalization scheme which is equivalent to applying shallow fusion in training. The proposed methods show good improvements over standard model combination (shallow fusion) on our state-of-the-art Librispeech system. Furthermore, the improvements are persistent even if the LM is exchanged for a more powerful one after training.
Wilfried Michel, Ralf Schlüter, Hermann Ney
INTERSPEECH2
2020 Context-Dependent Acoustic Modeling Without Explicit Phone Clustering
abstract
Phoneme-based acoustic modeling of large vocabulary automatic speech recognition takes advantage of phoneme context. The large number of context-dependent (CD) phonemes and their highly varying statistics require tying or smoothing to enable robust training. Usually, classification and regression trees are used for phonetic clustering, which is standard in hidden Markov model (HMM)-based systems. However, this solution introduces a secondary training objective and does not allow for end-to-end training. In this work, we address a direct phonetic context modeling for the hybrid deep neural network (DNN)/HMM, that does not build on any phone clustering algorithm for the determination of the HMM state inventory. By performing different decompositions of the joint probability of the center phoneme state and its left and right contexts, we obtain a factorized network consisting of different components, trained jointly. Moreover, the representation of the phonetic context for the network relies on phoneme embeddings. The recognition accuracy of our proposed models on the Switchboard task is comparable and outperforms slightly the hybrid model using the standard state-tying decision trees.
Tina Raissi, Eugen Beck, Ralf Schlüter, Hermann Ney
INTERSPEECH3
2020 A New Training Pipeline for an Improved Neural Transducer
abstract
The RNN transducer is a promising end-to-end model candidate. We compare the original training criterion with the full marginalization over all alignments, to the commonly used maximum approximation, which simplifies, improves and speeds up our training. We also generalize from the original neural network model and study more powerful models, made possible due to the maximum approximation. We further generalize the output label topology to cover RNN-T, RNA and CTC. We perform several studies among all these aspects, including a study on the effect of external alignments. We find that the transducer model generalizes much better on longer sequences than the attention model. Our final transducer model outperforms our attention model on Switchboard 300h by over 6% relative WER.
Albert Zeyer, André Merboldt, Ralf Schlüter, Hermann Ney
INTERSPEECH3
2020 Robust Beam Search for Encoder-Decoder Attention Based Speech Recognition Without Length Bias
abstract
As one popular modeling approach for end-to-end speech recognition, attention-based encoder-decoder models are known to suffer the length bias and corresponding beam problem.Different approaches have been applied in simple beam search to ease the problem, most of which are heuristic-based and require considerable tuning.We show that heuristics are not proper modeling refinement, which results in severe performance degradation with largely increased beam sizes.We propose a novel beam search derived from reinterpreting the sequence posterior with an explicit length modeling.By applying the reinterpreted probability together with beam pruning, the obtained final probability leads to a robust model modification, which allows reliable comparison among output sequences of different lengths.Experimental verification on the LibriSpeech corpus shows that the proposed approach solves the length bias problem without heuristics or additional tuning effort.It provides robust decision making and consistently good performance under both small and very large beam sizes.Compared with the best results of the heuristic baseline, the proposed approach achieves the same WER on the 'clean' sets and 4% relative improvement on the 'other' sets.We also show that it is more efficient with the additional derived early stopping criterion.
Wei Zhou 0043, Ralf Schlüter, Hermann Ney
INTERSPEECH2
2019 Training Language Models for Long-Span Cross-Sentence Evaluation
abstract
While recurrent neural networks can motivate cross-sentence language modeling and its application to automatic speech recognition (ASR), corresponding modifications of the training method for that end are rarely discussed. In fact, even more generally, the impact of training sequence construction strategy in language modeling for different evaluation conditions is typically ignored. In this work, we revisit this basic but fundamental question. We train language models based on long short-term memory recurrent neural networks and Transformers using various types of training sequences and study their robustness with respect to different evaluation modes. Our experiments on 300h Switchboard and Quaero English datasets show that models trained with back-propagation over sequences consisting of concatenation of multiple sentences with state carry-over across sequences effectively outperform those trained with the sentence-level training, both in terms of perplexity and word error rates for cross-utterance ASR.
Kazuki Irie, Albert Zeyer, Ralf Schlüter, Hermann Ney
ASRU3
2019 A Comparison of Transformer and LSTM Encoder Decoder Models for ASR
abstract
We present competitive results using a Transformer encoder-decoder-attention model for end-to-end speech recognition needing less training time compared to a similarly performing LSTM model. We observe that the Transformer training is in general more stable compared to the LSTM, although it also seems to overfit more, and thus shows more problems with generalization. We also find that two initial LSTM layers in the Transformer encoder provide a much better positional encoding. Data-augmentation, a variant of SpecAugment, helps to improve both the Transformer by 33% and the LSTM by 15% relative. We analyze several pretraining and scheduling schemes, which is crucial for both the Transformer and the LSTM models. We improve our LSTM model by additional convolutional layers. We perform our experiments on Lib-riSpeech 1000h, Switchboard 300h and TED-LIUM-v2 200h, and we show state-of-the-art performance on TED-LIUM-v2 for attention based end-to-end models. We deliberately limit the training on LibriSpeech to 12.5 epochs of the training data for comparisons, to keep the results of practical interest, although we show that longer training time still improves more. We publish all the code and setups to run our experiments.
Albert Zeyer, Parnia Bahar, Kazuki Irie, Ralf Schlüter, Hermann Ney
ASRU4
2019 On Using 2D Sequence-to-sequence Models for Speech Recognition
abstract
Attention-based sequence-to-sequence models have shown promising results in automatic speech recognition. Using these architectures, one-dimensional input and output sequences are related by an attention approach, thereby replacing more explicit alignment processes, like in classical HMM-based modeling. In contrast, here we apply a novel two-dimensional long short-term memory (2DLSTM) architecture to directly model the input/output relation between audio/feature vector sequences and word sequences. The proposed model is an alternative model such that instead of using any type of attention components, we apply a 2DLSTM layer to assimilate the context from both input observations and output transcriptions. The experimental evaluation on the Switchboard 300h automatic speech recognition task shows word error rates for the 2DLSTM model that are competitive to end-to-end attention-based model.
Parnia Bahar, Albert Zeyer, Ralf Schlüter, Hermann Ney
ICASSP3
2019 Investigation into Joint Optimization of Single Channel Speech Enhancement and Acoustic Modeling for Robust ASR
abstract
This paper investigates the joint optimization of single channel speech enhancement and the acoustic model of a hybrid DNN-HMM system for noise robust ASR. Two enhancement methods are investigated. A masking of the noisy speech signal with a speech mask estimated by a DNN based mask estimator, as well as a parametric Wiener filter employing a DNN based noise estimator and a DNN based frame wise estimation of the filter parameters. Those components are jointly optimized with the acoustic model of the ASR system. It is shown that the Wiener filter approach can be used to improve the performance of a state-of-the-art single-channel ASR system on the single channel track of the CHiME-4 data, where the WER of the real evaluation set is reduced from 11.6 % to 10.5 %.
Tobias Menne, Ralf Schlüter, Hermann Ney
ICASSP2
2019 Language Modeling with Deep Transformers
abstract
We explore deep autoregressive Transformer models in language modeling for speech recognition. We focus on two aspects. First, we revisit Transformer model configurations specifically for language modeling. We show that well configured Transformer models outperform our baseline models based on the shallow stack of LSTM recurrent neural network layers. We carry out experiments on the open-source LibriSpeech 960hr task, for both 200K vocabulary word-level and 10K byte-pair encoding subword-level language modeling. We apply our word-level models to conventional hybrid speech recognition by lattice rescoring, and the subword-level models to attention based encoder-decoder models by shallow fusion. Second, we show that deep Transformer language models do not require positional encoding. The positional encoding is an essential augmentation for the self-attention mechanism which is invariant to sequence ordering. However, in autoregressive setup, as is the case for language modeling, the amount of information increases along the position dimension, which is a positional signal by its own. The analysis of attention weights shows that deep autoregressive self-attention models can automatically make use of such positional information. We find that removing the positional encoding even slightly improves the performance of these models.
Kazuki Irie, Albert Zeyer, Ralf Schlüter, Hermann Ney
INTERSPEECH3
2019 Cumulative Adaptation for BLSTM Acoustic Models
abstract
This paper addresses the robust speech recognition problem as an adaptation task. Specifically, we investigate the cumulative application of adaptation methods. A bidirectional Long Short-Term Memory (BLSTM) based neural network, capable of learning temporal relationships and translation invariant representations, is used for robust acoustic modelling. Further, i-vectors were used as an input to the neural network to perform instantaneous speaker and environment adaptation, providing 8\% relative improvement in word error rate on the NIST Hub5 2000 evaluation test set. By enhancing the first-pass i-vector based adaptation with a second-pass adaptation using speaker and environment dependent transformations within the network, a further relative improvement of 5\% in word error rate was achieved. We have reevaluated the features used to estimate i-vectors and their normalization to achieve the best performance in a modern large scale automatic speech recognition system.
Markus Kitza, Pavel Golik, Ralf Schlüter, Hermann Ney
INTERSPEECH3
2019 RWTH ASR Systems for LibriSpeech: Hybrid vs Attention
abstract
We present state-of-the-art automatic speech recognition (ASR) systems employing a standard hybrid DNN/HMM architecture compared to an attention-based encoder-decoder design for the LibriSpeech task. Detailed descriptions of the system development, including model design, pretraining schemes, training schedules, and optimization approaches are provided for both system architectures. Both hybrid DNN/HMM and attention-based systems employ bi-directional LSTMs for acoustic modeling/encoding. For language modeling, we employ both LSTM and Transformer based architectures. All our systems are built using RWTHs open-source toolkits RASR and RETURNN. To the best knowledge of the authors, the results obtained when training on the full LibriSpeech training set, are the best published currently, both for the hybrid DNN/HMM and the attention-based systems. Our single hybrid system even outperforms previous results obtained from combining eight single systems. Our comparison shows that on the LibriSpeech 960h task, the hybrid DNN/HMM system outperforms the attention-based system by 15% relative on the clean and 40% relative on the other test sets in terms of word error rate. Moreover, experiments on a reduced 100h-subset of the LibriSpeech training corpus even show a more pronounced margin between the hybrid DNN/HMM and attention-based architectures.
Christoph Lüscher, Eugen Beck, Kazuki Irie, Markus Kitza, Wilfried Michel, Albert Zeyer, Ralf Schlüter, Hermann Ney
INTERSPEECH7
2019 Analysis of Deep Clustering as Preprocessing for Automatic Speech Recognition of Sparsely Overlapping Speech
abstract
Significant performance degradation of automatic speech recognition (ASR) systems is observed when the audio signal contains cross-talk. One of the recently proposed approaches to solve the problem of multi-speaker ASR is the deep clustering (DPCL) approach. Combining DPCL with a state-of-the-art hybrid acoustic model, we obtain a word error rate (WER) of 16.5 % on the commonly used wsj0-2mix dataset, which is the best performance reported thus far to the best of our knowledge. The wsj0-2mix dataset contains simulated cross-talk where the speech of multiple speakers overlaps for almost the entire utterance. In a more realistic ASR scenario the audio signal contains significant portions of single-speaker speech and only part of the signal contains speech of multiple competing speakers. This paper investigates obstacles of applying DPCL as a preprocessing method for ASR in such a scenario of sparsely overlapping speech. To this end we present a data simulation approach, closely related to the wsj0-2mix dataset, generating sparsely overlapping speech datasets of arbitrary overlap ratio. The analysis of applying DPCL to sparsely overlapping speech is an important interim step between the fully overlapping datasets like wsj0-2mix and more realistic ASR datasets, such as CHiME-5 or AMI.
Tobias Menne, Ilya Sklyar, Ralf Schlüter, Hermann Ney
INTERSPEECH3
2019 An Analysis of Local Monotonic Attention Variants
André Merboldt, Albert Zeyer, Ralf Schlüter, Hermann Ney
INTERSPEECH3
2019 Comparison of Lattice-Free and Lattice-Based Sequence Discriminative Training Criteria for LVCSR
abstract
Sequence discriminative training criteria have long been a standard tool in automatic speech recognition for improving the performance of acoustic models over their maximum likelihood / cross entropy trained counterparts. While previously a lattice approximation of the search space has been necessary to reduce computational complexity, recently proposed methods use other approximations to dispense of the need for the computationally expensive step of separate lattice creation. In this work we present a memory efficient implementation of the forward-backward computation that allows us to use uni-gram word-level language models in the denominator calculation while still doing a full summation on GPU. This allows for a direct comparison of lattice-based and lattice-free sequence discriminative training criteria such as MMI and sMBR, both using the same language model during training. We compared performance, speed of convergence, and stability on large vocabulary continuous speech recognition tasks like Switchboard and Quaero. We found that silence modeling seriously impacts the performance in the lattice-free case and needs special treatment. In our experiments lattice-free MMI comes on par with its lattice-based counterpart. Lattice-based sMBR still outperforms all lattice-free training criteria.
Wilfried Michel, Ralf Schlüter, Hermann Ney
INTERSPEECH2
2019 Rescoring Keyword Search Confidence Estimates with Graph-Based Re-Ranking Using Acoustic Word Embeddings
Anna Piunova, Eugen Beck, Ralf Schlüter, Hermann Ney
INTERSPEECH3
2019 Survey Talk: Modeling in Automatic Speech Recognition: Beyond Hidden Markov Models
Ralf Schlüter
INTERSPEECH1
2019 Upper and Lower Tight Error Bounds for Feature Omission with an Extension to Context Reduction
abstract
In this work, fundamental analytic results in the form of error bounds are presented that quantify the effect of feature omission and selection for pattern classification in general, as well as the effect of context reduction in string classification, like automatic speech recognition, printed/handwritten character recognition, or statistical machine translation. A general simulation framework is introduced that supports discovery and proof of error bounds, which lead to the error bounds presented here. Initially derived tight lower and upper bounds for feature omission are generalized to feature selection, followed by another extension to context reduction of string class priors (aka language models) in string classification. For string classification, the quantitative effect of string class prior context reduction on symbol-level Bayes error is presented. The tightness of the original feature omission bounds seems lost in this case, as further simulations indicate. However, combining both feature omission andcontext reduction, the tightness of the bounds is retained. A central result of this work is the proof of the existence, and the amount of a statistical threshold w.r.t. the introduction of additional features in general pattern classification, or the increase of context in string classification beyond which a decrease in Bayes error is guaranteed.
Ralf Schlüter, Eugen Beck, Hermann Ney
IEEE Trans. Pattern Anal. Mach. Intell.1
2019 Training of reduced-rank linear transformations for multi-layer polynomial acoustic features for speech recognition
Muhammad Ali Tahir, Heyun Huang, Albert Zeyer, Ralf Schlüter, Hermann Ney
Speech Commun.4
2018 Prediction of LSTM-RNN Full Context States as a Subtask for N-Gram Feedforward Language Models
abstract
Long short-term memory (LSTM) recurrent neural network language models compress the full context of variable lengths into a fixed size vector. In this work, we investigate the task of predicting the LSTM hidden representation of the full context from a truncated n-gram context as a subtask for training an n-gram feedforward language model. Since this approach is a form of knowledge distillation, we compare two methods. First, we investigate the standard transfer based on the Kullback-Leibler divergence of the output distribution of the feedforward model from that of the LSTM. Second, we minimize the mean squared error between the hidden state of the LSTM and that of the n-gram feedforward model. We carry out experiments on different subsets of the Switchboard speech recognition dataset for feedforward models with a short (5-gram) and a medium (10-gram) context length. We show that we get improvements in perplexity and word error rate of up to 8% and 4% relative for the medium model, while the improvements are only marginal for the short model.
Kazuki Irie, Zhihong Lei, Ralf Schlüter, Hermann Ney
ICASSP3
2018 Acoustic Modeling of Speech Waveform Based on Multi-Resolution, Neural Network Signal Processing
abstract
Recently, several papers have demonstrated that neural networks (NN) are able to perform the feature extraction as part of the acoustic model. Motivated by the Gammatone feature extraction pipeline, in this paper we extend the waveform based NN model by a second level of time-convolutional element. The proposed extension generalizes the envelope extraction block, and allows the model to learn multi-resolutional representations. Automatic speech recognition (ASR) experiments show significant word error rate reduction over our previous best acoustic model trained in the signal domain directly. Although we use only 250 hours of speech, the data-driven NN based speech signal processing performs nearly equally to traditional handcrafted feature extractors. In additional experiments, we also test segment-level feature normalization techniques on NN derived features, which improve the results further. However, the porting of speech representations derived by a feed-forward NN to a LSTM back-end model indicates much less robustness of the NN front-end compared to the standard feature extractors. Analysis of the weights in the proposed new layer reveals that the NN prefers both multi-resolution and modulation spectrum representations.
Zoltán Tüske, Ralf Schlüter, Hermann Ney
ICASSP2
2018 Segmental Encoder-Decoder Models for Large Vocabulary Automatic Speech Recognition
abstract
It has been known for a long time that the classic Hidden-Markov-Model (HMM) derivation for speech recognition contains assumptions such as independence of observation vectors and weak duration modeling that are practical but unrealistic.When using the hybrid approach this is amplified by trying to fit a discriminative model into a generative one.Hidden Conditional Random Fields (CRFs) and segmental models (e.g.Semi-Markov CRFs / Segmental CRFs) have been proposed as an alternative, but for a long time have failed to get traction until recently.In this paper we explore different length modeling approaches for segmental models, their relation to attention-based systems.Furthermore we show experimental results on a handwriting recognition task and to the best of our knowledge the first reported results on the Switchboard 300h speech recognition corpus using this approach.
Eugen Beck, Mirko Hannemann, Patrick Doetsch, Ralf Schlüter, Hermann Ney
INTERSPEECH4
2018 Investigation on Estimation of Sentence Probability by Combining Forward, Backward and Bi-directional LSTM-RNNs
Kazuki Irie, Zhihong Lei, Liuhui Deng, Ralf Schlüter, Hermann Ney
INTERSPEECH4
2018 Comparison of BLSTM-Layer-Specific Affine Transformations for Speaker Adaptation
Markus Kitza, Ralf Schlüter, Hermann Ney
INTERSPEECH2
2018 Investigation on LSTM Recurrent N-gram Language Models for Speech Recognition
abstract
Recurrent neural networks (NN) with long short-term memory (LSTM) are the current state of the art to model long term dependencies.However, recent studies indicate that NN language models (LM) need only limited length of history to achieve excellent performance.In this paper, we extend the previous investigation on LSTM network based n-gram modeling to the domain of automatic speech recognition (ASR).First, applying recent optimization techniques and up to 6-layer LSTM networks, we improve LM perplexities by nearly 50% relative compared to classic count models on three different domains.Then, we demonstrate by experimental results that perplexities improve significantly only up to 40-grams when limiting the LM history.Nevertheless, the ASR performance saturates already around 20-grams despite across sentence modeling.Analysis indicates that the performance gain of LSTM NNLM over count models results only partially from the longer context and cross sentence modeling capabilities.Using equal context, we show that deep 4-gram LSTM can significantly outperform large interpolated count models by performing the backing off and smoothing significantly better.This observation also underlines the decreasing importance to combine state-of-the-art deep NNLM with count based model.
Zoltán Tüske, Ralf Schlüter, Hermann Ney
INTERSPEECH2
2018 Improved Training of End-to-end Attention Models for Speech Recognition
abstract
Sequence-to-sequence attention-based models on subword units allow simple open-vocabulary end-to-end speech recognition. In this work, we show that such models can achieve competitive results on the Switchboard 300h and LibriSpeech 1000h tasks. In particular, we report the state-of-the-art word error rates (WER) of 3.54% on the dev-clean and 3.82% on the test-clean evaluation subsets of LibriSpeech. We introduce a new pretraining scheme by starting with a high time reduction factor and lowering it during training, which is crucial both for convergence and final performance. In some experiments, we also use an auxiliary CTC loss function to help the convergence. In addition, we train long short-term memory (LSTM) language models on subword units. By shallow fusion, we report up to 27% relative improvements in WER over the attention baseline without a language model.
Albert Zeyer, Kazuki Irie, Ralf Schlüter, Hermann Ney
INTERSPEECH3
2018 Speaker Adapted Beamforming for Multi-Channel Automatic Speech Recognition
abstract
This paper presents, in the context of multi-channel ASR, a method to adapt a mask based, statistically optimal beamforming approach to a speaker of interest. The beamforming vector of the statistically optimal beamformer is computed by utilizing speech and noise masks, which are estimated by a neural network. The proposed adaptation approach is based on the integration of the beamformer, which includes the mask estimation network, and the acoustic model of the ASR system. This allows for the propagation of the training error, from the acoustic modeling cost function, all the way through the beamforming operation and through the mask estimation network. By using the results of a first pass recognition and by keeping all other parameters fixed, the mask estimation network can therefore be fine tuned by retraining. Utterances of a speaker of interest can thus be used in a two pass approach, to optimize the beamforming for the speech characteristics of that specific speaker. It is shown that this approach improves the ASR performance of a state-of-the-art multi-channel ASR system on the CHiME-4 data. Furthermore the effect of the adaptation on the estimated speech masks is discussed.
Tobias Menne, Ralf Schlüter, Hermann Ney
SLT2
2017 Returnn: The RWTH extensible training framework for universal recurrent neural networks
abstract
In this work we release our extensible and easily configurable neural network training software. It provides a rich set of functional layers with a particular focus on efficient training of recurrent neural network topologies on multiple GPUs. The source of the software package is public and freely available for academic research purposes and can be used as a framework or as a standalone tool which supports a flexible configuration. The software allows to train state-of-the-art deep bidirectional long short-term memory (LSTM) models on both one dimensional data like speech or two dimensional data like handwritten text and was used to develop successful submission systems in several evaluation campaigns.
Patrick Doetsch, Albert Zeyer, Paul Voigtlaender, Ilia Kulikov, Ralf Schlüter, Hermann Ney
ICASSP5
2017 Investigations on byte-level convolutional neural networks for language modeling in low resource speech recognition
abstract
In this paper, we present an investigation on technical details of the byte-level convolutional layer which replaces the conventional linear word projection layer in the neural language model. In particular, we discuss and compare the effective filter configurations, pooling types and the use of bytes instead of characters. We carry out experiments on language packs released by the IARPA Babel project and measure the performance in terms of perplexity and word error rate. Introducing a convolutional layer consistently improves the results on all languages. Also, there is no degradation from using raw bytes instead of proper Unicode characters, even on syllabic alphabets like Amharic. In addition, we report improvements in word error rate from rescoring lattices and evaluate keyword search performance on several languages.
Kazuki Irie, Pavel Golik, Ralf Schlüter, Hermann Ney
ICASSP3
2017 Noisy objective functions based on the f-divergence
abstract
Dropout, the random dropping out of activations according to a specified rate, is a very simple but effective method to avoid over-fitting of deep neural networks to the training data.
Markus Nußbaum-Thom, Ralf Schlüter, Vaibhava Goel, Hermann Ney
ICASSP2
2017 A comprehensive study of deep bidirectional LSTM RNNS for acoustic modeling in speech recognition
abstract
Recent experiments show that deep bidirectional long short-term memory (BLSTM) recurrent neural network acoustic models outperform feedforward neural networks for automatic speech recognition (ASR). However, their training requires a lot of tuning and experience. In this work, we provide a comprehensive overview over various BLSTM training aspects and their interplay within ASR, which has been missing so far in the literature. We investigate on different variants of optimization methods, batching, truncated backpropagation, and regularization techniques such as dropout, and we study the effect of size and depth, training models of up to 10 layers. This includes a comparison of computation times vs. recognition performance. Furthermore, we introduce a pretraining scheme for LSTMs with layer-wise construction of the network showing good improvements especially for deep networks. The experimental analysis mainly was performed on the Quaero task, with additional results on Switchboard. The best BLSTM model gave a relative improvement in word error rate of over 15% compared to our best feed-forward baseline on our Quaero 50h task. All experiments were done using RETURNN and RASR, RWTH's extensible training framework for universal recurrent neural networks and ASR toolkit. The training configuration files are publicly available.
Albert Zeyer, Patrick Doetsch, Paul Voigtlaender, Ralf Schlüter, Hermann Ney
ICASSP4
2017 Faster sequence training
abstract
It has been shown that sequence-discriminative training can improve the performance for large vocabulary continuous speech recognition. Our main contribution is a novel method for reducing the computation time of any sort of sequence training while only slightly decreasing the overall performance. The method allows to parallelize the forward propagation through the network, the loss and loss gradient calculation which will provide a frame-wise error signal, and an independent forward and back propagation using that error signal. That last step can be calculated in a frame-wise manner and thus allows to use frame chunking to further improve the runtime. The loss calculation can itself be parallelized over many sequences. In addition to several experiments which outline the runtime gains, we also provide a convergence proof sketch. We extend on the research of sequence training of bidirectional long-short term memory ((B)LSTM) networks and provide an overview and comparison over different criteria. We have published all the code as part of our RETURNN and RASR framework including our training setup configurations.
Albert Zeyer, Ilia Kulikov, Ralf Schlüter, Hermann Ney
ICASSP3
2017 Parallel Neural Network Features for Improved Tandem Acoustic Modeling
abstract
The combination of acoustic models or features is a standard approach to exploit various knowledge sources.This paper investigates the concatenation of different bottleneck (BN) neural network (NN) outputs for tandem acoustic modeling.Thus, combination of NN features is performed via Gaussian mixture models (GMM).Complementarity between the NN feature representations is attained by using various network topologies: LSTM recurrent, feed-forward, and hierarchical, as well as different non-linearities: hyperbolic tangent, sigmoid, and rectified linear units.Speech recognition experiments are carried out on various tasks: telephone conversations, Skype calls, as well as broadcast news and conversations.Results indicate that LSTM based tandem approach is still competitive, and such tandem model can challenge comparable hybrid systems.The traditional steps of tandem modeling, speaker adaptive and sequence discriminative GMM training, improve the tandem results further.Furthermore, these "old-fashioned" steps remain applicable after the concatenation of multiple neural network feature streams.Exploiting the parallel processing of input feature streams, it is shown that 2-5% relative improvement could be achieved over the single best BN feature set.Finally, we also report results after neural network based language model rescoring and examine the system combination possibilities using such complex tandem models.
Zoltán Tüske, Wilfried Michel, Ralf Schlüter, Hermann Ney
INTERSPEECH3
2017 CTC in the Context of Generalized Full-Sum HMM Training
abstract
We formulate a generalized hybrid HMM-NN training procedure using the full-sum over the hidden state-sequence and identify CTC as a special case of it.We present an analysis of the alignment behavior of such a training procedure and explain the strong localization of label output behavior of full-sum training (also referred to as peaky or spiky behavior).We show how to avoid that behavior by using a state prior.We discuss the temporal decoupling between output label position/time-frame, and the corresponding evidence in the input observations when this is trained with BLSTM models.We also show a way how to overcome this by jointly training a FFNN.We implemented the Baum-Welch alignment algorithm in CUDA to be able to do fast soft realignments on GPU.We have published this code along with some of our experiments as part of RETURNN, RWTH's extensible training framework for universal recurrent neural networks.We finish with experimental validation of our study on WSJ and Switchboard.
Albert Zeyer, Eugen Beck, Ralf Schlüter, Hermann Ney
INTERSPEECH3
2016 Investigation on log-linear interpolation of multi-domain neural network language model
abstract
Inspired by the success of multi-task training in acoustic modeling, this paper investigates a new architecture for a multi-domain neural network based language model (NNLM). The proposed model has several shared hidden layers and domain-specific output layers. As will be shown, the log-linear interpolation of the multi-domain outputs and the optimization of interpolation weights fit naturally in the framework of NNLM. The resulting model can be expressed as a single NNLM. As an initial study of such an architecture, this paper focuses on deep feed-forward neural networks (DNNs). We also re-investigate the potential of long context up to 30-grams, and depth up to 5 hidden layers in DNN-LM. Our final feed-forward multidomain NNLM is trained on 3.1B running words across 11 domains for English broadcast news and conversations large vocabulary continuous speech recognition task. After log-linear interpolation and fine-tuning, we measured improvements in terms of perplexity and word error rate over the models trained on 50M running words of in-domain news resources. The final multi-domain feed-forward LM outperformed our previous best LSTM-RNN LM trained on the 50M in-domain corpus, even after linear interpolation with large count models.
Zoltán Tüske, Kazuki Irie, Ralf Schlüter, Hermann Ney
ICASSP3
2016 LSTM, GRU, Highway and a Bit of Attention: An Empirical Overview for Language Modeling in Speech Recognition
abstract
Popularized by the long short-term memory (LSTM), multiplicative gates have become a standard means to design artificial neural networks with intentionally organized information flow.Notable examples of such architectures include gated recurrent units (GRU) and highway networks.In this work, we first focus on the evaluation of each of the classical gated architectures for language modeling for large vocabulary speech recognition.Namely, we evaluate the highway network, lateral network, LSTM and GRU.Furthermore, the motivation underlying the highway network also applies to LSTM and GRU.An extension specific to the LSTM has been recently proposed with an additional highway connection between the memory cells of adjacent LSTM layers.In contrast, we investigate an approach which can be used with both LSTM and GRU: a highway network in which the LSTM or GRU is used as the transformation function.We found that the highway connections enable both standalone feedforward and recurrent neural language models to benefit better from the deep structure and provide a slight improvement of recognition accuracy after interpolation with count models.To complete the overview, we include our initial investigations on the use of the attention mechanism for learning word triggers.
Kazuki Irie, Zoltán Tüske, Tamer Alkhouli, Ralf Schlüter, Hermann Ney
INTERSPEECH4
2016 Towards Online-Recognition with Deep Bidirectional LSTM Acoustic Models
abstract
Online-Recognition requires the acoustic model to provide posterior probabilities after a limited time delay given the online input audio data.This necessitates unidirectional modeling and the standard solution is to use unidirectional long short-term memory (LSTM) recurrent neural networks (RNN) or feedforward neural networks (FFNN).It is known that bidirectional LSTMs are more powerful and perform better than unidirectional LSTMs.To demonstrate the performance difference, we start by comparing several different bidirectional and unidirectional LSTM topologies.Furthermore, we apply a modification to bidirectional RNNs to enable online-recognition by moving a window over the input stream and perform one forwarding through the RNN on each window.Then, we combine the posteriors of each forwarding and we renormalize them.We show in experiments that the performance of this online-enabled bidirectional LSTM performs as good as the offline bidirectional LSTM and much better than the unidirectional LSTM.
Albert Zeyer, Ralf Schlüter, Hermann Ney
INTERSPEECH2
2015 Multilingual representations for low resource speech recognition and keyword search
abstract
This paper examines the impact of multilingual (ML) acoustic representations on Automatic Speech Recognition (ASR) and keyword search (KWS) for low resource languages in the context of the OpenKWS15 evaluation of the IARPA Babel program. The task is to develop Swahili ASR and KWS systems within two weeks using as little as 3 hours of transcribed data. Multilingual acoustic representations proved to be crucial for building these systems under strict time constraints. The paper discusses several key insights on how these representations are derived and used. First, we present a data sampling strategy that can speed up the training of multilingual representations without appreciable loss in ASR performance. Second, we show that fusion of diverse multilingual representations developed at different LORELEI sites yields substantial ASR and KWS gains. Speaker adaptation and data augmentation of these representations improves both ASR and KWS performance (up to 8.7% relative). Third, incorporating un-transcribed data through semi-supervised learning, improves WER and KWS performance. Finally, we show that these multilingual representations significantly improve ASR and KWS performance (relative 9% for WER and 5% for MTWV) even when forty hours of transcribed audio in the target language is available. Multilingual representations significantly contributed to the LORELEI KWS systems winning the OpenKWS15 evaluation.
Jia Cui, Brian Kingsbury, Bhuvana Ramabhadran, Abhinav Sethy, Kartik Audhkhasi, Ellen Eide, Lidia Mangu, Markus Nußbaum-Thom, Michael Picheny, Zoltán Tüske, Pavel Golik, Ralf Schlüter, Hermann Ney, Mark J. F. Gales, Kate M. Knill, Anton Ragni, Philip C. Woodland
ASRU13
2015 Speaker adaptive joint training of Gaussian mixture models and bottleneck features
abstract
In the tandem approach, the output of a neural network (NN) serves as input features to a Gaussian mixture model (GMM) aiming to improve the emission probability estimates. As has been shown in our previous work, GMM with pooled covariance matrix can be integrated into a neural network framework as a softmax layer with hidden variables, which allows for joint estimation of both neural network and Gaussian mixture parameters. Here, this approach is extended to include speaker adaptive training (SAT) by introducing a speaker dependent neural network layer. Error backpropagation beyond this speaker dependent layer realizes the adaptive training of the Gaussian parameters as well as the optimization of the bottleneck (BN) tandem features of the underlying acoustic model, simultaneously. In this study, after the initialization by constrained maximum likelihood linear regression (CMLLR) the speaker dependent layer itself is kept constant during the joint training. Experiments show that the deeper backpropagation through the speaker dependent layer is necessary for improved recognition performance. The speaker adaptively and jointly trained BN-GMM results in 5% relative improvement over very strong speaker-independent hybrid baseline on the Quaero English broadcast news and conversations task, and on the 300-hour Switchboard task.
Zoltán Tüske, Pavel Golik, Ralf Schlüter, Hermann Ney
ASRU3
2015 Unsupervised adaptation of a denoising autoencoder by Bayesian Feature Enhancement for reverberant asr under mismatch conditions
abstract
The parametric Bayesian Feature Enhancement (BFE) and a datadriven Denoising Autoencoder (DA) both bring performance gains in severe single-channel speech recognition conditions. The first can be adjusted to different conditions by an appropriate parameter setting, while the latter needs to be trained on conditions similar to the ones expected at decoding time, making it vulnerable to a mismatch between training and test conditions. We use a DNN backend and study reverberant ASR under three types of mismatch conditions: different room reverberation times, different speaker to microphone distances and the difference between artificially reverberated data and the recordings in a reverberant environment. We show that for these mismatch conditions BFE can provide the targets for a DA. This unsupervised adaptation provides a performance gain over the direct use of BFE and even enables to compensate for the mismatch of real and simulated reverberant data.
Jahn Heymann, Reinhold Häb-Umbach, Pavel Golik, Ralf Schlüter
ICASSP4
2015 Improved strategies for a zero oov rate LVCSR system
abstract
In this work, multiple hierarchical language modeling strategies for a zero OOV rate large vocabulary continuous speech recognition system are investigated. In our previously proposed hierarchical approach, a full-word language model and a context independent character-level LM (CLM) are directly used during search. The novelty of this work is to jointly model the character-level prior and the pronunciation probabilities, to introduce across-word context into the characterlevel LM, and to properly normalize the character-level LM using prefix-tree based normalization for the hierarchical approach. Significant reductions in-terms of word error rates (WER) on the best full-word Quaero Polish LVCSR system are reported.
M. Ali Basha Shaik, Amr El-Desoky Mousa, Stefan Hahn, Ralf Schlüter, Hermann Ney
ICASSP4
2015 Investigation of mixture splitting concept for training linear bottlenecks of deep neural network acoustic models
abstract
A Gaussian or log-linear mixture model trained by maximum likelihood may be trained further using discriminative training. It is desirable that the mixture splitting is also done during the discriminative training, to achieve better mixture density distribution. In previous work such a discriminative splitting approach was presented. Similarly, the resolution of a deep neural network may also be increased by splitting. In this paper, discriminative splitting is applied as a way of initializing a linear bottleneck between two layers of a DNN. Experiments for a single hidden layer and six hidden layer cases show the potential of this approach as an alternative method of pre-training for linear bottlenecks for MLP hidden layers.
Muhammad Ali Tahir, Simon Wiesler, Ralf Schlüter, Hermann Ney
ICASSP3
2015 Integrating Gaussian mixtures into deep neural networks: Softmax layer with hidden variables
abstract
In the hybrid approach, neural network output directly serves as hidden Markov model (HMM) state posterior probability estimates. In contrast to this, in the tandem approach neural network output is used as input features to improve classic Gaussian mixture model (GMM) based emission probability estimates. This paper shows that GMM can be easily integrated into the deep neural network framework. By exploiting its equivalence with the log-linear mixture model (LMM), GMM can be transformed to a large softmax layer followed by a summation pooling layer. Theoretical and experimental results indicate that the jointly trained and optimally chosen GMM and bottleneck tandem features cannot perform worse than a hybrid model. Thus, the question “hybrid vs. tandem” simplifies to optimizing the output layer of a neural network. Speech recognition experiments are carried out on a broadcast news and conversations task using up to 12 feed-forward hidden layers with sigmoid and rectified linear unit activation functions. The evaluation of the LMM layer shows recognition gains over the classic softmax output.
Zoltán Tüske, Muhammad Ali Tahir, Ralf Schlüter, Hermann Ney
ICASSP3
2015 Sequence-discriminative training of recurrent neural networks
abstract
We investigate sequence-discriminative training of long shortterm memory recurrent neural networks using the maximum mutual information criterion. We show that although recurrent neural networks already make use of the whole observation sequence and are able to incorporate more contextual information than feed forward networks, their performance can be improved with sequence-discriminative training. Experiments are performed on two publicly available handwriting recognition tasks containing English and French handwriting. On the English corpus, we obtain a relative improvement in WER of over 11% with maximum mutual information (MMI) training compared to cross-entropy training. On the French corpus, we observed that it is necessary to interpolate the MMI objective function with cross-entropy.
Paul Voigtlaender, Patrick Doetsch, Simon Wiesler, Ralf Schlüter, Hermann Ney
ICASSP4
2015 Investigations on sequence training of neural networks
abstract
In this paper we present an investigation of sequence-discriminative training of deep neural networks for automatic speech recognition. We evaluate different sequence-discriminative training criteria (MMI and MPE) and optimization algorithms (including SGD and Rprop) using the RASR toolkit. Further, we compare the training of the whole network with that of the output layer only. Technical details necessary for a robust training are studied, since there is no consensus yet on the ultimate training recipe. The investigation extends our previous work on training linear bottleneck networks from scratch showing the consistently positive effect of sequence training.
Simon Wiesler, Pavel Golik, Ralf Schlüter, Hermann Ney
ICASSP3
2015 Error bounds for context reduction and feature omission
Eugen Beck, Ralf Schlüter, Hermann Ney
INTERSPEECH2
2015 Convolutional neural networks for acoustic modeling of raw time signal in LVCSR
abstract
In this paper we continue to investigate how the deep neural network (DNN) based acoustic models for automatic speech recognition can be trained without hand-crafted feature extraction. Previously, we have shown that a simple fully connected feedforward DNN performs surprisingly well when trained directly on the raw time signal. The analysis of the weights revealed that the DNN has learned a kind of short-time time-frequency decomposition of the speech signal. In conventional feature extraction pipelines this is done manually by means of a filter bank that is shared between the neighboring analysis windows. Following this idea, we show that the performance gap between DNNs trained on spliced hand-crafted features and DNNs trained on raw time signal can be strongly reduced by introducing 1D-convolutional layers. Thus, the DNN is forced to learn a short-time filter bank shared over a longer time span. This also allows us to interpret the weights of the second convolutional layer in the same way as 2D patches learned on critical band energies by typical convolutional neural networks. The evaluation is performed on an English LVCSR task. Trained on the raw time signal, the convolutional layers allow to reduce the WER on the test set from 25.5% to 23.4%, compared to an MFCC based result of 22.1% using fully connected layers. Index Terms: acoustic modeling, raw time signal, convolutional neural networks
Pavel Golik, Zoltán Tüske, Ralf Schlüter, Hermann Ney
INTERSPEECH3
2015 Multilingual features based keyword search for very low-resource languages
abstract
In this paper we describe RWTH Aachen’s system for keyword search (KWS) with very limited amount of transcribed audio data available in the target language. This setting has become this year’s primary condition within the Babel project [1], seeking to minimize the amount of human effort while retaining a reasonable KWS performance. Thus the highlights presented in this paper include graphemic acoustic modeling; multilingual features trained on language data from the previous project periods; comparison of tandem and hybrid DNN-HMM acoustic models; processing of large amounts of text data available on the web and the morphological KWS based on automatically derived word fragments. The evaluation is performed using two training sets for each of the six current project period’s languages ‐ full language pack (FLP), consisting of 30 hours and very limited language pack (VLLP), comprising less than 3 hours of transcribed audio data. We put our focus on the latter of the two, which is clearly more challenging. The methods described in this work allowed us to exceed 0.3 MTWV on five out of six languages using development queries. Index Terms: acoustic modeling, keyword search, graphemic, multilingual, neural networks, semi-supervised learning
Pavel Golik, Zoltán Tüske, Ralf Schlüter, Hermann Ney
INTERSPEECH3
2015 Bag-of-words input for long history representation in neural network-based language models for speech recognition
abstract
In most of previous works on neural network based language models (NNLMs), the words are represented as 1-of-N encoded feature vectors. In this paper we investigate an alternative encoding of the word history, known as bag-of-words (BOW) representation of a word sequence, and use it as an additional input feature to the NNLM. Both the feedforward neural network (FFNN) and the long short-term memory recurrent neural network (LSTM-RNN) language models (LMs) with additional BOW input are evaluated on an English large vocabulary automatic speech recognition (ASR) task. We show that the BOW features significantly improve both the perplexity (PP) and the word error rate (WER) of a standard FFNN LM. In contrast, the LSTM-RNN LM does not benefit from such an explicit long context feature. Therefore the performance gap between feedforward and recurrent architectures for language modeling is reduced. In addition, we revisit the cache based LM, a seeming analog of the BOW for the count based LM, which was unsuccessful for ASR in the past. Although the cache is able to improve the perplexity, we only observe a very small reduction in WER. Index Terms: language modeling, speech recognition, bag-ofwords, feedforward neural networks, recurrent neural networks, long short-term memory, cache language model
Kazuki Irie, Ralf Schlüter, Hermann Ney
INTERSPEECH2
2015 Improvements in RWTH LVCSR evaluation systems for Polish, Portuguese, English, urdu, and Arabic
abstract
In this work, Portuguese, Polish, English, Urdu, and Arabic automatic speech recognition evaluation systems developed by the RWTH Aachen University are presented. Our LVCSR systems focus on various domains like broadcast news, spontaneous speech, and podcasts. All these systems but Urdu are used for Euronews and Skynews evaluations as part of the EUBridge project. Our previously developed LVCSR systems were improved using different techniques for the aforementioned languages. Significant improvements are obtained using multilingual tandem and hybrid approaches, minimum phone error training, lexical adaptation, open vocabulary long short term memory language models, maximum entropy language models and confusion-network based system combination. Index Terms: LVCSR, LSTM, open-vocabulary, EU-Bridge
M. Ali Basha Shaik, Zoltán Tüske, Muhammad Ali Tahir, Markus Nußbaum-Thom, Ralf Schlüter, Hermann Ney
INTERSPEECH5
2015 From Feedforward to Recurrent LSTM Neural Networks for Language Modeling
abstract
Language models have traditionally been estimated based on relative frequencies, using count statistics that can be extracted from huge amounts of text data. More recently, it has been found that neural networks are particularly powerful at estimating probability distributions over word sequences, giving substantial improvements over state-of-the-art count models. However, the performance of neural network language models strongly depends on their architectural structure. This paper compares count models to feedforward, recurrent, and long short-term memory (LSTM) neural network variants on two large-vocabulary speech recognition tasks. We evaluate the models in terms of perplexity and word error rate, experimentally validating the strong correlation of the two quantities, which we find to hold regardless of the underlying type of the language model. Furthermore, neural networks incur an increased computational complexity compared to count models, and they differently model context dependences, often exceeding the number of words that are taken into account by count based approaches. These differences require efficient search methods for neural networks, and we analyze the potential improvements that can be obtained when applying advanced algorithms to the rescoring of word lattices on large-scale setups.
Martin Sundermeyer, Hermann Ney, Ralf Schlüter
IEEE ACM Trans. Audio Speech Lang. Process.3
2014 A family of discriminative training criteria based on the F-divergence for deep neural networks
abstract
We present novel bounds on the classification error which are based on the f-Divergence and, at the same time, can be used as practical training criteria. There exist virtually no studies which investigate the link between the f-Divergence, the classification error and practical training criteria. So far only the Kullback-Leibler f-Divergence has been examined in this context to formulate a bound on the classification error and to derive the cross-entropy criterion. We extend this concept to a larger class of f-Divergences. We also successfully investigate if the novel training criteria based on the f-Divergence are suited for frame-wise training of deep neural networks on the Babel Vietnamese and Bengali speech recognition tasks.
Markus Nußbaum-Thom, Ralf Schlüter, Vaibhava Goel, Hermann Ney
ICASSP3
2014 Multilingual MRASTA features for low-resource keyword search and speech recognition systems
abstract
This paper investigates the application of hierarchical MRASTA bottleneck (BN) features for under-resourced languages within the IARPA Babel project. Through multilingual training of Multilayer Perceptron (MLP) BN features on five languages (Cantonese, Pashto, Tagalog, Turkish, and Vietnamese), we could end up in a single feature stream which is more beneficial to all languages than the unilingual features. In the case of balanced corpus sizes, the multilingual BN features improve the automatic speech recognition (ASR) performance by 3-5% and the keyword search (KWS) by 3-10% relative for both limited (LLP) and full language packs (FLP). Borrowing orders of magnitude more data from non-target FLPs, the recognition error rate is reduced by 8-10%, and the spoken term detection is improved by over 40% relative on Vietnamese and Pashto LLP. Aiming at the fast development of acoustic models, cross-lingual transfer of multilingually ”pretrained” BN features for a new language is also investigated. Without the need of any MLP training on the new language, the ported BN features performed similarly to the unilingual features on FLP and significantly better on LLP. Results also show that a simple fine-tuning step on the new language is enough to achieve comparable KWS and ASR performance to that system where the target language is also involved in the time-consuming multilingual training.
Zoltán Tüske, David Nolden, Ralf Schlüter, Hermann Ney
ICASSP3
2014 The RWTH English lecture recognition system
abstract
In this paper, we describe the RWTH speech recognition system for English lectures developed within the Translectures project. A difficulty in the development of an English lectures recognition system, is the high ratio of non-native speakers. We address this problem by using very effective deep bottleneck features trained on multilingual data. The acoustic model is trained on large amounts of data from different domains and with different dialects. Large improvements are obtained from unsupervised acoustic adaptation. Another challenge is the frequent use of technical terms and the wide range of topics. In our recognition system, slides, which are attached to most lectures, are used for improving lexical coverage and language model adaptation.
Simon Wiesler, Kazuki Irie, Zoltán Tüske, Ralf Schlüter, Hermann Ney
ICASSP4
2014 RASR/NN: The RWTH neural network toolkit for speech recognition
abstract
This paper describes the new release of RASR — the open source version of the well-proven speech recognition toolkit developed and used at RWTH Aachen University. The focus is put on the implementation of the NN module for training neural network acoustic models. We describe code design, configuration, and features of the NN module. The key feature is a high flexibility regarding the network topology, choice of activation functions, training criteria, and optimization algorithm, as well as a built-in support for efficient GPU computing. The evaluation of run-time performance and recognition accuracy is performed exemplary with a deep neural network as acoustic model in a hybrid NN/HMM system. The results show that RASR achieves a state-of-the-art performance on a real-world large vocabulary task, while offering a complete pipeline for building and applying large scale speech recognition systems.
Simon Wiesler, Alexander Richard, Pavel Golik, Ralf Schlüter, Hermann Ney
ICASSP4
2014 Mean-normalized stochastic gradient for large-scale deep learning
abstract
Deep neural networks are typically optimized with stochastic gradient descent (SGD). In this work, we propose a novel second-order stochastic optimization algorithm. The algorithm is based on analytic results showing that a non-zero mean of features is harmful for the optimization. We prove convergence of our algorithm in a convex setting. In our experiments we show that our proposed algorithm converges faster than SGD. Further, in contrast to earlier work, our algorithm allows for training models with a factorized structure from scratch. We found this structure to be very useful not only because it accelerates training and decoding, but also because it is a very effective means against overfitting. Combining our proposed optimization algorithm with this model structure, model size can be reduced by a factor of eight and still improvements in recognition error rate are obtained. Additional gains are obtained by improving the Newbob learning rate strategy.
Simon Wiesler, Alexander Richard, Ralf Schlüter, Hermann Ney
ICASSP3
2014 Open-Lexicon Language Modeling Combining Word and Character Levels
abstract
In this paper we investigate different n-gram language models that are defined over an open lexicon. We introduce a character-level language model and combine it with a standard word-level language model in a back off fashion. The character-level language model is redefined and renormalized to assign zero probability to words from a fixed vocabulary. Furthermore we present a way to interpolate language models created at the word and character levels. The computation of character-level probabilities incorporates the across-word context. We compare perplexities on all words from the test set and on in-lexicon and OOV words separately on corpora of English and Arabic text.
Michal Kozielski, Martin Matysiak, Patrick Doetsch, Ralf Schlüter, Hermann Ney
ICFHR4
2014 Word pair approximation for more efficient decoding with high-order language models
David Nolden, Ralf Schlüter, Hermann Ney
INTERSPEECH2
2014 RWTH LVCSR systems for quaero and EU-bridge: German, Polish, Spanish and Portuguese
abstract
In this paper, German, Polish, Spanish, and Portuguese large vocabulary continuous speech recognition (LVCSR) systems developed by the RWTH Aachen University are presented.All the above mentioned systems for the aforementioned languages are used for the Quaero and EU-Bridge project evaluations.The LVCSR systems developed for these competitive evaluations focus on various domains like broadcast news, podcasts and lecture domain.Transcription of the speech for these tasks is challenging due to huge variability in the acoustic conditions and a significant portion of audio data includes spontaneous speech.Good improvements are obtained using stateof-the-art multilingual bottleneck features, minimum phone error trained acoustic models, language model (LM) adaptation and confusion-network based system combination.In addition, an open vocabulary approach using morphemic units is investigated along with the LM adaptation for the German LVCSR.
M. Ali Basha Shaik, Zoltán Tüske, Muhammad Ali Tahir, Markus Nußbaum-Thom, Ralf Schlüter, Hermann Ney
INTERSPEECH5
2014 rwthlm - the RWTH aachen university neural network language modeling toolkit
abstract
We present a novel toolkit that implements the long short-term memory (LSTM) neural network concept for language modeling. The main goal is to provide a software which is easy to use, and which allows fast training of standard recurrent and LSTM neural network language models. The toolkit obtains state-of-the-art performance on the standard Treebank corpus. To reduce the training time, BLAS and related libraries are supported, and it is possible to evaluate multiple word sequences in parallel. In addition, arbitrary word classes can be used to speed up the computation in case of large vocabulary sizes. Finally, the software allows easy integration with SRILM, and it supports direct decoding and rescoring of HTK lattices. The toolkit is available for download under an open source license.
Martin Sundermeyer, Ralf Schlüter, Hermann Ney
INTERSPEECH2
2014 Lattice decoding and rescoring with long-Span neural network language models
abstract
With long-span neural network language models, considerable improvements have been obtained in speech recognition. However, it is difficult to apply these models if the underlying search space is large. In this paper, we combine previous work on lattice decoding with long short-term memory (LSTM) neural network language models. By adding refined pruning techniques, we are able to reduce the search effort by a factor of three. Furthermore, we introduce two novel approximations for full lattice rescoring, which opens the potential of lattice-based speech recognition techniques. Compared to 1000-best lists, we find that we can increase the word error rate improvements obtained with LSTMs from 8.2 % to 10.7 % relative over a stateof-the-art baseline, while the resulting lattices are even considerably smaller. In addition, we investigate the use of LSTMs for Babel Assamese keyword search, obtaining significant improvements of 2.5 % relative.
Martin Sundermeyer, Zoltán Tüske, Ralf Schlüter, Hermann Ney
INTERSPEECH3
2014 Data augmentation, feature combination, and multilingual neural networks to improve ASR and KWS performance for low-resource languages
abstract
This paper presents the progress of acoustic models for lowresourced languages (Assamese, Bengali, Haitian Creole, Lao, Zulu) developed within the second evaluation campaign of the IARPA Babel project.This year, the main focus of the project is put on training high-performing automatic speech recognition (ASR) and keyword search (KWS) systems from language resources limited to about 10 hours of transcribed speech data.Optimizing the structure of Multilayer Perceptron (MLP) based feature extraction and switching from the sigmoid activation function to rectified linear units results in about 5% relative improvement over baseline MLP features.Further improvements are obtained when the MLPs are trained on multiple feature streams and by exploiting label preserving data augmentation techniques like vocal tract length perturbation.Systematic application of these methods allows to improve the unilingual systems by 4-6% absolute in WER and 0.064-0.105absolute in MTWV.Transfer and adaptation of multilingually trained MLPs lead to additional gains, clearly exceeding the project goal of 0.3 MTWV even when only the limited language pack of the target language is used.
Zoltán Tüske, Pavel Golik, David Nolden, Ralf Schlüter, Hermann Ney
INTERSPEECH4
2014 Acoustic modeling with deep neural networks using raw time signal for LVCSR
Zoltán Tüske, Pavel Golik, Ralf Schlüter, Hermann Ney
INTERSPEECH3
2013 Efficient nearly error-less LVCSR decoding based on incremental forward and backward passes
abstract
We show that most search errors can be identified by aligning the results of a symmetric forward and backward decoding pass. Based on this knowledge, we introduce an efficient high-level decoding architecture which yields virtually no search errors, and requires virtually no manual tuning. We perform an initial forward- and backward decoding with tight initial beams, then we identify search errors, and then we recursively increment the beam sizes and perform new forward and backward decodings for erroneous intervals until no more search errors are detected. Consequently, each utterance and even each single word is decoded with the smallest beam size required to decode it correctly. On all tested systems we achieve an error rate equal or very close to classical decoding with ideally tuned beam size, but unsupervisedly without specific tuning, and at around 2 times faster runtime. An additional speedup by factor 2 can be achieved by decoding the forward and backward pass in separate threads.
David Nolden, Ralf Schlüter, Hermann Ney
ASRU2
2013 A high-performance Cantonese keyword search system
abstract
We present a system for keyword search on Cantonese conversational telephony audio, collected for the IARPA Babel program, that achieves good performance by combining postings lists produced by diverse speech recognition systems from three different research groups. We describe the keyword search task, the data on which the work was done, four different speech recognition systems, and our approach to system combination for keyword search. We show that the combination of four systems outperforms the best single system by 7%, achieving an actual term-weighted value of 0.517.
Brian Kingsbury, Jia Cui, Mark J. F. Gales, Kate M. Knill, Jonathan Mamou, Lidia Mangu, David Nolden, Michael Picheny, Bhuvana Ramabhadran, Ralf Schlüter, Abhinav Sethy, Philip C. Woodland
ICASSP11
2013 Open vocabulary handwriting recognition using combined word-level and character-level language models
abstract
In this paper, we present a unified search strategy for open vocabulary handwriting recognition using weighted finite state transducers. Additionally to a standard word-level language model we introduce a separate n-gram character-level language model for out-of-vocabulary word detection and recognition. The probabilities assigned by those two models are combined into one Bayes decision rule. We evaluate the proposed method on the IAM database of English handwriting. An improvement from 22.2% word error rate to 17.3% is achieved comparing to the closed-vocabulary scenario and the best published result.
Michal Kozielski, David Rybach, Stefan Hahn, Ralf Schlüter, Hermann Ney
ICASSP4
2013 System combination and score normalization for spoken term detection
abstract
Spoken content in languages of emerging importance needs to be searchable to provide access to the underlying information. In this paper, we investigate the problem of extending data fusion methodologies from Information Retrieval for Spoken Term Detection on low-resource languages in the framework of the IARPA Babel program. We describe a number of alternative methods improving keyword search performance. We apply these methods to Cantonese, a language that presents some new issues in terms of reduced resources and shorter query lengths. First, we show score normalization methodology that improves in average by 20% keyword search performance. Second, we show that properly combining the outputs of diverse ASR systems performs 14% better than the best normalized ASR system.
Jonathan Mamou, Jia Cui, Mark J. F. Gales, Brian Kingsbury, Kate M. Knill, Lidia Mangu, David Nolden, Michael Picheny, Bhuvana Ramabhadran, Ralf Schlüter, Abhinav Sethy, Philip C. Woodland
ICASSP11
2013 Advanced search space pruning with acoustic look-ahead for WFST based LVCSR
abstract
In this work we show how some concepts already known from dynamic network decoding can be used to improve the efficiency of WFST based decoders. First we apply the concept of acoustic look-ahead to a WFST based decoder, and then we analyze the applicability of LM state pruning, a well motivated pruning method which is fundamental to token-passing decoders. The structure of the composed WFST search network makes it difficult to motivate advanced pruning methods, and consequently it is difficult to achieve a real reduction in search space. Nonetheless, we show how LM state pruning can be applied to WFST based decoders to improve their efficiency. The search space can be reduced by up to 50% at equal precision through acoustic look-ahead. Since our decoder follows a dynamic composition approach, the advantage in search space does not fully transfer to the RTF, which can be reduced by around 20% through acoustic look-ahead, and additional 5% through LM state pruning.
David Nolden, Ralf Schlüter, Hermann Ney
ICASSP2
2013 Feature combination and stacking of recurrent and non-recurrent neural networks for LVCSR
abstract
This paper investigates the combination of different short-term features and the combination of recurrent and non-recurrent neural networks (NNs) on a Spanish speech recognition task. Several methods exist to combine different feature sets such as concatenation or linear discriminant analysis (LDA). Even though all these techniques achieve reasonable improvements, feature combination by multi-layer perceptrons (MLPs) outperforms all known approaches. We develop the concept of MLP based feature combination further using recurrent neural networks (RNNs). The phoneme posterior estimates derived from an RNN lead to a significant improvement over the result of the MLPs and achieve a 5% relative better word error rate (WER) with much less parameters. Moreover, we improve the system performance further by combining an MLP and an RNN in a hierarchical framework. The MLP benefits from the preprocessing of the RNN. All NNs are trained on phonemes. Nevertheless, the same concepts could be applied using context-dependent states. In addition to the improvements in recognition performance w.r.t. WER, NN based feature combination methods reduce both, the training and the testing complexity. Overall, the systems are based on a single set of acoustic models, together with the training of different NNs.
Christian Plahl, Michal Kozielski, Ralf Schlüter, Hermann Ney
ICASSP3
2013 Comparison of feedforward and recurrent neural network language models
abstract
Research on language modeling for speech recognition has increasingly focused on the application of neural networks. Two competing concepts have been developed: On the one hand, feedforward neural networks representing an n-gram approach, on the other hand recurrent neural networks that may learn context dependencies spanning more than a fixed number of predecessor words. To the best of our knowledge, no comparison has been carried out between feedforward and state-of-the-art recurrent networks when applied to speech recognition. This paper analyzes this aspect in detail on a well-tuned French speech recognition task. In addition, we propose a simple and efficient method to normalize language model probabilities across different vocabularies, and we show how to speed up training of recurrent neural networks by parallelization.
Martin Sundermeyer, Ilya Oparin, Jean-Luc Gauvain, B. Freiberg, Ralf Schlüter, Hermann Ney
ICASSP5
2013 Investigation on cross- and multilingual MLP features under matched and mismatched acoustical conditions
abstract
In this paper, Multi Layer Perceptron (MLP) based multilingual bottleneck features are investigated for acoustic modeling in three languages - German, French, and US English. We use a modified training algorithm to handle the multilingual training scenario without having to explicitly map the phonemes to a common phoneme set. Furthermore, the cross-lingual portability of bottleneck features between the three languages are also investigated. Single pass recognition experiments on large vocabulary SMS dictation task indicate that (1) multilingual bottleneck features yield significantly lower word error rates compared to standard MFCC features (2) multilingual bottleneck features are superior to monolingual bottleneck features trained for the target language with limited training data, and (3) multilingual bottleneck features are beneficial in training acoustic models in a low resource language where only mismatched training data is available-by exploiting the more matched training data from other languages.
Zoltán Tüske, Joel Pinto, Daniel Willett, Ralf Schlüter
ICASSP4
2013 Deep hierarchical bottleneck MRASTA features for LVCSR
abstract
Hierarchical Multi Layer Perceptron (MLP) based long-term feature extraction is optimized for TANDEM connectionist large vocabulary continuous speech recognition (LVCSR) system within the QUAERO project. Training the bottleneck MLP on multi-resolutional RASTA filtered critical band energies, more than 20% relative word error rate (WER) reduction over standard MFCC system is observed after optimizing the number of target labels. Furthermore, introducing a deeper structure in the hierarchical bottleneck processing the relative gain increases to 25%. The final system based on deep bottleneck TANDEM features clearly outperforms the hybrid approach, even if the long-term features are also presented to the deep MLP acoustic model. The results are also verified on evaluation data of the year 2012, and about 20% relative WER improvement over classical cepstral system is measured even after speaker adaptive training.
Zoltán Tüske, Ralf Schlüter, Hermann Ney
ICASSP2
2013 A critical evaluation of stochastic algorithms for convex optimization
abstract
Log-linear models find a wide range of applications in pattern recognition. The training of log-linear models is a convex optimization problem. In this work, we compare the performance of stochastic and batch optimization algorithms. Stochastic algorithms are fast on large data sets but can not be parallelized well. In our experiments on a broadcast conversations recognition task, stochastic methods yield competitive results after only a short training period, but when spending enough computational resources for parallelization, batch algorithms are competitive with stochastic algorithms. We obtained slight improvements by using a stochastic second order algorithm. Our best log-linear model outperforms the maximum likelihood trained Gaussian mixture model baseline although being ten times smaller.
Simon Wiesler, Alexander Richard, Ralf Schlüter, Hermann Ney
ICASSP3
2013 Development of the RWTH transcription system for slovenian
abstract
In this paper we describe the RWTH automatic speech recognition system for Slovenian developed within the transLectures project.The project aims at supporting the transcription and translation of video lectures freely available on the web.Difficulties arise on all levels of modeling: Slovenian is a morphologically rich language with a high level of inflection (pronunciation model), and a large variety of dialects and recording conditions brings uncertainty into the audio signal (acoustic model).Moreover, the video lectures cover a wide spectrum of topics with a high share of spontaneous speech and technical terms (language model).These issues require application of robust and adaptive methods.Besides the system description, this study mainly focuses on robust acoustic modeling.Building acoustic models from various resources, we also compare the influence of speaker adaptation to different neural network based acoustic features.Systematic application of these methods allows us to reduce the word error rate on the evaluation corpus from 59.2% to 43.4%.We also give a motivation for Slovenian open vocabulary recognition and perform some first steps.
Pavel Golik, Zoltán Tüske, Ralf Schlüter, Hermann Ney
INTERSPEECH3
2013 Improving LVCSR with hidden conditional random fields for grapheme-to-phoneme conversion
abstract
In virtually every state-of-the-art large vocabulary continuous speech recognition (LVCSR) system, grapheme-to-phoneme (G2P) conversion is applied to generalize beyond a fixed set of words given by a background lexicon. The overall performance of the G2P system has a strong effect on the recognition qual-ity. Typically, generative models based on joint-n-grams are used, although some discriminative models have a competitive performance but the training time may be quite large. In this work, the effect of using discriminative G2P modeling based on hidden conditional random fields (HCRFs) is ana-lyzed. Besides measuring and comparing the G2P qualities on a textual level, one focus is the performance of LVCSR systems. Although the HCRF model does not outperform the generative one on text data, we could improve our English QUAERO ASR system by 1-3 % relative on a couple of test corpora over a strong baseline by only replacing the G2P strategy. Index Terms: grapheme-to-phoneme conversion, G2P, LVCSR, HCRF, hidden conditional random fields
Stefan Hahn, Patrick Lehnen, Simon Wiesler, Ralf Schlüter, Hermann Ney
INTERSPEECH4
2013 Morpheme level hierarchical pitman-yor class-based language models for LVCSR of morphologically rich languages
abstract
Performing large vocabulary continuous speech recognition (LVCSR) for morphologically rich languages is considered a challenging task.The morphological richness of such languages leads to high out-of-vocabulary (OOV) rates and poor language model (LM) probabilities.In this case, the use of morphemes has been shown to increase the lexical coverage and lower the LM perplexity.Another approach used to improve the LM probability estimates is to incorporate additional knowledge sources in the LM estimation process using classbased LMs (CLMs).Recently, the hierarchical Pitman-Yor LMs (HPYLMs) have shown superiority over the modified Kneser-Ney (MKN) smoothed N-gram LMs in terms of both perplexity (PPL) and word error rate (WER) on word-based LVCSR tasks.In this paper, hierarchical Pitman-Yor class-based LMs (HPY-CLMs) are combined with morpheme level language modeling.This enables the application of the proposed models on top of morpheme-based systems.Experiments are conducted on Arabic and German LVCSR tasks.Consistent performance improvements are obtained for all the available corpora compared to the conventional morpheme-based and class-based LMs.
Amr El-Desoky Mousa, M. Ali Basha Shaik, Ralf Schlüter, Hermann Ney
INTERSPEECH3
2013 Relative error bounds for statistical classifiers based on the f-divergence
abstract
In language classification, measures like perplexity and Kullback-Leibler divergence are used to compare language models. While this bears the advantage of isolating the effect of the language model in speech and language processing problems, the measures have no clear relation to the corresponding classification error. In practice, an improvement in terms of perplexity does not necessarily correspond to an improvement in the error rate. It is well-known that Bayes decision rule is optimal if the true distribution is used for classification. Since the true distribution is unknown in practice, a model distribution is used instead, introducing suboptimality. We focus on the degradation introduced by a model distribution, and provide an upper bound on the error difference between Bayes decision and a modelbased decision rule in terms of the f-Divergence between the true and model distributions. Simulations are first presented to reveal a special case of the bound, followed by an analytic proof of the generalized bound and its tightness. In addition, the conditions that result in the boundary cases will be discussed. Several instances of the bound will be verified using simulations, and the bound will be used to study the effect of the language model on the classification error. Index Terms: generalization bounds, language modeling, perplexity, confidence measures, f-Divergence, error mismatch.
Markus Nußbaum-Thom, Eugen Beck, Tamer Alkhouli, Ralf Schlüter, Hermann Ney
INTERSPEECH4
2013 Feature-rich sub-lexical language models using a maximum entropy approach for German LVCSR
abstract
German is a morphologically rich language having a high degree of word inflections, derivations and compounding. This leads to high out-of-vocabulary (OOV) rates and poor language model (LM) probabilities in the large vocabulary continuous speech recognition (LVCSR) systems. One of the main challenges in the German LVCSR is the recognition of the OOV words. For this purpose, data-driven morphemes are used to provide higher lexical coverage. On the other hand, the probability estimates of a sub-lexical LM could be further improved using feature-rich LMs like maximum entropy (MaxEnt) and class-based LMs. In this work, for a sub-lexical level German LVCSR task, we investigate the use of the multiple morpheme level features as classes for building class-based LMs that are estimated using the state-of-the-art MaxEnt approach. Thus, the benefits of both the MaxEnt LMs and the traditional class-based LMs are effectively combined. Furthermore, we experiment the use of Maximum a-posteriori adaptation over the MaxEnt class-based LMs. We show consistent reductions in both the OOV recognition error rate and the word error rate (WER) on a German LVCSR task from the Quaero project, compared to the traditional class-based and theN -gram morpheme based LM. Index Terms: open-vocabulary, German LVCSR, features, maximum entropy, class-based
M. Ali Basha Shaik, Amr El-Desoky Mousa, Ralf Schlüter, Hermann Ney
INTERSPEECH3
2013 Training log-linear acoustic models in higher-order polynomial feature space for speech recognition
abstract
The use of higher-order polynomial acoustic features can improve the performance of automatic speech recognition.However, the dimensionality of the polynomial representation can be prohibitively large, making the training of acoustic models using polynomial features for large vocabulary ASR systems infeasible.This paper presents an iterative polynomial training framework for acoustic modeling, which recursively expands the current acoustic features into their second-order polynomial feature space.In each recursion the dimensionality is reduced by a linear projection, such that increasingly higher order polynomial information is incorporated while keeping the dimensionality of the acoustic models constant.Experimental results obtained for a large-vocabulary continuous speech recognition task show that the proposed method outperforms conventional mixture models.
Muhammad Ali Tahir, Heyun Huang, Ralf Schlüter, Hermann Ney, Louis ten Bosch, Bert Cranen, Lou Boves
INTERSPEECH3
2013 Multilingual hierarchical MRASTA features for ASR
abstract
Recently, a multilingual Multi Layer Perceptron (MLP) training method was introduced without having to explicitly map the phonetic units of multiple languages to a common set.This paper further investigates this method using bottleneck (BN) tandem connectionist acoustic modeling for four high-resourced languages -English, French, German, and Polish.Aiming at the improvement of already existing high performing automatic speech recognition (ASR) systems, the multilingual training of the BN-MLP is extended from short-term to hierarchical longterm (multi-resolutional RASTA) feature extraction.Furthermore, deeper structures and context-dependent target labels are also examined.We experimentally demonstrate that a single state-of-the-art BN feature set can be trained for multiple languages, which is superior to the monolingual feature set, and results in significant gains in all the four languages.Studying the scalability of the multilingual BN features, a similar gain is observed in small (50 hours) and in larger scale (300 hours) ASR experiments regardless of the distribution of the data amount between the languages.Using deeper structures, context-dependent targets, and speaker adaptation, the multilingual BN reduces the word error rates by 3-7% relative over the target language BN features and 25-30% over the conventional MFCC system.
Zoltán Tüske, Ralf Schlüter, Hermann Ney
INTERSPEECH2
2013 Novel tight classification error bounds under mismatch conditions based on f-Divergence
abstract
By default, statistical classification/multiple hypothesis testing is faced with the model mismatch introduced by replacing the true distributions in Bayes decision rule by model distributions estimated on training samples. Although a large number of statistical measures exist w.r.t. to the mismatch introduced, these works rarely relate to the mismatch in accuracy, i.e. the difference between model error and Bayes error. In this work, the accuracy mismatch between the ideal Bayes decision rule/Bayes test and a mismatched decision rule in statistical classification/multiple hypothesis testing is investigated explicitly. A proof of a novel generalized tight statistical bound on the accuracy mismatch is presented. This result is compared to existing statistical bounds related to the total variational distance that can be extended to bounds of the accuracy mismatch. The analytic results are supported by distribution simulations.
Ralf Schlüter, Markus Nußbaum-Thom, Eugen Beck, Tamer Alkhouli, Hermann Ney
ITW1
2013 Investigations on an EM-Style Optimization Algorithm for Discriminative Training of HMMs
abstract
Today's speech recognition systems are based on hidden Markov models (HMMs) with Gaussian mixture models whose parameters are estimated using a discriminative training criterion such as Maximum Mutual Information (MMI) or Minimum Phone Error (MPE). Currently, the optimization is almost always done with (empirical variants of) Extended Baum-Welch (EBW). This type of optimization requires sophisticated update schemes for the step sizes and a considerable amount of parameter tuning, and only little is known about its convergence behavior. In this paper, we derive an EM-style algorithm for discriminative training of HMMs. Like Expectation-Maximization (EM) for the generative training of HMMs, the proposed algorithm improves the training criterion on each iteration, converges to a local optimum, and is completely parameter-free. We investigate the feasibility of the proposed EM-style algorithm for discriminative training of two tasks, namely grapheme-to-phoneme conversion and spoken digit string recognition.
Georg Heigold, Hermann Ney, Ralf Schlüter
IEEE ACM Trans. Audio Speech Lang. Process.3
2013 Lexical Prefix Tree and WFST: A Comparison of Two Dynamic Search Concepts for LVCSR
abstract
Dynamic network decoders have the advantage of significantly lower memory consumption compared to static network decoders, especially when huge vocabularies and complex language models are required. This paper compares the properties of two well-known search strategies for dynamic network decoding, namely history conditioned lexical tree search and weighted finite-state transducer-based search using on-the-fly transducer composition. The two search strategies share many common principles like the use of dynamic programming, beam search, and many more. We point out the similarities of both approaches and investigate the implications of their differing features, both formally and experimentally, with a focus on implementation independent properties. Therefore, experimental results are obtained with a single decoder by representing the history conditioned lexical tree search strategy in the transducer framework. The properties analyzed cover structure and size of the search space, differences in hypotheses recombination, language model look-ahead techniques, and lattice generation.
David Rybach, Hermann Ney, Ralf Schlüter
IEEE Trans. Speech Audio Process.3
2012 Basis vector orthogonalization for an improved kernel gradient matching pursuit method
abstract
With the aim of achieving a computationally efficient optimization of kernel-based probabilistic models for various problems, such as sequential pattern recognition, we have already developed the kernel gradient matching pursuit method as an approximation technique for kernel-based classification. The conventional kernel gradient matching pursuit method approximates the optimal parameter vector by using a linear combination of a small number of basis vectors. In this paper, we propose an improved kernel gradient matching pursuit method that introduces orthogonality constraints to the obtained basis vector set. We verified the efficiency of the proposed method by conducting recognition experiments based on handwritten image datasets and speech datasets. We realized a scalable kernel optimization that incorporated various models, handled very high-dimensional features (>;100 K features), and enabled the use of large scale datasets (>; 10 M samples).
Yotaro Kubo, Shinji Watanabe 0001, Atsushi Nakamura, Simon Wiesler, Ralf Schlüter, Hermann Ney
ICASSP5
2012 Investigations on the use of morpheme level features in Language Models for Arabic LVCSR
abstract
A major challenge for Arabic Large Vocabulary Continuous Speech Recognition (LVCSR) is the rich morphology of Arabic, which leads to high Out-of-vocabulary (OOV) rates, and poor Language Model (LM) probabilities. In such cases, the use of morphemes rather than full-words is considered a better choice for LMs. Thereby, higher lexical coverage and less LM perplexities are achieved. On the other side, an effective way to increase the robustness of LMs is to incorporate features of words into LMs. In this paper, we investigate the use of features derived for morphemes rather than words. Thus, we combine the benefits of both morpheme level and feature rich modeling. We compare the performance of stream-based, class-based and Factored LMs (FLMs) estimated over sequences of morphemes and their features for performing Arabic LVCSR. A relative reduction of 3.9% in Word Error Rate (WER) is achieved compared to a word-based system.
Amr El-Desoky Mousa, Ralf Schlüter, Hermann Ney
ICASSP2
2012 Joining advantages of word-conditioned and token-passing decoding
abstract
We compare the families of token-passing and word conditioned decoders, and derive a more efficient dynamic decoder. The advantage of the word conditioned approach is a trivially simple hypothesis recombination, while the advantage of the token-passing approach is a straight-forward minimization of the search network with compressed word tails. We derive a dynamic decoder which joins the advantages of both decoding architectures by minimizing the search network of a word conditioned decoder. We describe the decoder and analyze its efficiency regarding acoustic look-ahead, network minimization, and facilitated pruning methods. Finally, we compare the new decoder with a WFST based decoder extended by acoustic look-ahead.
David Nolden, David Rybach, Ralf Schlüter, Hermann Ney
ICASSP3
2012 Extended search space pruning in LVCSR
abstract
We compare the most important pruning methods which are common in different LVCSR decoding architectures and lead them back to a theoretical motivation. Based on this motivation, we propose a new pruning method which fades the word end pruning over a large part of the search network. We analyze the methods regarding their relationship between search-space and word error rate, and regarding their mutual dependence. We show that the different pruning methods are mutually dependent and difficult to combine, and that our new pruning method is the most effective method regarding both the search space and runtime efficiency.
David Nolden, Ralf Schlüter, Hermann Ney
ICASSP2
2012 Silence is golden: Modeling non-speech events in WFST-based dynamic network decoders
abstract
Models for silence are a fundamental part of continuous speech recognition systems. Depending on application requirements, audio data segmentation, and availability of detailed training data annotations, it may be necessary or beneficial to differentiate between other non-speech events, for example breath and background noise. The integration of multiple non-speech models in a WFST-based dynamic network decoder is not straightforward, because these models do not perfectly fit in the transducer framework. This paper describes several options for the transducer construction with multiple non-speech models, shows their considerable different characteristics in memory and runtime efficiency, and analyzes the impact on the recognition performance.
David Rybach, Ralf Schlüter, Hermann Ney
ICASSP2
2012 Comparison and combination of different CRBE based MLP features for LVCSR
abstract
Multi Layer Perceptron (MLP) features extracted from different types of critical band energies (CRBE) - derived from MFCC, GT, and PLP pipeline - are compared on French broadcast news and conversational speech recognition task. Though the MLP structure is kept fixed, ROVER combination of different CRBE based systems leads to 4% relative improvement. Furthermore, aiming at the combination of state-of-the-art features based on various signal analysis methods into one single stream, posterior feature space based combination technique is proposed. The speaker normalized features originated from different CRBEs are merged after additional MLP training by Dempster-Shafer rule. The performance of these posterior features unifying the different CRBE based features is superior to the best single CRBE based posterior features by 6% relative. Further results reveal that the concatenated cepstral and unified posterior features perform nearly as well as the ROVER combination of the different CRBE based systems.
Zoltán Tüske, Ralf Schlüter, Hermann Ney
ICASSP2
2012 Morpheme Level Feature-based Language Models for German LVCSR
abstract
One of the challenges for Large Vocabulary Continuous Speech Recognition (LVCSR) of German is its complex morphology and high level of compounding.It leads to high Out-of-vocabulary (OOV) rates, and poor Language Model (LM) probabilities.In such cases, building LMs on morpheme level can be considered a better choice.Thereby, higher lexical coverage and lower LM perplexities are achieved.On the other side, a successful approach to improve the LM probability estimation is to incorporate features of words using feature-based LMs.In this paper, we use features derived for morphemes as well as words.Thus, we combine the benefits of both morpheme level and feature rich modeling.We compare the performance of stream-based, class-based and factored LMs (FLMs).Relative reductions of around 1.5% in Word Error Rate (WER) are achieved compared to the best previous results obtained using FLMs.
Amr El-Desoky Mousa, M. Ali Basha Shaik, Ralf Schlüter, Hermann Ney
INTERSPEECH3
2012 Search Space Pruning Based on Anticipated Path Recombination in LVCSR
abstract
In this paper we introduce a well-motivated abstract pruning criterion for LVCSR decoders based on the anticipated recombination of HMM state alignment paths.We show that several heuristical pruning methods common in dynamic network decoders are approximations of this pruning criterion.The abstract criterion is too complex to be applied directly in an efficient manner, so we derive approximations which can be applied efficiently.Our new pruning methods allow much more exhaustive pruning of the search space than previous methods.We show that the size of the search space can be reduced by up to 50% at equal precision over the previous state of the art, and the RTF by 20%.The abstract pruning criterion can be considered a guide to derive effective pruning methods for any kind of time synchronous decoder.
David Nolden, Ralf Schlüter, Hermann Ney
INTERSPEECH2
2012 Posterior-Scaled MPE: Novel Discriminative Training Criteria
abstract
We recently discovered novel discriminative training criteria following a principled approach. In this approach training criteria are developed from error bounds on the global error for pattern classification tasks that depend on non-trivial loss functions. Automatic speech recognition (ASR) is a prominent example for such a task depending on the non-trivial Levenshtein loss. In this context, the posterior-scaled Minimum Phoneme Error (MPE) training criterion, which is the state-of-the-art discriminative training criterion in ASR, was shown to be an approximation to one of the novel criteria. Here, we describe the implementation of the posterior-scaled MPE criterion in a transducer-based framework, and compare this criterion to other discriminative training criteria on an ASR task. This comparison indicates that the posterior-scaled MPE criterion performs better than other discriminative criteria including MPE. Index Terms: error bounds, discriminative training criteria, margin, MPE
Markus Nußbaum-Thom, Zoltán Tüske, Georg Heigold, Ralf Schlüter, Hermann Ney
INTERSPEECH4
2012 Investigation of Maximum Entropy Hybrid Language Models for Open Vocabulary German and Polish LVCSR
abstract
For languages like German and Polish, higher numbers of word inflections lead to high out-of-vocabulary (OOV) rates and high language model (LM) perplexities.Thus, one of the main challenges in large vocabulary continuous speech recognition (LVCSR) is recognizing an open vocabulary.In this paper, we investigate the use of mixed type of sub-word units in the same recognition lexicon.Namely, morphemic or syllabic units combined with pronunciations called graphones, normal graphemic morphemes or syllables, along with full-words.In addition, we investigate the suitability of hybrid mixed-unit Ngrams as features for Maximum Entropy LM along with adaptation.We achieve significant improvements in recognizing OOVs and word error rate reductions for German and Polish LVCSR compared to the conventional full-word approach and state-of-the-art N-gram mixed type hybrid LM.
M. Ali Basha Shaik, Amr El-Desoky Mousa, Ralf Schlüter, Hermann Ney
INTERSPEECH3
2012 LSTM Neural Networks for Language Modeling
abstract
Neural networks have become increasingly popular for the task of language modeling.Whereas feed-forward networks only exploit a fixed context length to predict the next word of a sequence, conceptually, standard recurrent neural networks can take into account all of the predecessor words.On the other hand, it is well known that recurrent networks are difficult to train and therefore are unlikely to show the full potential of recurrent models.These problems are addressed by a the Long Short-Term Memory neural network architecture.In this work, we analyze this type of network on an English and a large French language modeling task.Experiments show improvements of about 8 % relative in perplexity over standard recurrent neural network LMs.In addition, we gain considerable improvements in WER on top of a state-of-the-art speech recognition system.
Martin Sundermeyer, Ralf Schlüter, Hermann Ney
INTERSPEECH2
2012 Simultaneous Discriminative Training and Mixture Splitting of HMMs for Speech Recognition
abstract
A method is proposed to incorporate mixture density splitting into the acoustic model discriminative training for speech recognition.The standard method is to obtain a high resolution acoustic model by maximum likelihood training and density splitting, and then improving this model by discriminative training.We choose a log-linear form of acoustic model because for a single Gaussian density per triphone state the log-linear MMI optimization is a convex optimization problem, and by further splitting and discriminative training of this model we can get a higher complexity model.Previously it was shown that we achieve large gains in the objective function and corresponding moderate gains in the word error rate on a large vocabulary corpus.This paper incorporates the state of the art minimum phone error training criterion into the framework, and shows that after discriminative splitting, a subsequent log-linear MPE training achieves better results than Gaussian mixture model MPE optimization alone.
Muhammad Ali Tahir, Markus Nußbaum-Thom, Ralf Schlüter, Hermann Ney
INTERSPEECH3
2012 Context-Dependent MLPs for LVCSR: TANDEM, Hybrid or Both?
abstract
Gaussian Mixture Model (GMM) and Multi Layer Perceptron (MLP) based acoustic models are compared on a French large vocabulary continuous speech recognition (LVCSR) task.In addition to optimizing the output layer size of the MLP, the effect of the deep neural network structure is also investigated.Moreover, using different linear transformations (time derivatives, LDA, CMLLR) on conventional MFCC, the study is also extended to MLP based probabilistic and bottle-neck TANDEM features.Results show that using either the hybrid or bottleneck TANDEM approach leads to similar recognition performance.However, the best performance is achieved when deep MLP acoustic models are trained on concatenated cepstral and context-dependent bottle-neck features.Further experiments reveal the importance of the neighbouring frames in case of MLP based modeling, and that its gain over GMM acoustic models is strongly reduced by more complex features.
Zoltán Tüske, Ralf Schlüter, Hermann Ney, Martin Sundermeyer
INTERSPEECH2
2012 Accelerated Batch Learning of Convex Log-linear Models for LVCSR
abstract
This paper describes a log-linear modeling framework suitable for large-scale speech recognition tasks. We introduce modifications to our training procedure that are required for extending our previous work on log-linear models to larger tasks. We give a detailed description of the training procedure with a focus on aspects that impact computational efficiency. The performance of our approach is evaluated on the English Quaero corpus, a challenging broadcast conversations task. The log-linear model consistently outperforms the maximum likelihood baseline system. Comparable performance to a system with minimum-phone-error training is achieved. Index Terms: acoustic modeling, discriminative models 1.
Simon Wiesler, Ralf Schlüter, Hermann Ney
INTERSPEECH2
2012 Does the Cost Function Matter in Bayes Decision Rule?
abstract
In many tasks in pattern recognition, such as automatic speech recognition (ASR), optical character recognition (OCR), part-of-speech (POS) tagging, and other string recognition tasks, we are faced with a well-known inconsistency: The Bayes decision rule is usually used to minimize string (symbol sequence) error, whereas, in practice, we want to minimize symbol (word, character, tag, etc.) error. When comparing different recognition systems, we do indeed use symbol error rate as an evaluation measure. The topic of this work is to analyze the relation between string (i.e., 0-1) and symbol error (i.e., metric, integer valued) cost functions in the Bayes decision rule, for which fundamental analytic results are derived. Simple conditions are derived for which the Bayes decision rule with integer-valued metric cost function and with 0-1 cost gives the same decisions or leads to classes with limited cost. The corresponding conditions can be tested with complexity linear in the number of classes. The results obtained do not make any assumption w.r.t. the structure of the underlying distributions or the classification problem. Nevertheless, the general analytic results are analyzed via simulations of string recognition problems with Levenshtein (edit) distance cost function. The results support earlier findings that considerable improvements are to be expected when initial error rates are high.
Ralf Schlüter, Markus Nußbaum-Thom, Hermann Ney
IEEE Trans. Pattern Anal. Mach. Intell.1
2012 WFST Enabled Solutions to ASR Problems: Beyond HMM Decoding
abstract
During the last decade, weighted finite-state transducers (WFSTs) have become popular in speech recognition. While their main field of application remains hidden Markov model (HMM) decoding, the WFST framework is now also seen as a brick in solutions to many other central problems in automatic speech recognition (ASR). These solutions are less known, and this work aims at giving an overview of the applications of WFSTs in large-vocabulary continuous speech recognition (LVCSR) besides HMM decoding: discriminative acoustic model training, Bayes risk decoding, and system combination. The application of the WFST framework has a big practical impact: we show how the framework helps to structure problems, to develop generic solutions, and to delegate complex computations to WFST toolkits. In this paper, we review the literature, discuss existing approaches, and provide new insights into WFST enabled solutions. We also present a novel, purely WFST-based algorithm for computing the exact Bayes risk hypothesis from a lattice with the Levenshtein distance as loss function. We present the problems and their solutions in a unified framework and discuss the advantages and limits of using WFSTs. We do not provide new experimental results, but refer to the existing literature. Our work helps to identify where and how the transducer framework can contribute to a compact and generic solution to LVCSR problems.
Björn Hoffmeister, Georg Heigold, David Rybach, Ralf Schlüter, Hermann Ney
IEEE Trans. Speech Audio Process.4
2011 Cross-lingual portability of Chinese and english neural network features for French and German LVCSR
abstract
This paper investigates neural network (NN) based cross-lingual probabilistic features. Earlier work reports that intra-lingual features consistently outperform the corresponding cross-lingual features. We show that this may not generalize. Depending on the complexity of the NN features, cross-lingual features reduce the resources used for training -the NN has to be trained on one language only- without any loss in performance w.r.t. word error rate (WER). To further investigate this inconsistency concerning intra- vs. cross-lingual neural network features, we analyze the performance of these features w.r.t. the degree of kinship between training and testing language, and the amount of training data used. Whenever the same amount of data is used for NN training, a close relationship between training and testing language is required to achieve similar results. By increasing the training data the relationship becomes less, as well as changing the topology of the NN to the bottle neck structure. Moreover, cross-lingual features trained on English or Chinese improve the best intra-lingual system for German up to 2% relative in WER and up to 3% relative for French and achieve the same improvement as for discriminative training. Moreover, we gain again up to 8% relative in WER by combining intra- and cross-lingual systems.
Christian Plahl, Ralf Schlüter, Hermann Ney
ASRU2
2011 Discriminative splitting of Gaussian/log-linear mixture HMMs for speech recognition
abstract
This paper presents a method to incorporate mixture density splitting into the acoustic model discriminative log-linear training. The standard method is to obtain a high resolution model by maximum likelihood training and density splitting, and then further training this model discriminatively. For a single Gaussian density per state the log-linear MMI optimization is a global maximum problem, and by further splitting and discriminative training of this model we can get a higher complexity model. The mixture training is not a global maximum problem, nevertheless experimentally we achieve large gains in the objective function and corresponding moderate gains in the word error rate on a large vocabulary corpus.
Muhammad Ali Tahir, Ralf Schlüter, Hermann Ney
ASRU2
2011 A convergence analysis of log-linear training and its application to speech recognition
abstract
Log-linear models are a promising approach for speech recognition. Typically, log-linear models are trained according to a strictly convex criterion. Optimization algorithms are guaranteed to converge to the unique global optimum of the objective function from any initialization. For large-scale applications, considerations in the limit of infinite iterations are not sufficient. We show that log-linear training can be a highly ill-conditioned optimization problem, resulting in extremely slow convergence. Conversely, the optimization problem can be preconditioned by feature transformations. Making use of our convergence analysis, we improve our log-linear speech recognition system and achieve a strong reduction of its training time. In addition, we validate our analysis on a continuous handwriting recognition task.
Simon Wiesler, Ralf Schlüter, Hermann Ney
ASRU2
2011 Subspace pursuit method for kernel-log-linear models
abstract
This paper presents a novel method for reducing the dimensionality of kernel spaces. Recently, to maintain the convexity of training, log linear models without mixtures have been used as emission probability density functions in hidden Markov models for automatic speech recognition. In that framework, nonlinearly-transformed high-dimensional features are used to achieve the nonlinear classification of the original observation vectors without using mixtures. In this paper, with the goal of using high-dimensional features in kernel spaces, the cutting plane subspace pursuit method proposed for support vector machines is generalized and applied to log-linear models. The experimental results show that the proposed method achieved an efficient approximation of the feature space by using a limited number of basis vectors.
Yotaro Kubo, Simon Wiesler, Ralf Schlüter, Hermann Ney, Shinji Watanabe 0001, Atsushi Nakamura, Tetsunori Kobayashi
ICASSP3
2011 Exploiting sparseness of backing-off language models for efficient look-ahead in LVCSR
abstract
In this paper, we propose a new method for computing and applying language model look-ahead in a dynamic network decoder, exploiting the sparseness of backing-off n-gram language models. Only partial (sparse) look-ahead tables are computed, with a size that depends on the number of words that have an n-gram score in the language model for a specific context, rather than a constant, vocabulary dependent size. Since high order backing-off language models are inherently sparse, this mechanism reduces the runtime- and memory effort of computing the look-ahead tables by magnitudes. A modified decoding algorithm is required to apply these sparse LM look-ahead tables efficiently. We show that sparse LM look-ahead is much more efficient than the classical method, and that full n-gram look-ahead becomes favorable over lower order look-ahead even when many distinct LM contexts appear during decoding.
David Nolden, Hermann Ney, Ralf Schlüter
ICASSP3
2011 A comparative analysis of dynamic network decoding
abstract
The use of statically compiled search networks for ASR systems using huge vocabularies and complex language models often becomes challenging in terms of memory requirements. Dynamic network decoders introduce additional computations in favor of significantly lower memory consumption. In this paper we investigate the properties of two well-known search strategies for dynamic network decoding, namely history conditioned tree search and WFST-based search using dynamic transducer composition. We analyze the impact of the differences in search graph representation, search space structure, and language model look-ahead techniques. Experiments on an LVCSR task illustrate the influence of the compared properties.
David Rybach, Ralf Schlüter, Hermann Ney
ICASSP2
2011 Using morpheme and syllable based sub-words for polish LVCSR
abstract
Polish is a synthetic language with a high morpheme-per-word ratio. It makes use of a high degree of inflection leading to high out-of-vocabulary (OOV) rates, and high Language Model (LM) perplexities. This poses a challenge for Large Vocabulary and Continuous Speech Recognition (LVCSR) systems. Here, the use of morpheme and syllable based units is investigated for building sub-lexical LMs. A different type of sub-lexical units is proposed based on combining morphemic or syllabic units with corresponding pronunciations. Thereby, a set of grapheme-phoneme pairs called graphones are used for building LMs. A relative reduction of 3.5% in Word Error Rate (WER) is obtained with respect to a traditional system based on full-words.
M. Ali Basha Shaik, Amr El-Desoky Mousa, Ralf Schlüter, Hermann Ney
ICASSP3
2011 The RWTH 2010 Quaero ASR evaluation system for English, French, and German
abstract
Recognizing Broadcast Conversational (BC) speech data is a difficult task, which can be regarded as one of the major challenges beyond the recognition of Broadcast News (BN).
Martin Sundermeyer, Markus Nußbaum-Thom, Simon Wiesler, Christian Plahl, Amr El-Desoky Mousa, Stefan Hahn, David Nolden, Ralf Schlüter, Hermann Ney
ICASSP8
2011 Non-stationary feature extraction for automatic speech recognition
abstract
In current speech recognition systems mainly Short-Time Fourier Transform based features like MFCC are applied. Dropping the short-time stationarity assumption of the voiced speech, this paper introduces the non-stationary signal analysis into the ASR framework. We present new acoustic features extracted by a pitch-adaptive Gammatone filter bank. The noise robustness was proved on AURORA 2 and 4 tasks, where the proposed features outperform the standard MFCC. Furthermore, successful combination experiments via ROVER indicate the differences between the new features and MFCC.
Zoltán Tüske, Pavel Golik, Ralf Schlüter, Friedhelm R. Drepper
ICASSP3
2011 Feature selection for log-linear acoustic models
abstract
Log-linear acoustic models have been shown to be competitive with Gaussian mixture models in speech recognition. Their high training time can be reduced by feature selection. We compare a simple univariate feature selection algorithm with ReliefF - an efficient multivariate algorithm. An alternative to feature selection is ℓ1-regularized training, which leads to sparse models. We observe that this gives no speedup when sparse features are used, hence feature selection methods are preferable. For dense features, ℓ1-regularization can reduce training and recognition time. We generalize the well known Rprop algorithm for the optimization of ℓ1-regularized functions. Experiments on the Wall Street Journal corpus showed that a large number of sparse features could be discarded without loss of performance. A strong regularization led to slight performance degradations, but can be useful on large tasks, where training the full model is not tractable.
Simon Wiesler, Alexander Richard, Yotaro Kubo, Ralf Schlüter, Hermann Ney
ICASSP4
2011 Morpheme Based Factored Language Models for German LVCSR
abstract
German is a highly inflectional language, where a large number of words can be generated from the same root.It makes a liberal use of compounding leading to high Out-of-vocabulary (OOV) rates, and poor Language Model (LM) probability estimates.Therefore, the use of morphemes for language modeling is considered a better choice for Large Vocabulary Continuous Speech Recognition (LVCSR) than the full-words.Thereby, better lexical coverage and less LM perplexities are achieved.On the other side, the use of Factored Language Models (FLMs) is considered a successful approach that allows the integration of many information sources to get better LM probability estimates.In this paper, we try a combined methodology for language modeling where both morphological decomposition and factored language modeling are used in one model called morpheme based FLM.Finally, we obtain around 2.5% relative reduction in Word Error Rate (WER) with respect to a traditional full-words system.
Amr El-Desoky Mousa, M. Ali Basha Shaik, Ralf Schlüter, Hermann Ney
INTERSPEECH3
2011 Acoustic Look-Ahead for More Efficient Decoding in LVCSR
abstract
In this paper we propose novel approximations of a generalized acoustic look-ahead to speed up the search process in large vocabulary continuous speech recognition (LVCSR).Unlike earlier methods, we do not employ any phoneme-or syllable level heuristics.First we define and analyze the perfect acoustic look-ahead as a simple pre-evaluation of the original acoustic models into the future.This method is very slow, but reveals the best possible impact on the search space that can be achieved through acoustic look-ahead.In a second step, we derive efficient and simple approximative look-ahead models from the perfect models.We show that the approximative models compare well to the perfect models regarding the search space, and that the approximative models significantly improve the efficiency in comparison to the baseline, without any negative effect on the precision.
David Nolden, Ralf Schlüter, Hermann Ney
INTERSPEECH2
2011 Compound Word Recombination for German LVCSR
abstract
Compound words are a difficulty for German speech recognition systems since they cause high out-of-vocabulary and word error rates.State of the art approaches augment the language model by the fragments of compounds in order to increase lexical coverage, lower the perplexity and out-of-vocabulary rate.The fragments are tagged in order to concatenate subsequent equally tagged fragments in the recognition result, but this does not guarantee the recombination of proper words.Such recombination techniques neglect the large vocabulary of the language model training data for recombination although most compounds are covered by it.In this paper, we investigate the use of this vocabulary for the recombination of compound words from the recognition result.The approach is tested on two large vocabulary tasks on top of full-word and fragment based language models and achieves good improvements of 3-7% relative over the baseline compound-sensitive word error rate.
Markus Nußbaum-Thom, Amr El-Desoky Mousa, Ralf Schlüter, Hermann Ney
INTERSPEECH3
2011 Improved Acoustic Feature Combination for LVCSR by Neural Networks
abstract
This paper investigates the combination of different acoustic features.Several methods to combine these features such as concatenation or LDA are well known.Even though LDA improves the system, feature combination by LDA has been shown to be suboptimal.We introduce a new method based on neural networks.The posterior estimates derived from the NN lead to a significant improvement and achieve a 6% relative better word error rate (WER).Results are also compared to system combination.While system combination has been reported to outperform all other combination techniques, in this work the proposed NN-based combination outperforms system combination.We achieve a 2% relative better WER, resulting in an improvement of 7% relative to the baseline system.In addition to giving better recognition performance w.r.t.WER, NN-based combination reduces both, training and testing complexity.Overall, we use a single set of acoustic models, together with the training of the NN.
Christian Plahl, Ralf Schlüter, Hermann Ney
INTERSPEECH2
2011 Hybrid Language Models Using Mixed Types of Sub-Lexical Units for Open Vocabulary German LVCSR
abstract
German is a highly inflected language with a large number of words derived from the same root.It makes use of a high degree of word compounding leading to high Out-of-vocabulary (OOV) rates, and Language Model (LM) perplexities.For such languages the use of sub-lexical units for Large Vocabulary Continuous Speech Recognition (LVCSR) becomes a natural choice.In this paper, we investigate the use of mixed types of sub-lexical units in the same recognition lexicon.Namely, morphemic or syllabic units combined with pronunciations called graphones, normal graphemic morphemes or syllables along with full-words.This mixture of units is used for building hybrid LMs suitable for open vocabulary LVCSR where the system operates over an open, constantly changing vocabulary like in broadcast news, political debates, etc.A relative reduction of around 5.0% in Word Error Rate (WER) is obtained compared to a traditional full-words system.Moreover, around 40% of the OOVs are recognized.
M. Ali Basha Shaik, Amr El-Desoky Mousa, Ralf Schlüter, Hermann Ney
INTERSPEECH3
2011 On the Estimation of Discount Parameters for Language Model Smoothing
abstract
The goal of statistical language modeling is to find probability estimates for arbitrary word sequences. To obtain non-zero values, the probability distributions found in the training data need to be smoothed. In the widely-used Kneser-Ney family of smoothing algorithms, this is achieved by absolute discounting. The discount parameters can be computed directly using some approximation formulas minimizing the leaving-one-out log-likelihood of the training data. In this work, we outline several shortcomings of the standard estimators for the discount parameters. We propose an efficient method for computing the discount values on heldout data and analyze the resulting parameter estimates. Experiments on large English and French corpora show consistent improvements in perplexity and word error rate over the baseline method. At the same time, this approach can be used for language model pruning, leading to slightly better results than standard pruning algorithms. Index Terms: language model smoothing, absolute discounting, Kneser-Ney method, language model pruning
Martin Sundermeyer, Ralf Schlüter, Hermann Ney
INTERSPEECH2
2011 Log-Linear Optimization of Second-Order Polynomial Features with Subsequent Dimension Reduction for Speech Recognition
abstract
Second order polynomial features are useful for speech recognition because they can be used to model class specific covariance even with a pooled covariance acoustic model.Previous experiments with second order features have shown word error rate improvements.However, the improvement comes at the price of a large increase in the number of parameters.This paper investigates the discriminative training of second order features, with a subsequent dimension reduction transform to limit the increase in number of parameters.The acoustic model parameters and the transformation matrix parameters are modeled log-linearly and optimized using maximum mutual information criterion.The advantage of log-linear optimization lies in its ability to robustly combine different kinds of features.Experiments are performed for second order MFCC features on the EPPS large vocabulary task and have resulted in a decrease in word error rate.
Muhammad Ali Tahir, Ralf Schlüter, Hermann Ney
INTERSPEECH2
2011 A Study on Speaker Normalized MLP Features in LVCSR
abstract
Different normalization methods are applied in recent Large Vocabulary Continuous Speech Recognition Systems (LVCSR) to reduce the influence of speaker variability on the acoustic models.In this paper we investigate the use of Vocal Tract Length Normalization (VTLN) and Speaker Adaptive Training (SAT) in Multi Layer Perceptron (MLP) feature extraction on an English task.We achieve significant improvements by each normalization method and we gain further by stacking the normalizations.Studying features transformed by Constrained Maximum Likelihood Linear Regression (CMLLR) based SAT as possible input for MLP, further experiments show that MLP could not consistently take advantage of SAT as it does in case of VTLN.
Zoltán Tüske, Christian Plahl, Ralf Schlüter
INTERSPEECH3
2011 Equivalence of Generative and Log-Linear Models
abstract
Conventional speech recognition systems are based on hidden Markov models (HMMs) with Gaussian mixture models (GHMMs). Discriminative log-linear models are an alternative modeling approach and have been investigated recently in speech recognition. GHMMs are directed models with constraints, e.g., positivity of variances and normalization of conditional probabilities, while log-linear models do not use such constraints. This paper compares the posterior form of typical generative models related to speech recognition with their log-linear model counterparts. The key result will be the derivation of the equivalence of these two different approaches under weak assumptions. In particular, we study Gaussian mixture models, part-of-speech bigram tagging models, and eventually, the GHMMs. This result unifies two important but fundamentally different modeling paradigms in speech recognition on the functional level. Furthermore, this paper will present comparative experimental results for various speech tasks of different complexity, including a digit string and large-vocabulary continuous speech recognition tasks.
Georg Heigold, Hermann Ney, Patrick Lehnen, Tobias Gass, Ralf Schlüter
IEEE Trans. Speech Audio Process.5
2011 On the Relationship Between Bayes Risk and Word Error Rate in ASR
abstract
Recently, a number of (approximate) approaches emerged in speech processing, which try to overcome the known lack of match between symbol level evaluation measures (e.g., word error rate) and the standard string (symbol sequence) cost (e.g., sentence error)-based Bayes decision rule, by using symbol level cost functions for Bayes decision rule. Nevertheless, experiments show that for a majority of test samples both decision rules still give equal decisions, especially at lower error rates. In this paper, analytic evidence for these observations is provided. A set of conditions is presented, for which Bayes decision rule with symbol level and string level cost function leads to the same decisions. Furthermore, the case of word error cost represented by the Levenshtein (edit) distance is investigated, which upon others covers the important case of speech recognition. A Hamming distance-based upper bound to the Levenshtein cost function is discussed. This cost function relates to former, word-posterior based decision rules, and the corresponding efficient decision rule is shown to be strongly related to Bayes decision rule with the Levenshtein cost. The analytic results are verified experimentally, and their quantitative effect is studied by experiments on four different well-known large vocabulary automatic speech recognition tasks.
Ralf Schlüter, Markus Nußbaum-Thom, Hermann Ney
IEEE Trans. Speech Audio Process.1
2010 Discriminative HMMS, log-linear models, and CRFS: What is the difference?
abstract
Recently, there have been many papers studying discriminative acoustic modeling techniques like conditional random fields or discriminative training of conventional Gaussian HMMs. This paper will give an overview of the recent work and progress. We will strictly distinguish between the type of acoustic models on the one hand and the training criterion on the other hand. We will address two issues in more detail: the relation between conventional Gaussian HMMs and conditional random fields and the advantages of formulating the training criterion as a convex optimization problem. Experimental results for various speech tasks will be presented to carefully evaluate the different concepts and approaches, including both a digit string and large vocabulary continuous speech recognition tasks.
Georg Heigold, Simon Wiesler, Markus Nußbaum-Thom, Patrick Lehnen, Ralf Schlüter, Hermann Ney
ICASSP5
2010 Discriminative adaptation for log-linear acoustic models
abstract
Log-linear models have recently been used in acoustic modeling for speech recognition systems.This has been motivated by competitive results compared to systems based on Gaussian models, and a more direct parametrisation of the posterior model.To competitively use log-linear models for speech recognition, important methods, such as speaker adaptation, have to be reformulated in a log-linear framework.In this work, an approach to log-linear affine feature transforms for speaker adaptation is described.Experiments for both supervised and unsupervised adaptation are presented, showing improvements over a maximum likelihood baseline in the form of feature space maximum likelihood linear regression for the case of supervised adaptation.
Jonas Lööf, Ralf Schlüter, Hermann Ney
INTERSPEECH2
2010 Time conditioned search in automatic speech recognition reconsidered
abstract
In this paper we re-investigate the time conditioned search (TCS) method in comparison to the well known word conditioned search (WCS), and analyze its applicability on state-ofthe-art large vocabulary continuous speech recognition tasks. In contrast to current standard approaches, time conditioned search offers theoretical advantages particularly in combination with huge vocabularies and huge language models, but it is difficult to combine with across word modelling, which was proven to be an important technique in automatic speech recognition. Our novel contributions for TCS are a pruning step during the recombination called Early Word End Pruning, an additional recombination technique called Context Recombination, the idea of a Startup Interval to reduce the number of started trees, and a mechanism to combine TCS with across word modelling. We show that, with these techniques, TCS can outperform WCS on current ASR tasks. Index Terms: speech recognition, search, word conditioned, time conditioned
David Nolden, Hermann Ney, Ralf Schlüter
INTERSPEECH3
2010 The RWTH 2009 quaero ASR evaluation system for English and German
abstract
In this work, the RWTH automatic speech recognition systems for English and German for the second Quaero evaluation campaign 2009 are presented. The systems are designed to transcribe web data, European parliament plenary sessions and broadcast news data. Another challenge in the 2009 evaluation is that almost no in-domain training data is provided and the test data contains a large variety of speech types. The RWTH participates for the English and German languages with the best results for German and competitive results for the English. Contributing to the enhancements are the systematic use of hierarchical neural network based posterior features, system combination, speaker adaptation, cross speaker adaptation, domain dependent modeling and the usage of additional training data.
Markus Nußbaum-Thom, Simon Wiesler, Martin Sundermeyer, Christian Plahl, Stefan Hahn, Ralf Schlüter, Hermann Ney
INTERSPEECH6
2010 Parallel lexical-tree based LVCSR on multi-core processors
Naveen Parihar, Ralf Schlüter, David Rybach, Eric A. Hansen
INTERSPEECH2
2010 Hierarchical bottle neck features for LVCSR
abstract
This paper investigates the combination of different neural network topologies for probabilistic feature extraction.On one hand, a five-layer neural network used in bottle neck feature extraction allows to obtain arbitrary feature size without dimensionality reduction by transform, independently of the training targets.On the other hand, a hierarchical processing technique is effective and robust over several conditions.Even though the hierarchical and bottle neck processing performs equally well, the combination of both topologies improves the system by 5% relative.Furthermore, the MFCC baseline system is improved by up to 20% relative.This behaviour could be confirmed on two different tasks.In addition, we analyse the influence of multi-resolution RASTA filtering and long-term spectral features as input for the neural network feature extraction.
Christian Plahl, Ralf Schlüter, Hermann Ney
INTERSPEECH2
2010 Revisiting VTLN using linear transformation on conventional MFCC
abstract
In this paper, we revisit the linear transformation for VTLN on conventional MFCC proposed by Sanand et al. in [1], using the idea of band-limited interpolation.The filter-bank is modified to include half-filters at zero and nyquist frequencies, as the full symmetric spectrum is required for performing bandlimited interpolation.In this paper, we show that the filter-bank with half-filters does not affect the recognition performance on clean speech (also shown in [1]), but does affect the recognition performance on noisy speech.This motivated us to revisit the linear transformation for VTLN in [1] and propose modifications to undo the affect of half-filters during the feature extraction.We show through recognition experiments that the proposed modifications to the linear transformation have comparable performance as the conventional VTLN approach, still enabling us to perform VTLN using a linear transformation on conventional MFCC.
Rama Sanand Doddipatla, Ralf Schlüter, Hermann Ney
INTERSPEECH2
2010 On the relation of Bayes risk, word error, and word posteriors in ASR
abstract
In automatic speech recognition, we are faced with a wellknown inconsistency: Bayes decision rule is usually used to minimize sentence (word sequence) error, whereas in practice we want to minimize word error, which also is the usual evaluation measure.Recently, a number of speech recognition approaches to approximate Bayes decision rule with word error (Levenshtein/edit distance) cost were proposed.Nevertheless, experiments show that the decisions often remain the same and that the effect on the word error rate is limited, especially at low error rates.In this work, further analytic evidence for these observations is provided.A set of conditions is presented, for which Bayes decision rule with sentence and word error cost function leads to the same decisions.Furthermore, the case of word error cost is investigated and related to word posterior probabilities.The analytic results are verified experimentally on several large vocabulary speech recognition tasks.
Ralf Schlüter, Markus Nußbaum-Thom, Hermann Ney
INTERSPEECH1
2010 A discriminative splitting criterion for phonetic decision trees
abstract
Phonetic decision trees are a key concept in acoustic modeling for large vocabulary continuous speech recognition.Although discriminative training has become a major line of research in speech recognition and all state-of-the-art acoustic models are trained discriminatively, the conventional phonetic decision tree approach still relies on the maximum likelihood principle.In this paper we develop a splitting criterion based on the minimization of the classification error.An improvement of more than 10% relative over a discriminatively trained baseline system on the Wall Street Journal corpus suggests that the proposed approach is promising.
Simon Wiesler, Georg Heigold, Markus Nußbaum-Thom, Ralf Schlüter, Hermann Ney
INTERSPEECH4
2010 A Hybrid Morphologically Decomposed Factored Language Models for Arabic LVCSR
Amr El-Desoky Mousa, Ralf Schlüter, Hermann Ney
HLT-NAACL2
2010 Evaluation of automatic transcription systems for the judicial domain
abstract
This paper describes two different automatic transcription systems developed for judicial application domains for the Polish and Italian languages. The judicial domain requires to cope with several factors which are known to be critical for automatic speech recognition, such as: background noise, reverberation, spontaneous and accented speech, overlapped speech, cross channel effects, etc. The two automatic speech recognition (ASR) systems have been developed independently starting from out-of-domain data and, then, they have been adapted using a certain amount of in-domain audio and text data. The ASR performance have been measured on audio data acquired in the courtrooms of Naples and Wroclaw. The resulting word error rates are around 40%, for Italian, and around between 30% and 50% for Polish. This performance, similar to that reported for other comparable ASR tasks (e.g. meeting transcriptions with distant microphone), suggests that possible applications can address tasks such as indexing and/or information retrieval in multimedia documents recorded during judicial debates.
Jonas Lööf, Daniele Falavigna, Ralf Schlüter, Diego Giuliani, Roberto Gretter, Hermann Ney
SLT3
2010 Sub-lexical language models for German LVCSR
abstract
One of the major difficulties related to German LVCSR is the rich morphology nature of German, leading to high out-of-vocabulary (OOV) rates, and high language model (LM) perplexities. Normally, compound words make up an essential fraction of the German vocabulary. Most compound OOVs are composed of frequent in-vocabulary words. Here, we investigate the use of sub-lexical LMs based on different approaches for word decomposition, namely supervised and unsupervised decomposition, as well as decomposition derived from grapheme-to-phoneme (G2P) conversion. In the later approach, we augment a normal word model with a set of grapheme-phoneme pairs called graphones used to model the OOV words. A novel approach is proposed to select the representative graphone sequences for OOVs based on unsupervised decomposition and word-pronunciation alignment. We obtain relative reductions in word error rate (WER) from 4.2% to 6.5% with respect to a comparable full-words system.
Amr El-Desoky Mousa, M. Ali Basha Shaik, Ralf Schlüter, Hermann Ney
SLT3
2009 Generalized likelihood ratio discriminant analysis
abstract
Linear Discriminant Analysis (LDA) has been established as an important means for dimension reduction and decorrelation in speech recognition. The major points of criticism of LDA are that it uses an ad hoc and non-discriminative training criterion, and that the estimation is performed in a separate preprocessing step. This paper presents a new discriminative training method for the estimation of (projecting) linear feature transforms. More precisely, the problem is formulated in the loglinear framework, resulting in a convex optimization problem. Experimental results are provided for a digit string recognition task to compare the performance and robustness of the proposed approach (in combination with ML or MMI optimized acoustic models) with conventional LDA. Also, first experiments for a large vocabulary task are presented.
Muhammad Ali Tahir, Georg Heigold, Christian Plahl, Ralf Schlüter, Hermann Ney
ASRU4
2009 Investigations on features for log-linear acoustic models in continuous speech recognition
abstract
Hidden Markov Models with Gaussian Mixture Models as emission probabilities (GHMMs) are the underlying structure of all state-of-the-art speech recognition systems. Using Gaussian mixture distributions follows the generative approach where the class-conditional probability is modeled, although for classification only the posterior probability is needed. Though being very successful in related tasks like Natural Language Processing (NLP), in speech recognition direct modeling of posterior probabilities with log-linear models has rarely been used and has not been applied successfully to continuous speech recognition. In this paper we report competitive results for a speech recognizer with a log-linear acoustic model on the Wall Street Journal corpus, a Large Vocabulary Continuous Speech Recognition (LVCSR) task. We trained this model from scratch, i.e. without relying on an existing GHMM system. Previously the use of data dependent sparse features for log-linear models has been proposed. We compare them with polynomial features and show that the combination of polynomial and data dependent sparse features leads to better results.
Simon Wiesler, Markus Nußbaum-Thom, Georg Heigold, Ralf Schlüter, Hermann Ney
ASRU4
2009 Modified MPE/MMI in a transducer-based framework
abstract
In this paper we show how common training criteria like for example MPE or MMI can be extended to incorporate a margin term. In addition, a transducer-based training implementation is presented, which covers a large variety of discriminative training criteria for ASR, including the standard MMI, MPE, and MCE criteria, as well as the modifications to these criteria presented here. The modified criteria are directly related with the conventional large margin formulation of SVMs. In the proposed approach, we can take advantage of the generalization guarantees of large margin classifiers while keeping the existing framework for the discriminative training, including the efficient algorithms for conventional MPE or MMI. On the conceptual side, this allows for a direct evaluation of the margin term. Finally, experimental results are presented for different large vocabulary continuous speech recognition tasks (one of which is trained on a very large amount of training data) using these modified criteria.
Georg Heigold, Ralf Schlüter, Hermann Ney
ICASSP2
2009 Audio segmentation for speech recognition using segment features
abstract
Audio segmentation is an essential preprocessing step in several audio processing applications with a significant impact e.g. on speech recognition performance. We introduce a novel framework which combines the advantages of different well known segmentation methods. An automatically estimated log-linear segment model is used to determine the segmentation of an audio stream in a holistic way by a maximum a posteriori decoding strategy, instead of classifying change points locally. A comparison to other segmentation techniques in terms of speech recognition performance is presented, showing a promising segmentation quality of our approach.
David Rybach, Christian Gollan, Ralf Schlüter, Hermann Ney
ICASSP3
2009 Investigating the use of morphological decomposition and diacritization for improving Arabic LVCSR
abstract
One of the challenges related to large vocabulary Arabic speech recognition is the rich morphology nature of Arabic language which leads to both high out-of-vocabulary (OOV) rates and high language model (LM) perplexities.Another challenge is the absence of the short vowels (diacritics) from the Arabic written transcripts which causes a large difference between spoken and written language and thus a weaker connection between the acoustic and language models.In this work, we try to address these two important challenges by introducing both morphological decomposition and diacritization in Arabic language modeling.Finally, we are able to obtain about 3.7% relative reduction in word error rate (WER) with respect to a comparable non-diacritized full-words system running on our test set.
Amr El-Desoky Mousa, Christian Gollan, David Rybach, Ralf Schlüter, Hermann Ney
INTERSPEECH4
2009 Investigations on convex optimization using log-linear HMMs for digit string recognition
abstract
Discriminative methods are an important technique to refine the acoustic model in speech recognition.Conventional discriminative training is initialized with some baseline model and the parameters are re-estimated in a separate step.This approach has proven to be successful, but it includes many heuristics, approximations, and parameters to be tuned.This tuning involves much engineering and makes it difficult to reproduce and compare experiments.In contrast to the conventional training, convex optimization techniques provide a sound approach to estimate all model parameters from scratch.Such a straight approach hopefully dispense with additional heuristics, e.g.scaling of posteriors.This paper addresses the question how well this concept using log-linear models carries over to practice.Experimental results are reported for a digit string recognition task, which allows for the investigation of this issue without approximations.
Georg Heigold, David Rybach, Ralf Schlüter, Hermann Ney
INTERSPEECH3
2009 Log-linear model combination with word-dependent scaling factors
abstract
Log-linear model combination is the standard approach in LVCSR to combine several knowledge sources, usually an acoustic and a language model.Instead of using a single scaling factor per knowledge source, we make the scaling factor wordand pronunciation-dependent.In this work, we combine three acoustic models, a pronunciation model, and a language model for a Mandarin BN/BC task.The achieved error rate reduction of 2% relative is small but consistent for two test sets.An analysis of the results shows that the major contribution comes from the improved interdependency of language and acoustic model.
Björn Hoffmeister, Ruoying Liang, Ralf Schlüter, Hermann Ney
INTERSPEECH3
2009 Bayes risk approximations using time overlap with an application to system combination
abstract
The computation of the Minimum Bayes Risk (MBR) decoding rule for word lattices needs approximations. We investigate a class of approximations where the Levenshtein alignment is approximated under the condition that competing lattice arcs overlap in time. The approximations have their origins in MBR decoding and in discriminative training. We develop modified versions and propose a new, conceptually extremely simple confusion network algorithm. The MBR decoding rule is extended to scope with several lattices, which enables us to apply all the investigated approximations to system combination. All approximations are tested on a Mandarin and on an English LVCSR task for a single system and for system combination. The new methods are competitive in error rate and show some advantages over the standard approaches to MBR decoding. Index Terms: speech recognition, minimum bayes risk, confusion network, system combination, discriminative training
Björn Hoffmeister, Ralf Schlüter, Hermann Ney
INTERSPEECH2
2009 Parallel fast likelihood computation for LVCSR using mixture decomposition
Naveen Parihar, Ralf Schlüter, David Rybach, Eric A. Hansen
INTERSPEECH2
2009 Development of the GALE 2008 Mandarin LVCSR system
abstract
This paper describes the current improvements of the RWTH Mandarin LVCSR system.We introduce vocal tract length normalization for the Gammatone features and present comparable results for Gammatone based feature extraction and classical feature extraction.In order to benefit from the huge amount of data of 1600h available in the GALE project we have trained the acoustic models up to 8M Gaussians.We present detailed character error rates for the different number of Gaussians.Different kinds of systems are developed and a two stage decoding framework is applied, which uses cross-adaptation and a subsequent lattice-based system combination.In addition to various acoustic front-ends, these systems use different kinds of neural network toneme posterior features.We present detailed recognition results of the development cycle and the different acoustic front-ends of the systems.Finally, we compare the ultimate evaluation system to our last years system and can report a 10% relative improvement.
Christian Plahl, Björn Hoffmeister, Georg Heigold, Jonas Lööf, Ralf Schlüter, Hermann Ney
INTERSPEECH5
2009 The RWTH aachen university open source speech recognition system
abstract
We announce the public availability of the RWTH Aachen University speech recognition toolkit.The toolkit includes state of the art speech recognition technology for acoustic model training and decoding.Speaker adaptation, speaker adaptive training, unsupervised training, a finite state automata library, and an efficient tree search decoder are notable components.Comprehensive documentation, example setups for training and recognition, and a tutorial are provided to support newcomers.
David Rybach, Christian Gollan, Georg Heigold, Björn Hoffmeister, Jonas Lööf, Ralf Schlüter, Hermann Ney
INTERSPEECH6
2008 A GIS-like training algorithm for log-linear models with hidden variables
abstract
Conditional random fields (CRFs) are often estimated using an entropy based criterion in combination with generalized iterative scaling (GIS). GIS offers, upon others, the immediate advantages that it is locally convergent, completely parameter free, and guarantees an improvement of the criterion in each step. GIS, however, is limited in two aspects. GIS cannot be applied when the model incorporates hidden variables, and it can only be applied to optimize the maximum mutual information criterion (MMI). Here, we extend the GIS algorithm to resolve these two limitations. The new approach allows for training log-linear models with hidden variables and optimizes discriminative training criteria different from maximum mutual information (MMI), including minimum phone error (MPE). The proposed GIS-like method shares the above-mentioned theoretical properties of GIS. The framework is tested for optical character recognition on the USPS task, and for speech recognition on the Sietill task for continuous digit string recognition.
Georg Heigold, Thomas Deselaers, Ralf Schlüter, Hermann Ney
ICASSP3
2008 Modified MMI/MPE: a direct evaluation of the margin in speech recognition
abstract
In this paper we show how common speech recognition training criteria such as the Minimum Phone Error criterion or the Maximum Mutual Information criterion can be extended to incorporate a margin term. Different margin-based training algorithms have been proposed to refine existing training algorithms for general machine learning problems. However, for speech recognition, some special problems have to be addressed and all approaches proposed either lack practical applicability or the inclusion of a margin term enforces significant changes to the underlying model, e.g. the optimization algorithm, the loss function, or the parameterization of the model. In our approach, the conventional training criteria are modified to incorporate a margin term. This allows us to do large-margin training in speech recognition using the same efficient algorithms for accumulation and optimization and to use the same software as for conventional discriminative training. We show that the proposed criteria are equivalent to Support Vector Machines with suitable smooth loss functions, approximating the non-smooth hinge loss function or the hard error (e.g. phone error). Experimental results are given for two different tasks: the rather simple digit string recognition task Sietill which severely suffers from overfitting and the large vocabulary European Parliament Plenary Sessions English task which is supposed to be dominated by the risk and the generalization does not seem to be such an issue.
Georg Heigold, Thomas Deselaers, Ralf Schlüter, Hermann Ney
ICML3
2008 On the equivalence of Gaussian and log-linear HMMs
abstract
The acoustic models of conventional state-of-the-art speech recognition systems use generative Gaussian HMMs. In the past few years, discriminative models like for example Conditional Random Fields (CRFs) have been proposed to refine the acoustic models. CRFs directly model the class posteriors, the quantities of interest in recognition. CRFs are undirected models, and CRFs do not assume local normalization constraints as HMMs do. This paper addresses the issue to what extent such less restricted models add flexiblity to the model compared with the generative counterpart. This work extends our previous work in that it provides the technical details used for showing the equivalence of Gaussian and log-linear HMMs. The correctness of the proposed equivalence transformation for conditional probabilities is demonstrated on a simple concept tagging task.
Georg Heigold, Patrick Lehnen, Ralf Schlüter, Hermann Ney
INTERSPEECH3
2008 iCNC and iROVER: the limits of improving system combination with classification?
abstract
We show how ROVER and confusion network combination (CNC) can be improved with classification. The general idea of improving combination with classification is that each word is assigned to a certain location and at each location a classifier decides which of the provided alternatives is most likely correct. We investigate four variations of this idea and three different classifiers, which are trained on various features derived from ASR lattices. For our experiments, we use highly optimized ROVER and CNC systems as baseline, which already give a relative reduction in WER of more than 20 % for the TC-STAR 2007 English task. With our methods we can further improve the result of the corresponding standard combination method. Index Terms: speech recognition, system combination 1.
Björn Hoffmeister, Ralf Schlüter, Hermann Ney
INTERSPEECH2
2008 Recent improvements of the RWTH GALE Mandarin LVCSR system
abstract
This paper describes the current improvements of the RWTH Mandarin LVCSR system. We introduce a new reduced toneme set developed at RWTH. We are using different toneme sets and pronunciation lexica. For the purpose of discriminative training we will show a fast way to transform word lattices between systems using different toneme sets and pronunciation lexica. In addition to various acoustic front-ends, the current systems use different kinds of neural network toneme posterior features. While different kinds of systems are developed, a two stage decoding framework for combining these systems is applied. We show detailed recognition results of the development cycle of the systems. Finally, two methods to integrate tonal features are compared.
Christian Plahl, Björn Hoffmeister, Mei-Yuh Hwang, Danju Lu, Georg Heigold, Jonas Lööf, Ralf Schlüter, Hermann Ney
INTERSPEECH7
2008 Development of the SRI/nightingale Arabic ASR system
abstract
We describe the large vocabulary automatic speech recognition system developed for Modern Standard Arabic by the SRI/Nightingale team, and used for the 2007 GALE evaluation as part of the speech translation system. We show how system performance is affected by different development choices, ranging from text processing and lexicon to decoding system architecture design. Word error rate results are reported on broadcast news and conversational data from the GALE development and evaluation test sets. Index Terms: speech recognition, large vocabulary, Arabic 1.
Dimitra Vergyri, Arindam Mandal, Wen Wang 0001, Andreas Stolcke, Jing Zheng 0001, Martin Graciarena, David Rybach, Christian Gollan, Ralf Schlüter, Katrin Kirchhoff, Arlo Faria, Nelson Morgan
INTERSPEECH9
2007 Development of the 2007 RWTH Mandarin LVCSR system
abstract
This paper describes the development of the RWTH Mandarin LVCSR system. Different acoustic front-ends together with multiple system cross-adaptation are used in a two stage decoding framework. We describe the system in detail and present systematic recognition results. Especially, we compare a variety of approaches for cross-adapting to multiple systems. During the development we did a comparative study on different methods for integrating tone and phoneme posterior features. Furthermore, we apply lattice based consensus decoding and system combination methods. In these methods, the effect of minimizing character instead of word errors is compared. The final system obtains a character error rate of 17.7% on the GALE 2006 evaluation data.
Björn Hoffmeister, Christian Plahl, Peter Fritz, Georg Heigold, Jonas Lööf, Ralf Schlüter, Hermann Ney
ASRU6
2007 Advances in Arabic broadcast news transcription at RWTH
abstract
This paper describes the RWTH speech recognition system for Arabic. Several design aspects of the system, including cross-adaptation, multiple system design and combination, are analyzed. We summarize the semi-automatic lexicon generation for Arabic using a statistical approach to grapheme-to-phoneme conversion and pronunciation statistics. Furthermore, a novel ASR-based audio segmentation algorithm is presented. Finally, we discuss practical approaches for parallelized acoustic training and memory efficient lattice rescoring. Systematic results are reported on recent GALE evaluation corpora.
David Rybach, Stefan Hahn, Christian Gollan, Ralf Schlüter, Hermann Ney
ASRU4
2007 Cross-Site and Intra-Site ASR System Combination: Comparisons on Lattice and 1-Best Methods
abstract
We evaluate system combination techniques for automatic speech recognition using systems from multiple sites who participated in the TC-STAR 2006 evaluation. Both lattice and 1-best combination techniques are tested for cross-site and intra-site tasks. For pairwise combinations the lattice based approaches can outperform 1-best ROVER with confidence scores, but 1-best ROVER results are equal (or even better) when combining three or four systems.
Björn Hoffmeister, Dustin Hillard, Stefan Hahn, Ralf Schlüter, Mari Ostendorf, Hermann Ney
ICASSP (4)4
2007 Gammatone Features and Feature Combination for Large Vocabulary Speech Recognition
abstract
In this work, an acoustic feature set based on a gammatone filterbank is introduced for large vocabulary speech recognition. The gammatone features presented here lead to competitive results on the EPPS English task, and considerable improvements were obtained by subsequent combination to a number of standard acoustic features, i.e. MFCC, PLP, MF-PLP, and VTLN plus voicedness. Best results were obtained when combining gammatone features to all other features using weighted ROVER, resulting in a relative improvement of about 12% in word error rate compared to the best single feature system. We also found that ROVER gives better results for feature combination than both log-linear model combination and LDA.
Ralf Schlüter, Ilja Bezrukov, Hermann Wagner, Hermann Ney
ICASSP (4)1
2007 An improved method for unsupervised training of LVCSR systems
abstract
In this paper, we introduce an improved method for unsupervised training where the data selection or filtering process is done on state level. We describe in detail the setup of the experiments and introduce the state confidence scores on word and allophone state level for performing the data selection for mixture training on state level. Although we are using a relatively small amount of 180 hours of untranscribed recordings in addition to the available carefully manually transcribed transcriptions of 100 hours, we are able to significantly improve our final speaker adaptive acoustic model. Furthermore, we present promising results by doing system combination using the acoustic models trained on different confidence thresholds. These methods are evaluated on the EPPS corpus starting from the RWTH European English parliamentary speech transcription system. A significant improvement of 7 % relative is achieved using less data for unsupervised training than conventional systems require.
Christian Gollan, Stefan Hahn, Ralf Schlüter, Hermann Ney
INTERSPEECH3
2007 On the equivalence of Gaussian HMM and Gaussian HMM-like hidden conditional random fields
abstract
In this work we show that Gaussian HMMs (GHMMs) are equivalent to GHMM-like Hidden Conditional Random Fields (HCRFs). Hence, improvements of HCRFs over GHMMs found in literature are not due to a refined acoustic modeling but rather come from the more robust formulation of the underlying optimization problem or spurious local optima. Conventional GHMMs are usually estimated with a criterion on segment level whereas hybrid approaches are based on a formulation of the criterion on frame level. In contrast to CRFs, these approaches do not provide scores or do not support more than two classes in a natural way. In this work we analyze these two classes of criteria and propose a refined frame based criterion, which is shown to be an approximation of the associated criterion on segment level. Experimental results concerning these issues are reported for the German digit string recognition task Sietill and the large vocabulary English European Parliament Plenary Sessions (EPPS) task. Index Terms: speech recognition, parameter estimation, maximum entropy methods
Georg Heigold, Ralf Schlüter, Hermann Ney
INTERSPEECH2
2007 The RWTH 2007 TC-STAR evaluation system for european English and Spanish
abstract
In this work, the RWTH automatic speech recognition systems developed for the third TC-STAR evaluation campaign 2007 are presented.The RWTH systems make systematic use of internal system combination, combining systems with differences in feature extraction, adaptation methods, and training data used.To take advantage of this, novel feature extraction methods were employed; this year saw the introduction of Gammatone features and MLP based phone posterior features.Further improvements were achieved using unsupervised training, and it is notable that these improvements were achieved using a fairly low amount of automatically transcribed data.Also contributing to the improvements over last year was the switch to MPE training, and the introduction of projecting SAT transforms.
Jonas Lööf, Christian Gollan, Stefan Hahn, Georg Heigold, Björn Hoffmeister, Christian Plahl, David Rybach, Ralf Schlüter, Hermann Ney
INTERSPEECH8
2007 Efficient estimation of speaker-specific projecting feature transforms
abstract
This paper introduces a new, efficient approach for estimating projecting feature transforms for speech recognition.It is based on the MMI criterion, a likelihood ratio criterion motivated by a simplification of the MMI criterion, and is shown to be closely related to HLDA.In comparison to current methods, the new method is faster, making it more suitable for speaker adaptive training, where the number of speakers, and therefore the number of transforms are substantial.The proposed method was integrated into the RWTH parliamentary speeches transcription system.Experimental results are presented using speaker specific projecting transforms, both when used in recognition only and when used for speaker adaptive training, showing consistent improvements.Furthermore, the observed improvements are shown to be additive to the improvement of MLLR.Comparisons to DLT are presented, and results are presented for a new projecting DLT method.
Jonas Lööf, Ralf Schlüter, Hermann Ney
INTERSPEECH2
2007 Hierarchical neural networks feature extraction for LVCSR system
abstract
This paper investigates the use of a hierarchy of Neural Networks for performing data driven feature extraction.Two different hierarchical structures based on long and short temporal context are considered.Features are tested on two different LVCSR systems for Meetings data (RT05 evaluation data) and for Arabic Broadcast News (BNAT05 evaluation data).The hierarchical NNs outperforms the single NN features consistently on different type of data and tasks and provides significant improvements w.r.t.respective baselines systems.Best results are obtained when different time resolutions are used at different level of the hierarchy.
Fabio Valente, Jithendra Vepa, Christian Plahl, Christian Gollan, Hynek Hermansky, Ralf Schlüter
INTERSPEECH6
2007 Using multiple acoustic feature sets for speech recognition
András Zolnay, Daniil Kocharov, Ralf Schlüter, Hermann Ney
Speech Commun.3
2006 Frame based system combination and a comparison with weighted ROVER and CNC
abstract
In this paper we present a novel ASR system combination technique able to combine systems producing word graphs of different structure and with different segmentations. The new method is based on the definition of a time frame-wise word error cost function in a minimum Bayes risk framework. In contrast to confusion network combination (CNC), it preserves both the word graph structure and the word boundaries. First experimental results are presented on the European Parliament Plenary Sessions (EPPS) task for European Spanish and British English. The new approach to system combination is compared to both ROVER and CNC. In addition, we also apply datadriven weighting schemes for all system combination approaches addressed in this work. For the experiments presented, a variety of internal systems as well as an additional external system were combined. Index Terms: speech recognition, system combination, word posteriors. 1.
Björn Hoffmeister, Tobias Klein, Ralf Schlüter, Hermann Ney
INTERSPEECH3
2006 The 2006 RWTH parliamentary speeches transcription system
abstract
In this work, investigations in the course of the developement of RWTH automatic speech recognition systems developed for the second TC-STAR evaluation campaign 2006 are presented.The systems were designed to transcribe parliamentary speeches taken from the European Parliament Plenary Sessions (EPPS) in European English and Spanish, as well as speeches from the Spanish Parliament.The RWTH systems apply a two pass search strategy with a fourgram one-pass decoder including a fast vocal tract length normalization variant as first pass.The systems further include several adaptation and normalization methods, minimum classification error trained models, and bayes risk minimization.For all relevant individual components contrastive results are presented on the EPPS Spanish and English data, including investigations which did not yet enter the evaluation systems.
Jonas Lööf, Maximilian Bisani, Christian Gollan, Georg Heigold, Björn Hoffmeister, Christian Plahl, Ralf Schlüter, Hermann Ney
INTERSPEECH7
2006 Feature combination using linear discriminant analysis and its pitfalls
abstract
In this paper, Linear Discriminant Analysis (LDA) is investigated with respect to the combination of different acoustic features for automatic speech recognition. It is shown that the combination of acoustic features using LDA does not consistently lead to improvements in word error rate. A detailed analysis of the recognition results on the Verbmobil (VM II) and on the English portion of the European Parliament Plenary Sessions (EPPS) corpus is given. This includes an independent analysis of the effect of the dimension of the input to LDA, the effect of strongly correlated input features, as well as a detailed numerical analysis of the generalized eigenvalue problem underlying LDA. Relative improvements in word error rate of up to 5 % were observed for LDA-based combination of multiple acoustic features. 1.
Ralf Schlüter, András Zolnay, Hermann Ney
INTERSPEECH1
2005 Cross Domain Automatic Transcription on the TC-STAR EPPS Corpus
abstract
This paper describes the ongoing development of the British English European Parliament Plenary Session corpus. This corpus will be part of the speech-to-speech translation evaluation infrastructure of the European TC-STAR project. Furthermore, we present first recognition results on the English speech recordings. The transcription system has been derived from an older speech recognition system built for the North-American broadcast news task. We report on the measures taken for rapid cross-domain porting and present encouraging results.
Christian Gollan, Maximilian Bisani, Stephan Kanthak, Ralf Schlüter, Hermann Ney
ICASSP (1)4
2005 Acoustic Feature Combination for Robust Speech Recognition
abstract
In this paper, we consider the use of multiple acoustic features of the speech signal for robust speech recognition. We investigate the combination of various auditory based (mel frequency cepstrum coefficients, perceptual linear prediction, etc.) and articulatory based (voicedness) features. Features are combined by linear discriminant analysis and log-linear model combination based techniques. We describe the two feature combination techniques and compare the experimental results. Experiments performed on the large-vocabulary task VerbMobil II (German conversational speech) show that the accuracy of automatic speech recognition systems can be improved by the combination of different acoustic features.
András Zolnay, Ralf Schlüter, Hermann Ney
ICASSP (1)2
2005 Articulatory motivated acoustic features for speech recognition
abstract
In this paper, we consider the use of multiple acoustic features of the speech signal for continuous speech recognition. A novel articulatory motivated acoustic feature is introduced, namely the spectrum derivative feature. The new feature is tested in combination with the standard Mel Frequency Cepstral Coefficients (MFCC) and the voicedness features. Linear Discriminant Analysis is applied to find the optimal combination of different acoustic features. Experiments have been performed on small and large vocabulary tasks. Significant improvements in word error rate have been obtained by combining the MFCC feature with the articulatory motivated voicedness and spectrum derivative features: improvements of up to 25 % on the small-vocabulary task and improvements of up to 4 % on the large-vocabulary task relative to using MFCC alone with the same overall number of parameters in the system. 1.
Daniil Kocharov, András Zolnay, Ralf Schlüter, Hermann Ney
INTERSPEECH3
2005 Investigations on error minimizing training criteria for discriminative training in automatic speech recognition
abstract
Discriminative training criteria have been shown to consistently outperform maximum likelihood trained speech recognition systems. In this paper we employ the Minimum Classification Error (MCE) criterion to optimize the parameters of the acoustic model of a large scale speech recognition system. The statistics for both the correct and the competing model are solely collected on word lattices without the use of N-best lists. Thus, particularly for long utterances, the number of sentence alternatives taken into account is significantly larger compared to N-best lists. The MCE criterion is embedded in an extended unifying approach for a class of discriminative training criteria which allows for direct comparison of the performance gain obtained with the improvements of other commonly used criteria such as Maximum Mutual Information (MMI) and Minimum Word Error (MWE). Experiments conducted on large vocabulary tasks show a consistent performance gain for MCE over MMI. Moreover, the improvements obtained with MCE turn out to be in the same order of magnitude as the performance gains obtained with the MWE criterion. 1.
Wolfgang Macherey, Lars Haferkamp, Ralf Schlüter, Hermann Ney
INTERSPEECH3
2005 Bayes risk minimization using metric loss functions
abstract
In this work, fundamental properties of Bayes decision rule using general loss functions are derived analytically and are verified experimentally for automatic speech recognition. It is shown that, for maximum posterior probabilities larger than 1/2, Bayes decision rule with a metric loss function always decides on the posterior maximizing class independent of the specific choice of (metric) loss function. Also for maximum posterior probabilities less than 1/2, a condition is derived under which the Bayes risk using a general metric loss function is still minimized by the posterior maximizing class. For a speech recognition task with low initial word error rate, it is shown that nearly 2/3 of the test utterances fulfil these conditions and need not be considered for Bayes risk minimization with Levenshtein loss, which reduces the computational complexity of Bayes risk minimization. In addition, bounds for the difference between the Bayes risk for the posterior maximizing class and minimum Bayes risk are derived, which can serve as cost estimates for Bayes risk minimization approaches.
Ralf Schlüter, T. Scharrenbach, Volker Steinbiss, Hermann Ney
INTERSPEECH1
2004 Discriminative training with tied covariance matrices
abstract
Discriminative training techniques have proved to be a powerful method for improving large vocabulary speech recognition systems based on Gaussian mixture hidden Markov models.Typically, the optimization of discriminative objective functions is done using the extended Baum algorithm.Since for continuous distributions no proof of fast and stable convergence is known up to now, parameter re-estimation depends on setting the iteration constants in the update rules heuristically, ensuring that the new variances are positive definite.In case of density specific variances this leads to a system of quadratic inequalities.However, if tied variances are used, the inequalities become more complicated and often the resulting constants are too large to be appropriate for discriminative training.In this paper we present an alternative approach to setting the iteration constants to alleviate this problem.First experimental results show that the new method leads to improved convergence speed and test set performance.
Wolfgang Macherey, Ralf Schlüter, Hermann Ney
INTERSPEECH2
2003 Extraction methods of voicing feature for robust speech recognition
abstract
In this paper, three different voicing features are studied as additional acoustic features for continuous speech recognition.The harmonic product spectrum based feature is extracted in frequency domain while the autocorrelation and the average magnitude difference based methods work in time domain.The algorithms produce a measure of voicing for each time frame.The voicing measure was combined with the standard Mel Frequency Cepstral Coefficients (MFCC) using linear discriminant analysis to choose the most relevant features.Experiments have been performed on small and large vocabulary tasks.The three different voicing measures combined with MFCCs resulted in similar improvements in word error rate: improvements of up to 14% on the smallvocabulary task and improvements of up to 6% on the largevocabulary task relative to using MFCC alone with the same overall number of parameters in the system.
András Zolnay, Ralf Schlüter, Hermann Ney
INTERSPEECH2
2002 Robust speech recognition using a voiced-unvoiced feature
abstract
In this paper, a voiced-unvoiced measure is used as acoustic feature for continuous speech recognition. The voiced-unvoiced measure was combined with the standard Mel Frequency Cepstral Coefficients (MFCC) using linear discriminant analysis (LDA) to choose the most relevant features. Experiments were performed on the SieTill (German digit strings recorded over telephone line) and on the SPINE (English spontaneous speech under different simulated noisy environments) corpus. The additional voiced-unvoiced measure results in improvements in word error rate (WER) of up to 11% relative to using MFCC alone with the same overall number of parameters in the system.
András Zolnay, Ralf Schlüter, Hermann Ney
INTERSPEECH2
2001 Computing Mel-frequency cepstral coefficients on the power spectrum
abstract
We present a method to derive Mel-frequency cepstral coefficients directly from the power spectrum of a speech signal. We show that omitting the filterbank in signal analysis does not affect the word error rate. The presented approach simplifies the speech recognizers front end by merging subsequent signal analysis steps into a single one. It avoids possible interpolation and discretization problems and results in a compact implementation. We show that frequency warping schemes like vocal tract normalization can be integrated easily in our concept without additional computational efforts. Recognition test results obtained with the RWTH large vocabulary speech recognition system are presented for two different corpora: The German VerbMobil II dev99 corpus, and the English North American Business News 94 20k development corpus.
Sirko Molau, Michael Pitz, Ralf Schlüter, Hermann Ney
ICASSP3
2001 Using phase spectrum information for improved speech recognition performance
abstract
New acoustic features for continuous speech recognition based on the short-term Fourier phase spectrum are introduced for mono (telephone) recordings. The new phase based features were combined with standard Mel Frequency Cepstral Coefficients (MFCC), and results were produced with and without using additional linear discriminant analysis (LDA) to choose the most relevant features. Experiments were performed on the SieTill corpus for telephone line recorded German digit strings. Using LDA to combine purely phase based features with MFCCs, we obtained improvements in word error rate of up to 25% relative to using MFCCs alone with the same overall number of parameters in the system.
Ralf Schlüter, Hermann Ney
ICASSP1
2001 Explicit word error minimization using word hypothesis posterior probabilities
abstract
We introduce a new concept, the time frame error rate. We show that this error rate is closely correlated with the word error rate and use it to overcome the mismatch between Bayes' decision rule which aims at minimizing the expected sentence error rate and the word error rate which is used to assess the performance of speech recognition systems. Based on the time frame errors we derive a new decision rule and show that the word error rate can be reduced consistently with it on various recognition tasks. All stochastic models are left completely unchanged. We present experimental results on five corpora, the Dutch Arise corpus, the German Verbmobil '98 corpus, the English North American Business '94 20k and 64k development corpora, and the English Broadcast News '96 corpus. The relative reduction of the word error rate ranges from 2.3% to 5.1%.
Frank Wessel, Ralf Schlüter, Hermann Ney
ICASSP2
2001 Vocal tract normalization equals linear transformation in cepstral space
Michael Pitz, Sirko Molau, Ralf Schlüter, Hermann Ney
INTERSPEECH3
2001 Comparison of discriminative training criteria and optimization methods for speech recognition
Ralf Schlüter, Wolfgang Macherey, Boris Müller, Hermann Ney
Speech Commun.1
2001 Model-based MCE bound to the true Bayes' error
abstract
We show that the minimum classification error (MCE) criterion gives an upper bound to the true Bayes' error rate independent of the corresponding model distribution. In addition, we show that model-free optimization of the MCE criterion leads to a closed form solution in the asymptotic case of infinite training data. While leading to the Bayes' error rate, the resulting model distribution differs from the true distribution. This suggests that the structure of model distributions trained with the MCE criterion should differ from the structure of the true distributions, as they are usually used in statistical pattern recognition.
Ralf Schlüter, Hermann Ney
IEEE Signal Process. Lett.1
2001 Confidence measures for large vocabulary continuous speech recognition
abstract
In this paper, we present several confidence measures for large vocabulary continuous speech recognition. We propose to estimate the confidence of a hypothesized word directly as its posterior probability, given all acoustic observations of the utterance. These probabilities are computed on word graphs using a forward-backward algorithm. We also study the estimation of posterior probabilities on N-best lists instead of word graphs and compare both algorithms in detail. In addition, we compare the posterior probabilities with two alternative confidence measures, i.e., the acoustic stability and the hypothesis density. We present experimental results on five different corpora: the Dutch ARISE 1k evaluation corpus, the German Verbmobil '98 7k evaluation corpus, the English North American Business '94 20k and 64k development corpora, and the English Broadcast News '96 65k evaluation corpus. We show that the posterior probabilities computed on word graphs outperform all other confidence measures. The relative reduction in confidence error rate ranges between 19% and 35% compared to the baseline confidence error rate.
Frank Wessel, Ralf Schlüter, Klaus Macherey, Hermann Ney
IEEE Trans. Speech Audio Process.2
2000 Recent improvements of the RWTH large vocabulary speech recognition system on spontaneous speech
abstract
The paper presents recent improvements of the RWTH large vocabulary continuous speech recognition system (LVCSR). In particular, we report on the integration of across-word models into the first recognition pass, and describe better algorithms for fast vocal tract normalization (VTN). We focus both on improvements in word error rate and how to speed up the recognizer with only minimal loss of recognition accuracy. Implementation details and experimental results are given for the VerbMobil task, a German spontaneous speech corpus. The 25.0% word error rate (WER) of our within-word baseline system was reduced to 21.4% with VTN and across-word models. Decreasing the real-time factor (RTF) by up to 85% resulted in only a small degradation in recognition performance of 2% relative on average.
Achim Sixtus, Sirko Molau, Stephan Kanthak, Ralf Schlüter, Hermann Ney
ICASSP4
2000 Using posterior word probabilities for improved speech recognition
abstract
In this paper we present a new scoring scheme for speech recognition. Instead of using the joint probability of a word sequence and a sequence of acoustic observations, we determine the best path through a word graph using posterior word probabilities. These probabilities are computed beforehand with a modified forward-backward algorithm. It is important to note that during the search for the best path no language model is needed because it is already considered in the posterior word probabilities. Subsequent modules can thus process these word graphs very efficiently. Also, confidence measures can be computed from the posterior word probabilities with no additional cost. We present experimental results on five corpora, the Dutch Arise corpus, the German Verbmobil '98 corpus, the English North American Business '94 20 k and 64 k development corpora, and the English Broadcast News '96 corpus. The relative reduction in word error rate ranges between 1.5% and 5.0%.
Frank Wessel, Ralf Schlüter, Hermann Ney
ICASSP2
2000 Speech recognition using context conditional word posterior probabilities
abstract
, and all����. an acoustic observation sequence sequence: In this paper two new scoring schemes for large vocabulary continuous speech recognition are compared. Instead of using the joint probability of a word sequence and a sequence of acoustic observations, we determine the best path through a word graph using posterior word probabilities with or without word context. The exact calculation of the posterior probability for a word sequence implies a sum over all possible word boundaries, which is approximated by a maximum operation in the standard scoring approach. The new scoring scheme using word posterior probabilities could be expected to lead to improved recognition performance, because it involves partial summation over word boundaries. We present experimental results on five different corpora, the Dutch Arise corpus, the German Verbmobil ’98 corpus, the English North American Business ’94 20k and 64k development corpora, and the English Broadcast News ’96 corpus. It is shown that the Viterbi approximation within words has no effect on standard and word posterior based recognition. Using word posterior probabilities with and without word context, the relative reduction in word error rate is comparable and ranges between 1.5% and 5%. A reason why the additional consideration of word context does not further improve the recognition performance might be that the increase in word context information is traded against a decrease in the number of word sequences that contributes to a particular word posterior probability. 1.
Ralf Schlüter, Frank Wessel, Hermann Ney
INTERSPEECH1
1999 A combined maximum mutual information and maximum likelihood approach for mixture density splitting
abstract
In this paper, we present a markovian system, which takes into account pronunciation variation of French. After doing a brief overview of different methods allowing to deal with pronunciation variation in ASR, we describe our approach (which is based on the MHAT (Markovian Harmonic Adaptation and Transduction) model), as well as the lexical and phonological materials defined in order to implement MHAT into a classic ASR system based on HMM models. We finally compare two approaches (both issue from the MHAT model) that differ each other by the level of pronunciation modeling : at the lexicon level and at the language model level (by introducing an intermediate level of words representations depending on the context of words in the sentence). Results show an improvement of French continuous speech recognition when taking into account the context of words in the sentence within the language model.
Ralf Schlüter, Wolfgang Macherey, Boris Müller, Hermann Ney
EUROSPEECH1
1998 Comparison of discriminative training criteria
abstract
A formally unifying approach for a class of discriminative training criteria including maximum mutual information (MMI) and minimum classification error (MCE) criterion is presented, together with the optimization methods of the gradient descent (GD) and extended Baum-Welch (EB) algorithm. Comparisons are discussed for the MMI and the MCE criterion, including the determination of the sets of word sequence hypotheses for discrimination using word graphs. Experiments have been carried out on the SieTill corpus for telephone line recorded German continuous digit strings. Using several approaches for acoustic modeling, the word error rates obtained by MMI training using single densities always were better than those for maximum likelihood (ML) using mixture densities. Finally, the results obtained for corrective training (CT), i.e. using only the best recognized word sequence in addition to the spoken word sequence, could not be improved by using the word graph based discriminative training.
Ralf Schlüter, Wolfgang Macherey
ICASSP1
1998 Using word probabilities as confidence measures
abstract
Estimates of confidence for the output of a speech recognition system can be used in many practical applications of speech recognition technology. They can be employed for detecting possible errors and can help to avoid undesirable verification turns in automatic inquiry systems. We propose to estimate the confidence in a hypothesized word as its posterior probability, given all acoustic feature vectors of the speaker utterance. The basic idea of our approach is to estimate the posterior word probabilities as the sum of all word hypothesis probabilities which represent the occurrence of the same word in more or less the same segment of time. The word hypothesis probabilities are approximated by paths in a wordgraph and are computed using a simplified forward-backward algorithm. We present experimental results on the North American Business (NAB'94) and the German Verbmobil recognition task.
Frank Wessel, Klaus Macherey, Ralf Schlüter
ICASSP3
1997 Comparison of optimization methods for discriminative training criteria
abstract
In this work we compare two parameter optimization techniques for discriminative training using the MMI criterion: the extended Baum-Welch (EBW) algorithm and the generalized probabilistic descent (GPD) method.Using Gaussian emission densities we found special expressions for the step sizes in GPD, leading to reestimation formula very similar to those derived for the EBW algorithm.Results were produced for both the TI digitstring and the SieTill corpus for continuously spoken American English and German digitstrings.The results for both techniques do not show signi cant dierences.This experimental results support the strong link between EBW and GPD as expected from the analytic comparison.
Ralf Schlüter, Wolfgang Macherey, Stephan Kanthak, Hermann Ney, Lutz Welling
EUROSPEECH1