VLDB 2026 Research / reviewers in the wild / expert
Hermann Ney
dblp:n/HermannNey
· DBLP profile ↗
559ranked-venue papers
39as first author
40since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 387 · 18 first-author · 23 since 2021Graphics, computer vision, multimedia, augmented reality and games · 377 · 27 first-author · 36 since 2021Databases, data management, data science and information retrieval · 14Human-computer interaction and ubiquitous computing · 2 · 2 first-authorTheory of computation · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Right Label Context in End-to-End Training of Time-Synchronous ASR ModelsabstractCurrent time-synchronous sequence-to-sequence automatic speech recognition (ASR) models are trained by using sequence level cross-entropy that sums over all alignments. Due to the discriminative formulation, incorporating the right label context into the training criterion’s gradient causes normalization problems and is not mathematically well-defined. The classic hybrid neural network hidden Markov model (NN-HMM) with its inherent generative formulation enables conditioning on the right label context. However, due to the HMM state-tying the identity of the right label context is never modeled explicitly. In this work, we propose a factored loss with auxiliary left and right label contexts that sums over all alignments. We show that the inclusion of the right label context is particularly beneficial when training data resources are limited. Moreover, we also show that it is possible to build a factored hybrid HMM system by relying exclusively on the full-sum criterion. Experiments were conducted on Switchboard 300h and LibriSpeech 960h. Tina Raissi, Ralf Schlüter, Hermann Ney |
ICASSP | 3 |
| 2025 | The Conformer Encoder May Reverse the Time DimensionabstractWe sometimes observe monotonically decreasing cross-attention weights in our Conformer-based global attention-based encoder-decoder (AED) models, negatively affecting performance compared to monotonically increasing attention weights. Further investigation shows that the Conformer encoder reverses the sequence in the time dimension. We analyze the initial behavior of the decoder cross-attention mechanism and find that it encourages the Conformer encoder self-attention to build a connection between the initial frames and all other informative frames. Furthermore, we show that, at some point in training, the self-attention module of the Conformer starts dominating the output over the preceding feed-forward module, which then only allows the reversed information to pass through. We propose methods and ideas of how this flipping can be avoided and investigate a novel method to obtain label-frame-position alignments by using the gradients of the label log probabilities w.r.t. the encoder input frames. Robin Schmitt, Albert Zeyer, Mohammad Zeineldeen, Ralf Schlüter, Hermann Ney |
ICASSP | 5 |
| 2025 | Classification Error Bound for Low Bayes Error Conditions in Machine LearningabstractIn statistical classification and machine learning, classification error is an important performance measure, which is minimized by the Bayes decision rule. In practice, the unknown true distribution is usually replaced with a model distribution estimated from the training data in the Bayes decision rule. This substitution introduces a mismatch between the Bayes error and the model-based classification error. In this work, we apply classification error bounds to study the relationship between the error mismatch and the Kullback-Leibler divergence in machine learning. Motivated by recent observations of low model-based classification errors in many machine learning tasks, bounding the Bayes error to be lower, we propose a linear approximation of the classification error bound for low Bayes error conditions. Then, the bound for class priors are discussed. Moreover, we extend the classification error bound for sequences. Using automatic speech recognition as a representative example of machine learning applications, this work analytically discusses the correlations among different performance measures with extended bounds, including cross-entropy loss, language model perplexity, and word error rate. Vahe Eminyan, Ralf Schlüter, Hermann Ney |
ICASSP | 4 |
| 2025 | Dynamic Acoustic Model Architecture Optimization in Training for ASR
Jingjing Xu 0002, Albert Zeyer, Eugen Beck, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 6 |
| 2025 | Regularizing Learnable Feature Extraction for Automatic Speech RecognitionabstractNeural front-ends are an appealing alternative to traditional, fixed feature extraction pipelines for automatic speech recognition (ASR) systems since they can be directly trained to fit the acoustic model. However, their performance often falls short compared to classical methods, which we show is largely due to their increased susceptibility to overfitting. This work therefore investigates regularization methods for training ASR models with learnable feature extraction front-ends. First, we examine audio perturbation methods and show that larger relative improvements can be obtained for learnable features. Additionally, we identify two limitations in the standard use of SpecAugment for these front-ends and propose masking in the short time Fourier transform (STFT)-domain as a simple but effective modification to address these challenges. Finally, integrating both regularization approaches effectively closes the performance gap between traditional and learnable features. Peter Vieting, Maximilian Kannen, Benedikt Hilmes, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 5 |
| 2025 | Label-Context-Dependent Internal Language Model Estimation for CTC
Minh-Nghia Phan, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 4 |
| 2024 | On the Relation Between Internal Language Model and Sequence Discriminative Training for Neural TransducersabstractInternal language model (ILM) subtraction has been widely applied to improve the performance of the RNN-Transducer with external language model (LM) fusion for speech recognition. In this work, we show that sequence discriminative training has a strong correlation with ILM subtraction from both theoretical and empirical points of view. Theoretically, we derive that the global optimum of maximum mutual information (MMI) training shares a similar formula as ILM subtraction. Empirically, we show that ILM subtraction and sequence discriminative training achieve similar effects across a wide range of experiments on Librispeech, including both MMI and minimum Bayes risk (MBR) criteria, as well as neural transducers and LMs of both full and limited context. The benefit of ILM subtraction also becomes much smaller after sequence discriminative training. We also provide an indepth study to show that sequence discriminative training has a minimal effect on the commonly used zero-encoder ILM estimation, but a joint effect on both encoder and prediction + joint network for posterior probability reshaping including both ILM and blank suppression. Wei Zhou 0043, Ralf Schlüter, Hermann Ney |
ICASSP | 4 |
| 2024 | Chunked Attention-Based Encoder-Decoder Model for Streaming Speech RecognitionabstractWe study a streamable attention-based encoder-decoder model in which either the decoder, or both the encoder and decoder, operate on pre-defined, fixed-size windows called chunks. A special end-of-chunk (EOC) symbol advances from one chunk to the next chunk, effectively replacing the conventional end-of-sequence symbol. This modification, while minor, situates our model as equivalent to a transducer model that operates on chunks instead of frames, where EOC corresponds to the blank symbol. We further explore the remaining differences between a standard transducer and our model. Additionally, we examine relevant aspects such as long-form speech generalization, beam size, and length normalization. Through experiments on Librispeech and TED-LIUM-v2, and by concatenating consecutive sequences for long-form trials, we find that our streamable model maintains competitive performance compared to the non-streamable variant and generalizes very well to long-form speech. Mohammad Zeineldeen, Albert Zeyer, Ralf Schlüter, Hermann Ney |
ICASSP | 4 |
| 2024 | Investigating the Effect of Label Topology and Training Criterion on ASR Performance and Alignment Quality
Tina Raissi, Christoph Lüscher, Simon Berger, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 5 |
| 2024 | Refined Statistical Bounds for Classification Error Mismatches with Constrained Bayes ErrorabstractIn statistical classification/multiple hypothesis testing and machine learning, a model distribution estimated from the training data is usually applied to replace the unknown true distribution in the Bayes decision rule, which introduces a mismatch between the Bayes error and the model-based classification error. In this work, we derive the classification error bound to study the relationship between the Kullback-Leibler divergence and the classification error mismatch. We first reconsider the statistical bounds based on classification error mismatch derived in previous works, employing a different method of derivation. Then, motivated by the observation that the Bayes error is typically low in machine learning tasks like speech recognition and pattern recognition, we derive a refined Kullback-Leibler-divergence-based bound on the error mismatch with the constraint that the Bayes error is lower than a threshold. Vahe Eminyan, Ralf Schlüter, Hermann Ney |
ITW | 4 |
| 2024 | Task-Oriented Document-Grounded Dialog Systems by HLTPR@RWTH for DSTC9 and DSTC10abstractThis paper summarizes our contributions to the document-grounded dialog tasks at the 9th and 10th Dialog System Technology Challenges (DSTC9 and DSTC10). In both iterations the task consists of three subtasks: first detect whether the current turn is knowledge seeking, second select a relevant knowledge document, and third generate a response grounded on the selected document. For DSTC9 we proposed different approaches to make the selection task more efficient. The best method, Hierarchical Selection, actually improves the results compared to the original baseline and gives a speedup of 24x. In the DSTC10 iteration of the task, the challenge was to adapt systems trained on written dialogs to perform well on noisy automatic speech recognition transcripts. Therefore, we proposed data augmentation techniques to increase the robustness of the models as well as methods to adapt the style of generated responses to fit well into the proceeding dialog. Additionally, we proposed a noisy channel model that allows for increasing the factuality of the generated responses. In addition to summarizing our previous contributions, in this work, we also report on a few small improvements and reconsider the automatic evaluation metrics for the generation task which have shown a low correlation to human judgments. David Thulke, Nico Daheim, Christian Dugast, Hermann Ney |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2023 | End-To-End Training of a Neural HMM with Label and Transition ProbabilitiesabstractWe investigate a novel modeling approach for end-to-end neural network training using hidden Markov models (HMM) where the transition probabilities between hidden states are modeled and learned explicitly. Most contemporary sequence-to-sequence models allow for from-scratch training by summing over all possible label segmentations in a given topology. In our approach there are explicit, learnable probabilities for transitions between segments as opposed to a blank label that implicitly encodes duration statistics.We implement a GPU-based forward-backward algorithm that enables the simultaneous training of label and transition probabilities.We investigate recognition results and additionally Viterbi alignments of our models. We find that while the transition model training does not improve recognition performance, it has a positive impact on the alignment quality. The generated alignments are shown to be viable targets in state-of-the-art Viterbi trainings. Daniel Mann, Tina Raissi, Wilfried Michel, Ralf Schlüter, Hermann Ney |
ASRU | 5 |
| 2023 | Investigating The Effect of Language Models in Sequence Discriminative Training For Neural TransducersabstractIn this work, we investigate the effect of language models (LMs) with different context lengths and label units (phoneme vs. word) used in sequence discriminative training for phoneme-based neural transducers. Both lattice-free and N-best-list approaches are examined. For lattice-free methods with phoneme-level LMs, we propose a method to approximate the context history to employ LMs with full-context dependency. This approximation can be extended to arbitrary context length and enables the usage of word-level LMs in lattice-free methods. Moreover, a systematic comparison is conducted across lattice-free and N-best-list-based methods. Experimental results on Librispeech show that using the word-level LM in training outperforms the phoneme-level LM. Besides, we find that the context size of the LM used for probability computation has a limited effect on performance. Moreover, our results reveal the pivotal importance of the hypothesis space quality in sequence discriminative training. Wei Zhou 0043, Ralf Schlüter, Hermann Ney |
ASRU | 4 |
| 2023 | Lattice-Free Sequence Discriminative Training for Phoneme-Based Neural TransducersabstractRecently, RNN-Transducers have achieved remarkable results on various automatic speech recognition tasks. However, lattice-free sequence discriminative training methods, which obtain superior performance in hybrid models, are rarely investigated in RNN-Transducers. In this work, we propose three lattice-free training objectives, namely lattice-free maximum mutual information, lattice-free segment-level minimum Bayes risk, and lattice-free minimum Bayes risk, which are used for the final posterior output of the phoneme-based neural transducer with a limited context dependency. Compared to criteria using N-best lists, lattice-free methods eliminate the decoding step for hypotheses generation during training, which leads to more efficient training. Experimental results show that lattice-free methods gain up to 6.5% relative improvement in word error rate compared to a sequence-level cross-entropy trained model. Compared to the N-best-list based minimum Bayes risk objectives, lattice-free methods gain 40% - 70% relative training time speedup with a small degradation in performance. Wei Zhou 0043, Ralf Schlüter, Hermann Ney |
ICASSP | 4 |
| 2023 | Enhancing and Adversarial: Improve ASR with Speaker LabelsabstractASR can be improved by multi-task learning (MTL) with domain enhancing or domain adversarial training, which are two opposite objectives with the aim to increase/decrease domain variance towards domain-aware/agnostic ASR, respectively. In this work, we study how to best apply these two opposite objectives with speaker labels to improve conformer-based ASR. We also propose a novel adaptive gradient reversal layer for stable and effective adversarial training without tuning effort. Detailed analysis and experimental verification are conducted to show the optimal positions in the ASR neural network (NN) to apply speaker enhancing and adversarial training. We also explore their combination for further improvement, achieving the same performance as i-vectors plus adversarial training. Our best speaker-based MTL achieves 7% relative improvement on the Switchboard Hub5’00 set. We also investigate the effect of such speaker-based MTL w.r.t. cleaner dataset and weaker ASR NN. Wei Zhou 0043, Jingjing Xu 0002, Mohammad Zeineldeen, Christoph Lüscher, Ralf Schlüter, Hermann Ney |
ICASSP | 7 |
| 2023 | RASR2: The RWTH ASR Toolkit for Generic Sequence-to-sequence Speech RecognitionabstractModern public ASR tools usually provide rich support for training various sequence-to-sequence (S2S) models, but rather simple support for decoding open-vocabulary scenarios only.For closed-vocabulary scenarios, public tools supporting lexicalconstrained decoding are usually only for classical ASR, or do not support all S2S models.To eliminate this restriction on research possibilities such as modeling unit choice, we present RASR2 in this work, a research-oriented generic S2S decoder implemented in C++.It offers a strong flexibility/compatibility for various S2S models, language models, label units/topologies and neural network architectures.It provides efficient decoding for both open-and closed-vocabulary scenarios based on a generalized search framework with rich support for different search modes and settings.We evaluate RASR2 with a wide range of experiments on both switchboard and Librispeech corpora.Our source code is public online. Wei Zhou 0043, Eugen Beck, Simon Berger, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 5 |
| 2023 | Competitive and Resource Efficient Factored Hybrid HMM Systems are Simpler Than You Think
Tina Raissi, Christoph Lüscher, Moritz Gunz, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 5 |
| 2022 | Does Joint Training Really Help Cascaded Speech Translation?abstractCurrently, in speech translation, the straightforward approach -cascading a recognition system with a translation system -delivers state-of-theart results.However, fundamental challenges such as error propagation from the automatic speech recognition system still remain.To mitigate these problems, recently, people turn their attention to direct data and propose various joint training methods.In this work, we seek to answer the question of whether joint training really helps cascaded speech translation.We review recent papers on the topic and also investigate a joint training criterion by marginalizing the transcription posterior probabilities.Our findings show that a strong cascaded baseline can diminish any improvements obtained using joint training, and we suggest alternatives to joint training.We hope this work can serve as a refresher of the current speech translation landscape, and motivate research in finding more efficient and creative ways to utilize the direct data for speech translation. Viet Anh Khoa Tran, David Thulke, Yingbo Gao, Christian Herold, Hermann Ney |
EMNLP | 5 |
| 2022 | Improving Factored Hybrid HMM Acoustic Modeling without State TyingabstractIn this work, we show that a factored hybrid hidden Markov model (FH-HMM) which is defined without any phonetic state-tying outperforms a state-of-the-art hybrid HMM. The factored hybrid HMM provides a link to transducer models in the way it models phonetic (label) context while preserving the strict separation of acoustic and language model of the hybrid HMM approach. Furthermore, we show that the factored hybrid model can be trained from scratch without using phonetic state-tying in any of the training steps. Our modeling approach enables triphone context while avoiding phonetic state-tying by a decomposition into locally normalized factored posteriors for monophones/HMM states in phoneme context. Experimental results are provided for Switchboard 300h and LibriSpeech. On the former task we also show that by avoiding the phonetic state-tying step, the factored hybrid can take better advantage of regularization techniques during training, compared to the standard hybrid HMM with phonetic state-tying based on classification and regression trees (CART). Tina Raissi, Eugen Beck, Ralf Schlüter, Hermann Ney |
ICASSP | 4 |
| 2022 | Efficient Sequence Training of Attention Models Using Approximative RecombinationabstractSequence discriminative training is a great tool to improve the performance of an automatic speech recognition system. It does, however, necessitate a sum over all possible word sequences, which is intractable to compute in practice. Current state-of-the-art systems with unlimited label context circumvent this problem by limiting the summation to an n-best list of relevant competing hypotheses obtained from beam search.This work proposes to perform (approximative) recombinations of hypotheses during beam search, if they share a common local history. The error that is incurred by the approximation is analyzed and it is shown that using this technique the effective beam size can be increased by several orders of magnitude without significantly increasing the computational requirements. Lastly, it is shown that this technique can be used to effectively perform sequence discriminative training for attention-based encoder-decoder acoustic models on the LibriSpeech task. Nils-Philipp Wynands, Wilfried Michel, Jan Rosendahl, Ralf Schlüter, Hermann Ney |
ICASSP | 5 |
| 2022 | Conformer-Based Hybrid ASR System For Switchboard DatasetabstractThe recently proposed conformer architecture has been successfully used for end-to-end automatic speech recognition (ASR) architectures achieving state-of-the-art performance on different datasets. To our best knowledge, the impact of using conformer acoustic model for hybrid ASR is not investigated. In this paper, we present and evaluate a competitive conformer-based hybrid model training recipe. We study different training aspects and methods to improve worderror-rate as well as to increase training speed. We apply time downsampling methods for efficient training and use transposed convolutions to upsample the output sequence again. We conduct experiments on Switchboard 300h dataset and our conformer-based hybrid model achieves competitive results compared to other architectures. It generalizes very well on Hub5’01 test set and outperforms the BLSTM-based hybrid model significantly. Mohammad Zeineldeen, Jingjing Xu 0002, Christoph Lüscher, Wilfried Michel, Alexander Gerstenberger, Ralf Schlüter, Hermann Ney |
ICASSP | 7 |
| 2022 | On Language Model Integration for RNN Transducer Based Speech RecognitionabstractThe mismatch between an external language model (LM) and the implicitly learned internal LM (ILM) of RNN-Transducer (RNN-T) can limit the performance of LM integration such as simple shallow fusion. A Bayesian interpretation suggests to remove this sequence prior as ILM correction. In this work, we study various ILM correction-based LM integration methods formulated in a common RNN-T framework. We provide a decoding interpretation on two major reasons for performance improvement with ILM correction, which is further experimentally verified with detailed analysis. We also propose an exact-ILM training framework by extending the proof given in the hybrid autoregressive transducer, which enables a theoretical justification for other ILM approaches. Systematic comparison is conducted for both in-domain and cross-domain evaluation on the Librispeech and TED-LIUM Release 2 corpora, respectively. Our proposed exact-ILM training can further improve the best ILM method. Wei Zhou 0043, Zuoyun Zheng, Ralf Schlüter, Hermann Ney |
ICASSP | 4 |
| 2022 | Automatic Learning of Subword Dependent Model ScalesabstractTo improve the performance of state-of-the-art automatic speech recognition systems it is common practice to include external knowledge sources such as language models or prior corrections.This is usually done via log-linear model combination using separate scaling parameters for each model.Typically these parameters are manually optimized on some held-out data.In this work we propose to optimize these scaling parameters via automatic differentiation and stochastic gradient decent similar to the neural network model parameters.We show on the LibriSpeech (LBS) and Switchboard (SWB) corpora that the model scales for a combination of attentionbased encoder-decoder acoustic model and language model can be learned as effectively as with manual tuning.We further extend this approach to subword dependent model scales which could not be tuned manually which leads to 7% improvement on LBS and 3% on SWB.We also show that joint training of scales and model parameters is possible and gives additional 6% improvement on LBS. Felix Meyer, Wilfried Michel, Mohammad Zeineldeen, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 5 |
| 2022 | Self-Normalized Importance Sampling for Neural Language ModelingabstractTo mitigate the problem of having to traverse over the full vocabulary in the softmax normalization of a neural language model, sampling-based training criteria are proposed and investigated in the context of large vocabulary word-based neural language models. These training criteria typically enjoy the benefit of faster training and testing, at a cost of slightly degraded performance in terms of perplexity and almost no visible drop in word error rate. While noise contrastive estimation is one of the most popular choices, recently we show that other sampling-based criteria can also perform well, as long as an extra correction step is done, where the intended class posterior probability is recovered from the raw model outputs. In this work, we propose self-normalized importance sampling. Compared to our previous work, the criteria considered in this work are self-normalized and there is no need to further conduct a correction step. Through self-normalized language model training as well as lattice rescoring experiments, we show that our proposed self-normalized importance sampling is competitive in both research-oriented and production-oriented automatic speech recognition tasks. Yingbo Gao, Alexander Gerstenberger, Jintao Jiang, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 6 |
| 2022 | Improving the Training Recipe for a Robust Conformer-based Hybrid Model
Mohammad Zeineldeen, Jingjing Xu 0002, Christoph Lüscher, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 5 |
| 2022 | Efficient Training of Neural Transducer for Speech RecognitionabstractAs one of the most popular sequence-to-sequence modeling approaches for speech recognition, the RNN-Transducer has achieved evolving performance with more and more sophisticated neural network models of growing size and increasing training epochs. While strong computation resources seem to be the prerequisite of training superior models, we try to overcome it by carefully designing a more efficient training pipeline. In this work, we propose an efficient 3-stage progressive training pipeline to build highly-performing neural transducer models from scratch with very limited computation resources in a reasonable short time period. The effectiveness of each stage is experimentally verified on both Librispeech and Switchboard corpora. The proposed pipeline is able to train transducer models approaching state-of-the-art performance with a single GPU in just 2-3 weeks. Our best conformer transducer achieves 4.1% WER on Librispeech test-other with only 35 epochs of training. Wei Zhou 0043, Wilfried Michel, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 4 |
| 2022 | HMM vs. CTC for Automatic Speech Recognition: Comparison Based on Full-Sum Training from ScratchabstractIn this work, we compare from-scratch sequence-level cross-entropy (full-sum) training of Hidden Markov Model (HMM) and Connectionist Temporal Classification (CTC) topologies for automatic speech recognition (ASR). Besides accuracy, we further analyze their capability for generating high-quality time alignment between the speech signal and the transcription, which can be crucial for many subsequent applications. Moreover, we propose several methods to improve convergence of from-scratch full-sum training by addressing the alignment modeling issue. Systematic comparison is conducted on both Switchboard and LibriSpeech corpora across CTC, posterior HMM with and w/o transition probabilities, and standard hybrid HMM. We also provide a detailed analysis of both Viterbi forced-alignment and Baum-Welch full-sum occupation probabilities. Tina Raissi, Wei Zhou 0043, Simon Berger, Ralf Schlüter, Hermann Ney |
SLT | 5 |
| 2022 | Monotonic Segmental Attention for Automatic Speech RecognitionabstractWe introduce a novel segmental-attention model for automatic speech recognition. We restrict the decoder attention to segments to avoid quadratic runtime of global attention, better generalize to long sequences, and eventually enable streaming. We directly compare global-attention and different segmental-attention modeling variants. We develop and compare two separate time-synchronous decoders, one specifically taking the segmental nature into account, yielding further improvements. Using time-synchronous decoding for segmental models is novel and a step towards streaming applications. Our experiments show the importance of a length model to predict the segment boundaries. The final best segmental-attention model using segmental decoding performs better than global-attention, in contrast to other monotonic attention approaches in the literature. Further, we observe that the segmental model generalizes much better to long sequences of up to several minutes. Albert Zeyer, Robin Schmitt, Wei Zhou 0043, Ralf Schlüter, Hermann Ney |
SLT | 5 |
| 2021 | Comparing the Benefit of Synthetic Training Data for Various Automatic Speech Recognition ArchitecturesabstractRecent publications on automatic-speech-recognition (ASR) have a strong focus on attention encoder-decoder (AED) architectures which tend to suffer from over-fitting in low resource scenarios. One solution to tackle this issue is to generate synthetic data with a trained text-to-speech system (TTS) if additional text is available. This was successfully applied in many publications with AED systems, but only very limited in the context of other ASR architectures. We investigate the effect of varying pre-processing, the speaker embedding and input encoding of the TTS system w.r.t. the effectiveness of the synthesized data for AED-ASR training. Additionally, we also consider internal language model subtraction for the first time, resulting in up to 38% relative improvement. We compare the AED results to a state-of-the-art hybrid ASR system, a monophone based system using connectionist-temporal-classification (CTC) and a monotonic transducer based system. We show that for the later systems the addition of synthetic data has no relevant effect, but they still outperform the AED systems on LibriSpeech-100h. We achieve a final word-error-rate of 3.3%/10.0% with a hybrid system on the clean/noisy test-sets, surpassing any previous state-of-the-art systems on Librispeech-100h that do not include unlabeled audio data. Nick Rossenbach, Mohammad Zeineldeen, Benedikt Hilmes, Ralf Schlüter, Hermann Ney |
ASRU | 5 |
| 2021 | On Architectures and Training for Raw Waveform Feature Extraction in ASRabstractWith the success of neural network based modeling in auto-matic speech recognition (ASR), many studies investigated acoustic modeling and learning of feature extractors directly based on the raw waveform. Recently, one line of research has focused on unsupervised pre-training of feature extractors on audio-only data to improve downstream ASR performance. In this work, we investigate the usefulness of one of these front-end frameworks, namely wav2vec, in a setting without additional untranscribed data for hybrid ASR systems. We compare this framework both to the manually defined stan-dard Gammatone feature set, as well as to features extracted as part of the acoustic model of an ASR system trained su-pervised. We study the benefits of using the pre-trained feature extractor and explore how to additionally exploit an ex-isting acoustic model trained with different features. Finally, we systematically examine combinations of the described features in order to further advance the performance. Peter Vieting, Christoph Lüscher, Wilfried Michel, Ralf Schlüter, Hermann Ney |
ASRU | 5 |
| 2021 | Phoneme Based Neural Transducer for Large Vocabulary Speech RecognitionabstractTo join the advantages of classical and end-to-end approaches for speech recognition, we present a simple, novel and competitive approach for phoneme-based neural transducer modeling. Different alignment label topologies are compared and word-end-based phoneme label augmentation is proposed to improve performance. Utilizing the local dependency of phonemes, we adopt a simplified neural network structure and a straightforward integration with the external word-level language model to preserve the consistency of seq-to-seq modeling. We also present a simple, stable and efficient training procedure using frame-wise cross-entropy loss. A phonetic context size of one is shown to be sufficient for the best performance. A simplified scheduled sampling approach is applied for further improvement and different decoding approaches are briefly compared. The overall performance of our best model is comparable to state-of-the-art (SOTA) results for the TED-LIUM Release 2 and Switchboard corpora. Wei Zhou 0043, Simon Berger, Ralf Schlüter, Hermann Ney |
ICASSP | 4 |
| 2021 | On Sampling-Based Training Criteria for Neural Language ModelingabstractAs the vocabulary size of modern word-based language models becomes ever larger, many sampling-based training criteria are proposed and investigated.The essence of these sampling methods is that the softmax-related traversal over the entire vocabulary can be simplified, giving speedups compared to the baseline.A problem we notice about the current landscape of such sampling methods is the lack of a systematic comparison and some myths about preferring one over another.In this work, we consider Monte Carlo sampling, importance sampling, a novel method we call compensated partial summation, and noise contrastive estimation.Linking back to the three traditional criteria, namely mean squared error, binary cross-entropy, and crossentropy, we derive the theoretical solutions to the training problems.Contrary to some common belief, we show that all these sampling methods can perform equally well, as long as we correct for the intended class posterior probabilities.Experimental results in language modeling and automatic speech recognition on Switchboard and LibriSpeech support our claim, with all sampling-based methods showing similar perplexities and word error rates while giving the expected speedups. Yingbo Gao, David Thulke, Alexander Gerstenberger, Khoa Viet Tran, Ralf Schlüter, Hermann Ney |
Interspeech | 6 |
| 2021 | Forty Years of Speech and Language Processing: From Bayes Decision Rule to Deep Learning
Hermann Ney |
Interspeech | 1 |
| 2021 | Investigating Methods to Improve Language Model Integration for Attention-Based Encoder-Decoder ASR ModelsabstractAttention-based encoder-decoder (AED) models learn an implicit internal language model (ILM) from the training transcriptions. The integration with an external LM trained on much more unpaired text usually leads to better performance. A Bayesian interpretation as in the hybrid autoregressive transducer (HAT) suggests dividing by the prior of the discriminative acoustic model, which corresponds to this implicit LM, similarly as in the hybrid hidden Markov model approach. The implicit LM cannot be calculated efficiently in general and it is yet unclear what are the best methods to estimate it. In this work, we compare different approaches from the literature and propose several novel methods to estimate the ILM directly from the AED model. Our proposed methods outperform all previous approaches. We also investigate other methods to suppress the ILM mainly by decreasing the capacity of the AED model, limiting the label context, and also by training the AED model together with a pre-existing LM. Mohammad Zeineldeen, Aleksandr Glushko, Wilfried Michel, Albert Zeyer, Ralf Schlüter, Hermann Ney |
Interspeech | 6 |
| 2021 | Librispeech Transducer Model with Internal Language Model Prior CorrectionabstractWe present our transducer model on Librispeech. We study variants to include an external language model (LM) with shallow fusion and subtract an estimated internal LM. This is justified by a Bayesian interpretation where the transducer model prior is given by the estimated internal LM. The subtraction of the internal LM gives us over 14% relative improvement over normal shallow fusion. Our transducer has a separate probability distribution for the non-blank labels which allows for easier combination with the external LM, and easier estimation of the internal LM. We additionally take care of including the end-of-sentence (EOS) probability of the external LM in the last blank probability which further improves the performance. All our code and setups are published. Albert Zeyer, André Merboldt, Wilfried Michel, Ralf Schlüter, Hermann Ney |
Interspeech | 5 |
| 2021 | Equivalence of Segmental and Neural Transducer Modeling: A Proof of ConceptabstractWith the advent of direct models in automatic speech recognition (ASR), the formerly prevalent frame-wise acoustic modeling based on hidden Markov models (HMM) diversified into a number of modeling architectures like encoder-decoder attention models, transducer models and segmental models (direct HMM).While transducer models stay with a frame-level model definition, segmental models are defined on the level of label segments directly.While (soft-)attention-based models avoid explicit alignment, transducer and segmental approach internally do model alignment, either by segment hypotheses or, more implicitly, by emitting so-called blank symbols.In this work, we prove that the widely used class of RNN-Transducer models and segmental models (direct HMM) are equivalent and therefore show equal modeling power.It is shown that blank probabilities translate into segment length probabilities and vice versa.In addition, we provide initial experiments investigating decoding and beam-pruning, comparing time-synchronous and label-/segment-synchronous search strategies and their properties using the same underlying model. Wei Zhou 0043, Albert Zeyer, André Merboldt, Ralf Schlüter, Hermann Ney |
Interspeech | 5 |
| 2021 | Acoustic Data-Driven Subword Modeling for End-to-End Speech RecognitionabstractSubword units are commonly used for end-to-end automatic speech recognition (ASR), while a fully acoustic-oriented subword modeling approach is somewhat missing.We propose an acoustic data-driven subword modeling (ADSM) approach that adapts the advantages of several text-based and acousticbased subword methods into one pipeline.With a fully acousticoriented label design and learning process, ADSM produces acoustic-structured subword units and acoustic-matched target sequence for further ASR training.The obtained ADSM labels are evaluated with different end-to-end ASR approaches including CTC, RNN-Transducer and attention models.Experiments on the LibriSpeech corpus show that ADSM clearly outperforms both byte pair encoding (BPE) and pronunciationassisted subword modeling (PASM) in all cases.Detailed analysis shows that ADSM achieves acoustically more logical word segmentation and more balanced sequence length, and thus, is suitable for both time-synchronous and label-synchronous models.We also briefly describe how to apply acoustic-based subword regularization and unseen text segmentation using ADSM. Wei Zhou 0043, Mohammad Zeineldeen, Zuoyun Zheng, Ralf Schlüter, Hermann Ney |
Interspeech | 5 |
| 2021 | Data Filtering using Cross-Lingual Word EmbeddingsabstractChristian Herold, Jan Rosendahl, Joris Vanvinckenroye, Hermann Ney. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Christian Herold, Jan Rosendahl, Joris Vanvinckenroye, Hermann Ney |
NAACL-HLT | 4 |
| 2021 | Two-Way Neural Machine Translation: A Proof of Concept for Bidirectional Translation Modeling Using a Two-Dimensional GridabstractNeural translation models have proven to be effective in capturing sufficient information from a source sentence and generating a high-quality target sentence. However, it is not easy to get the best effect for bidirectional translation, i.e., both source-to-target and target-to-source translation using a single model. If we exclude some pioneering attempts, such as multilingual systems, all other bidirectional translation approaches are required to train two individual models. This paper proposes to build a single end-to-end bidirectional translation model using a two-dimensional grid, where the left-to-right decoding generates source-to-target, and the bottom-to-up decoding creates target-to-source output. Instead of training two models independently, our approach encourages a single network to jointly learn to translate in both directions. Experiments on the WMT2018 German↔English and Turkish↔English translation tasks show that the proposed model is capable of generating a good translation quality and has sufficient potential to direct the research. Parnia Bahar, Christopher Brix, Hermann Ney |
SLT | 3 |
| 2021 | Tight Integrated End-to-End Training for Cascaded Speech TranslationabstractA cascaded speech translation model relies on discrete and non-differentiable transcription, which provides a supervision signal from the source side and helps the transformation between source speech and target text. Such modeling suffers from error propagation between ASR and MT models. Direct speech translation is an alternative method to avoid error propagation; however, its performance is often behind the cascade system. To use an intermediate representation and preserve the end-to-end trainability, previous studies have proposed using two-stage models by passing the hidden vectors of the recognizer into the decoder of the MT model and ignoring the MT encoder. This work explores the feasibility of collapsing the entire cascade components into a single end-to-end trainable model by optimizing all parameters of ASR and MT models jointly without ignoring any learned parameters. It is a tightly integrated method that passes renormalized source word posterior distributions as a soft decision instead of one-hot vectors and enables backpropagation. Therefore, it provides both transcriptions and translations and achieves strong consistency between them. Our experiments on four tasks with different data scenarios show that the model outperforms cascade models up to 1.8% in BLEU and 2.0% in TER and is superior compared to direct models. Parnia Bahar, Tobias Bieschke, Ralf Schlüter, Hermann Ney |
SLT | 4 |
| 2020 | Successfully Applying the Stabilized Lottery Ticket Hypothesis to the Transformer ArchitectureabstractSparse models require less memory for storage and enable a faster inference by reducing the necessary number of FLOPs.This is relevant both for time-critical and on-device computations using neural networks.The stabilized lottery ticket hypothesis states that networks can be pruned after none or few training iterations, using a mask computed based on the unpruned converged model.On the transformer architecture and the WMT 2014 English→German and English→French tasks, we show that stabilized lottery ticket pruning performs similar to magnitude pruning for sparsity levels of up to 85%, and propose a new combination of pruning techniques that outperforms all other techniques for even higher levels of sparsity.Furthermore, we confirm that the parameter's initial sign and not its specific value is the primary factor for successful training, and show that magnitude pruning cannot be used to find winning lottery tickets. Christopher Brix, Parnia Bahar, Hermann Ney |
ACL | 3 |
| 2020 | Unifying Input and Output Smoothing in Neural Machine TranslationabstractSoft contextualized data augmentation is a recent method that replaces one-hot representation of words with soft posterior distributions of an external language model, smoothing the input of neural machine translation systems.Label smoothing is another effective method that penalizes over-confident model outputs by discounting some probability mass from the true target word, smoothing the output of neural machine translation systems.Having the benefit of updating all word vectors in each optimization step and better regularizing the models, the two smoothing methods are shown to bring significant improvements in translation performance.In this work, we study how to best combine the methods and stack the improvements.Specifically, we vary the prior distributions to smooth with, the hyperparameters that control the smoothing strength, and the token selection procedures.We conduct extensive experiments on small datasets, evaluate the recipes on larger datasets, and examine the implications when back-translation is further used.Our results confirm cumulative improvements when input and output smoothing are used in combination, giving up to +1.9 BLEU scores on standard machine translation tasks and reveal reasons why these smoothing methods should be preferred. Yingbo Gao, Baohao Liao, Hermann Ney |
COLING | 3 |
| 2020 | Neural Language Modeling for Named Entity RecognitionabstractNamed entity recognition is a key component in various natural language processing systems, and neural architectures provide significant improvements over conventional approaches. Regardless of different word embedding and hidden layer structures of the networks, a conditional random field layer is commonly used for the output. This work proposes to use a neural language model as an alternative to the conditional random field layer, which is more flexible for the size of the corpus. Experimental results show that the proposed system has a significant advantage in terms of training speed, with a marginal performance degradation. Zhihong Lei, Weiyue Wang 0001, Christian Dugast, Hermann Ney |
COLING | 4 |
| 2020 | When and Why is Unsupervised Neural Machine Translation Useless?abstractThis paper studies the practicality of the current state-of-the-art unsupervised methods in neural machine translation (NMT). In ten translation tasks with various data settings, we analyze the conditions under which the unsupervised methods fail to produce reasonable translations. We show that their performance is severely affected by linguistic dissimilarity and domain mismatch between source and target monolingual data. Such conditions are common for low-resource language pairs, where unsupervised learning works poorly. In all of our experiments, supervised and semi-supervised baselines with 50k-sentence bilingual data outperform the best unsupervised results. Our analyses pinpoint the limits of the current unsupervised NMT and also suggest immediate research directions. Yunsu Kim 0001, Miguel Graça, Hermann Ney |
EAMT | 3 |
| 2020 | Exploring A Zero-Order Direct Hmm Based on Latent Attention for Automatic Speech RecognitionabstractIn this paper, we study a simple yet elegant latent variable attention model for automatic speech recognition (ASR) which enables an integration of attention sequence modeling into the direct hidden Markov model (HMM) concept. We use a sequence of hidden variables that establishes a mapping from output labels to input frames. Inspired by the direct HMM model, we assume a decomposition of the label sequence posterior into emission and transition probabilities using zero-order assumption and incorporate both Transformer and LSTM attention models into it. The method keeps the explicit alignment as part of the stochastic model and combines the ease of the end-to-end training of the attention model as well as an efficient and simple beam search. To study the effect of the latent model, we qualitatively analyze the alignment behavior of the different approaches. Our experiments on three ASR tasks show promising results in WER with more focused alignments in comparison to the attention models. Parnia Bahar, Nikita Makarov, Albert Zeyer, Ralf Schlüter, Hermann Ney |
ICASSP | 5 |
| 2020 | A Comprehensive Study of Residual CNNS for Acoustic Modeling in ASRabstractLong short-term memory (LSTM) networks are the dominant architecture for large vocabulary continuous speech recognition (LVCSR) acoustic modeling due to their good performance. However, LSTMs are hard to tune and computationally expensive. To build a system with lower computational costs and which allows online streaming applications, we explore convolutional neural networks (CNN). To the best of our knowledge there is no overview on CNN hyper-parameter tuning for LVCSR in the literature, so we present our results explicitly. Apart from recognition performance, we focus on the training and evaluation speed and provide a time-efficient setup for CNNs. We faced an overfitting problem in training and solved it with data augmentation, namely SpecAugment. The system achieves results competitive with the top LSTM results. We significantly increased the speed of CNN in training and decoding approaching the speed of the offline LSTM. Vitalii Bozheniuk, Albert Zeyer, Ralf Schlüter, Hermann Ney |
ICASSP | 4 |
| 2020 | Domain Robust, Fast, and Compact Neural Language ModelsabstractDespite advances in neural language modeling, obtaining a good model on a large scale multi-domain dataset still remains a difficult task. We propose training methods for building neural language models for such a task, which are not only domain robust, but reasonable in model size and fast for evaluation. We combine knowledge distillation from pretrained domain expert language models with the noise contrastive estimation (NCE) loss. Knowledge distillation allows to train a single student model which is both compact and domain robust, while the use of NCE loss makes the model self-normalized, which enables fast evaluation. We conduct experiments on a large English multi-domain speech recognition dataset provided by AppTek. The resulting student model is of the size of one domain expert, while it gives similar perplexities as various teacher models on their expert domain; the model is self-normalized, allowing for 30% faster first pass decoding than the naive models which require the full soft- max computation, and finally it gives improvements of more than 8% relative in terms of word error rate over a large multidomain 4-gram count model trained on more than 10 B words. Alexander Gerstenberger, Kazuki Irie, Pavel Golik, Eugen Beck, Hermann Ney |
ICASSP | 5 |
| 2020 | How Much Self-Attention Do We Need? Trading Attention for Feed-Forward LayersabstractWe propose simple architectural modifications in the standard Transformer with the goal to reduce its total state size (defined as the number of self-attention layers times the sum of the key and value dimensions, times position) without loss of performance. Large scale Transformer language models have been empirically proved to give very good performance. However, scaling up results in a model that needs to store large states at evaluation time. This can increase the memory requirement dramatically for search e.g., in speech recognition (first pass decoding, lattice rescoring, or shallow fusion). In order to efficiently increase the model capacity without increasing the state size, we replace the single-layer feed-forward module in the Transformer layer by a deeper network, and decrease the total number of layers. In addition, we also evaluate the effect of key-value tying which directly divides the state size in half. On TED-LIUM 2, we obtain a model of state size 4 times smaller than the standard Transformer, with only 2% relative loss in terms of perplexity, which makes the deployment of Transformer language models more convenient. Kazuki Irie, Alexander Gerstenberger, Ralf Schlüter, Hermann Ney |
ICASSP | 4 |
| 2020 | Frame-Level MMI as A Sequence Discriminative Training Criterion for LVCSRabstractIn this work we present frame-level maximum mutual information (MMI) as a novel sequence discriminative training criterion for hybrid HMM-DNN acoustic models. Compared to the standard, sequence-level MMI criterion we show that frame-level MMI has increased robustness towards missing cross-entropy (CE) smoothing and can converge even without interpolation. Using model free optimization, we show that in the asymptotic case of an infinite amount of training data models trained using this criterion are equal to the true class posterior distribution, whereas training using the state-level minimum Bayes risk (sMBR) criterion leads to a distorted function of the true class posterior distribution. This analytical result is backed by experimental evidence. We further propose a generalized class of training criteria, that continuously interpolates between frame-level MMI and sMBR criterion. Wilfried Michel, Ralf Schlüter, Hermann Ney |
ICASSP | 3 |
| 2020 | Generating Synthetic Audio Data for Attention-Based Speech Recognition SystemsabstractRecent advances in text-to-speech (TTS) led to the development of flexible multi-speaker end-to-end TTS systems. We extend state-of-the-art attention-based automatic speech recognition (ASR) systems with synthetic audio generated by a TTS system trained only on the ASR corpora itself. ASR and TTS systems are built separately to show that text-only data can be used to enhance existing end-to-end ASR systems without the necessity of parameter or architecture changes. We compare our method with language model integration of the same text data and with simple data augmentation methods like SpecAugment and show that performance improvements are mostly independent. We achieve improvements of up to 33% relative in word-error-rate (WER) over a strong baseline with data-augmentation in a low-resource environment (LibriSpeech-100h), closing the gap to a comparable oracle experiment by more than 50%. We also show improvements of up to 5% relative WER over our most recent ASR baseline on LibriSpeech-960h. Nick Rossenbach, Albert Zeyer, Ralf Schlüter, Hermann Ney |
ICASSP | 4 |
| 2020 | Layer-Normalized LSTM for Hybrid-Hmm and End-To-End ASRabstractTraining deep neural networks is often challenging in terms of training stability. It often requires careful hyperparameter tuning or a pretraining scheme to converge. Layer normalization (LN) has shown to be a crucial ingredient in training deep encoder-decoder models. We explore various LN long short-term memory (LSTM) recurrent neural networks (RNN) variants by applying LN to different parts of the internal recurrency of LSTMs. There is no previous work that investigates this. We carry out experiments on the Switchboard 300h task for both hybrid and end-to-end ASR models and we show that LN improves the final word error rate (WER), the stability during training, allows to train even deeper models, requires less hyperparameter tuning, and works well even without pre-training. We find that applying LN to both forward and recurrent inputs globally, which we denoted by Global Joined Norm variant, gives a 10% relative improvement in WER. Mohammad Zeineldeen, Albert Zeyer, Ralf Schlüter, Hermann Ney |
ICASSP | 4 |
| 2020 | The Rwth Asr System for Ted-Lium Release 2: Improving Hybrid Hmm With SpecaugmentabstractWe present a complete training pipeline to build a state-of-the-art hybrid HMM-based ASR system on the 2nd release of the TED-LIUM corpus. Data augmentation using SpecAugment is successfully applied to improve performance on top of our best SAT model using i-vectors. By investigating the effect of different maskings, we achieve improvements from SpecAugment on hybrid HMM models without increasing model size and training time. A subsequent sMBR training is applied to fine-tune the final acoustic model, and both LSTM and Transformer language models are trained and evaluated. Our best system achieves a 5.6% WER on the test set, which outperforms the previous state-of-the-art by 27% relative. Wei Zhou 0043, Wilfried Michel, Kazuki Irie, Markus Kitza, Ralf Schlüter, Hermann Ney |
ICASSP | 6 |
| 2020 | Full-Sum Decoding for Hybrid Hmm Based Speech Recognition Using LSTM Language ModelabstractIn hybrid HMM based speech recognition, LSTM language models have been widely applied and achieved large improvements. The theoretical capability of modeling any unlimited context suggests that no recombination should be applied in decoding. This motivates to reconsider full summation over the HMM-state sequences instead of Viterbi approximation in decoding. We explore the potential gain from more accurate probabilities in terms of decision making and apply the full-sum decoding with a modified prefix-tree search framework. The proposed full-sum decoding is evaluated on both Switchboard and Librispeech corpora. Different models using CE and sMBR training criteria are used. Additionally, both MAP and confusion network decoding as approximated variants of general Bayes decision rule are evaluated. Consistent improvements over strong baselines are achieved in almost all cases without extra cost. We also discuss tuning effort, efficiency and some limitations of full-sum decoding. Wei Zhou 0043, Ralf Schlüter, Hermann Ney |
ICASSP | 3 |
| 2020 | LVCSR with Transformer Language Models
Eugen Beck, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 3 |
| 2020 | Investigation of Large-Margin Softmax in Neural Language Modeling
Jingjing Huo, Yingbo Gao, Weiyue Wang 0001, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 5 |
| 2020 | Early Stage LM Integration Using Local and Global Log-Linear CombinationabstractSequence-to-sequence models with an implicit alignment mechanism (e.g. attention) are closing the performance gap towards traditional hybrid hidden Markov models (HMM) for the task of automatic speech recognition. One important factor to improve word error rate in both cases is the use of an external language model (LM) trained on large text-only corpora. Language model integration is straightforward with the clear separation of acoustic model and language model in classical HMM-based modeling. In contrast, multiple integration schemes have been proposed for attention models. In this work, we present a novel method for language model integration into implicit-alignment based sequence-to-sequence models. Log-linear model combination of acoustic and language model is performed with a per-token renormalization. This allows us to compute the full normalization term efficiently both in training and in testing. This is compared to a global renormalization scheme which is equivalent to applying shallow fusion in training. The proposed methods show good improvements over standard model combination (shallow fusion) on our state-of-the-art Librispeech system. Furthermore, the improvements are persistent even if the LM is exchanged for a more powerful one after training. Wilfried Michel, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 3 |
| 2020 | Context-Dependent Acoustic Modeling Without Explicit Phone ClusteringabstractPhoneme-based acoustic modeling of large vocabulary automatic speech recognition takes advantage of phoneme context. The large number of context-dependent (CD) phonemes and their highly varying statistics require tying or smoothing to enable robust training. Usually, classification and regression trees are used for phonetic clustering, which is standard in hidden Markov model (HMM)-based systems. However, this solution introduces a secondary training objective and does not allow for end-to-end training. In this work, we address a direct phonetic context modeling for the hybrid deep neural network (DNN)/HMM, that does not build on any phone clustering algorithm for the determination of the HMM state inventory. By performing different decompositions of the joint probability of the center phoneme state and its left and right contexts, we obtain a factorized network consisting of different components, trained jointly. Moreover, the representation of the phonetic context for the network relies on phoneme embeddings. The recognition accuracy of our proposed models on the Switchboard task is comparable and outperforms slightly the hybrid model using the standard state-tying decision trees. Tina Raissi, Eugen Beck, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 4 |
| 2020 | A New Training Pipeline for an Improved Neural TransducerabstractThe RNN transducer is a promising end-to-end model candidate. We compare the original training criterion with the full marginalization over all alignments, to the commonly used maximum approximation, which simplifies, improves and speeds up our training. We also generalize from the original neural network model and study more powerful models, made possible due to the maximum approximation. We further generalize the output label topology to cover RNN-T, RNA and CTC. We perform several studies among all these aspects, including a study on the effect of external alignments. We find that the transducer model generalizes much better on longer sequences than the attention model. Our final transducer model outperforms our attention model on Switchboard 300h by over 6% relative WER. Albert Zeyer, André Merboldt, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 4 |
| 2020 | Robust Beam Search for Encoder-Decoder Attention Based Speech Recognition Without Length BiasabstractAs one popular modeling approach for end-to-end speech recognition, attention-based encoder-decoder models are known to suffer the length bias and corresponding beam problem.Different approaches have been applied in simple beam search to ease the problem, most of which are heuristic-based and require considerable tuning.We show that heuristics are not proper modeling refinement, which results in severe performance degradation with largely increased beam sizes.We propose a novel beam search derived from reinterpreting the sequence posterior with an explicit length modeling.By applying the reinterpreted probability together with beam pruning, the obtained final probability leads to a robust model modification, which allows reliable comparison among output sequences of different lengths.Experimental verification on the LibriSpeech corpus shows that the proposed approach solves the length bias problem without heuristics or additional tuning effort.It provides robust decision making and consistently good performance under both small and very large beam sizes.Compared with the best results of the heuristic baseline, the proposed approach achieves the same WER on the 'clean' sets and 4% relative improvement on the 'other' sets.We also show that it is more efficient with the additional derived early stopping criterion. Wei Zhou 0043, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 3 |
| 2020 | ELoPE: Fine-Grained Visual Classification with Efficient Localization, Pooling and EmbeddingabstractThe task of fine-grained visual classification (FGVC) deals with classification problems that display a small inter-class variance such as distinguishing between different bird species or car models. State-of-the-art approaches typically tackle this problem by integrating an elaborate attention mechanism or (part-) localization method into a standard convolutional neural network (CNN). Also in this work the aim is to enhance the performance of a backbone CNN such as ResNet by including three efficient and lightweight components specifically designed for FGVC. This is achieved by using global k-max pooling, a discriminative embedding layer trained by optimizing class means and an efficient localization module that estimates bounding boxes using only class labels for training. The resulting model achieves state-of-the-art recognition accuracies on multiple FGVC benchmark datasets. Harald Hanselmann, Hermann Ney |
WACV | 2 |
| 2020 | Weakly Supervised Learning with Multi-Stream CNN-LSTM-HMMs to Discover Sequential Parallelism in Sign Language VideosabstractIn this work we present a new approach to the field of weakly supervised learning in the video domain. Our method is relevant to sequence learning problems which can be split up into sub-problems that occur in parallel. Here, we experiment with sign language data. The approach exploits sequence constraints within each independent stream and combines them by explicitly imposing synchronisation points to make use of parallelism that all sub-problems share. We do this with multi-stream HMMs while adding intermediate synchronisation constraints among the streams. We embed powerful CNN-LSTM models in each HMM stream following the hybrid approach. This allows the discovery of attributes which on their own lack sufficient discriminative power to be identified. We apply the approach to the domain of sign language recognition exploiting the sequential parallelism to learn sign language, mouth shape and hand shape classifiers. We evaluate the classifiers on three publicly available benchmark data sets featuring challenging real-life sign language with over 1,000 classes, full sentence based lip-reading and articulated hand shape recognition on a fine-grained hand shape taxonomy featuring over 60 different hand shapes. We clearly outperform the state-of-the-art on all data sets and observe significantly faster convergence using the parallel alignment approach. Oscar Koller, Necati Cihan Camgöz, Hermann Ney, Richard Bowden |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2019 | Effective Cross-lingual Transfer of Neural Machine Translation Models without Shared VocabulariesabstractTransfer learning or multilingual model is essential for low-resource neural machine translation (NMT), but the applicability is limited to cognate languages by sharing their vocabularies.This paper shows effective techniques to transfer a pre-trained NMT model to a new, unrelated language without shared vocabularies.We relieve the vocabulary mismatch by using cross-lingual word embedding, train a more language-agnostic encoder by injecting artificial noises, and generate synthetic data easily from the pre-training data without back-translation.Our methods do not require restructuring the vocabulary or retraining the model.We improve plain NMT transfer by up to +5.1% BLEU in five low-resource translation tasks, outperforming multilingual joint training by a large margin.We also provide extensive ablation studies on pre-trained embedding, synthetic data, vocabulary size, and parameter freezing for a better understanding of NMT transfer. Yunsu Kim 0001, Yingbo Gao, Hermann Ney |
ACL (1) | 3 |
| 2019 | A Comparative Study on End-to-End Speech to Text TranslationabstractRecent advances in deep learning show that end-to-end speech to text translation model is a promising approach to direct the speech translation field. In this work, we provide an overview of different end-to-end architectures, as well as the usage of an auxiliary connectionist temporal classification (CTC) loss for better convergence. We also investigate on pre-training variants such as initializing different components of a model using pretrained models, and their impact on the final performance, which gives boosts up to 4% in Bleu and 5% in Ter. Our experiments are performed on 270h IWSLT TED-talks En→De, and 100h LibriSpeech Audio-books En→Fr. We also show improvements over the current end-to-end state-of-the-art systems on both tasks. Parnia Bahar, Tobias Bieschke, Hermann Ney |
ASRU | 3 |
| 2019 | Training Language Models for Long-Span Cross-Sentence EvaluationabstractWhile recurrent neural networks can motivate cross-sentence language modeling and its application to automatic speech recognition (ASR), corresponding modifications of the training method for that end are rarely discussed. In fact, even more generally, the impact of training sequence construction strategy in language modeling for different evaluation conditions is typically ignored. In this work, we revisit this basic but fundamental question. We train language models based on long short-term memory recurrent neural networks and Transformers using various types of training sequences and study their robustness with respect to different evaluation modes. Our experiments on 300h Switchboard and Quaero English datasets show that models trained with back-propagation over sequences consisting of concatenation of multiple sentences with state carry-over across sequences effectively outperform those trained with the sentence-level training, both in terms of perplexity and word error rates for cross-utterance ASR. Kazuki Irie, Albert Zeyer, Ralf Schlüter, Hermann Ney |
ASRU | 4 |
| 2019 | A Comparison of Transformer and LSTM Encoder Decoder Models for ASRabstractWe present competitive results using a Transformer encoder-decoder-attention model for end-to-end speech recognition needing less training time compared to a similarly performing LSTM model. We observe that the Transformer training is in general more stable compared to the LSTM, although it also seems to overfit more, and thus shows more problems with generalization. We also find that two initial LSTM layers in the Transformer encoder provide a much better positional encoding. Data-augmentation, a variant of SpecAugment, helps to improve both the Transformer by 33% and the LSTM by 15% relative. We analyze several pretraining and scheduling schemes, which is crucial for both the Transformer and the LSTM models. We improve our LSTM model by additional convolutional layers. We perform our experiments on Lib-riSpeech 1000h, Switchboard 300h and TED-LIUM-v2 200h, and we show state-of-the-art performance on TED-LIUM-v2 for attention based end-to-end models. We deliberately limit the training on LibriSpeech to 12.5 epochs of the training data for comparisons, to keep the results of practical interest, although we show that longer training time still improves more. We publish all the code and setups to run our experiments. Albert Zeyer, Parnia Bahar, Kazuki Irie, Ralf Schlüter, Hermann Ney |
ASRU | 5 |
| 2019 | uniblock: Scoring and Filtering Corpus with Unicode Block InformationabstractYingbo Gao, Weiyue Wang, Hermann Ney. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Yingbo Gao, Weiyue Wang 0001, Hermann Ney |
EMNLP/IJCNLP (1) | 3 |
| 2019 | Pivot-based Transfer Learning for Neural Machine Translation between Non-English LanguagesabstractYunsu Kim, Petre Petrov, Pavel Petrushkov, Shahram Khadivi, Hermann Ney. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Yunsu Kim 0001, Petre Petrov, Pavel Petrushkov, Shahram Khadivi, Hermann Ney |
EMNLP/IJCNLP (1) | 5 |
| 2019 | On Using 2D Sequence-to-sequence Models for Speech RecognitionabstractAttention-based sequence-to-sequence models have shown promising results in automatic speech recognition. Using these architectures, one-dimensional input and output sequences are related by an attention approach, thereby replacing more explicit alignment processes, like in classical HMM-based modeling. In contrast, here we apply a novel two-dimensional long short-term memory (2DLSTM) architecture to directly model the input/output relation between audio/feature vector sequences and word sequences. The proposed model is an alternative model such that instead of using any type of attention components, we apply a 2DLSTM layer to assimilate the context from both input observations and output transcriptions. The experimental evaluation on the Switchboard 300h automatic speech recognition task shows word error rates for the 2DLSTM model that are competitive to end-to-end attention-based model. Parnia Bahar, Albert Zeyer, Ralf Schlüter, Hermann Ney |
ICASSP | 4 |
| 2019 | Investigation into Joint Optimization of Single Channel Speech Enhancement and Acoustic Modeling for Robust ASRabstractThis paper investigates the joint optimization of single channel speech enhancement and the acoustic model of a hybrid DNN-HMM system for noise robust ASR. Two enhancement methods are investigated. A masking of the noisy speech signal with a speech mask estimated by a DNN based mask estimator, as well as a parametric Wiener filter employing a DNN based noise estimator and a DNN based frame wise estimation of the filter parameters. Those components are jointly optimized with the acoustic model of the ASR system. It is shown that the Wiener filter approach can be used to improve the performance of a state-of-the-art single-channel ASR system on the single channel track of the CHiME-4 data, where the WER of the real evaluation set is reduced from 11.6 % to 10.5 %. Tobias Menne, Ralf Schlüter, Hermann Ney |
ICASSP | 3 |
| 2019 | Language Modeling with Deep TransformersabstractWe explore deep autoregressive Transformer models in language modeling for speech recognition. We focus on two aspects. First, we revisit Transformer model configurations specifically for language modeling. We show that well configured Transformer models outperform our baseline models based on the shallow stack of LSTM recurrent neural network layers. We carry out experiments on the open-source LibriSpeech 960hr task, for both 200K vocabulary word-level and 10K byte-pair encoding subword-level language modeling. We apply our word-level models to conventional hybrid speech recognition by lattice rescoring, and the subword-level models to attention based encoder-decoder models by shallow fusion. Second, we show that deep Transformer language models do not require positional encoding. The positional encoding is an essential augmentation for the self-attention mechanism which is invariant to sequence ordering. However, in autoregressive setup, as is the case for language modeling, the amount of information increases along the position dimension, which is a positional signal by its own. The analysis of attention weights shows that deep autoregressive self-attention models can automatically make use of such positional information. We find that removing the positional encoding even slightly improves the performance of these models. Kazuki Irie, Albert Zeyer, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 4 |
| 2019 | Cumulative Adaptation for BLSTM Acoustic ModelsabstractThis paper addresses the robust speech recognition problem as an adaptation task. Specifically, we investigate the cumulative application of adaptation methods. A bidirectional Long Short-Term Memory (BLSTM) based neural network, capable of learning temporal relationships and translation invariant representations, is used for robust acoustic modelling. Further, i-vectors were used as an input to the neural network to perform instantaneous speaker and environment adaptation, providing 8\% relative improvement in word error rate on the NIST Hub5 2000 evaluation test set. By enhancing the first-pass i-vector based adaptation with a second-pass adaptation using speaker and environment dependent transformations within the network, a further relative improvement of 5\% in word error rate was achieved. We have reevaluated the features used to estimate i-vectors and their normalization to achieve the best performance in a modern large scale automatic speech recognition system. Markus Kitza, Pavel Golik, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 4 |
| 2019 | RWTH ASR Systems for LibriSpeech: Hybrid vs AttentionabstractWe present state-of-the-art automatic speech recognition (ASR) systems employing a standard hybrid DNN/HMM architecture compared to an attention-based encoder-decoder design for the LibriSpeech task. Detailed descriptions of the system development, including model design, pretraining schemes, training schedules, and optimization approaches are provided for both system architectures. Both hybrid DNN/HMM and attention-based systems employ bi-directional LSTMs for acoustic modeling/encoding. For language modeling, we employ both LSTM and Transformer based architectures. All our systems are built using RWTHs open-source toolkits RASR and RETURNN. To the best knowledge of the authors, the results obtained when training on the full LibriSpeech training set, are the best published currently, both for the hybrid DNN/HMM and the attention-based systems. Our single hybrid system even outperforms previous results obtained from combining eight single systems. Our comparison shows that on the LibriSpeech 960h task, the hybrid DNN/HMM system outperforms the attention-based system by 15% relative on the clean and 40% relative on the other test sets in terms of word error rate. Moreover, experiments on a reduced 100h-subset of the LibriSpeech training corpus even show a more pronounced margin between the hybrid DNN/HMM and attention-based architectures. Christoph Lüscher, Eugen Beck, Kazuki Irie, Markus Kitza, Wilfried Michel, Albert Zeyer, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 8 |
| 2019 | Analysis of Deep Clustering as Preprocessing for Automatic Speech Recognition of Sparsely Overlapping SpeechabstractSignificant performance degradation of automatic speech recognition (ASR) systems is observed when the audio signal contains cross-talk. One of the recently proposed approaches to solve the problem of multi-speaker ASR is the deep clustering (DPCL) approach. Combining DPCL with a state-of-the-art hybrid acoustic model, we obtain a word error rate (WER) of 16.5 % on the commonly used wsj0-2mix dataset, which is the best performance reported thus far to the best of our knowledge. The wsj0-2mix dataset contains simulated cross-talk where the speech of multiple speakers overlaps for almost the entire utterance. In a more realistic ASR scenario the audio signal contains significant portions of single-speaker speech and only part of the signal contains speech of multiple competing speakers. This paper investigates obstacles of applying DPCL as a preprocessing method for ASR in such a scenario of sparsely overlapping speech. To this end we present a data simulation approach, closely related to the wsj0-2mix dataset, generating sparsely overlapping speech datasets of arbitrary overlap ratio. The analysis of applying DPCL to sparsely overlapping speech is an important interim step between the fully overlapping datasets like wsj0-2mix and more realistic ASR datasets, such as CHiME-5 or AMI. Tobias Menne, Ilya Sklyar, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 4 |
| 2019 | An Analysis of Local Monotonic Attention Variants
André Merboldt, Albert Zeyer, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 4 |
| 2019 | Comparison of Lattice-Free and Lattice-Based Sequence Discriminative Training Criteria for LVCSRabstractSequence discriminative training criteria have long been a standard tool in automatic speech recognition for improving the performance of acoustic models over their maximum likelihood / cross entropy trained counterparts. While previously a lattice approximation of the search space has been necessary to reduce computational complexity, recently proposed methods use other approximations to dispense of the need for the computationally expensive step of separate lattice creation. In this work we present a memory efficient implementation of the forward-backward computation that allows us to use uni-gram word-level language models in the denominator calculation while still doing a full summation on GPU. This allows for a direct comparison of lattice-based and lattice-free sequence discriminative training criteria such as MMI and sMBR, both using the same language model during training. We compared performance, speed of convergence, and stability on large vocabulary continuous speech recognition tasks like Switchboard and Quaero. We found that silence modeling seriously impacts the performance in the lattice-free case and needs special treatment. In our experiments lattice-free MMI comes on par with its lattice-based counterpart. Lattice-based sMBR still outperforms all lattice-free training criteria. Wilfried Michel, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 3 |
| 2019 | Rescoring Keyword Search Confidence Estimates with Graph-Based Re-Ranking Using Acoustic Word Embeddings
Anna Piunova, Eugen Beck, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 4 |
| 2019 | Upper and Lower Tight Error Bounds for Feature Omission with an Extension to Context ReductionabstractIn this work, fundamental analytic results in the form of error bounds are presented that quantify the effect of feature omission and selection for pattern classification in general, as well as the effect of context reduction in string classification, like automatic speech recognition, printed/handwritten character recognition, or statistical machine translation. A general simulation framework is introduced that supports discovery and proof of error bounds, which lead to the error bounds presented here. Initially derived tight lower and upper bounds for feature omission are generalized to feature selection, followed by another extension to context reduction of string class priors (aka language models) in string classification. For string classification, the quantitative effect of string class prior context reduction on symbol-level Bayes error is presented. The tightness of the original feature omission bounds seems lost in this case, as further simulations indicate. However, combining both feature omission andcontext reduction, the tightness of the bounds is retained. A central result of this work is the proof of the existence, and the amount of a statistical threshold w.r.t. the introduction of additional features in general pattern classification, or the increase of context in string classification beyond which a decrease in Bayes error is guaranteed. Ralf Schlüter, Eugen Beck, Hermann Ney |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2019 | Training of reduced-rank linear transformations for multi-layer polynomial acoustic features for speech recognition
Muhammad Ali Tahir, Heyun Huang, Albert Zeyer, Ralf Schlüter, Hermann Ney |
Speech Commun. | 5 |
| 2018 | Neural Sign Language TranslationabstractSign Language Recognition (SLR) has been an active research field for the last two decades. However, most research to date has considered SLR as a naive gesture recognition problem. SLR seeks to recognize a sequence of continuous signs but neglects the underlying rich grammatical and linguistic structures of sign language that differ from spoken language. In contrast, we introduce the Sign Language Translation (SLT) problem. Here, the objective is to generate spoken language translations from sign language videos, taking into account the different word orders and grammar. We formalize SLT in the framework of Neural Machine Translation (NMT) for both end-to-end and pretrained settings (using expert knowledge). This allows us to jointly learn the spatial representations, the underlying language model, and the mapping between sign and spoken language. To evaluate the performance of Neural SLT, we collected the first publicly available Continuous SLT dataset, RWTH-PHOENIX-Weather 2014T1. It provides spoken language translations and gloss level annotations for German Sign Language videos of weather broadcasts. Our dataset contains over .95M frames with >67K signs from a sign vocabulary of >1K and >99K words from a German vocabulary of >2.8K. We report quantitative and qualitative results for various SLT setups to underpin future research in this newly established field. The upper bound for translation performance is calculated at 19.26 BLEU-4, while our end-to-end frame-level and gloss-level tokenization networks were able to achieve 9.58 and 18.13 respectively. Necati Cihan Camgöz, Simon Hadfield, Oscar Koller, Hermann Ney, Richard Bowden |
CVPR | 4 |
| 2018 | Towards Two-Dimensional Sequence to Sequence Model in Neural Machine TranslationabstractThis work investigates an alternative model for neural machine translation (NMT) and proposes a novel architecture, where we employ a multi-dimensional long short-term memory (MDLSTM) for translation modeling.In the state-of-the-art methods, source and target sentences are treated as one-dimensional sequences over time, while we view translation as a two-dimensional (2D) mapping using an MDLSTM layer to define the correspondence between source and target words.We extend beyond the current sequence to sequence backbone NMT models to a 2D structure in which the source and target sentences are aligned with each other in a 2D grid.Our proposed topology shows consistent improvements over attention-based sequence to sequence model on two WMT 2017 tasks, German↔English. Parnia Bahar, Christopher Brix, Hermann Ney |
EMNLP | 3 |
| 2018 | Improving Unsupervised Word-by-Word Translation with Language Model and Denoising AutoencoderabstractUnsupervised learning of cross-lingual word embedding offers elegant matching of words across languages, but has fundamental limitations in translating sentences.In this paper, we propose simple yet effective methods to improve word-by-word translation of crosslingual embeddings, using only monolingual corpora but without any back-translation.We integrate a language model for context-aware search, and use a novel denoising autoencoder to handle reordering.Our system surpasses state-of-the-art unsupervised neural translation systems without costly iterative training.We also analyze the effect of vocabulary size and denoising type on the translation performance, which provides better understanding of learning the cross-lingual word embedding and its usage in translation. Yunsu Kim 0001, Jiahui Geng, Hermann Ney |
EMNLP | 3 |
| 2018 | Prediction of LSTM-RNN Full Context States as a Subtask for N-Gram Feedforward Language ModelsabstractLong short-term memory (LSTM) recurrent neural network language models compress the full context of variable lengths into a fixed size vector. In this work, we investigate the task of predicting the LSTM hidden representation of the full context from a truncated n-gram context as a subtask for training an n-gram feedforward language model. Since this approach is a form of knowledge distillation, we compare two methods. First, we investigate the standard transfer based on the Kullback-Leibler divergence of the output distribution of the feedforward model from that of the LSTM. Second, we minimize the mean squared error between the hidden state of the LSTM and that of the n-gram feedforward model. We carry out experiments on different subsets of the Switchboard speech recognition dataset for feedforward models with a short (5-gram) and a medium (10-gram) context length. We show that we get improvements in perplexity and word error rate of up to 8% and 4% relative for the medium model, while the improvements are only marginal for the short model. Kazuki Irie, Zhihong Lei, Ralf Schlüter, Hermann Ney |
ICASSP | 4 |
| 2018 | Acoustic Modeling of Speech Waveform Based on Multi-Resolution, Neural Network Signal ProcessingabstractRecently, several papers have demonstrated that neural networks (NN) are able to perform the feature extraction as part of the acoustic model. Motivated by the Gammatone feature extraction pipeline, in this paper we extend the waveform based NN model by a second level of time-convolutional element. The proposed extension generalizes the envelope extraction block, and allows the model to learn multi-resolutional representations. Automatic speech recognition (ASR) experiments show significant word error rate reduction over our previous best acoustic model trained in the signal domain directly. Although we use only 250 hours of speech, the data-driven NN based speech signal processing performs nearly equally to traditional handcrafted feature extractors. In additional experiments, we also test segment-level feature normalization techniques on NN derived features, which improve the results further. However, the porting of speech representations derived by a feed-forward NN to a LSTM back-end model indicates much less robustness of the NN front-end compared to the standard feature extractors. Analysis of the weights in the proposed new layer reveals that the NN prefers both multi-resolution and modulation spectrum representations. Zoltán Tüske, Ralf Schlüter, Hermann Ney |
ICASSP | 3 |
| 2018 | Optimizing Energies for Pose-Invariant Face RecognitionabstractOne of the most difficult challenges in face recognition is the large variation in pose. One approach to handle this problem is to use a 2D-Warping algorithm in a nearest-neighbor classifier. The 2D-Warping algorithm optimizes an energy function that captures the cost of matching pixels between two images while respecting the 2D dependencies defined by local pixel neighborhoods. Optimizing this energy function is an NP-complete problem and is therefore approached with algorithms that aim to approximate the optimal solution. In this paper we compare two algorithms that do this without discarding any 2D dependencies and we study the effect of the quality of the approximate solutions on the classification performance. Additionally, we propose a new algorithm that is capable of finding better solutions and obtaining better energies than the other methods. The experimental evaluation on the CMU-MultiPIE database shows that the proposed algorithm also achieves state-of-the-art recognition accuracies. Harald Hanselmann, Hermann Ney |
ICPR | 2 |
| 2018 | Segmental Encoder-Decoder Models for Large Vocabulary Automatic Speech RecognitionabstractIt has been known for a long time that the classic Hidden-Markov-Model (HMM) derivation for speech recognition contains assumptions such as independence of observation vectors and weak duration modeling that are practical but unrealistic.When using the hybrid approach this is amplified by trying to fit a discriminative model into a generative one.Hidden Conditional Random Fields (CRFs) and segmental models (e.g.Semi-Markov CRFs / Segmental CRFs) have been proposed as an alternative, but for a long time have failed to get traction until recently.In this paper we explore different length modeling approaches for segmental models, their relation to attention-based systems.Furthermore we show experimental results on a handwriting recognition task and to the best of our knowledge the first reported results on the Switchboard 300h speech recognition corpus using this approach. Eugen Beck, Mirko Hannemann, Patrick Doetsch, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 5 |
| 2018 | Investigation on Estimation of Sentence Probability by Combining Forward, Backward and Bi-directional LSTM-RNNs
Kazuki Irie, Zhihong Lei, Liuhui Deng, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 5 |
| 2018 | Comparison of BLSTM-Layer-Specific Affine Transformations for Speaker Adaptation
Markus Kitza, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 3 |
| 2018 | Investigation on LSTM Recurrent N-gram Language Models for Speech RecognitionabstractRecurrent neural networks (NN) with long short-term memory (LSTM) are the current state of the art to model long term dependencies.However, recent studies indicate that NN language models (LM) need only limited length of history to achieve excellent performance.In this paper, we extend the previous investigation on LSTM network based n-gram modeling to the domain of automatic speech recognition (ASR).First, applying recent optimization techniques and up to 6-layer LSTM networks, we improve LM perplexities by nearly 50% relative compared to classic count models on three different domains.Then, we demonstrate by experimental results that perplexities improve significantly only up to 40-grams when limiting the LM history.Nevertheless, the ASR performance saturates already around 20-grams despite across sentence modeling.Analysis indicates that the performance gain of LSTM NNLM over count models results only partially from the longer context and cross sentence modeling capabilities.Using equal context, we show that deep 4-gram LSTM can significantly outperform large interpolated count models by performing the backing off and smoothing significantly better.This observation also underlines the decreasing importance to combine state-of-the-art deep NNLM with count based model. Zoltán Tüske, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 3 |
| 2018 | Improved Training of End-to-end Attention Models for Speech RecognitionabstractSequence-to-sequence attention-based models on subword units allow simple open-vocabulary end-to-end speech recognition. In this work, we show that such models can achieve competitive results on the Switchboard 300h and LibriSpeech 1000h tasks. In particular, we report the state-of-the-art word error rates (WER) of 3.54% on the dev-clean and 3.82% on the test-clean evaluation subsets of LibriSpeech. We introduce a new pretraining scheme by starting with a high time reduction factor and lowering it during training, which is crucial both for convergence and final performance. In some experiments, we also use an auxiliary CTC loss function to help the convergence. In addition, we train long short-term memory (LSTM) language models on subword units. By shallow fusion, we report up to 27% relative improvements in WER over the attention baseline without a language model. Albert Zeyer, Kazuki Irie, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 4 |
| 2018 | Speaker Adapted Beamforming for Multi-Channel Automatic Speech RecognitionabstractThis paper presents, in the context of multi-channel ASR, a method to adapt a mask based, statistically optimal beamforming approach to a speaker of interest. The beamforming vector of the statistically optimal beamformer is computed by utilizing speech and noise masks, which are estimated by a neural network. The proposed adaptation approach is based on the integration of the beamformer, which includes the mask estimation network, and the acoustic model of the ASR system. This allows for the propagation of the training error, from the acoustic modeling cost function, all the way through the beamforming operation and through the mask estimation network. By using the results of a first pass recognition and by keeping all other parameters fixed, the mask estimation network can therefore be fine tuned by retraining. Utterances of a speaker of interest can thus be used in a two pass approach, to optimize the beamforming for the speech characteristics of that specific speaker. It is shown that this approach improves the ASR performance of a state-of-the-art multi-channel ASR system on the CHiME-4 data. Furthermore the effect of the adaptation on the estimated speech masks is discussed. Tobias Menne, Ralf Schlüter, Hermann Ney |
SLT | 3 |
| 2018 | Deep Sign: Enabling Robust Statistical Continuous Sign Language Recognition via Hybrid CNN-HMMsabstractThis manuscript introduces the end-to-end embedding of a CNN into a HMM, while interpreting the outputs of the CNN in a Bayesian framework. The hybrid CNN-HMM combines the strong discriminative abilities of CNNs with the sequence modelling capabilities of HMMs. Most current approaches in the field of gesture and sign language recognition disregard the necessity of dealing with sequence data both for training and evaluation. With our presented end-to-end embedding we are able to improve over the state-of-the-art on three challenging benchmark continuous sign language recognition tasks by between 15 and 38% relative reduction in word error rate and up to 20% absolute. We analyse the effect of the CNN structure, network pretraining and number of hidden states. We compare the hybrid modelling to a tandem approach and evaluate the gain of model combination. Oscar Koller, Sepehr Zargaran, Hermann Ney, Richard Bowden |
Int. J. Comput. Vis. | 3 |
| 2017 | Deep fisher faces
Harald Hanselmann, Hermann Ney |
BMVC | 3 |
| 2017 | Re-Sign: Re-Aligned End-to-End Sequence Modelling with Deep Recurrent CNN-HMMsabstractThis work presents an iterative re-alignment approach applicable to visual sequence labelling tasks such as gesture recognition, activity recognition and continuous sign language recognition. Previous methods dealing with video data usually rely on given frame labels to train their classifiers. However, looking at recent data sets, these labels often tend to be noisy which is commonly overseen. We propose an algorithm that treats the provided training labels as weak labels and refines the label-to-image alignment on-the-fly in a weakly supervised fashion. Given a series of frames and sequence-level labels, a deep recurrent CNN-BLSTM network is trained end-to-end. Embedded into an HMM the resulting deep model corrects the frame labels and continuously improves its performance in several re-alignments. We evaluate on two challenging publicly available sign recognition benchmark data sets featuring over 1000 classes. We outperform the state-of-the-art by up to 10% absolute and 30% relative. Oscar Koller, Sepehr Zargaran, Hermann Ney |
CVPR | 3 |
| 2017 | Returnn: The RWTH extensible training framework for universal recurrent neural networksabstractIn this work we release our extensible and easily configurable neural network training software. It provides a rich set of functional layers with a particular focus on efficient training of recurrent neural network topologies on multiple GPUs. The source of the software package is public and freely available for academic research purposes and can be used as a framework or as a standalone tool which supports a flexible configuration. The software allows to train state-of-the-art deep bidirectional long short-term memory (LSTM) models on both one dimensional data like speech or two dimensional data like handwritten text and was used to develop successful submission systems in several evaluation campaigns. Patrick Doetsch, Albert Zeyer, Paul Voigtlaender, Ilia Kulikov, Ralf Schlüter, Hermann Ney |
ICASSP | 6 |
| 2017 | Investigations on byte-level convolutional neural networks for language modeling in low resource speech recognitionabstractIn this paper, we present an investigation on technical details of the byte-level convolutional layer which replaces the conventional linear word projection layer in the neural language model. In particular, we discuss and compare the effective filter configurations, pooling types and the use of bytes instead of characters. We carry out experiments on language packs released by the IARPA Babel project and measure the performance in terms of perplexity and word error rate. Introducing a convolutional layer consistently improves the results on all languages. Also, there is no degradation from using raw bytes instead of proper Unicode characters, even on syllabic alphabets like Amharic. In addition, we report improvements in word error rate from rescoring lattices and evaluate keyword search performance on several languages. Kazuki Irie, Pavel Golik, Ralf Schlüter, Hermann Ney |
ICASSP | 4 |
| 2017 | Noisy objective functions based on the f-divergenceabstractDropout, the random dropping out of activations according to a specified rate, is a very simple but effective method to avoid over-fitting of deep neural networks to the training data. Markus Nußbaum-Thom, Ralf Schlüter, Vaibhava Goel, Hermann Ney |
ICASSP | 4 |
| 2017 | A comprehensive study of deep bidirectional LSTM RNNS for acoustic modeling in speech recognitionabstractRecent experiments show that deep bidirectional long short-term memory (BLSTM) recurrent neural network acoustic models outperform feedforward neural networks for automatic speech recognition (ASR). However, their training requires a lot of tuning and experience. In this work, we provide a comprehensive overview over various BLSTM training aspects and their interplay within ASR, which has been missing so far in the literature. We investigate on different variants of optimization methods, batching, truncated backpropagation, and regularization techniques such as dropout, and we study the effect of size and depth, training models of up to 10 layers. This includes a comparison of computation times vs. recognition performance. Furthermore, we introduce a pretraining scheme for LSTMs with layer-wise construction of the network showing good improvements especially for deep networks. The experimental analysis mainly was performed on the Quaero task, with additional results on Switchboard. The best BLSTM model gave a relative improvement in word error rate of over 15% compared to our best feed-forward baseline on our Quaero 50h task. All experiments were done using RETURNN and RASR, RWTH's extensible training framework for universal recurrent neural networks and ASR toolkit. The training configuration files are publicly available. Albert Zeyer, Patrick Doetsch, Paul Voigtlaender, Ralf Schlüter, Hermann Ney |
ICASSP | 5 |
| 2017 | Faster sequence trainingabstractIt has been shown that sequence-discriminative training can improve the performance for large vocabulary continuous speech recognition. Our main contribution is a novel method for reducing the computation time of any sort of sequence training while only slightly decreasing the overall performance. The method allows to parallelize the forward propagation through the network, the loss and loss gradient calculation which will provide a frame-wise error signal, and an independent forward and back propagation using that error signal. That last step can be calculated in a frame-wise manner and thus allows to use frame chunking to further improve the runtime. The loss calculation can itself be parallelized over many sequences. In addition to several experiments which outline the runtime gains, we also provide a convergence proof sketch. We extend on the research of sequence training of bidirectional long-short term memory ((B)LSTM) networks and provide an overview and comparison over different criteria. We have published all the code as part of our RETURNN and RASR framework including our training setup configurations. Albert Zeyer, Ilia Kulikov, Ralf Schlüter, Hermann Ney |
ICASSP | 4 |
| 2017 | Parallel Neural Network Features for Improved Tandem Acoustic ModelingabstractThe combination of acoustic models or features is a standard approach to exploit various knowledge sources.This paper investigates the concatenation of different bottleneck (BN) neural network (NN) outputs for tandem acoustic modeling.Thus, combination of NN features is performed via Gaussian mixture models (GMM).Complementarity between the NN feature representations is attained by using various network topologies: LSTM recurrent, feed-forward, and hierarchical, as well as different non-linearities: hyperbolic tangent, sigmoid, and rectified linear units.Speech recognition experiments are carried out on various tasks: telephone conversations, Skype calls, as well as broadcast news and conversations.Results indicate that LSTM based tandem approach is still competitive, and such tandem model can challenge comparable hybrid systems.The traditional steps of tandem modeling, speaker adaptive and sequence discriminative GMM training, improve the tandem results further.Furthermore, these "old-fashioned" steps remain applicable after the concatenation of multiple neural network feature streams.Exploiting the parallel processing of input feature streams, it is shown that 2-5% relative improvement could be achieved over the single best BN feature set.Finally, we also report results after neural network based language model rescoring and examine the system combination possibilities using such complex tandem models. Zoltán Tüske, Wilfried Michel, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 4 |
| 2017 | CTC in the Context of Generalized Full-Sum HMM TrainingabstractWe formulate a generalized hybrid HMM-NN training procedure using the full-sum over the hidden state-sequence and identify CTC as a special case of it.We present an analysis of the alignment behavior of such a training procedure and explain the strong localization of label output behavior of full-sum training (also referred to as peaky or spiky behavior).We show how to avoid that behavior by using a state prior.We discuss the temporal decoupling between output label position/time-frame, and the corresponding evidence in the input observations when this is trained with BLSTM models.We also show a way how to overcome this by jointly training a FFNN.We implemented the Baum-Welch alignment algorithm in CUDA to be able to do fast soft realignments on GPU.We have published this code along with some of our experiments as part of RETURNN, RWTH's extensible training framework for universal recurrent neural networks.We finish with experimental validation of our study on WSJ and Switchboard. Albert Zeyer, Eugen Beck, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 4 |
| 2017 | Editorial
Qun Liu 0001, Xiaodong He 0001, Hermann Ney |
Mach. Transl. | 3 |
| 2016 | Deep Sign: Hybrid CNN-HMM for Continuous Sign Language Recognition
Oscar Koller, Sepehr Zargaran, Hermann Ney, Richard Bowden |
BMVC | 3 |
| 2016 | Deep Hand: How to Train a CNN on 1 Million Hand Images When Your Data is Continuous and Weakly LabelledabstractThis work presents a new approach to learning a framebased classifier on weakly labelled sequence data by embedding a CNN within an iterative EM algorithm. This allows the CNN to be trained on a vast number of example images when only loose sequence level information is available for the source videos. Although we demonstrate this in the context of hand shape recognition, the approach has wider application to any video recognition task where frame level labelling is not available. The iterative EM algorithm leverages the discriminative ability of the CNN to iteratively refine the frame level annotation and subsequent training of the CNN. By embedding the classifier within an EM framework the CNN can easily be trained on 1 million hand images. We demonstrate that the final classifier generalises over both individuals and data sets. The algorithm is evaluated on over 3000 manually labelled hand shape images of 60 different classes which will be released to the community. Furthermore, we demonstrate its use in continuous sign language recognition on two publicly available large sign language data sets, where it outperforms the current state-of-the-art by a large margin. To our knowledge no previous work has explored expectation maximization without Gaussian mixture models to exploit weak sequence labels for sign language recognition. Oscar Koller, Hermann Ney, Richard Bowden |
CVPR | 2 |
| 2016 | Investigation on log-linear interpolation of multi-domain neural network language modelabstractInspired by the success of multi-task training in acoustic modeling, this paper investigates a new architecture for a multi-domain neural network based language model (NNLM). The proposed model has several shared hidden layers and domain-specific output layers. As will be shown, the log-linear interpolation of the multi-domain outputs and the optimization of interpolation weights fit naturally in the framework of NNLM. The resulting model can be expressed as a single NNLM. As an initial study of such an architecture, this paper focuses on deep feed-forward neural networks (DNNs). We also re-investigate the potential of long context up to 30-grams, and depth up to 5 hidden layers in DNN-LM. Our final feed-forward multidomain NNLM is trained on 3.1B running words across 11 domains for English broadcast news and conversations large vocabulary continuous speech recognition task. After log-linear interpolation and fine-tuning, we measured improvements in terms of perplexity and word error rate over the models trained on 50M running words of in-domain news resources. The final multi-domain feed-forward LM outperformed our previous best LSTM-RNN LM trained on the 50M in-domain corpus, even after linear interpolation with large count models. Zoltán Tüske, Kazuki Irie, Ralf Schlüter, Hermann Ney |
ICASSP | 4 |
| 2016 | Bidirectional Decoder Networks for Attention-Based End-to-End Offline Handwriting RecognitionabstractRecurrent neural networks that can be trained end-to-end on sequence learning tasks provide promising benefits over traditional recognition systems. In this paper, we demonstrate the application of an attention-based long short-term memory decoder network for offline handwriting recognition and analyze the segmentation, classification and decoding errors produced by the model. We further extend the decoding network by a bidirectional topology together with an integrated length estimation procedure and show that it is superior to unidirectional decoder networks. Results are presented for the word and text line recognition tasks of the RIMES handwriting recognition database. The software used in the experiments is freely available for academic research purposes. Patrick Doetsch, Albert Zeyer, Hermann Ney |
ICFHR | 3 |
| 2016 | On the Benefits of Convolutional Neural Network Combinations in Offline Handwriting RecognitionabstractIn this paper, we elaborate the advantages of combining two neural network methodologies, convolutional neural networks (CNN) and long short-term memory (LSTM) recurrent neural networks, with the framework of hybrid hidden Markov models (HMM) for recognizing offline handwriting text. CNNs employ shift-invariant filters to generate discriminative features within neural networks. We show that CNNs are powerful tools to extract general purpose features that even work well for unknown classes. We evaluate our system on a Chinese handwritten text database and provide a GPU-based implementation that can be used to reproduce the experiments. All experiments were conducted with RWTH OCR, an open-source system developed at our institute. Dewi Suryani, Patrick Doetsch, Hermann Ney |
ICFHR | 3 |
| 2016 | Handwriting Recognition with Large Multidimensional Long Short-Term Memory Recurrent Neural NetworksabstractMultidimensional long short-term memory recurrent neural networks achieve impressive results for handwriting recognition. However, with current CPU-based implementations, their training is very expensive and thus their capacity has so far been limited. We release an efficient GPU-based implementation which greatly reduces training times by processing the input in a diagonal-wise fashion. We use this implementation to explore deeper and wider architectures than previously used for handwriting recognition and show that especially the depth plays an important role. We outperform state of the art results on two databases with a deep multidimensional network. Paul Voigtlaender, Patrick Doetsch, Hermann Ney |
ICFHR | 3 |
| 2016 | LSTM, GRU, Highway and a Bit of Attention: An Empirical Overview for Language Modeling in Speech RecognitionabstractPopularized by the long short-term memory (LSTM), multiplicative gates have become a standard means to design artificial neural networks with intentionally organized information flow.Notable examples of such architectures include gated recurrent units (GRU) and highway networks.In this work, we first focus on the evaluation of each of the classical gated architectures for language modeling for large vocabulary speech recognition.Namely, we evaluate the highway network, lateral network, LSTM and GRU.Furthermore, the motivation underlying the highway network also applies to LSTM and GRU.An extension specific to the LSTM has been recently proposed with an additional highway connection between the memory cells of adjacent LSTM layers.In contrast, we investigate an approach which can be used with both LSTM and GRU: a highway network in which the LSTM or GRU is used as the transformation function.We found that the highway connections enable both standalone feedforward and recurrent neural language models to benefit better from the deep structure and provide a slight improvement of recognition accuracy after interpolation with count models.To complete the overview, we include our initial investigations on the use of the attention mechanism for learning word triggers. Kazuki Irie, Zoltán Tüske, Tamer Alkhouli, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 5 |
| 2016 | Towards Online-Recognition with Deep Bidirectional LSTM Acoustic ModelsabstractOnline-Recognition requires the acoustic model to provide posterior probabilities after a limited time delay given the online input audio data.This necessitates unidirectional modeling and the standard solution is to use unidirectional long short-term memory (LSTM) recurrent neural networks (RNN) or feedforward neural networks (FFNN).It is known that bidirectional LSTMs are more powerful and perform better than unidirectional LSTMs.To demonstrate the performance difference, we start by comparing several different bidirectional and unidirectional LSTM topologies.Furthermore, we apply a modification to bidirectional RNNs to enable online-recognition by moving a window over the input stream and perform one forwarding through the RNN on each window.Then, we combine the posteriors of each forwarding and we renormalize them.We show in experiments that the performance of this online-enabled bidirectional LSTM performs as good as the offline bidirectional LSTM and much better than the unidirectional LSTM. Albert Zeyer, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 3 |
| 2015 | Multilingual representations for low resource speech recognition and keyword searchabstractThis paper examines the impact of multilingual (ML) acoustic representations on Automatic Speech Recognition (ASR) and keyword search (KWS) for low resource languages in the context of the OpenKWS15 evaluation of the IARPA Babel program. The task is to develop Swahili ASR and KWS systems within two weeks using as little as 3 hours of transcribed data. Multilingual acoustic representations proved to be crucial for building these systems under strict time constraints. The paper discusses several key insights on how these representations are derived and used. First, we present a data sampling strategy that can speed up the training of multilingual representations without appreciable loss in ASR performance. Second, we show that fusion of diverse multilingual representations developed at different LORELEI sites yields substantial ASR and KWS gains. Speaker adaptation and data augmentation of these representations improves both ASR and KWS performance (up to 8.7% relative). Third, incorporating un-transcribed data through semi-supervised learning, improves WER and KWS performance. Finally, we show that these multilingual representations significantly improve ASR and KWS performance (relative 9% for WER and 5% for MTWV) even when forty hours of transcribed audio in the target language is available. Multilingual representations significantly contributed to the LORELEI KWS systems winning the OpenKWS15 evaluation. Jia Cui, Brian Kingsbury, Bhuvana Ramabhadran, Abhinav Sethy, Kartik Audhkhasi, Ellen Eide, Lidia Mangu, Markus Nußbaum-Thom, Michael Picheny, Zoltán Tüske, Pavel Golik, Ralf Schlüter, Hermann Ney, Mark J. F. Gales, Kate M. Knill, Anton Ragni, Philip C. Woodland |
ASRU | 14 |
| 2015 | Speaker adaptive joint training of Gaussian mixture models and bottleneck featuresabstractIn the tandem approach, the output of a neural network (NN) serves as input features to a Gaussian mixture model (GMM) aiming to improve the emission probability estimates. As has been shown in our previous work, GMM with pooled covariance matrix can be integrated into a neural network framework as a softmax layer with hidden variables, which allows for joint estimation of both neural network and Gaussian mixture parameters. Here, this approach is extended to include speaker adaptive training (SAT) by introducing a speaker dependent neural network layer. Error backpropagation beyond this speaker dependent layer realizes the adaptive training of the Gaussian parameters as well as the optimization of the bottleneck (BN) tandem features of the underlying acoustic model, simultaneously. In this study, after the initialization by constrained maximum likelihood linear regression (CMLLR) the speaker dependent layer itself is kept constant during the joint training. Experiments show that the deeper backpropagation through the speaker dependent layer is necessary for improved recognition performance. The speaker adaptively and jointly trained BN-GMM results in 5% relative improvement over very strong speaker-independent hybrid baseline on the Quaero English broadcast news and conversations task, and on the 300-hour Switchboard task. Zoltán Tüske, Pavel Golik, Ralf Schlüter, Hermann Ney |
ASRU | 4 |
| 2015 | A Comparison between Count and Neural Network Models Based on Joint Translation and Reordering SequencesabstractWe propose a conversion of bilingual sentence pairs and the corresponding word alignments into novel linear sequences.These are joint translation and reordering (JTR) uniquely defined sequences, combining interdepending lexical and alignment dependencies on the word level into a single framework.They are constructed in a simple manner while capturing multiple alignments and empty words.JTR sequences can be used to train a variety of models.We investigate the performances of ngram models with modified Kneser-Ney smoothing, feed-forward and recurrent neural network architectures when estimated on JTR sequences, and compare them to the operation sequence model (Durrani et al., 2013b).Evaluations on the IWSLT German→English, WMT German→English and BOLT Chinese→English tasks show that JTR models improve state-of-the-art phrasebased systems by up to 2.2 BLEU. Andreas Guta, Tamer Alkhouli, Jan-Thorsten Peter, Jörn Wübker, Hermann Ney |
EMNLP | 5 |
| 2015 | Improved strategies for a zero oov rate LVCSR systemabstractIn this work, multiple hierarchical language modeling strategies for a zero OOV rate large vocabulary continuous speech recognition system are investigated. In our previously proposed hierarchical approach, a full-word language model and a context independent character-level LM (CLM) are directly used during search. The novelty of this work is to jointly model the character-level prior and the pronunciation probabilities, to introduce across-word context into the characterlevel LM, and to properly normalize the character-level LM using prefix-tree based normalization for the hierarchical approach. Significant reductions in-terms of word error rates (WER) on the best full-word Quaero Polish LVCSR system are reported. M. Ali Basha Shaik, Amr El-Desoky Mousa, Stefan Hahn, Ralf Schlüter, Hermann Ney |
ICASSP | 5 |
| 2015 | Investigation of mixture splitting concept for training linear bottlenecks of deep neural network acoustic modelsabstractA Gaussian or log-linear mixture model trained by maximum likelihood may be trained further using discriminative training. It is desirable that the mixture splitting is also done during the discriminative training, to achieve better mixture density distribution. In previous work such a discriminative splitting approach was presented. Similarly, the resolution of a deep neural network may also be increased by splitting. In this paper, discriminative splitting is applied as a way of initializing a linear bottleneck between two layers of a DNN. Experiments for a single hidden layer and six hidden layer cases show the potential of this approach as an alternative method of pre-training for linear bottlenecks for MLP hidden layers. Muhammad Ali Tahir, Simon Wiesler, Ralf Schlüter, Hermann Ney |
ICASSP | 4 |
| 2015 | Integrating Gaussian mixtures into deep neural networks: Softmax layer with hidden variablesabstractIn the hybrid approach, neural network output directly serves as hidden Markov model (HMM) state posterior probability estimates. In contrast to this, in the tandem approach neural network output is used as input features to improve classic Gaussian mixture model (GMM) based emission probability estimates. This paper shows that GMM can be easily integrated into the deep neural network framework. By exploiting its equivalence with the log-linear mixture model (LMM), GMM can be transformed to a large softmax layer followed by a summation pooling layer. Theoretical and experimental results indicate that the jointly trained and optimally chosen GMM and bottleneck tandem features cannot perform worse than a hybrid model. Thus, the question “hybrid vs. tandem” simplifies to optimizing the output layer of a neural network. Speech recognition experiments are carried out on a broadcast news and conversations task using up to 12 feed-forward hidden layers with sigmoid and rectified linear unit activation functions. The evaluation of the LMM layer shows recognition gains over the classic softmax output. Zoltán Tüske, Muhammad Ali Tahir, Ralf Schlüter, Hermann Ney |
ICASSP | 4 |
| 2015 | Sequence-discriminative training of recurrent neural networksabstractWe investigate sequence-discriminative training of long shortterm memory recurrent neural networks using the maximum mutual information criterion. We show that although recurrent neural networks already make use of the whole observation sequence and are able to incorporate more contextual information than feed forward networks, their performance can be improved with sequence-discriminative training. Experiments are performed on two publicly available handwriting recognition tasks containing English and French handwriting. On the English corpus, we obtain a relative improvement in WER of over 11% with maximum mutual information (MMI) training compared to cross-entropy training. On the French corpus, we observed that it is necessary to interpolate the MMI objective function with cross-entropy. Paul Voigtlaender, Patrick Doetsch, Simon Wiesler, Ralf Schlüter, Hermann Ney |
ICASSP | 5 |
| 2015 | Investigations on sequence training of neural networksabstractIn this paper we present an investigation of sequence-discriminative training of deep neural networks for automatic speech recognition. We evaluate different sequence-discriminative training criteria (MMI and MPE) and optimization algorithms (including SGD and Rprop) using the RASR toolkit. Further, we compare the training of the whole network with that of the output layer only. Technical details necessary for a robust training are studied, since there is no consensus yet on the ultimate training recipe. The investigation extends our previous work on training linear bottleneck networks from scratch showing the consistently positive effect of sequence training. Simon Wiesler, Pavel Golik, Ralf Schlüter, Hermann Ney |
ICASSP | 4 |
| 2015 | The LIMSI handwriting recognition system for the HTRtS 2014 contestabstractIn this paper we present the handwriting recognition systems submitted by the LIMSI to the HTRtS 2014 contest. The systems for both the restricted and unrestricted tracks consisted of combination of several optical models. We extracted handcrafted features as well as pixels values with a sliding window. We trained Deep Neural Networks (DNNs) and Bidirectional Long Short-Term Memory Recurrent Neural Networks (BLSTM-RNNs), which where plugged as the optical model in Hidden Markov Models (HMMs). We propose a novel method to build language models that can cope with hyphenation in the text. The combination was performed from lattices generated from the different systems. We were the only team participating in both tracks and ranked second in each. The final Word Error Rates were 15.0% and 11.0% for the restricted (resp. unrestricted) track. We studied the impact of adding data for optical and language modeling. After the evaluation, we also used the same corpus for the language model as the winning team and obtained comparable results. Théodore Bluche, Hermann Ney, Christopher Kermorvant |
ICDAR | 2 |
| 2015 | Framewise and CTC training of Neural Networks for handwriting recognitionabstractIn recent years, Long Short-Term Memory Recurrent Neural Networks (LSTM-RNNs) trained with the Connectionist Temporal Classification (CTC) objective won many international handwriting recognition evaluations. The CTC algorithm is based on a forward-backward procedure, avoiding the need of a segmentation of the input before training. The network outputs are characters labels, and a special non-character label. On the other hand, in the hybrid Neural Network / Hidden Markov Models (NN/HMM) framework, networks are trained with framewise criteria to predict state labels. In this paper, we show that CTC training is close to forward-backward training of NN/HMMs, and can be extended to more standard HMM topologies. We apply this method to Multi-Layer Perceptrons (MLPs), and investigate the properties of CTC, namely the modeling of character by single labels and the role of the special label. Théodore Bluche, Hermann Ney, Jérôme Louradour, Christopher Kermorvant |
ICDAR | 2 |
| 2015 | Bagging by design for continuous Handwriting Recognition using multi-objective particle swarm optimizationabstractMultiple classifier systems are used to improve baseline results using different strategies. Bagging by design improves standard bagging by the minimization of intersection between the different ensembles. This work proposes the use of design bagging for continuous handwriting recognition. The design is performed using a multi-objective particle swarm optimizer. Hidden Markov Models and Long-Short Term Memory Recurrent Neural Networks are used to validate the proposed design. Experiments on English and French Handwriting Recognition with different setups show significant improvements. Mahdi Hamdani, Patrick Doetsch, Hermann Ney |
ICDAR | 3 |
| 2015 | Investigation of Segmental Conditional Random Fields for large vocabulary handwriting recognitionabstractMultiple types of models are used in handwriting recognition and can be broadly categorized into generative and discriminative models. Gaussian Hidden Markov Models are used successfully in most of the systems. Discriminative training can be applied to these models to improve them further. Alternatively, Segmental Conditional Random Fields have the advantage of being discriminative as well as segmental. The novelty of this work is the investigation of Segmental Conditional Random Fields for handwriting recognition. In addition, Multi-Layer Perceptrons and Long Short Term Memory Recurrent Neural Networks are compared for the observations generation in this framework. Various types of features are investigated in the segmental models for handwriting recognition. Furthermore, class-based language model features are proposed to extend this model. Visual features based on moments are extracted at a word level to make the model more robust. Experimental results on English handwriting show a relative reduction of 13.7% in terms of word error rate w.r.t. the baseline system. The proposed system also outperforms the Gaussian Hidden Markov Models trained discriminatively using the minimum phone error criterion by a relative reduction of 6.9% in terms of word error rate. Mahdi Hamdani, M. Ali Basha Shaik, Patrick Doetsch, Hermann Ney |
ICDAR | 4 |
| 2015 | Error bounds for context reduction and feature omission
Eugen Beck, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 3 |
| 2015 | On efficient training of word classes and their application to recurrent neural network language modelsabstractIn this paper, we investigated various word clustering methods, by studying two clustering algorithms: Brown clustering and exchange algorithm, and three objective functions derived from different class-based language models (CBLM): two-sided, predictive and conditional models. In particular, we focused on the implementation of the exchange algorithm with improved speed. In total, we compared six clustering methods in terms of runtime and perplexity (PP) of the CBLM on a French corpus, and show that our accelerated implementation of exchange algorithm is up to 114 times faster than the original and around 6 times faster than the best implementation of Brown clustering we could find, while performing about the same (slightly better) in PP. In addition, we conducted a keyword search experiment on the Babel Lithuanian task (IARPA-babel304b-v1.0b), which showed that CBLM improves the word error rate (WER) but not the keyword search performance. Furthermore, we used these clustering techniques for the output layer of a recurrent neural network (RNN) language model (LM) and we show that in terms of PP of the RNN LM, word classes trained under the predictive model perform slightly better than those trained under other criteria we considered. Index Terms: word clustering, language modeling, neural network based language model, recurrent neural network, long short-term memory Rami Botros, Kazuki Irie, Martin Sundermeyer, Hermann Ney |
INTERSPEECH | 4 |
| 2015 | Convolutional neural networks for acoustic modeling of raw time signal in LVCSRabstractIn this paper we continue to investigate how the deep neural network (DNN) based acoustic models for automatic speech recognition can be trained without hand-crafted feature extraction. Previously, we have shown that a simple fully connected feedforward DNN performs surprisingly well when trained directly on the raw time signal. The analysis of the weights revealed that the DNN has learned a kind of short-time time-frequency decomposition of the speech signal. In conventional feature extraction pipelines this is done manually by means of a filter bank that is shared between the neighboring analysis windows. Following this idea, we show that the performance gap between DNNs trained on spliced hand-crafted features and DNNs trained on raw time signal can be strongly reduced by introducing 1D-convolutional layers. Thus, the DNN is forced to learn a short-time filter bank shared over a longer time span. This also allows us to interpret the weights of the second convolutional layer in the same way as 2D patches learned on critical band energies by typical convolutional neural networks. The evaluation is performed on an English LVCSR task. Trained on the raw time signal, the convolutional layers allow to reduce the WER on the test set from 25.5% to 23.4%, compared to an MFCC based result of 22.1% using fully connected layers. Index Terms: acoustic modeling, raw time signal, convolutional neural networks Pavel Golik, Zoltán Tüske, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 4 |
| 2015 | Multilingual features based keyword search for very low-resource languagesabstractIn this paper we describe RWTH Aachen’s system for keyword search (KWS) with very limited amount of transcribed audio data available in the target language. This setting has become this year’s primary condition within the Babel project [1], seeking to minimize the amount of human effort while retaining a reasonable KWS performance. Thus the highlights presented in this paper include graphemic acoustic modeling; multilingual features trained on language data from the previous project periods; comparison of tandem and hybrid DNN-HMM acoustic models; processing of large amounts of text data available on the web and the morphological KWS based on automatically derived word fragments. The evaluation is performed using two training sets for each of the six current project period’s languages ‐ full language pack (FLP), consisting of 30 hours and very limited language pack (VLLP), comprising less than 3 hours of transcribed audio data. We put our focus on the latter of the two, which is clearly more challenging. The methods described in this work allowed us to exceed 0.3 MTWV on five out of six languages using development queries. Index Terms: acoustic modeling, keyword search, graphemic, multilingual, neural networks, semi-supervised learning Pavel Golik, Zoltán Tüske, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 4 |
| 2015 | Bag-of-words input for long history representation in neural network-based language models for speech recognitionabstractIn most of previous works on neural network based language models (NNLMs), the words are represented as 1-of-N encoded feature vectors. In this paper we investigate an alternative encoding of the word history, known as bag-of-words (BOW) representation of a word sequence, and use it as an additional input feature to the NNLM. Both the feedforward neural network (FFNN) and the long short-term memory recurrent neural network (LSTM-RNN) language models (LMs) with additional BOW input are evaluated on an English large vocabulary automatic speech recognition (ASR) task. We show that the BOW features significantly improve both the perplexity (PP) and the word error rate (WER) of a standard FFNN LM. In contrast, the LSTM-RNN LM does not benefit from such an explicit long context feature. Therefore the performance gap between feedforward and recurrent architectures for language modeling is reduced. In addition, we revisit the cache based LM, a seeming analog of the BOW for the count based LM, which was unsuccessful for ASR in the past. Although the cache is able to improve the perplexity, we only observe a very small reduction in WER. Index Terms: language modeling, speech recognition, bag-ofwords, feedforward neural networks, recurrent neural networks, long short-term memory, cache language model Kazuki Irie, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 3 |
| 2015 | Improvements in RWTH LVCSR evaluation systems for Polish, Portuguese, English, urdu, and ArabicabstractIn this work, Portuguese, Polish, English, Urdu, and Arabic automatic speech recognition evaluation systems developed by the RWTH Aachen University are presented. Our LVCSR systems focus on various domains like broadcast news, spontaneous speech, and podcasts. All these systems but Urdu are used for Euronews and Skynews evaluations as part of the EUBridge project. Our previously developed LVCSR systems were improved using different techniques for the aforementioned languages. Significant improvements are obtained using multilingual tandem and hybrid approaches, minimum phone error training, lexical adaptation, open vocabulary long short term memory language models, maximum entropy language models and confusion-network based system combination. Index Terms: LVCSR, LSTM, open-vocabulary, EU-Bridge M. Ali Basha Shaik, Zoltán Tüske, Muhammad Ali Tahir, Markus Nußbaum-Thom, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 6 |
| 2015 | A Comparison of Update Strategies for Large-Scale Maximum Expected BLEU TrainingabstractJoern Wuebker, Sebastian Muehr, Patrick Lehnen, Stephan Peitz, Hermann Ney. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015. Jörn Wübker, Sebastian Muehr, Patrick Lehnen, Stephan Peitz, Hermann Ney |
HLT-NAACL | 5 |
| 2015 | Continuous sign language recognition: Towards large vocabulary statistical recognition systems handling multiple signers
Oscar Koller, Jens Forster, Hermann Ney |
Comput. Vis. Image Underst. | 3 |
| 2015 | From Feedforward to Recurrent LSTM Neural Networks for Language ModelingabstractLanguage models have traditionally been estimated based on relative frequencies, using count statistics that can be extracted from huge amounts of text data. More recently, it has been found that neural networks are particularly powerful at estimating probability distributions over word sequences, giving substantial improvements over state-of-the-art count models. However, the performance of neural network language models strongly depends on their architectural structure. This paper compares count models to feedforward, recurrent, and long short-term memory (LSTM) neural network variants on two large-vocabulary speech recognition tasks. We evaluate the models in terms of perplexity and word error rate, experimentally validating the strong correlation of the two quantities, which we find to hold regardless of the underlying type of the language model. Furthermore, neural networks incur an increased computational complexity compared to count models, and they differently model context dependences, often exceeding the number of words that are taken into account by count based approaches. These differences require efficient search methods for neural networks, and we analyze the potential improvements that can be obtained when applying advanced algorithms to the rescoring of word lattices on large-scale setups. Martin Sundermeyer, Hermann Ney, Ralf Schlüter |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2014 | The RWTH Large Vocabulary Arabic Handwriting Recognition SystemabstractThis paper describes the RWTH system for large vocabulary Arabic handwriting recognition. The recognizer is based on Hidden Markov Models (HMMs) with state of the art methods for visual/language modeling and decoding. The feature extraction is based on Recurrent Neural Networks (RNNs) which estimate the posterior distribution over the character labels for each observation. Discriminative training using the Minimum Phone Error (MPE) criterion is used to train the HMMs. The recognition is done with the help of n-gram Language Models (LMs) trained using in-domain text data. Unsupervised writer adaptation is also performed using the Constrained Maximum Likelihood Linear Regression (CMLLR) feature adaptation. The RWTH Arabic handwriting recognition system gave competitive results in previous handwriting recognition competitions. The used techniques allows to improve the performance of the system participating in the OpenHaRT 2013 evaluation. Mahdi Hamdani, Patrick Doetsch, Michal Kozielski, Amr El-Desoky Mousa, Hermann Ney |
Document Analysis Systems | 5 |
| 2014 | Multilingual Off-Line Handwriting Recognition in Real-World ImagesabstractWe propose a state-of-the-art system for recognizing real-world handwritten images exposing a huge degree of noise and a high out-of-vocabulary rate. We describe methods for successful image demising, line removal, deskewing, deslanting, and text line segmentation. We demonstrate how to use a HMM-based recognition system to obtain competitive results, and how to further improve it using LSTM neural networks in the tandem approach. The final system outperforms other approaches on a new dataset for English and French handwriting. The presented framework scales well across other standard datasets. Michal Kozielski, Patrick Doetsch, Mahdi Hamdani, Hermann Ney |
Document Analysis Systems | 4 |
| 2014 | Jane: Open Source Machine Translation System CombinationabstractDifferent machine translation engines can be remarkably dissimilar not only with respect to their technical paradigm, but also with respect to the translation output they yield.System combination is a method for combining the output of multiple machine translation engines in order to take benefit of the strengths of each of the individual engines.In this work we introduce a novel system combination implementation which is integrated into Jane, RWTH's open source statistical machine translation toolkit.On the most recent Workshop on Statistical Machine Translation system combination shared task, we achieve improvements of up to 0.7 points in BLEU over the best system combination hypotheses which were submitted for the official evaluation.Moreover, we enhance our system combination pipeline with additional n-gram language models and lexical translation models. Markus Freitag, Matthias Huck, Hermann Ney |
EACL | 3 |
| 2014 | Simple and Effective Approach for Consistent Training of Hierarchical Phrase-based Translation ModelsabstractIn this paper, we present a simple approach for consistent training of hierarchical phrase-based translation models.In order to consistently train a translation model, we perform hierarchical phrasebased decoding on training data to find derivations between the source and target sentences.This is done by synchronous parsing the given sentence pairs.After extracting k-best derivations, we reestimate the translation model probabilities based on collected rule counts.We show the effectiveness of our procedure on the IWSLT German→English and English→French translation tasks.Our results show improvements of up to 1.6 points BLEU. Stephan Peitz, David Vilar, Hermann Ney |
EACL | 3 |
| 2014 | Translation model based weighting for phrase extraction
Saab Mansour, Hermann Ney |
EAMT | 2 |
| 2014 | Read My Lips: Continuous Signer Independent Weakly Supervised Viseme Recognition
Oscar Koller, Hermann Ney, Richard Bowden |
ECCV (1) | 2 |
| 2014 | Improved Decipherment of Homophonic CiphersabstractIn this paper, we present two improvements to the beam search approach for solving homophonic substitution ciphers presented in Nuhn et al. (2013): An improved rest cost estimation together with an optimized strategy for obtaining the order in which the symbols of the cipher are deciphered reduces the beam size needed to successfully decipher the Zodiac-408 cipher from several million down to less than one hundred: The search effort is reduced from several hours of computation time to just a few seconds on a single CPU.These improvements allow us to successfully decipher the second part of the famous Beale cipher (see (Ward et al., 1885) and e.g.(King, 1993)): Having 182 different cipher symbols while having a length of just 762 symbols, the decipherment is way more challenging than the decipherment of the previously deciphered Zodiac-408 cipher (length 408, 54 different symbols).To the best of our knowledge, this cipher has not been deciphered automatically before. Malte Nuhn, Julian Schamper, Hermann Ney |
EMNLP | 3 |
| 2014 | Translation Modeling with Bidirectional Recurrent Neural NetworksabstractThis work presents two different trans-lation models using recurrent neural net-works. The first one is a word-based ap-proach using word alignments. Second, we present phrase-based translation mod-els that are more consistent with phrase-based decoding. Moreover, we introduce bidirectional recurrent neural models to the problem of machine translation, allow-ing us to use the full source sentence in our models, which is also of theoretical inter-est. We demonstrate that our translation models are capable of improving strong baselines already including recurrent neu-ral language models on three tasks: IWSLT 2013 German→English, BOLT Arabic→English and Chinese→English. We obtain gains up to 1.6 % BLEU and 1.7 % TER by rescoring 1000-best lists. 1 Martin Sundermeyer, Tamer Alkhouli, Jörn Wübker, Hermann Ney |
EMNLP | 4 |
| 2014 | Progress in dynamic network decodingabstractWe show how we boosted the efficiency of the dynamic network decoder in IBM's Attila speech recognition framework, by transforming the underlying concept from token-passing to word-conditioned, and adding speedup methods like sparse LM look-ahead. On several different tasks, we achieve improvements of 30 to 50% in efficiency at equal precision. We compare the efficiency to a state-of-the-art WFST based static decoder, and note that the added methods improve the dynamic decoder under conditions where it was lacking before in comparison, specifically when using a relatively small LM. Overall, the new dynamic decoder performs similarly to the static decoder, with a lead for the dynamic decoder on tasks with a larger LM, and a lead for the static decoder on tasks with a smaller LM. David Nolden, Hagen Soltau, Hermann Ney |
ICASSP | 3 |
| 2014 | A family of discriminative training criteria based on the F-divergence for deep neural networksabstractWe present novel bounds on the classification error which are based on the f-Divergence and, at the same time, can be used as practical training criteria. There exist virtually no studies which investigate the link between the f-Divergence, the classification error and practical training criteria. So far only the Kullback-Leibler f-Divergence has been examined in this context to formulate a bound on the classification error and to derive the cross-entropy criterion. We extend this concept to a larger class of f-Divergences. We also successfully investigate if the novel training criteria based on the f-Divergence are suited for frame-wise training of deep neural networks on the Babel Vietnamese and Bengali speech recognition tasks. Markus Nußbaum-Thom, Ralf Schlüter, Vaibhava Goel, Hermann Ney |
ICASSP | 5 |
| 2014 | Multilingual MRASTA features for low-resource keyword search and speech recognition systemsabstractThis paper investigates the application of hierarchical MRASTA bottleneck (BN) features for under-resourced languages within the IARPA Babel project. Through multilingual training of Multilayer Perceptron (MLP) BN features on five languages (Cantonese, Pashto, Tagalog, Turkish, and Vietnamese), we could end up in a single feature stream which is more beneficial to all languages than the unilingual features. In the case of balanced corpus sizes, the multilingual BN features improve the automatic speech recognition (ASR) performance by 3-5% and the keyword search (KWS) by 3-10% relative for both limited (LLP) and full language packs (FLP). Borrowing orders of magnitude more data from non-target FLPs, the recognition error rate is reduced by 8-10%, and the spoken term detection is improved by over 40% relative on Vietnamese and Pashto LLP. Aiming at the fast development of acoustic models, cross-lingual transfer of multilingually ”pretrained” BN features for a new language is also investigated. Without the need of any MLP training on the new language, the ported BN features performed similarly to the unilingual features on FLP and significantly better on LLP. Results also show that a simple fine-tuning step on the new language is enough to achieve comparable KWS and ASR performance to that system where the target language is also involved in the time-consuming multilingual training. Zoltán Tüske, David Nolden, Ralf Schlüter, Hermann Ney |
ICASSP | 4 |
| 2014 | The RWTH English lecture recognition systemabstractIn this paper, we describe the RWTH speech recognition system for English lectures developed within the Translectures project. A difficulty in the development of an English lectures recognition system, is the high ratio of non-native speakers. We address this problem by using very effective deep bottleneck features trained on multilingual data. The acoustic model is trained on large amounts of data from different domains and with different dialects. Large improvements are obtained from unsupervised acoustic adaptation. Another challenge is the frequent use of technical terms and the wide range of topics. In our recognition system, slides, which are attached to most lectures, are used for improving lexical coverage and language model adaptation. Simon Wiesler, Kazuki Irie, Zoltán Tüske, Ralf Schlüter, Hermann Ney |
ICASSP | 5 |
| 2014 | RASR/NN: The RWTH neural network toolkit for speech recognitionabstractThis paper describes the new release of RASR — the open source version of the well-proven speech recognition toolkit developed and used at RWTH Aachen University. The focus is put on the implementation of the NN module for training neural network acoustic models. We describe code design, configuration, and features of the NN module. The key feature is a high flexibility regarding the network topology, choice of activation functions, training criteria, and optimization algorithm, as well as a built-in support for efficient GPU computing. The evaluation of run-time performance and recognition accuracy is performed exemplary with a deep neural network as acoustic model in a hybrid NN/HMM system. The results show that RASR achieves a state-of-the-art performance on a real-world large vocabulary task, while offering a complete pipeline for building and applying large scale speech recognition systems. Simon Wiesler, Alexander Richard, Pavel Golik, Ralf Schlüter, Hermann Ney |
ICASSP | 5 |
| 2014 | Mean-normalized stochastic gradient for large-scale deep learningabstractDeep neural networks are typically optimized with stochastic gradient descent (SGD). In this work, we propose a novel second-order stochastic optimization algorithm. The algorithm is based on analytic results showing that a non-zero mean of features is harmful for the optimization. We prove convergence of our algorithm in a convex setting. In our experiments we show that our proposed algorithm converges faster than SGD. Further, in contrast to earlier work, our algorithm allows for training models with a factorized structure from scratch. We found this structure to be very useful not only because it accelerates training and decoding, but also because it is a very effective means against overfitting. Combining our proposed optimization algorithm with this model structure, model size can be reduced by a factor of eight and still improvements in recognition error rate are obtained. Additional gains are obtained by improving the Newbob learning rate strategy. Simon Wiesler, Alexander Richard, Ralf Schlüter, Hermann Ney |
ICASSP | 4 |
| 2014 | Fast and Robust Training of Recurrent Neural Networks for Offline Handwriting RecognitionabstractIn this paper we demonstrate a modified topology for long short-term memory recurrent neural networks that controls the shape of the squashing functions in gating units. We further propose an efficient training framework based on a mini-batch training on sequence level combined with a sequence chunking approach. The framework is evaluated on publicly available data sets containing English and French handwriting by utilizing a GPU based implementation. Speedups of more than 3x are achieved in training recurrent neural network models which outperform state of the art recognition results. Patrick Doetsch, Michal Kozielski, Hermann Ney |
ICFHR | 3 |
| 2014 | Improvement of Context Dependent Modeling for Arabic Handwriting RecognitionabstractThis paper proposes the improvement of context dependent modeling for Arabic handwriting recognition. Since the number of parameters in context dependent models is huge, CART trees are used for state tying. This work is based on a new set of questions for the CART tree construction based on a "lossy mapping" categorization of the Arabic shapes. The used system is a combination of Hidden Markov Models and Recurrent Neural Networks using the hybrid approach. A comparison between a Neural network trained using the baseline labels and another one based on the CART tree labels is done. The experimental results show that the use of the CART labels for the Neural Network training beneficial. The lossy mapping based CART tree performed better than the baseline system. An absolute improvement of 2.9% in terms of Word Error Rate is performed on the test set of the Open Hart database. Mahdi Hamdani, Patrick Doetsch, Hermann Ney |
ICFHR | 3 |
| 2014 | Open-Lexicon Language Modeling Combining Word and Character LevelsabstractIn this paper we investigate different n-gram language models that are defined over an open lexicon. We introduce a character-level language model and combine it with a standard word-level language model in a back off fashion. The character-level language model is redefined and renormalized to assign zero probability to words from a fixed vocabulary. Furthermore we present a way to interpolate language models created at the word and character levels. The computation of character-level probabilities incorporates the across-word context. We compare perplexities on all words from the test set and on in-lexicon and OOV words separately on corpora of English and Arabic text. Michal Kozielski, Martin Matysiak, Patrick Doetsch, Ralf Schlüter, Hermann Ney |
ICFHR | 5 |
| 2014 | Towards Unsupervised Learning for Handwriting RecognitionabstractWe present a method for training an off-line handwriting recognition system in an unsupervised manner. For an isolated word recognition task, we are able to bootstrap the system without any annotated data. We then retrain the system using the best hypothesis from a previous recognition pass in an iterative fashion. Our approach relies only on a prior language model and does not depend on an explicit segmentation of words into characters. The resulting system shows a promising performance on a standard dataset in comparison to a system trained in a supervised fashion for the same amount of training data. Michal Kozielski, Malte Nuhn, Patrick Doetsch, Hermann Ney |
ICFHR | 4 |
| 2014 | Fine-Grained Visual Categorization with 2D-WarpingabstractThe task of fine-grained visual categorization is related to both general object recognition and specialized tasks such as face recognition. Hence, we propose to combine two methods popular for general object recognition and face recognition to build a new model-free system for fine-grained visual categorization. Specifically, we use Local Naive-Bayes Nearest Neighbor as a pre-selection method and 2D-Warping as a refinement step. For the latter, we explore different ways to use the alignments computed by a 2D-Warping algorithm for classification. We demonstrate the performance of our approach on the CUB200-2011 database and show that our approach outperforms the recognition accuracy of current state-of-the-art methods. Harald Hanselmann, Hermann Ney |
ICPR | 2 |
| 2014 | Word pair approximation for more efficient decoding with high-order language models
David Nolden, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 3 |
| 2014 | Removing redundancy from lattices
David Nolden, Hagen Soltau, Daniel Povey, Pegah Ghahremani, Lidia Mangu, Hermann Ney |
INTERSPEECH | 6 |
| 2014 | RWTH LVCSR systems for quaero and EU-bridge: German, Polish, Spanish and PortugueseabstractIn this paper, German, Polish, Spanish, and Portuguese large vocabulary continuous speech recognition (LVCSR) systems developed by the RWTH Aachen University are presented.All the above mentioned systems for the aforementioned languages are used for the Quaero and EU-Bridge project evaluations.The LVCSR systems developed for these competitive evaluations focus on various domains like broadcast news, podcasts and lecture domain.Transcription of the speech for these tasks is challenging due to huge variability in the acoustic conditions and a significant portion of audio data includes spontaneous speech.Good improvements are obtained using stateof-the-art multilingual bottleneck features, minimum phone error trained acoustic models, language model (LM) adaptation and confusion-network based system combination.In addition, an open vocabulary approach using morphemic units is investigated along with the LM adaptation for the German LVCSR. M. Ali Basha Shaik, Zoltán Tüske, Muhammad Ali Tahir, Markus Nußbaum-Thom, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 6 |
| 2014 | rwthlm - the RWTH aachen university neural network language modeling toolkitabstractWe present a novel toolkit that implements the long short-term memory (LSTM) neural network concept for language modeling. The main goal is to provide a software which is easy to use, and which allows fast training of standard recurrent and LSTM neural network language models. The toolkit obtains state-of-the-art performance on the standard Treebank corpus. To reduce the training time, BLAS and related libraries are supported, and it is possible to evaluate multiple word sequences in parallel. In addition, arbitrary word classes can be used to speed up the computation in case of large vocabulary sizes. Finally, the software allows easy integration with SRILM, and it supports direct decoding and rescoring of HTK lattices. The toolkit is available for download under an open source license. Martin Sundermeyer, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 3 |
| 2014 | Lattice decoding and rescoring with long-Span neural network language modelsabstractWith long-span neural network language models, considerable improvements have been obtained in speech recognition. However, it is difficult to apply these models if the underlying search space is large. In this paper, we combine previous work on lattice decoding with long short-term memory (LSTM) neural network language models. By adding refined pruning techniques, we are able to reduce the search effort by a factor of three. Furthermore, we introduce two novel approximations for full lattice rescoring, which opens the potential of lattice-based speech recognition techniques. Compared to 1000-best lists, we find that we can increase the word error rate improvements obtained with LSTMs from 8.2 % to 10.7 % relative over a stateof-the-art baseline, while the resulting lattices are even considerably smaller. In addition, we investigate the use of LSTMs for Babel Assamese keyword search, obtaining significant improvements of 2.5 % relative. Martin Sundermeyer, Zoltán Tüske, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 4 |
| 2014 | Data augmentation, feature combination, and multilingual neural networks to improve ASR and KWS performance for low-resource languagesabstractThis paper presents the progress of acoustic models for lowresourced languages (Assamese, Bengali, Haitian Creole, Lao, Zulu) developed within the second evaluation campaign of the IARPA Babel project.This year, the main focus of the project is put on training high-performing automatic speech recognition (ASR) and keyword search (KWS) systems from language resources limited to about 10 hours of transcribed speech data.Optimizing the structure of Multilayer Perceptron (MLP) based feature extraction and switching from the sigmoid activation function to rectified linear units results in about 5% relative improvement over baseline MLP features.Further improvements are obtained when the MLPs are trained on multiple feature streams and by exploiting label preserving data augmentation techniques like vocal tract length perturbation.Systematic application of these methods allows to improve the unilingual systems by 4-6% absolute in WER and 0.064-0.105absolute in MTWV.Transfer and adaptation of multilingually trained MLPs lead to additional gains, clearly exceeding the project goal of 0.3 MTWV even when only the limited language pack of the target language is used. Zoltán Tüske, Pavel Golik, David Nolden, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 5 |
| 2014 | Acoustic modeling with deep neural networks using raw time signal for LVCSR
Zoltán Tüske, Pavel Golik, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 4 |
| 2014 | Extensions of the Sign Language Recognition and Translation Corpus RWTH-PHOENIX-Weather
Jens Forster, Oscar Koller, Martin Bellgardt, Hermann Ney |
LREC | 5 |
| 2013 | Advancements in Reordering Models for Statistical Machine Translation
Minwei Feng, Jan-Thorsten Peter, Hermann Ney |
ACL (1) | 3 |
| 2013 | Decipherment Complexity in 1: 1 Substitution Ciphers
Malte Nuhn, Hermann Ney |
ACL (1) | 2 |
| 2013 | Beam Search for Solving Substitution Ciphers
Malte Nuhn, Julian Schamper, Hermann Ney |
ACL (1) | 3 |
| 2013 | Efficient nearly error-less LVCSR decoding based on incremental forward and backward passesabstractWe show that most search errors can be identified by aligning the results of a symmetric forward and backward decoding pass. Based on this knowledge, we introduce an efficient high-level decoding architecture which yields virtually no search errors, and requires virtually no manual tuning. We perform an initial forward- and backward decoding with tight initial beams, then we identify search errors, and then we recursively increment the beam sizes and perform new forward and backward decodings for erroneous intervals until no more search errors are detected. Consequently, each utterance and even each single word is decoded with the smallest beam size required to decode it correctly. On all tested systems we achieve an error rate equal or very close to classical decoding with ideally tuned beam size, but unsupervisedly without specific tuning, and at around 2 times faster runtime. An additional speedup by factor 2 can be achieved by decoding the forward and backward pass in separate threads. David Nolden, Ralf Schlüter, Hermann Ney |
ASRU | 3 |
| 2013 | Improving Statistical Machine Translation with Word Class ModelsabstractAutomatically clustering words from a monolingual or bilingual training corpus into classes is a widely used technique in statistical natural language processing.We present a very simple and easy to implement method for using these word classes to improve translation quality.It can be applied across different machine translation paradigms and with arbitrary types of models.We show its efficacy on a small German→English and a larger French→German translation task with both standard phrase-based and hierarchical phrase-based translation systems for a common set of models.Our results show that with word class models, the baseline can be improved by up to 1.4% BLEU and 1.0% TER on the French→German task and 0.3% BLEU and 1.1% TER on the German→English task. Jörn Wübker, Stephan Peitz, Felix Rietig, Hermann Ney |
EMNLP | 4 |
| 2013 | Tandem HMM with convolutional neural network for handwritten word recognitionabstractIn this paper, we investigate the combination of hidden Markov models and convolutional neural networks for handwritten word recognition. The convolutional neural networks have been successfully applied to various computer vision tasks, including handwritten character recognition. In this work, we show that they can replace Gaussian mixtures to compute emission probabilities in hidden Markov models (hybrid combination), or serve as feature extractor for a standard Gaussian HMM system (tandem combination). The proposed systems outperform a basic HMM based on either decorrelated pixels or handcrafted features. We validated the approach on two publicly available databases, and we report up to 60% (Rimes) and 35% (IAM) relative improvement compared to a Gaussian HMM based on pixel values. The final systems give comparable results to recurrent neural networks, which are the best systems since 2009. Théodore Bluche, Hermann Ney, Christopher Kermorvant |
ICASSP | 2 |
| 2013 | Open vocabulary handwriting recognition using combined word-level and character-level language modelsabstractIn this paper, we present a unified search strategy for open vocabulary handwriting recognition using weighted finite state transducers. Additionally to a standard word-level language model we introduce a separate n-gram character-level language model for out-of-vocabulary word detection and recognition. The probabilities assigned by those two models are combined into one Bayes decision rule. We evaluate the proposed method on the IAM database of English handwriting. An improvement from 22.2% word error rate to 17.3% is achieved comparing to the closed-vocabulary scenario and the best published result. Michal Kozielski, David Rybach, Stefan Hahn, Ralf Schlüter, Hermann Ney |
ICASSP | 5 |
| 2013 | Advanced search space pruning with acoustic look-ahead for WFST based LVCSRabstractIn this work we show how some concepts already known from dynamic network decoding can be used to improve the efficiency of WFST based decoders. First we apply the concept of acoustic look-ahead to a WFST based decoder, and then we analyze the applicability of LM state pruning, a well motivated pruning method which is fundamental to token-passing decoders. The structure of the composed WFST search network makes it difficult to motivate advanced pruning methods, and consequently it is difficult to achieve a real reduction in search space. Nonetheless, we show how LM state pruning can be applied to WFST based decoders to improve their efficiency. The search space can be reduced by up to 50% at equal precision through acoustic look-ahead. Since our decoder follows a dynamic composition approach, the advantage in search space does not fully transfer to the RTF, which can be reduced by around 20% through acoustic look-ahead, and additional 5% through LM state pruning. David Nolden, Ralf Schlüter, Hermann Ney |
ICASSP | 3 |
| 2013 | Feature combination and stacking of recurrent and non-recurrent neural networks for LVCSRabstractThis paper investigates the combination of different short-term features and the combination of recurrent and non-recurrent neural networks (NNs) on a Spanish speech recognition task. Several methods exist to combine different feature sets such as concatenation or linear discriminant analysis (LDA). Even though all these techniques achieve reasonable improvements, feature combination by multi-layer perceptrons (MLPs) outperforms all known approaches. We develop the concept of MLP based feature combination further using recurrent neural networks (RNNs). The phoneme posterior estimates derived from an RNN lead to a significant improvement over the result of the MLPs and achieve a 5% relative better word error rate (WER) with much less parameters. Moreover, we improve the system performance further by combining an MLP and an RNN in a hierarchical framework. The MLP benefits from the preprocessing of the RNN. All NNs are trained on phonemes. Nevertheless, the same concepts could be applied using context-dependent states. In addition to the improvements in recognition performance w.r.t. WER, NN based feature combination methods reduce both, the training and the testing complexity. Overall, the systems are based on a single set of acoustic models, together with the training of different NNs. Christian Plahl, Michal Kozielski, Ralf Schlüter, Hermann Ney |
ICASSP | 4 |
| 2013 | Comparison of feedforward and recurrent neural network language modelsabstractResearch on language modeling for speech recognition has increasingly focused on the application of neural networks. Two competing concepts have been developed: On the one hand, feedforward neural networks representing an n-gram approach, on the other hand recurrent neural networks that may learn context dependencies spanning more than a fixed number of predecessor words. To the best of our knowledge, no comparison has been carried out between feedforward and state-of-the-art recurrent networks when applied to speech recognition. This paper analyzes this aspect in detail on a well-tuned French speech recognition task. In addition, we propose a simple and efficient method to normalize language model probabilities across different vocabularies, and we show how to speed up training of recurrent neural networks by parallelization. Martin Sundermeyer, Ilya Oparin, Jean-Luc Gauvain, B. Freiberg, Ralf Schlüter, Hermann Ney |
ICASSP | 6 |
| 2013 | Deep hierarchical bottleneck MRASTA features for LVCSRabstractHierarchical Multi Layer Perceptron (MLP) based long-term feature extraction is optimized for TANDEM connectionist large vocabulary continuous speech recognition (LVCSR) system within the QUAERO project. Training the bottleneck MLP on multi-resolutional RASTA filtered critical band energies, more than 20% relative word error rate (WER) reduction over standard MFCC system is observed after optimizing the number of target labels. Furthermore, introducing a deeper structure in the hierarchical bottleneck processing the relative gain increases to 25%. The final system based on deep bottleneck TANDEM features clearly outperforms the hybrid approach, even if the long-term features are also presented to the deep MLP acoustic model. The results are also verified on evaluation data of the year 2012, and about 20% relative WER improvement over classical cepstral system is measured even after speaker adaptive training. Zoltán Tüske, Ralf Schlüter, Hermann Ney |
ICASSP | 3 |
| 2013 | A critical evaluation of stochastic algorithms for convex optimizationabstractLog-linear models find a wide range of applications in pattern recognition. The training of log-linear models is a convex optimization problem. In this work, we compare the performance of stochastic and batch optimization algorithms. Stochastic algorithms are fast on large data sets but can not be parallelized well. In our experiments on a broadcast conversations recognition task, stochastic methods yield competitive results after only a short training period, but when spending enough computational resources for parallelization, batch algorithms are competitive with stochastic algorithms. We obtained slight improvements by using a stochastic second order algorithm. Our best log-linear model outperforms the maximum likelihood trained Gaussian mixture model baseline although being ten times smaller. Simon Wiesler, Alexander Richard, Ralf Schlüter, Hermann Ney |
ICASSP | 4 |
| 2013 | Feature Extraction with Convolutional Neural Networks for Handwritten Word RecognitionabstractIn this paper, we show that learning features with convolutional neural networks is better than using hand-crafted features for handwritten word recognition. We consider two kinds of systems: a grapheme based segmentation and a sliding window segmentation. In both cases, the combination of a convolutional neural network with a HMM outperform a state-of-the art HMM system based on explicit feature extraction. The experiments are conducted on the Rimes database. The systems obtained with the two kinds of segmentation are complementary: when they are combined, they outperform the systems in isolation. The system based on grapheme segmentation yields lower recognition rate but is very fast, which is suitable for specific applications such as document classification. Théodore Bluche, Hermann Ney, Christopher Kermorvant |
ICDAR | 2 |
| 2013 | Open Vocabulary Arabic Handwriting Recognition Using Morphological DecompositionabstractThe use of Language Models (LMs) is a very important component in large and open vocabulary recognition systems. This paper presents an open-vocabulary approach for Arabic handwriting recognition. The proposed approach makes use of Arabic word decomposition based on morphological analysis. The vocabulary is a combination of words and sub-words obtained by the decomposition process. Out Of Vocabulary (OOV) words can be recognized by combining different elements from the lexicon. The recognition system is based on Hidden Markov Models (HMMs) with position and context dependent character models. An n-gram LM trained on the decomposed text is used along with the HMMs during the search. The approach is evaluated using two Arabic handwriting datasets. The open vocabulary approach leads to a significant improvement in the system performance. Two different types experiments for two Arabic handwriting recognition tasks are conducted in this work. The proposed approach for open vocabulary allows to have an absolute improvement of up to 1% in the Word Error Rate (WER) for the constrained task and to keep the same performance of the baseline system for the unconstrained one. Mahdi Hamdani, Amr El-Desoky Mousa, Hermann Ney |
ICDAR | 3 |
| 2013 | Improvements in RWTH's System for Off-Line Handwriting RecognitionabstractIn this paper we describe a novel HMM-based system for off-line handwriting recognition. We adapt successful techniques from the domains of large vocabulary speech recognition and image object recognition: moment-based image normalization, writer adaptation, discriminative feature extraction and training, and open-vocabulary recognition. We evaluate those methods and examine their cumulative effect on the recognition performance. The final system outperforms current state-of-the-art approaches on two standard evaluation corpora for English and French handwriting. Michal Kozielski, Patrick Doetsch, Hermann Ney |
ICDAR | 3 |
| 2013 | Cross-entropy vs. squared error training: a theoretical and experimental comparisonabstractIn this paper we investigate the error criteria that are optimized during the training of artificial neural networks (ANN).We compare the bounds of the squared error (SE) and the crossentropy (CE) criteria being the most popular choices in stateof-the art implementations.The evaluation is performed on automatic speech recognition (ASR) and handwriting recognition (HWR) tasks using a hybrid HMM-ANN model.We find that with randomly initialized weights, the squared error based ANN does not converge to a good local optimum.However, with a good initialization by pre-training, the word error rate of our best CE trained system could be reduced from 30.9% to 30.5% on the ASR, and from 22.7% to 21.9% on the HWR task by performing a few additional "fine-tuning" iterations with the SE criterion. Pavel Golik, Patrick Doetsch, Hermann Ney |
INTERSPEECH | 3 |
| 2013 | Development of the RWTH transcription system for slovenianabstractIn this paper we describe the RWTH automatic speech recognition system for Slovenian developed within the transLectures project.The project aims at supporting the transcription and translation of video lectures freely available on the web.Difficulties arise on all levels of modeling: Slovenian is a morphologically rich language with a high level of inflection (pronunciation model), and a large variety of dialects and recording conditions brings uncertainty into the audio signal (acoustic model).Moreover, the video lectures cover a wide spectrum of topics with a high share of spontaneous speech and technical terms (language model).These issues require application of robust and adaptive methods.Besides the system description, this study mainly focuses on robust acoustic modeling.Building acoustic models from various resources, we also compare the influence of speaker adaptation to different neural network based acoustic features.Systematic application of these methods allows us to reduce the word error rate on the evaluation corpus from 59.2% to 43.4%.We also give a motivation for Slovenian open vocabulary recognition and perform some first steps. Pavel Golik, Zoltán Tüske, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 4 |
| 2013 | Improving LVCSR with hidden conditional random fields for grapheme-to-phoneme conversionabstractIn virtually every state-of-the-art large vocabulary continuous speech recognition (LVCSR) system, grapheme-to-phoneme (G2P) conversion is applied to generalize beyond a fixed set of words given by a background lexicon. The overall performance of the G2P system has a strong effect on the recognition qual-ity. Typically, generative models based on joint-n-grams are used, although some discriminative models have a competitive performance but the training time may be quite large. In this work, the effect of using discriminative G2P modeling based on hidden conditional random fields (HCRFs) is ana-lyzed. Besides measuring and comparing the G2P qualities on a textual level, one focus is the performance of LVCSR systems. Although the HCRF model does not outperform the generative one on text data, we could improve our English QUAERO ASR system by 1-3 % relative on a couple of test corpora over a strong baseline by only replacing the G2P strategy. Index Terms: grapheme-to-phoneme conversion, G2P, LVCSR, HCRF, hidden conditional random fields Stefan Hahn, Patrick Lehnen, Simon Wiesler, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 5 |
| 2013 | Structure learning in hidden conditional random fields for grapheme-to-phoneme conversionabstractAccurate grapheme-to-phoneme (g2p) conversion is needed for several speech processing applications, such as automatic speech synthesis and recognition.For some languages, notably English, improvements of g2p systems are very slow, due to the intricacy of the associations between letter and sounds.In recent years, several improvements have been obtained either by using variable-length associations in generative models (jointn-grams), or by recasting the problem as a conventional sequence labeling task, enabling to integrate rich dependencies in discriminative models.In this paper, we consider several ways to reconciliate these two approaches.Introducing hidden variable-length alignments through latent variables, our Hidden Conditional Random Field (HCRF) models are able to produce comparative performance compared to strong generative and discriminative models on the CELEX database. Patrick Lehnen, Alexandre Allauzen, Thomas Lavergne, François Yvon, Stefan Hahn, Hermann Ney |
INTERSPEECH | 6 |
| 2013 | Morpheme level hierarchical pitman-yor class-based language models for LVCSR of morphologically rich languagesabstractPerforming large vocabulary continuous speech recognition (LVCSR) for morphologically rich languages is considered a challenging task.The morphological richness of such languages leads to high out-of-vocabulary (OOV) rates and poor language model (LM) probabilities.In this case, the use of morphemes has been shown to increase the lexical coverage and lower the LM perplexity.Another approach used to improve the LM probability estimates is to incorporate additional knowledge sources in the LM estimation process using classbased LMs (CLMs).Recently, the hierarchical Pitman-Yor LMs (HPYLMs) have shown superiority over the modified Kneser-Ney (MKN) smoothed N-gram LMs in terms of both perplexity (PPL) and word error rate (WER) on word-based LVCSR tasks.In this paper, hierarchical Pitman-Yor class-based LMs (HPY-CLMs) are combined with morpheme level language modeling.This enables the application of the proposed models on top of morpheme-based systems.Experiments are conducted on Arabic and German LVCSR tasks.Consistent performance improvements are obtained for all the available corpora compared to the conventional morpheme-based and class-based LMs. Amr El-Desoky Mousa, M. Ali Basha Shaik, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 4 |
| 2013 | Relative error bounds for statistical classifiers based on the f-divergenceabstractIn language classification, measures like perplexity and Kullback-Leibler divergence are used to compare language models. While this bears the advantage of isolating the effect of the language model in speech and language processing problems, the measures have no clear relation to the corresponding classification error. In practice, an improvement in terms of perplexity does not necessarily correspond to an improvement in the error rate. It is well-known that Bayes decision rule is optimal if the true distribution is used for classification. Since the true distribution is unknown in practice, a model distribution is used instead, introducing suboptimality. We focus on the degradation introduced by a model distribution, and provide an upper bound on the error difference between Bayes decision and a modelbased decision rule in terms of the f-Divergence between the true and model distributions. Simulations are first presented to reveal a special case of the bound, followed by an analytic proof of the generalized bound and its tightness. In addition, the conditions that result in the boundary cases will be discussed. Several instances of the bound will be verified using simulations, and the bound will be used to study the effect of the language model on the classification error. Index Terms: generalization bounds, language modeling, perplexity, confidence measures, f-Divergence, error mismatch. Markus Nußbaum-Thom, Eugen Beck, Tamer Alkhouli, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 5 |
| 2013 | Feature-rich sub-lexical language models using a maximum entropy approach for German LVCSRabstractGerman is a morphologically rich language having a high degree of word inflections, derivations and compounding. This leads to high out-of-vocabulary (OOV) rates and poor language model (LM) probabilities in the large vocabulary continuous speech recognition (LVCSR) systems. One of the main challenges in the German LVCSR is the recognition of the OOV words. For this purpose, data-driven morphemes are used to provide higher lexical coverage. On the other hand, the probability estimates of a sub-lexical LM could be further improved using feature-rich LMs like maximum entropy (MaxEnt) and class-based LMs. In this work, for a sub-lexical level German LVCSR task, we investigate the use of the multiple morpheme level features as classes for building class-based LMs that are estimated using the state-of-the-art MaxEnt approach. Thus, the benefits of both the MaxEnt LMs and the traditional class-based LMs are effectively combined. Furthermore, we experiment the use of Maximum a-posteriori adaptation over the MaxEnt class-based LMs. We show consistent reductions in both the OOV recognition error rate and the word error rate (WER) on a German LVCSR task from the Quaero project, compared to the traditional class-based and theN -gram morpheme based LM. Index Terms: open-vocabulary, German LVCSR, features, maximum entropy, class-based M. Ali Basha Shaik, Amr El-Desoky Mousa, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 4 |
| 2013 | Training log-linear acoustic models in higher-order polynomial feature space for speech recognitionabstractThe use of higher-order polynomial acoustic features can improve the performance of automatic speech recognition.However, the dimensionality of the polynomial representation can be prohibitively large, making the training of acoustic models using polynomial features for large vocabulary ASR systems infeasible.This paper presents an iterative polynomial training framework for acoustic modeling, which recursively expands the current acoustic features into their second-order polynomial feature space.In each recursion the dimensionality is reduced by a linear projection, such that increasingly higher order polynomial information is incorporated while keeping the dimensionality of the acoustic models constant.Experimental results obtained for a large-vocabulary continuous speech recognition task show that the proposed method outperforms conventional mixture models. Muhammad Ali Tahir, Heyun Huang, Ralf Schlüter, Hermann Ney, Louis ten Bosch, Bert Cranen, Lou Boves |
INTERSPEECH | 4 |
| 2013 | Multilingual hierarchical MRASTA features for ASRabstractRecently, a multilingual Multi Layer Perceptron (MLP) training method was introduced without having to explicitly map the phonetic units of multiple languages to a common set.This paper further investigates this method using bottleneck (BN) tandem connectionist acoustic modeling for four high-resourced languages -English, French, German, and Polish.Aiming at the improvement of already existing high performing automatic speech recognition (ASR) systems, the multilingual training of the BN-MLP is extended from short-term to hierarchical longterm (multi-resolutional RASTA) feature extraction.Furthermore, deeper structures and context-dependent target labels are also examined.We experimentally demonstrate that a single state-of-the-art BN feature set can be trained for multiple languages, which is superior to the monolingual feature set, and results in significant gains in all the four languages.Studying the scalability of the multilingual BN features, a similar gain is observed in small (50 hours) and in larger scale (300 hours) ASR experiments regardless of the distribution of the data amount between the languages.Using deeper structures, context-dependent targets, and speaker adaptation, the multilingual BN reduces the word error rates by 3-7% relative over the target language BN features and 25-30% over the conventional MFCC system. Zoltán Tüske, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 3 |
| 2013 | Novel tight classification error bounds under mismatch conditions based on f-DivergenceabstractBy default, statistical classification/multiple hypothesis testing is faced with the model mismatch introduced by replacing the true distributions in Bayes decision rule by model distributions estimated on training samples. Although a large number of statistical measures exist w.r.t. to the mismatch introduced, these works rarely relate to the mismatch in accuracy, i.e. the difference between model error and Bayes error. In this work, the accuracy mismatch between the ideal Bayes decision rule/Bayes test and a mismatched decision rule in statistical classification/multiple hypothesis testing is investigated explicitly. A proof of a novel generalized tight statistical bound on the accuracy mismatch is presented. This result is compared to existing statistical bounds related to the total variational distance that can be extended to bounds of the accuracy mismatch. The analytic results are supported by distribution simulations. Ralf Schlüter, Markus Nußbaum-Thom, Eugen Beck, Tamer Alkhouli, Hermann Ney |
ITW | 5 |
| 2013 | Reverse Word Order Model
Markus Freitag, Minwei Feng, Matthias Huck, Stephan Peitz, Hermann Ney |
MTSummit | 5 |
| 2013 | (Hidden) Conditional Random Fields Using Intermediate Classes for Statistical Machine Translation
Patrick Lehnen, Jan-Thorsten Peter, Jörn Wübker, Stephan Peitz, Hermann Ney |
MTSummit | 5 |
| 2013 | Phrase Training Based Adaptation for Statistical Machine Translation
Saab Mansour, Hermann Ney |
HLT-NAACL | 2 |
| 2013 | Investigations on an EM-Style Optimization Algorithm for Discriminative Training of HMMsabstractToday's speech recognition systems are based on hidden Markov models (HMMs) with Gaussian mixture models whose parameters are estimated using a discriminative training criterion such as Maximum Mutual Information (MMI) or Minimum Phone Error (MPE). Currently, the optimization is almost always done with (empirical variants of) Extended Baum-Welch (EBW). This type of optimization requires sophisticated update schemes for the step sizes and a considerable amount of parameter tuning, and only little is known about its convergence behavior. In this paper, we derive an EM-style algorithm for discriminative training of HMMs. Like Expectation-Maximization (EM) for the generative training of HMMs, the proposed algorithm improves the training criterion on each iteration, converges to a local optimum, and is completely parameter-free. We investigate the feasibility of the proposed EM-style algorithm for discriminative training of two tasks, namely grapheme-to-phoneme conversion and spoken digit string recognition. Georg Heigold, Hermann Ney, Ralf Schlüter |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2013 | Lexical Prefix Tree and WFST: A Comparison of Two Dynamic Search Concepts for LVCSRabstractDynamic network decoders have the advantage of significantly lower memory consumption compared to static network decoders, especially when huge vocabularies and complex language models are required. This paper compares the properties of two well-known search strategies for dynamic network decoding, namely history conditioned lexical tree search and weighted finite-state transducer-based search using on-the-fly transducer composition. The two search strategies share many common principles like the use of dynamic programming, beam search, and many more. We point out the similarities of both approaches and investigate the implications of their differing features, both formally and experimentally, with a focus on implementation independent properties. Therefore, experimental results are obtained with a single decoder by representing the history conditioned lexical tree search strategy in the transducer framework. The properties analyzed cover structure and size of the search space, differences in hypotheses recombination, language model look-ahead techniques, and lattice generation. David Rybach, Hermann Ney, Ralf Schlüter |
IEEE Trans. Speech Audio Process. | 2 |
| 2012 | Deciphering Foreign Language by Combining Language Models and Context Vectors
Malte Nuhn, Arne Mauser, Hermann Ney |
ACL (1) | 3 |
| 2012 | Semantic Cohesion Model for Phrase-Based SMT
Minwei Feng, Weiwei Sun 0007, Hermann Ney |
COLING | 3 |
| 2012 | Discriminative Reordering Extensions for Hierarchical Phrase-Based Machine Translation
Matthias Huck, Stephan Peitz, Markus Freitag, Hermann Ney |
EAMT | 4 |
| 2012 | Conditional leaving-one-out and cross-validation for discount estimation in Kneser-Ney-like extensionsabstractThe smoothing of n-gram models is a core technique in language modelling (LM). Modified Kneser-Ney (mKN) ranges among one of the best smoothing techniques. This technique discounts a fixed quantity from the observed counts in order to approximate the Turing-Good (TG) counts. Despite the TG counts optimise the leaving-one-out (L1O) criterion, the discounting parameters introduced in mKN do not. Moreover, the approximation to the TG counts for large counts is heavily simplified. In this work, both ideas are addressed: the estimation of the discounting parameters by L1O and better functional forms to approximate larger TG counts. The L1O performance is compared with cross-validation (CV) and mKN baseline in two large vocabulary tasks. Jesús Andrés-Ferrer, Martin Sundermeyer, Hermann Ney |
ICASSP | 3 |
| 2012 | Overview of large scale optimization for discriminative training in speech recognitionabstractOver the past few decades, a variety of specialized approaches have been proposed to solve large problems in speech recognition. Conventional optimization techniques have not been widely applied, because the problems do not readily admit an objective for evaluating a given set of parameters and because of the large number of parameters. This situation is changing, due to recent developments in algorithmic optimization. In this paper, we review the specialized algorithms, including methods derived from the extended Baum-Welch (EBW) approach, Rprop, and GIS. We discuss optimization frameworks that could also potentially be applied, and outline some connections between the optimization methods and existing specialized methods. Dimitri Kanevsky, Georg Heigold, Stephen J. Wright 0001, Hermann Ney |
ICASSP | 4 |
| 2012 | Basis vector orthogonalization for an improved kernel gradient matching pursuit methodabstractWith the aim of achieving a computationally efficient optimization of kernel-based probabilistic models for various problems, such as sequential pattern recognition, we have already developed the kernel gradient matching pursuit method as an approximation technique for kernel-based classification. The conventional kernel gradient matching pursuit method approximates the optimal parameter vector by using a linear combination of a small number of basis vectors. In this paper, we propose an improved kernel gradient matching pursuit method that introduces orthogonality constraints to the obtained basis vector set. We verified the efficiency of the proposed method by conducting recognition experiments based on handwritten image datasets and speech datasets. We realized a scalable kernel optimization that incorporated various models, handled very high-dimensional features (>;100 K features), and enabled the use of large scale datasets (>; 10 M samples). Yotaro Kubo, Shinji Watanabe 0001, Atsushi Nakamura, Simon Wiesler, Ralf Schlüter, Hermann Ney |
ICASSP | 6 |
| 2012 | Investigations on the use of morpheme level features in Language Models for Arabic LVCSRabstractA major challenge for Arabic Large Vocabulary Continuous Speech Recognition (LVCSR) is the rich morphology of Arabic, which leads to high Out-of-vocabulary (OOV) rates, and poor Language Model (LM) probabilities. In such cases, the use of morphemes rather than full-words is considered a better choice for LMs. Thereby, higher lexical coverage and less LM perplexities are achieved. On the other side, an effective way to increase the robustness of LMs is to incorporate features of words into LMs. In this paper, we investigate the use of features derived for morphemes rather than words. Thus, we combine the benefits of both morpheme level and feature rich modeling. We compare the performance of stream-based, class-based and Factored LMs (FLMs) estimated over sequences of morphemes and their features for performing Arabic LVCSR. A relative reduction of 3.9% in Word Error Rate (WER) is achieved compared to a word-based system. Amr El-Desoky Mousa, Ralf Schlüter, Hermann Ney |
ICASSP | 3 |
| 2012 | Joining advantages of word-conditioned and token-passing decodingabstractWe compare the families of token-passing and word conditioned decoders, and derive a more efficient dynamic decoder. The advantage of the word conditioned approach is a trivially simple hypothesis recombination, while the advantage of the token-passing approach is a straight-forward minimization of the search network with compressed word tails. We derive a dynamic decoder which joins the advantages of both decoding architectures by minimizing the search network of a word conditioned decoder. We describe the decoder and analyze its efficiency regarding acoustic look-ahead, network minimization, and facilitated pruning methods. Finally, we compare the new decoder with a WFST based decoder extended by acoustic look-ahead. David Nolden, David Rybach, Ralf Schlüter, Hermann Ney |
ICASSP | 4 |
| 2012 | Extended search space pruning in LVCSRabstractWe compare the most important pruning methods which are common in different LVCSR decoding architectures and lead them back to a theoretical motivation. Based on this motivation, we propose a new pruning method which fades the word end pruning over a large part of the search network. We analyze the methods regarding their relationship between search-space and word error rate, and regarding their mutual dependence. We show that the different pruning methods are mutually dependent and difficult to combine, and that our new pruning method is the most effective method regarding both the search space and runtime efficiency. David Nolden, Ralf Schlüter, Hermann Ney |
ICASSP | 3 |
| 2012 | Performance analysis of Neural Networks in combination with n-gram language modelsabstractNeural Network language models (NNLMs) have recently become an important complement to conventional n-gram language models (LMs) in speech-to-text systems. However, little is known about the behavior of NNLMs. The analysis presented in this paper aims to understand which types of events are better modeled by NNLMs as compared to n-gram LMs, in what cases improvements are most substantial and why this is the case. Such an analysis is important to take further benefit from NNLMs used in combination with conventional n-gram models. The analysis is carried out for different types of neural network (feed-forward and recurrent) LMs. The results showing for which type of events NNLMs provide better probability estimates are validated on two setups that are different in their size and the degree of data homogeneity. Ilya Oparin, Martin Sundermeyer, Hermann Ney, Jean-Luc Gauvain |
ICASSP | 3 |
| 2012 | Silence is golden: Modeling non-speech events in WFST-based dynamic network decodersabstractModels for silence are a fundamental part of continuous speech recognition systems. Depending on application requirements, audio data segmentation, and availability of detailed training data annotations, it may be necessary or beneficial to differentiate between other non-speech events, for example breath and background noise. The integration of multiple non-speech models in a WFST-based dynamic network decoder is not straightforward, because these models do not perfectly fit in the transducer framework. This paper describes several options for the transducer construction with multiple non-speech models, shows their considerable different characteristics in memory and runtime efficiency, and analyzes the impact on the recognition performance. David Rybach, Ralf Schlüter, Hermann Ney |
ICASSP | 3 |
| 2012 | Comparison and combination of different CRBE based MLP features for LVCSRabstractMulti Layer Perceptron (MLP) features extracted from different types of critical band energies (CRBE) - derived from MFCC, GT, and PLP pipeline - are compared on French broadcast news and conversational speech recognition task. Though the MLP structure is kept fixed, ROVER combination of different CRBE based systems leads to 4% relative improvement. Furthermore, aiming at the combination of state-of-the-art features based on various signal analysis methods into one single stream, posterior feature space based combination technique is proposed. The speaker normalized features originated from different CRBEs are merged after additional MLP training by Dempster-Shafer rule. The performance of these posterior features unifying the different CRBE based features is superior to the best single CRBE based posterior features by 6% relative. Further results reveal that the concatenated cepstral and unified posterior features perform nearly as well as the ROVER combination of the different CRBE based systems. Zoltán Tüske, Ralf Schlüter, Hermann Ney |
ICASSP | 3 |
| 2012 | Comparison of Bernoulli and Gaussian HMMs Using a Vertical Repositioning Technique for Off-Line Handwriting RecognitionabstractIn this paper a vertical repositioning method based on the center of gravity is investigated for handwriting recognition systems and evaluated on databases containing Arabic and French handwriting. Experiments show that vertical distortion in images has a large impact on the performance of HMM based handwriting recognition systems. Recently good results were obtained with Bernoulli HMMs (BHMMs) using a preprocessing with vertical repositioning of binarized images. In order to isolate the effect of the preprocessing from the BHMM model, experiments were conducted with Gaussian HMMs and the LSTM-RNN tandem HMM approach with relative improvements of 33% WER on the Arabic and up to 62% on the French database. Patrick Doetsch, Mahdi Hamdani, Hermann Ney, Adrià Giménez, Jesús Andrés-Ferrer, Alfons Juan-Císcar |
ICFHR | 3 |
| 2012 | Moment-Based Image Normalization for Handwritten Text RecognitionabstractIn this paper, we extend the concept of moment-based normalization of images from digit recognition to the recognition of handwritten text. Image moments provide robust estimates for text characteristics such as size and position of words within an image. For handwriting recognition the normalization procedure is applied to image slices independently. Additionally, a novel moment-based algorithm for line-thickness normalization is presented. The proposed normalization methods are evaluated on the RIMES database of French handwriting and the IAM database of English handwriting. For RIMES we achieve an improvement from 16.7% word error rate to 13.4% and for IAM from 46.6% to 37.3%. Michal Kozielski, Jens Forster, Hermann Ney |
ICFHR | 3 |
| 2012 | Analysis of Preprocessing Techniques for Latin Handwriting RecognitionabstractIn this work we analyze the contribution of preprocessing steps for Latin handwriting recognition. A preprocessing pipeline based on geometric heuristics and image statistics is used. This pipeline is applied to French and English handwriting recognition in an HMM based framework. Results show that preprocessing improves recognition performance for the two tasks. The Maximum Likelihood (ML)-trained HMM system reaches a competitive WER of 16.7% and outperforms many sophisticated systems for the French handwriting recognition task. The results for English handwriting are comparable to other ML-trained HMM recognizers. Using MLP preprocessing a WER of 35.3% is achieved. Hendrik Pesch, Mahdi Hamdani, Jens Forster, Hermann Ney |
ICFHR | 4 |
| 2012 | Hidden Conditional Random Fields with M-to-N Alignments for Grapheme-to-Phoneme ConversionabstractConditional Random Fields have been successfully applied to a number of NLP tasks like concept tagging, named entity tagging, or grapheme-to-phoneme conversion.When no alignment between source and target side is provided with the training data, it is challenging to build a CRF system with state-of-the-art performance.In this work, we present an approach incorporating an Mto-N alignment as a hidden variable within a transducerbased implementation of CRFs.Including integrated estimation of transition penalties, it was possible to train a state-of-the-art hidden CRF system in reasonable time for an English grapheme-to-phoneme conversion task without using an external model to provide the alignment. Patrick Lehnen, Stefan Hahn, Vlad-Andrei Guta, Hermann Ney |
INTERSPEECH | 4 |
| 2012 | Morpheme Level Feature-based Language Models for German LVCSRabstractOne of the challenges for Large Vocabulary Continuous Speech Recognition (LVCSR) of German is its complex morphology and high level of compounding.It leads to high Out-of-vocabulary (OOV) rates, and poor Language Model (LM) probabilities.In such cases, building LMs on morpheme level can be considered a better choice.Thereby, higher lexical coverage and lower LM perplexities are achieved.On the other side, a successful approach to improve the LM probability estimation is to incorporate features of words using feature-based LMs.In this paper, we use features derived for morphemes as well as words.Thus, we combine the benefits of both morpheme level and feature rich modeling.We compare the performance of stream-based, class-based and factored LMs (FLMs).Relative reductions of around 1.5% in Word Error Rate (WER) are achieved compared to the best previous results obtained using FLMs. Amr El-Desoky Mousa, M. Ali Basha Shaik, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 4 |
| 2012 | Search Space Pruning Based on Anticipated Path Recombination in LVCSRabstractIn this paper we introduce a well-motivated abstract pruning criterion for LVCSR decoders based on the anticipated recombination of HMM state alignment paths.We show that several heuristical pruning methods common in dynamic network decoders are approximations of this pruning criterion.The abstract criterion is too complex to be applied directly in an efficient manner, so we derive approximations which can be applied efficiently.Our new pruning methods allow much more exhaustive pruning of the search space than previous methods.We show that the size of the search space can be reduced by up to 50% at equal precision over the previous state of the art, and the RTF by 20%.The abstract pruning criterion can be considered a guide to derive effective pruning methods for any kind of time synchronous decoder. David Nolden, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 3 |
| 2012 | Posterior-Scaled MPE: Novel Discriminative Training CriteriaabstractWe recently discovered novel discriminative training criteria following a principled approach. In this approach training criteria are developed from error bounds on the global error for pattern classification tasks that depend on non-trivial loss functions. Automatic speech recognition (ASR) is a prominent example for such a task depending on the non-trivial Levenshtein loss. In this context, the posterior-scaled Minimum Phoneme Error (MPE) training criterion, which is the state-of-the-art discriminative training criterion in ASR, was shown to be an approximation to one of the novel criteria. Here, we describe the implementation of the posterior-scaled MPE criterion in a transducer-based framework, and compare this criterion to other discriminative training criteria on an ASR task. This comparison indicates that the posterior-scaled MPE criterion performs better than other discriminative criteria including MPE. Index Terms: error bounds, discriminative training criteria, margin, MPE Markus Nußbaum-Thom, Zoltán Tüske, Georg Heigold, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 5 |
| 2012 | Investigation of Maximum Entropy Hybrid Language Models for Open Vocabulary German and Polish LVCSRabstractFor languages like German and Polish, higher numbers of word inflections lead to high out-of-vocabulary (OOV) rates and high language model (LM) perplexities.Thus, one of the main challenges in large vocabulary continuous speech recognition (LVCSR) is recognizing an open vocabulary.In this paper, we investigate the use of mixed type of sub-word units in the same recognition lexicon.Namely, morphemic or syllabic units combined with pronunciations called graphones, normal graphemic morphemes or syllables, along with full-words.In addition, we investigate the suitability of hybrid mixed-unit Ngrams as features for Maximum Entropy LM along with adaptation.We achieve significant improvements in recognizing OOVs and word error rate reductions for German and Polish LVCSR compared to the conventional full-word approach and state-of-the-art N-gram mixed type hybrid LM. M. Ali Basha Shaik, Amr El-Desoky Mousa, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 4 |
| 2012 | LSTM Neural Networks for Language ModelingabstractNeural networks have become increasingly popular for the task of language modeling.Whereas feed-forward networks only exploit a fixed context length to predict the next word of a sequence, conceptually, standard recurrent neural networks can take into account all of the predecessor words.On the other hand, it is well known that recurrent networks are difficult to train and therefore are unlikely to show the full potential of recurrent models.These problems are addressed by a the Long Short-Term Memory neural network architecture.In this work, we analyze this type of network on an English and a large French language modeling task.Experiments show improvements of about 8 % relative in perplexity over standard recurrent neural network LMs.In addition, we gain considerable improvements in WER on top of a state-of-the-art speech recognition system. Martin Sundermeyer, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 3 |
| 2012 | Simultaneous Discriminative Training and Mixture Splitting of HMMs for Speech RecognitionabstractA method is proposed to incorporate mixture density splitting into the acoustic model discriminative training for speech recognition.The standard method is to obtain a high resolution acoustic model by maximum likelihood training and density splitting, and then improving this model by discriminative training.We choose a log-linear form of acoustic model because for a single Gaussian density per triphone state the log-linear MMI optimization is a convex optimization problem, and by further splitting and discriminative training of this model we can get a higher complexity model.Previously it was shown that we achieve large gains in the objective function and corresponding moderate gains in the word error rate on a large vocabulary corpus.This paper incorporates the state of the art minimum phone error training criterion into the framework, and shows that after discriminative splitting, a subsequent log-linear MPE training achieves better results than Gaussian mixture model MPE optimization alone. Muhammad Ali Tahir, Markus Nußbaum-Thom, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 4 |
| 2012 | Context-Dependent MLPs for LVCSR: TANDEM, Hybrid or Both?abstractGaussian Mixture Model (GMM) and Multi Layer Perceptron (MLP) based acoustic models are compared on a French large vocabulary continuous speech recognition (LVCSR) task.In addition to optimizing the output layer size of the MLP, the effect of the deep neural network structure is also investigated.Moreover, using different linear transformations (time derivatives, LDA, CMLLR) on conventional MFCC, the study is also extended to MLP based probabilistic and bottle-neck TANDEM features.Results show that using either the hybrid or bottleneck TANDEM approach leads to similar recognition performance.However, the best performance is achieved when deep MLP acoustic models are trained on concatenated cepstral and context-dependent bottle-neck features.Further experiments reveal the importance of the neighbouring frames in case of MLP based modeling, and that its gain over GMM acoustic models is strongly reduced by more complex features. Zoltán Tüske, Ralf Schlüter, Hermann Ney, Martin Sundermeyer |
INTERSPEECH | 3 |
| 2012 | Accelerated Batch Learning of Convex Log-linear Models for LVCSRabstractThis paper describes a log-linear modeling framework suitable for large-scale speech recognition tasks. We introduce modifications to our training procedure that are required for extending our previous work on log-linear models to larger tasks. We give a detailed description of the training procedure with a focus on aspects that impact computational efficiency. The performance of our approach is evaluated on the English Quaero corpus, a challenging broadcast conversations task. The log-linear model consistently outperforms the maximum likelihood baseline system. Comparable performance to a system with minimum-phone-error training is achieved. Index Terms: acoustic modeling, discriminative models 1. Simon Wiesler, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 3 |
| 2012 | RWTH-PHOENIX-Weather: A Large Vocabulary Sign Language Recognition and Translation Corpus
Jens Forster, Thomas Hoyoux, Oscar Koller, Uwe Zelle, Justus H. Piater, Hermann Ney |
LREC | 7 |
| 2012 | Arabic-Segmentation Combination Strategies for Statistical Machine Translation
Saab Mansour, Hermann Ney |
LREC | 2 |
| 2012 | Insertion and Deletion Models for Statistical Machine Translation
Matthias Huck, Hermann Ney |
HLT-NAACL | 2 |
| 2012 | A comparison of segmentation methods and extended lexicon models for Arabic statistical machine translation
Sasa Hasan, Saab Mansour, Hermann Ney |
Mach. Transl. | 3 |
| 2012 | Analysis, preparation, and optimization of statistical sign language machine translation
Daniel Stein, Hermann Ney |
Mach. Transl. | 3 |
| 2012 | Cardinality pruning and language model heuristics for hierarchical phrase-based translation
David Vilar, Hermann Ney |
Mach. Transl. | 2 |
| 2012 | Jane: an advanced freely available hierarchical machine translation toolkit
David Vilar, Daniel Stein, Matthias Huck, Hermann Ney |
Mach. Transl. | 4 |
| 2012 | Latent Log-Linear Models for Handwritten Digit ClassificationabstractWe present latent log-linear models, an extension of log-linear models incorporating latent variables, and we propose two applications thereof: log-linear mixture models and image deformation-aware log-linear models. The resulting models are fully discriminative, can be trained efficiently, and the model complexity can be controlled. Log-linear mixture models offer additional flexibility within the log-linear modeling framework. Unlike previous approaches, the image deformation-aware model directly considers image deformations and allows for a discriminative training of the deformation parameters. Both are trained using alternating optimization. For certain variants, convergence to a stationary point is guaranteed and, in practice, even variants without this guarantee converge and find models that perform well. We tune the methods on the USPS data set and evaluate on the MNIST data set, demonstrating the generalization capabilities of our proposed models. Our models, although using significantly fewer parameters, are able to obtain competitive results with models proposed in the literature. Thomas Deselaers, Tobias Gass, Georg Heigold, Hermann Ney |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2012 | Does the Cost Function Matter in Bayes Decision Rule?abstractIn many tasks in pattern recognition, such as automatic speech recognition (ASR), optical character recognition (OCR), part-of-speech (POS) tagging, and other string recognition tasks, we are faced with a well-known inconsistency: The Bayes decision rule is usually used to minimize string (symbol sequence) error, whereas, in practice, we want to minimize symbol (word, character, tag, etc.) error. When comparing different recognition systems, we do indeed use symbol error rate as an evaluation measure. The topic of this work is to analyze the relation between string (i.e., 0-1) and symbol error (i.e., metric, integer valued) cost functions in the Bayes decision rule, for which fundamental analytic results are derived. Simple conditions are derived for which the Bayes decision rule with integer-valued metric cost function and with 0-1 cost gives the same decisions or leads to classes with limited cost. The corresponding conditions can be tested with complexity linear in the number of classes. The results obtained do not make any assumption w.r.t. the structure of the underlying distributions or the classification problem. Nevertheless, the general analytic results are analyzed via simulations of string recognition problems with Levenshtein (edit) distance cost function. The results support earlier findings that considerable improvements are to be expected when initial error rates are high. Ralf Schlüter, Markus Nußbaum-Thom, Hermann Ney |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2012 | Image warping for face recognition: From local optimality towards global optimization
Leonid Pishchulin, Tobias Gass, Philippe Dreuw, Hermann Ney |
Pattern Recognit. | 4 |
| 2012 | WFST Enabled Solutions to ASR Problems: Beyond HMM DecodingabstractDuring the last decade, weighted finite-state transducers (WFSTs) have become popular in speech recognition. While their main field of application remains hidden Markov model (HMM) decoding, the WFST framework is now also seen as a brick in solutions to many other central problems in automatic speech recognition (ASR). These solutions are less known, and this work aims at giving an overview of the applications of WFSTs in large-vocabulary continuous speech recognition (LVCSR) besides HMM decoding: discriminative acoustic model training, Bayes risk decoding, and system combination. The application of the WFST framework has a big practical impact: we show how the framework helps to structure problems, to develop generic solutions, and to delegate complex computations to WFST toolkits. In this paper, we review the literature, discuss existing approaches, and provide new insights into WFST enabled solutions. We also present a novel, purely WFST-based algorithm for computing the exact Bayes risk hypothesis from a lattice with the Levenshtein distance as loss function. We present the problems and their solutions in a unified framework and discuss the advantages and limits of using WFSTs. We do not provide new experimental results, but refer to the existing literature. Our work helps to identify where and how the transducer framework can contribute to a compact and generic solution to LVCSR problems. Björn Hoffmeister, Georg Heigold, David Rybach, Ralf Schlüter, Hermann Ney |
IEEE Trans. Speech Audio Process. | 5 |
| 2011 | Cross-lingual portability of Chinese and english neural network features for French and German LVCSRabstractThis paper investigates neural network (NN) based cross-lingual probabilistic features. Earlier work reports that intra-lingual features consistently outperform the corresponding cross-lingual features. We show that this may not generalize. Depending on the complexity of the NN features, cross-lingual features reduce the resources used for training -the NN has to be trained on one language only- without any loss in performance w.r.t. word error rate (WER). To further investigate this inconsistency concerning intra- vs. cross-lingual neural network features, we analyze the performance of these features w.r.t. the degree of kinship between training and testing language, and the amount of training data used. Whenever the same amount of data is used for NN training, a close relationship between training and testing language is required to achieve similar results. By increasing the training data the relationship becomes less, as well as changing the topology of the NN to the bottle neck structure. Moreover, cross-lingual features trained on English or Chinese improve the best intra-lingual system for German up to 2% relative in WER and up to 3% relative for French and achieve the same improvement as for discriminative training. Moreover, we gain again up to 8% relative in WER by combining intra- and cross-lingual systems. Christian Plahl, Ralf Schlüter, Hermann Ney |
ASRU | 3 |
| 2011 | Discriminative splitting of Gaussian/log-linear mixture HMMs for speech recognitionabstractThis paper presents a method to incorporate mixture density splitting into the acoustic model discriminative log-linear training. The standard method is to obtain a high resolution model by maximum likelihood training and density splitting, and then further training this model discriminatively. For a single Gaussian density per state the log-linear MMI optimization is a global maximum problem, and by further splitting and discriminative training of this model we can get a higher complexity model. The mixture training is not a global maximum problem, nevertheless experimentally we achieve large gains in the objective function and corresponding moderate gains in the word error rate on a large vocabulary corpus. Muhammad Ali Tahir, Ralf Schlüter, Hermann Ney |
ASRU | 3 |
| 2011 | A convergence analysis of log-linear training and its application to speech recognitionabstractLog-linear models are a promising approach for speech recognition. Typically, log-linear models are trained according to a strictly convex criterion. Optimization algorithms are guaranteed to converge to the unique global optimum of the objective function from any initialization. For large-scale applications, considerations in the limit of infinite iterations are not sufficient. We show that log-linear training can be a highly ill-conditioned optimization problem, resulting in extremely slow convergence. Conversely, the optimization problem can be preconditioned by feature transformations. Making use of our convergence analysis, we improve our log-linear speech recognition system and achieve a strong reduction of its training time. In addition, we validate our analysis on a continuous handwriting recognition task. Simon Wiesler, Ralf Schlüter, Hermann Ney |
ASRU | 3 |
| 2011 | Advancements in Arabic-to-English Hierarchical Machine Translation
Matthias Huck, David Vilar, Daniel Stein, Hermann Ney |
EAMT | 4 |
| 2011 | Warp that smile on your face: Optimal and smooth deformations for face recognitionabstractIn this work, we present novel warping algorithms for full 2D pixel-grid deformations for face recognition. Due to high variation in face appearance, face recognition is considered a very difficult task, especially if only a single reference image, for example a mug-shot, per face is available. Usually model-based approaches with additional training data are used to cope with several types of variation occurring in facial imaging. Image warping contrarily yields a distance measure which is invariant with regard to several types of variation. This allows for precise recognition even using only very few reference observations. Due to the computationally complex problem of optimal 2D warping, pseudo-2D warping-based approaches in the past represented strong approximations of the original problem, and were mainly successful on data with low variability or rectified images. We propose a novel 2D warping method which is globally optimal and makes no prior assumtions on the data variability besides two-dimensional smootheness constraints which both avoid local mirroring and gaps and significantly speed up the optimization. Furthermore, we show that occlusion handling is imperative to obtain smooth warpings in a variety of domains. We evaluate our novel algorithm on various well known databases, such as the AR-Face and CMU-PIE database, and provide a detailed comparison to existing warping approaches. We show that by using simple relative 2D constraints, strong local features and a kernel, which is robust w.r.t. occlusions, our computationally complex approaches outperform state-of-the-art results for recognizing faces under varying expressions, occlusions and poses. Most interestingly, we achieve higher accuracy using fewer training instances per class compared to methods learning a model of the 3D shape. Tobias Gass, Leonid Pishchulin, Philippe Dreuw, Hermann Ney |
FG | 4 |
| 2011 | Powerful extensions to CRFS for grapheme to phoneme conversionabstractConditional Random Fields (CRFs) have proven to per form well on natural language processing tasks like name transliteration, concept tagging or grapheme-to-phoneme (g2p) conversion. The aim of this paper is to propose some extension to the state-of-the-art CRF systems for these tasks. Since the number of features can grow rapidly, a method for features selection is very helpful to boost performance. A combination of L1 and L2 regularization (elastic net) has been adopted and implemented within the Rprop optimization algorithm. Usually, dependencies on the target side are limited to bigram dependencies since the computational complexity grows exponentially with the history length. We present a modified CRF decoding where a conventional language model on target side is integrated into the CRF search process. Thus, larger contexts can be taken into account. Besides these two main parts, the already published margin-extension to the CRF training criterion has been adopted. Stefan Hahn, Patrick Lehnen, Hermann Ney |
ICASSP | 3 |
| 2011 | EM-style optimization of hidden conditional random fields for grapheme-to-phoneme conversionabstractWe have recently proposed an EM-style algorithm to optimize log-linear models with hidden variables. In this paper, we use this algorithm to optimize a hidden conditional random field, i.e., a conditional random field with hidden variables. Similar to hidden Markov models, the alignments are the hidden variables in the examples considered. Here, EM-style algorithms are iterative optimization algorithms which are guaranteed to improve the training criterion in each iteration without the need for tuning step sizes, sophisticated update schemes or numerical line optimization (with hardly predictable complexity). This is a rather strong property which conventional gradient-based optimization algorithms do not have. We present experimental results for a grapheme-to-phoneme conversion task and compare the convergence behavior of the EM-style algorithm with L-BFGS and Rprop. Georg Heigold, Stefan Hahn, Patrick Lehnen, Hermann Ney |
ICASSP | 4 |
| 2011 | Subspace pursuit method for kernel-log-linear modelsabstractThis paper presents a novel method for reducing the dimensionality of kernel spaces. Recently, to maintain the convexity of training, log linear models without mixtures have been used as emission probability density functions in hidden Markov models for automatic speech recognition. In that framework, nonlinearly-transformed high-dimensional features are used to achieve the nonlinear classification of the original observation vectors without using mixtures. In this paper, with the goal of using high-dimensional features in kernel spaces, the cutting plane subspace pursuit method proposed for support vector machines is generalized and applied to log-linear models. The experimental results show that the proposed method achieved an efficient approximation of the feature space by using a limited number of basis vectors. Yotaro Kubo, Simon Wiesler, Ralf Schlüter, Hermann Ney, Shinji Watanabe 0001, Atsushi Nakamura, Tetsunori Kobayashi |
ICASSP | 4 |
| 2011 | Incorporating alignments into Conditional Random Fields for grapheme to phoneme conversionabstractConditional Random Fields (CRFs) are a state-of-the-art approach to natural language processing tasks like grapheme-to phoneme (g2p) conversion which is used to produce pronunciations or pronunciation variants for almost all ASR pronunciation lexica. One drawback of CRFs is that for training, an alignment is needed between graphemes and phonemes, usually even 1-to-l. The quality of the g2p result heavily depends on this alignment. Since these alignments are usually not annotated within the corpora, external models have to be used to produce such an alignment in a preprocessing step. In this work, we propose two approaches to integrate the alignment generation directly and efficiently into the CRF training process. Whereas the first approach relies on linear segmentation as starting point, the second approach considers all possible alignments given certain constraints. Both methods have been evaluated on two English g2p tasks, namely NETtalk and Celex, on which state-of-the-art results have been reported in the literature. The proposed approaches lead to results comparable to the state-of-the art. Patrick Lehnen, Stefan Hahn, Andreas Guta, Hermann Ney |
ICASSP | 4 |
| 2011 | Exploiting sparseness of backing-off language models for efficient look-ahead in LVCSRabstractIn this paper, we propose a new method for computing and applying language model look-ahead in a dynamic network decoder, exploiting the sparseness of backing-off n-gram language models. Only partial (sparse) look-ahead tables are computed, with a size that depends on the number of words that have an n-gram score in the language model for a specific context, rather than a constant, vocabulary dependent size. Since high order backing-off language models are inherently sparse, this mechanism reduces the runtime- and memory effort of computing the look-ahead tables by magnitudes. A modified decoding algorithm is required to apply these sparse LM look-ahead tables efficiently. We show that sparse LM look-ahead is much more efficient than the classical method, and that full n-gram look-ahead becomes favorable over lower order look-ahead even when many distinct LM contexts appear during decoding. David Nolden, Hermann Ney, Ralf Schlüter |
ICASSP | 2 |
| 2011 | A comparative analysis of dynamic network decodingabstractThe use of statically compiled search networks for ASR systems using huge vocabularies and complex language models often becomes challenging in terms of memory requirements. Dynamic network decoders introduce additional computations in favor of significantly lower memory consumption. In this paper we investigate the properties of two well-known search strategies for dynamic network decoding, namely history conditioned tree search and WFST-based search using dynamic transducer composition. We analyze the impact of the differences in search graph representation, search space structure, and language model look-ahead techniques. Experiments on an LVCSR task illustrate the influence of the compared properties. David Rybach, Ralf Schlüter, Hermann Ney |
ICASSP | 3 |
| 2011 | Using morpheme and syllable based sub-words for polish LVCSRabstractPolish is a synthetic language with a high morpheme-per-word ratio. It makes use of a high degree of inflection leading to high out-of-vocabulary (OOV) rates, and high Language Model (LM) perplexities. This poses a challenge for Large Vocabulary and Continuous Speech Recognition (LVCSR) systems. Here, the use of morpheme and syllable based units is investigated for building sub-lexical LMs. A different type of sub-lexical units is proposed based on combining morphemic or syllabic units with corresponding pronunciations. Thereby, a set of grapheme-phoneme pairs called graphones are used for building LMs. A relative reduction of 3.5% in Word Error Rate (WER) is obtained with respect to a traditional system based on full-words. M. Ali Basha Shaik, Amr El-Desoky Mousa, Ralf Schlüter, Hermann Ney |
ICASSP | 4 |
| 2011 | The RWTH 2010 Quaero ASR evaluation system for English, French, and GermanabstractRecognizing Broadcast Conversational (BC) speech data is a difficult task, which can be regarded as one of the major challenges beyond the recognition of Broadcast News (BN). Martin Sundermeyer, Markus Nußbaum-Thom, Simon Wiesler, Christian Plahl, Amr El-Desoky Mousa, Stefan Hahn, David Nolden, Ralf Schlüter, Hermann Ney |
ICASSP | 9 |
| 2011 | Feature selection for log-linear acoustic modelsabstractLog-linear acoustic models have been shown to be competitive with Gaussian mixture models in speech recognition. Their high training time can be reduced by feature selection. We compare a simple univariate feature selection algorithm with ReliefF - an efficient multivariate algorithm. An alternative to feature selection is ℓ1-regularized training, which leads to sparse models. We observe that this gives no speedup when sparse features are used, hence feature selection methods are preferable. For dense features, ℓ1-regularization can reduce training and recognition time. We generalize the well known Rprop algorithm for the optimization of ℓ1-regularized functions. Experiments on the Wall Street Journal corpus showed that a large number of sparse features could be discarded without loss of performance. A strong regularization led to slight performance degradations, but can be useful on large tasks, where training the full model is not tractable. Simon Wiesler, Alexander Richard, Yotaro Kubo, Ralf Schlüter, Hermann Ney |
ICASSP | 5 |
| 2011 | Hierarchical hybrid MLP/HMM or rather MLP features for a discriminatively trained Gaussian HMM: A comparison for offline handwriting recognitionabstractWe use neural network based features extracted by a hierarchical multilayer-perceptron (MLP) network either in a hybrid MLP/HMM approach or to discriminatively retrain a Gaussian hidden Markov model (GHMM) system in a tandem approach. MLP networks have been successfully used to model long-term and non-linear features dependencies in automatic speech and optical character recognition. In offline hand writing recognition, MLPs have been mostly used for isolated character and word recognition in hybrid approaches. Here we analyze MLPs within an LVCSR framework for continuous handwriting recognition using discriminative MMI/MPE training. Especially hybrid MLP/HMM and discriminatively retrained MLP-GHMM tandem approaches are evaluated. Significant improvements and competitive results are re ported for a closed-vocabulary task on the IfN/ENIT Arabic handwriting database and for a large-vocabulary task using the IAM English handwriting database. Philippe Dreuw, Patrick Doetsch, Christian Plahl, Hermann Ney |
ICIP | 4 |
| 2011 | N-Grams for Conditional Random Fields or a Failure-Transition(f) Posterior for Acyclic FSTs
Patrick Lehnen, Stefan Hahn, Hermann Ney |
INTERSPEECH | 3 |
| 2011 | Morpheme Based Factored Language Models for German LVCSRabstractGerman is a highly inflectional language, where a large number of words can be generated from the same root.It makes a liberal use of compounding leading to high Out-of-vocabulary (OOV) rates, and poor Language Model (LM) probability estimates.Therefore, the use of morphemes for language modeling is considered a better choice for Large Vocabulary Continuous Speech Recognition (LVCSR) than the full-words.Thereby, better lexical coverage and less LM perplexities are achieved.On the other side, the use of Factored Language Models (FLMs) is considered a successful approach that allows the integration of many information sources to get better LM probability estimates.In this paper, we try a combined methodology for language modeling where both morphological decomposition and factored language modeling are used in one model called morpheme based FLM.Finally, we obtain around 2.5% relative reduction in Word Error Rate (WER) with respect to a traditional full-words system. Amr El-Desoky Mousa, M. Ali Basha Shaik, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 4 |
| 2011 | Acoustic Look-Ahead for More Efficient Decoding in LVCSRabstractIn this paper we propose novel approximations of a generalized acoustic look-ahead to speed up the search process in large vocabulary continuous speech recognition (LVCSR).Unlike earlier methods, we do not employ any phoneme-or syllable level heuristics.First we define and analyze the perfect acoustic look-ahead as a simple pre-evaluation of the original acoustic models into the future.This method is very slow, but reveals the best possible impact on the search space that can be achieved through acoustic look-ahead.In a second step, we derive efficient and simple approximative look-ahead models from the perfect models.We show that the approximative models compare well to the perfect models regarding the search space, and that the approximative models significantly improve the efficiency in comparison to the baseline, without any negative effect on the precision. David Nolden, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 3 |
| 2011 | Compound Word Recombination for German LVCSRabstractCompound words are a difficulty for German speech recognition systems since they cause high out-of-vocabulary and word error rates.State of the art approaches augment the language model by the fragments of compounds in order to increase lexical coverage, lower the perplexity and out-of-vocabulary rate.The fragments are tagged in order to concatenate subsequent equally tagged fragments in the recognition result, but this does not guarantee the recombination of proper words.Such recombination techniques neglect the large vocabulary of the language model training data for recombination although most compounds are covered by it.In this paper, we investigate the use of this vocabulary for the recombination of compound words from the recognition result.The approach is tested on two large vocabulary tasks on top of full-word and fragment based language models and achieves good improvements of 3-7% relative over the baseline compound-sensitive word error rate. Markus Nußbaum-Thom, Amr El-Desoky Mousa, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 4 |
| 2011 | Improved Acoustic Feature Combination for LVCSR by Neural NetworksabstractThis paper investigates the combination of different acoustic features.Several methods to combine these features such as concatenation or LDA are well known.Even though LDA improves the system, feature combination by LDA has been shown to be suboptimal.We introduce a new method based on neural networks.The posterior estimates derived from the NN lead to a significant improvement and achieve a 6% relative better word error rate (WER).Results are also compared to system combination.While system combination has been reported to outperform all other combination techniques, in this work the proposed NN-based combination outperforms system combination.We achieve a 2% relative better WER, resulting in an improvement of 7% relative to the baseline system.In addition to giving better recognition performance w.r.t.WER, NN-based combination reduces both, training and testing complexity.Overall, we use a single set of acoustic models, together with the training of the NN. Christian Plahl, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 3 |
| 2011 | Hybrid Language Models Using Mixed Types of Sub-Lexical Units for Open Vocabulary German LVCSRabstractGerman is a highly inflected language with a large number of words derived from the same root.It makes use of a high degree of word compounding leading to high Out-of-vocabulary (OOV) rates, and Language Model (LM) perplexities.For such languages the use of sub-lexical units for Large Vocabulary Continuous Speech Recognition (LVCSR) becomes a natural choice.In this paper, we investigate the use of mixed types of sub-lexical units in the same recognition lexicon.Namely, morphemic or syllabic units combined with pronunciations called graphones, normal graphemic morphemes or syllables along with full-words.This mixture of units is used for building hybrid LMs suitable for open vocabulary LVCSR where the system operates over an open, constantly changing vocabulary like in broadcast news, political debates, etc.A relative reduction of around 5.0% in Word Error Rate (WER) is obtained compared to a traditional full-words system.Moreover, around 40% of the OOVs are recognized. M. Ali Basha Shaik, Amr El-Desoky Mousa, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 4 |
| 2011 | On the Estimation of Discount Parameters for Language Model SmoothingabstractThe goal of statistical language modeling is to find probability estimates for arbitrary word sequences. To obtain non-zero values, the probability distributions found in the training data need to be smoothed. In the widely-used Kneser-Ney family of smoothing algorithms, this is achieved by absolute discounting. The discount parameters can be computed directly using some approximation formulas minimizing the leaving-one-out log-likelihood of the training data. In this work, we outline several shortcomings of the standard estimators for the discount parameters. We propose an efficient method for computing the discount values on heldout data and analyze the resulting parameter estimates. Experiments on large English and French corpora show consistent improvements in perplexity and word error rate over the baseline method. At the same time, this approach can be used for language model pruning, leading to slightly better results than standard pruning algorithms. Index Terms: language model smoothing, absolute discounting, Kneser-Ney method, language model pruning Martin Sundermeyer, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 3 |
| 2011 | Log-Linear Optimization of Second-Order Polynomial Features with Subsequent Dimension Reduction for Speech RecognitionabstractSecond order polynomial features are useful for speech recognition because they can be used to model class specific covariance even with a pooled covariance acoustic model.Previous experiments with second order features have shown word error rate improvements.However, the improvement comes at the price of a large increase in the number of parameters.This paper investigates the discriminative training of second order features, with a subsequent dimension reduction transform to limit the increase in number of parameters.The acoustic model parameters and the transformation matrix parameters are modeled log-linearly and optimized using maximum mutual information criterion.The advantage of log-linear optimization lies in its ability to robustly combine different kinds of features.Experiments are performed for second order MFCC features on the EPPS large vocabulary task and have resulted in a decrease in word error rate. Muhammad Ali Tahir, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 3 |
| 2011 | A Convergence Analysis of Log-Linear TrainingabstractLog-linear models are widely used probability models for statistical pattern recognition. Typically, log-linear models are trained according to a convex criterion. In recent years, the interest in log-linear models has greatly increased. The optimization of log-linear model parameters is costly and therefore an important topic, in particular for large-scale applications. Different optimization algorithms have been evaluated empirically in many papers. In this work, we analyze the optimization problem analytically and show that the training of log-linear models can be highly ill-conditioned. We verify our findings on two handwriting tasks. By making use of our convergence analysis, we obtain good results on a large-scale continuous handwriting recognition task with a simple and generic approach. Simon Wiesler, Hermann Ney |
NIPS | 2 |
| 2011 | Towards Automatic Error Analysis of Machine Translation OutputabstractEvaluation and error analysis of machine translation output are important but difficult tasks. In this article, we propose a framework for automatic error analysis and classification based on the identification of actual erroneous words using the algorithms for computation of Word Error Rate (WER) and Position-independent word Error Rate (PER), which is just a very first step towards development of automatic evaluation measures that provide more specific information of certain translation problems. The proposed approach enables the use of various types of linguistic knowledge in order to classify translation errors in many different ways. This work focuses on one possible set-up, namely, on five error categories: inflectional errors, errors due to wrong word order, missing words, extra words, and incorrect lexical choices. For each of the categories, we analyze the contribution of various POS classes. We compared the results of automatic error analysis with the results of human error analysis in order to investigate two possible applications: estimating the contribution of each error type in a given translation output in order to identify the main sources of errors for a given translation system, and comparing different translation outputs using the introduced error categories in order to obtain more information about advantages and disadvantages of different systems and possibilites for improvements, as well as about advantages and disadvantages of applied methods for improvements. We used Arabic–English Newswire and Broadcast News and Chinese–English Newswire outputs created in the framework of the GALE project, several Spanish and English European Parliament outputs generated during the TC-Star project, and three German–English outputs generated in the framework of the fourth Machine Translation Workshop. We show that our results correlate very well with the results of a human error analysis, and that all our metrics except the extra words reflect well the differences between different versions of the same translation system as well as the differences between different translation systems. Maja Popovic, Hermann Ney |
Comput. Linguistics | 2 |
| 2011 | Confidence- and margin-based MMI/MPE discriminative training for off-line handwriting recognition
Philippe Dreuw, Georg Heigold, Hermann Ney |
Int. J. Document Anal. Recognit. | 3 |
| 2011 | Comparing Stochastic Approaches to Spoken Language Understanding in Multiple LanguagesabstractOne of the first steps in building a spoken language understanding (SLU) module for dialogue systems is the extraction of flat concepts out of a given word sequence, usually provided by an automatic speech recognition (ASR) system. In this paper, six different modeling approaches are investigated to tackle the task of concept tagging. These methods include classical, well-known generative and discriminative methods like Finite State Transducers (FSTs), Statistical Machine Translation (SMT), Maximum Entropy Markov Models (MEMMs), or Support Vector Machines (SVMs) as well as techniques recently applied to natural language processing such as Conditional Random Fields (CRFs) or Dynamic Bayesian Networks (DBNs). Following a detailed description of the models, experimental and comparative results are presented on three corpora in different languages and with different complexity. The French MEDIA corpus has already been exploited during an evaluation campaign and so a direct comparison with existing benchmarks is possible. Recently collected Italian and Polish corpora are used to test the robustness and portability of the modeling approaches. For all tasks, manual transcriptions as well as ASR inputs are considered. Additionally to single systems, methods for system combination are investigated. The best performing model on all tasks is based on conditional random fields. On the MEDIA evaluation corpus, a concept error rate of 12.6% could be achieved. Here, additionally to attribute names, attribute values have been extracted using a combination of a rule-based and a statistical approach. Applying system combination using weighted ROVER with all six systems, the concept error rate (CER) drops to 12.0%. Stefan Hahn, Marco Dinarelli, Christian Raymond, Fabrice Lefèvre, Patrick Lehnen, Renato De Mori, Alessandro Moschitti, Hermann Ney, Giuseppe Riccardi |
IEEE Trans. Speech Audio Process. | 8 |
| 2011 | Equivalence of Generative and Log-Linear ModelsabstractConventional speech recognition systems are based on hidden Markov models (HMMs) with Gaussian mixture models (GHMMs). Discriminative log-linear models are an alternative modeling approach and have been investigated recently in speech recognition. GHMMs are directed models with constraints, e.g., positivity of variances and normalization of conditional probabilities, while log-linear models do not use such constraints. This paper compares the posterior form of typical generative models related to speech recognition with their log-linear model counterparts. The key result will be the derivation of the equivalence of these two different approaches under weak assumptions. In particular, we study Gaussian mixture models, part-of-speech bigram tagging models, and eventually, the GHMMs. This result unifies two important but fundamentally different modeling paradigms in speech recognition on the functional level. Furthermore, this paper will present comparative experimental results for various speech tasks of different complexity, including a digit string and large-vocabulary continuous speech recognition tasks. Georg Heigold, Hermann Ney, Patrick Lehnen, Tobias Gass, Ralf Schlüter |
IEEE Trans. Speech Audio Process. | 2 |
| 2011 | Lattice-Based ASR-MT Interface for Speech TranslationabstractThe usual approach to improve the interface between automatic speech recognition (ASR) and machine translation (MT) is to use ASR word lattices for translation. In comparison with the previous research along this line, this paper presents an efficient algorithm for lattice-based search in MT. This algorithm utilizes confusion network information to enable phrase-level reordering, and is also able to process general lattices. The proposed search is not constrained to be monotonic; thus, it is able to perform the same type of reordering given lattice input as any statistical phrase-based search algorithm with a single sentence input. Using the concept described in this paper, we are able to significantly improve speech translation results on several small and large vocabulary tasks. The improvements of the MT quality as measured by BLEU are as high as 5% relative. We also show that the proposed lattice-based translation can outperform state-of-the-art translation of confusion networks and has advantages in terms of translation speed. Furthermore, we propose and evaluate a novel approach that shares the benefits of lattice-based translation with those translation systems which are not designed to process word lattices. Evgeny Matusov, Hermann Ney |
IEEE Trans. Speech Audio Process. | 2 |
| 2011 | On the Relationship Between Bayes Risk and Word Error Rate in ASRabstractRecently, a number of (approximate) approaches emerged in speech processing, which try to overcome the known lack of match between symbol level evaluation measures (e.g., word error rate) and the standard string (symbol sequence) cost (e.g., sentence error)-based Bayes decision rule, by using symbol level cost functions for Bayes decision rule. Nevertheless, experiments show that for a majority of test samples both decision rules still give equal decisions, especially at lower error rates. In this paper, analytic evidence for these observations is provided. A set of conditions is presented, for which Bayes decision rule with symbol level and string level cost function leads to the same decisions. Furthermore, the case of word error cost represented by the Levenshtein (edit) distance is investigated, which upon others covers the important case of speech recognition. A Hamming distance-based upper bound to the Levenshtein cost function is discussed. This cost function relates to former, word-posterior based decision rules, and the corresponding efficient decision rule is shown to be strongly related to Bayes decision rule with the Levenshtein cost. The analytic results are verified experimentally, and their quantitative effect is studied by experiments on four different well-known large vocabulary automatic speech recognition tasks. Ralf Schlüter, Markus Nußbaum-Thom, Hermann Ney |
IEEE Trans. Speech Audio Process. | 3 |
| 2010 | Training Phrase Translation Models with Leaving-One-Out
Jörn Wübker, Arne Mauser, Hermann Ney |
ACL | 3 |
| 2010 | Discriminative HMMS, log-linear models, and CRFS: What is the difference?abstractRecently, there have been many papers studying discriminative acoustic modeling techniques like conditional random fields or discriminative training of conventional Gaussian HMMs. This paper will give an overview of the recent work and progress. We will strictly distinguish between the type of acoustic models on the one hand and the training criterion on the other hand. We will address two issues in more detail: the relation between conventional Gaussian HMMs and conditional random fields and the advantages of formulating the training criterion as a convex optimization problem. Experimental results for various speech tasks will be presented to carefully evaluate the different concepts and approaches, including both a digit string and large vocabulary continuous speech recognition tasks. Georg Heigold, Simon Wiesler, Markus Nußbaum-Thom, Patrick Lehnen, Ralf Schlüter, Hermann Ney |
ICASSP | 6 |
| 2010 | Constrained Energy Minimization for Matching-Based Image RecognitionabstractWe propose to use energy minimization in MRFs for matching-based image recognition tasks. To this end, the Tree-Reweighted Message Passing algorithm is modified by geometric constraints and efficiently used by exploiting the guaranteed monotonicity of the lower bound within a nearest-neighbor based classification framework. The constraints allow for a speedup linear to the dimensionality of the reference image, and the lower bound allows to optimally prune the nearest-neighbor search without loosing accuracy, effectively allowing to increase the number of optimization iterations without an effect on runtime. We evaluate our approach on well-known OCR and face recognition tasks and on the latter outperform current state-of-the-art. Tobias Gass, Philippe Dreuw, Hermann Ney |
ICPR | 3 |
| 2010 | Discriminative adaptation for log-linear acoustic modelsabstractLog-linear models have recently been used in acoustic modeling for speech recognition systems.This has been motivated by competitive results compared to systems based on Gaussian models, and a more direct parametrisation of the posterior model.To competitively use log-linear models for speech recognition, important methods, such as speaker adaptation, have to be reformulated in a log-linear framework.In this work, an approach to log-linear affine feature transforms for speaker adaptation is described.Experiments for both supervised and unsupervised adaptation are presented, showing improvements over a maximum likelihood baseline in the form of feature space maximum likelihood linear regression for the case of supervised adaptation. Jonas Lööf, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 3 |
| 2010 | Time conditioned search in automatic speech recognition reconsideredabstractIn this paper we re-investigate the time conditioned search (TCS) method in comparison to the well known word conditioned search (WCS), and analyze its applicability on state-ofthe-art large vocabulary continuous speech recognition tasks. In contrast to current standard approaches, time conditioned search offers theoretical advantages particularly in combination with huge vocabularies and huge language models, but it is difficult to combine with across word modelling, which was proven to be an important technique in automatic speech recognition. Our novel contributions for TCS are a pruning step during the recombination called Early Word End Pruning, an additional recombination technique called Context Recombination, the idea of a Startup Interval to reduce the number of started trees, and a mechanism to combine TCS with across word modelling. We show that, with these techniques, TCS can outperform WCS on current ASR tasks. Index Terms: speech recognition, search, word conditioned, time conditioned David Nolden, Hermann Ney, Ralf Schlüter |
INTERSPEECH | 2 |
| 2010 | The RWTH 2009 quaero ASR evaluation system for English and GermanabstractIn this work, the RWTH automatic speech recognition systems for English and German for the second Quaero evaluation campaign 2009 are presented. The systems are designed to transcribe web data, European parliament plenary sessions and broadcast news data. Another challenge in the 2009 evaluation is that almost no in-domain training data is provided and the test data contains a large variety of speech types. The RWTH participates for the English and German languages with the best results for German and competitive results for the English. Contributing to the enhancements are the systematic use of hierarchical neural network based posterior features, system combination, speaker adaptation, cross speaker adaptation, domain dependent modeling and the usage of additional training data. Markus Nußbaum-Thom, Simon Wiesler, Martin Sundermeyer, Christian Plahl, Stefan Hahn, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 7 |
| 2010 | Hierarchical bottle neck features for LVCSRabstractThis paper investigates the combination of different neural network topologies for probabilistic feature extraction.On one hand, a five-layer neural network used in bottle neck feature extraction allows to obtain arbitrary feature size without dimensionality reduction by transform, independently of the training targets.On the other hand, a hierarchical processing technique is effective and robust over several conditions.Even though the hierarchical and bottle neck processing performs equally well, the combination of both topologies improves the system by 5% relative.Furthermore, the MFCC baseline system is improved by up to 20% relative.This behaviour could be confirmed on two different tasks.In addition, we analyse the influence of multi-resolution RASTA filtering and long-term spectral features as input for the neural network feature extraction. Christian Plahl, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 3 |
| 2010 | Revisiting VTLN using linear transformation on conventional MFCCabstractIn this paper, we revisit the linear transformation for VTLN on conventional MFCC proposed by Sanand et al. in [1], using the idea of band-limited interpolation.The filter-bank is modified to include half-filters at zero and nyquist frequencies, as the full symmetric spectrum is required for performing bandlimited interpolation.In this paper, we show that the filter-bank with half-filters does not affect the recognition performance on clean speech (also shown in [1]), but does affect the recognition performance on noisy speech.This motivated us to revisit the linear transformation for VTLN in [1] and propose modifications to undo the affect of half-filters during the feature extraction.We show through recognition experiments that the proposed modifications to the linear transformation have comparable performance as the conventional VTLN approach, still enabling us to perform VTLN using a linear transformation on conventional MFCC. Rama Sanand Doddipatla, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 3 |
| 2010 | On the relation of Bayes risk, word error, and word posteriors in ASRabstractIn automatic speech recognition, we are faced with a wellknown inconsistency: Bayes decision rule is usually used to minimize sentence (word sequence) error, whereas in practice we want to minimize word error, which also is the usual evaluation measure.Recently, a number of speech recognition approaches to approximate Bayes decision rule with word error (Levenshtein/edit distance) cost were proposed.Nevertheless, experiments show that the decisions often remain the same and that the effect on the word error rate is limited, especially at low error rates.In this work, further analytic evidence for these observations is provided.A set of conditions is presented, for which Bayes decision rule with sentence and word error cost function leads to the same decisions.Furthermore, the case of word error cost is investigated and related to word posterior probabilities.The analytic results are verified experimentally on several large vocabulary speech recognition tasks. Ralf Schlüter, Markus Nußbaum-Thom, Hermann Ney |
INTERSPEECH | 3 |
| 2010 | Direct observation of pruning errors (DOPE): a search analysis toolabstractThe search for the optimal word sequence can be performed efficiently even in a speech recognizer with a very large vocabulary and complex models. This is achieved using pruning methods with empirically chosen parameters and the willingness to accept a certain amount of pruning errors. Quite unsatisfying though, it is state-of-the-art that such pruning errors are not directly detected but, instead, indirect consequences of them, providing only a rough picture of what happens during search. With the tool Direct Observation of Pruning Errors (DOPE), described in this paper, pruning errors are detected on the state hypothesis level, which is a very fine level of granulation, several orders of magnitude finer than the sentence level. This allows much more exact analyses, including the analysis of pruning methods, or the effects of pruning parameters. 1. Volker Steinbiss, Martin Sundermeyer, Hermann Ney |
INTERSPEECH | 3 |
| 2010 | A discriminative splitting criterion for phonetic decision treesabstractPhonetic decision trees are a key concept in acoustic modeling for large vocabulary continuous speech recognition.Although discriminative training has become a major line of research in speech recognition and all state-of-the-art acoustic models are trained discriminatively, the conventional phonetic decision tree approach still relies on the maximum likelihood principle.In this paper we develop a splitting criterion based on the minimization of the classification error.An improvement of more than 10% relative over a discriminatively trained baseline system on the Wall Street Journal corpus suggests that the proposed approach is promising. Simon Wiesler, Georg Heigold, Markus Nußbaum-Thom, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 5 |
| 2010 | The SignSpeak Project - Bridging the Gap Between Signers and Speakers
Philippe Dreuw, Hermann Ney, Gregorio Martínez Pérez, Onno Crasborn, Justus H. Piater, Jose Miguel Moya, Mark Wheatley |
LREC | 2 |
| 2010 | A Hybrid Morphologically Decomposed Factored Language Models for Arabic LVCSR
Amr El-Desoky Mousa, Ralf Schlüter, Hermann Ney |
HLT-NAACL | 3 |
| 2010 | Evaluation of automatic transcription systems for the judicial domainabstractThis paper describes two different automatic transcription systems developed for judicial application domains for the Polish and Italian languages. The judicial domain requires to cope with several factors which are known to be critical for automatic speech recognition, such as: background noise, reverberation, spontaneous and accented speech, overlapped speech, cross channel effects, etc. The two automatic speech recognition (ASR) systems have been developed independently starting from out-of-domain data and, then, they have been adapted using a certain amount of in-domain audio and text data. The ASR performance have been measured on audio data acquired in the courtrooms of Naples and Wroclaw. The resulting word error rates are around 40%, for Italian, and around between 30% and 50% for Polish. This performance, similar to that reported for other comparable ASR tasks (e.g. meeting transcriptions with distant microphone), suggests that possible applications can address tasks such as indexing and/or information retrieval in multimedia documents recorded during judicial debates. Jonas Lööf, Daniele Falavigna, Ralf Schlüter, Diego Giuliani, Roberto Gretter, Hermann Ney |
SLT | 6 |
| 2010 | Sub-lexical language models for German LVCSRabstractOne of the major difficulties related to German LVCSR is the rich morphology nature of German, leading to high out-of-vocabulary (OOV) rates, and high language model (LM) perplexities. Normally, compound words make up an essential fraction of the German vocabulary. Most compound OOVs are composed of frequent in-vocabulary words. Here, we investigate the use of sub-lexical LMs based on different approaches for word decomposition, namely supervised and unsupervised decomposition, as well as decomposition derived from grapheme-to-phoneme (G2P) conversion. In the later approach, we augment a normal word model with a set of grapheme-phoneme pairs called graphones used to model the OOV words. A novel approach is proposed to select the representative graphone sequences for OOVs based on unsupervised decomposition and word-pronunciation alignment. We obtain relative reductions in word error rate (WER) from 4.2% to 6.5% with respect to a comparable full-words system. Amr El-Desoky Mousa, M. Ali Basha Shaik, Ralf Schlüter, Hermann Ney |
SLT | 4 |
| 2010 | Object classification by fusing SVMs and Gaussian mixtures
Thomas Deselaers, Georg Heigold, Hermann Ney |
Pattern Recognit. | 3 |
| 2009 | Generalized likelihood ratio discriminant analysisabstractLinear Discriminant Analysis (LDA) has been established as an important means for dimension reduction and decorrelation in speech recognition. The major points of criticism of LDA are that it uses an ad hoc and non-discriminative training criterion, and that the estimation is performed in a separate preprocessing step. This paper presents a new discriminative training method for the estimation of (projecting) linear feature transforms. More precisely, the problem is formulated in the loglinear framework, resulting in a convex optimization problem. Experimental results are provided for a digit string recognition task to compare the performance and robustness of the proposed approach (in combination with ML or MMI optimized acoustic models) with conventional LDA. Also, first experiments for a large vocabulary task are presented. Muhammad Ali Tahir, Georg Heigold, Christian Plahl, Ralf Schlüter, Hermann Ney |
ASRU | 5 |
| 2009 | Investigations on features for log-linear acoustic models in continuous speech recognitionabstractHidden Markov Models with Gaussian Mixture Models as emission probabilities (GHMMs) are the underlying structure of all state-of-the-art speech recognition systems. Using Gaussian mixture distributions follows the generative approach where the class-conditional probability is modeled, although for classification only the posterior probability is needed. Though being very successful in related tasks like Natural Language Processing (NLP), in speech recognition direct modeling of posterior probabilities with log-linear models has rarely been used and has not been applied successfully to continuous speech recognition. In this paper we report competitive results for a speech recognizer with a log-linear acoustic model on the Wall Street Journal corpus, a Large Vocabulary Continuous Speech Recognition (LVCSR) task. We trained this model from scratch, i.e. without relying on an existing GHMM system. Previously the use of data dependent sparse features for log-linear models has been proposed. We compare them with polynomial features and show that the combination of polynomial and data dependent sparse features leads to better results. Simon Wiesler, Markus Nußbaum-Thom, Georg Heigold, Ralf Schlüter, Hermann Ney |
ASRU | 5 |
| 2009 | SURF-Face: Face Recognition Under Viewpoint Consistency ConstraintsabstractWe analyze the usage of Speeded Up Robust Features (SURF) as local descriptors for face recognition.The effect of different feature extraction and viewpoint consistency constrained matching approaches are analyzed.Furthermore, a RANSAC based outlier removal for system combination is proposed.The proposed approach allows to match faces under partial occlusions, and even if they are not perfectly aligned or illuminated.Current approaches are sensitive to registration errors and usually rely on a very good initial alignment and illumination of the faces to be recognized.A grid-based and dense extraction of local features in combination with a block-based matching accounting for different viewpoint constraints is proposed, as interest-point based feature extraction approaches for face recognition often fail.The proposed SURF descriptors are compared to SIFT descriptors.Experimental results on the AR-Face and CMU-PIE database using manually aligned faces, unaligned faces, and partially occluded faces show that the proposed approach is robust and can outperform current generic approaches. Philippe Dreuw, Pascal Steingrube, Harald Hanselmann, Hermann Ney |
BMVC | 4 |
| 2009 | Log-Linear Mixtures for Object Class RecognitionabstractWe present log-linear mixture models as a fully discriminative approach to object cate-gory recognition which can, analogously to kernelised models, represent non-linear de-cision boundaries. We show that this model is the discriminative counterpart to Gaussian mixtures and that either one can be transformed into the respective other. However, the proposed model is easier to extend toward fusing multiple cues and numerically more stable to train and to evaluate. Experiments on the PASCAL VOC 2006 data show that the performance of our model compares favourably well to the state-of-the-art despite the model consisting of an order of magnitude fewer parameters, which suggests excellent generalisation capabilities. 1 Tobias Weyand, Thomas Deselaers, Hermann Ney |
BMVC | 3 |
| 2009 | On LM Heuristics for the Cube Growing Algorithm
David Vilar, Hermann Ney |
EAMT | 2 |
| 2009 | Are Unaligned Words Important for Machine Translation?
Evgeny Matusov, Hermann Ney |
EAMT | 3 |
| 2009 | Extending Statistical Machine Translation with Discriminative and Trigger-Based Lexicon Models
Arne Mauser, Sasa Hasan, Hermann Ney |
EMNLP | 3 |
| 2009 | Extensions of absolute discounting (Kneser-Ney method)abstractThe problem of estimating the parameters of an n-gram language model is a typical problem of estimating small probabilities. So far, two methods have been proposed and used to handle this problem: 1. the empirical Bayes method resulting in the Turing-Good estimates. Theses estimates do not have any constraints and tend to be very noisy. 2. discounting models like absolute (or linear) discounting. The discounting models are heavily constrained and typically have only a single free parameter. Both methods can be formulated in a leaving-one-out framework. In this paper, we study methods that lie between these two extremes. We design models with various types of constraints and derive efficient algorithms for estimating the parameters of these models. We propose two novel types of constraints or models: interval constraints and the exact extended Kneser-Ney model. The proposed methods are implemented and applied to language modelling in order to compare the methods in terms of perplexities. The results show that the new constrained methods outperform other unconstrained methods. Jesús Andrés-Ferrer, Hermann Ney |
ICASSP | 2 |
| 2009 | Modified MPE/MMI in a transducer-based frameworkabstractIn this paper we show how common training criteria like for example MPE or MMI can be extended to incorporate a margin term. In addition, a transducer-based training implementation is presented, which covers a large variety of discriminative training criteria for ASR, including the standard MMI, MPE, and MCE criteria, as well as the modifications to these criteria presented here. The modified criteria are directly related with the conventional large margin formulation of SVMs. In the proposed approach, we can take advantage of the generalization guarantees of large margin classifiers while keeping the existing framework for the discriminative training, including the efficient algorithms for conventional MPE or MMI. On the conceptual side, this allows for a direct evaluation of the margin term. Finally, experimental results are presented for different large vocabulary continuous speech recognition tasks (one of which is trained on a very large amount of training data) using these modified criteria. Georg Heigold, Ralf Schlüter, Hermann Ney |
ICASSP | 3 |
| 2009 | Audio segmentation for speech recognition using segment featuresabstractAudio segmentation is an essential preprocessing step in several audio processing applications with a significant impact e.g. on speech recognition performance. We introduce a novel framework which combines the advantages of different well known segmentation methods. An automatically estimated log-linear segment model is used to determine the segmentation of an audio stream in a holistic way by a maximum a posteriori decoding strategy, instead of classifying change points locally. A comparison to other segmentation techniques in terms of speech recognition performance is presented, showing a promising segmentation quality of our approach. David Rybach, Christian Gollan, Ralf Schlüter, Hermann Ney |
ICASSP | 4 |
| 2009 | Confidence-Based Discriminative Training for Model Adaptation in Offline Arabic Handwriting RecognitionabstractWe present a novel confidence-based discriminative training for model adaptation approach for an HMM based Arabic handwriting recognition system to handle different handwriting styles and their variations.Most current approaches are maximum-likelihood trained HMM systems and try to adapt their models to different writing styles using writer adaptive training, unsupervised clustering, or additional writer specific data.Discriminative training based on the maximum mutual information criterion is used to train writer independent handwriting models. For model adaptation during decoding, an unsupervised confidence-based discriminative training on a word and frame level within a two-pass decoding process is proposed. Additionally, the training criterion is extended to incorporate a margin term.The proposed methods are evaluated on the IFN/ENIT Arabic handwriting database, where the proposed novel adaptation approach can decrease the word-error-rate by 33% relative. Philippe Dreuw, Georg Heigold, Hermann Ney |
ICDAR | 3 |
| 2009 | Writer Adaptive Training and Writing Variant Model Refinement for Offline Arabic Handwriting RecognitionabstractWe present a writer adaptive training and writer clustering approach for an HMM based Arabic handwriting recognition system to handle different handwriting styles and their variations. Additionally, a writing variant model refinement for specific writing variants is proposed. Current approaches try to compensate the impact of different writing styles during preprocessing and normalization steps. Writer adaptive training with a CMLLR based feature adaptation is used to train writer dependent models. An unsupervised writer clustering with Bayesian information criterion based stopping condition for a CMLLR based feature adaptation during a two-pass decoding process is used to cluster different handwriting styles of unknown test writers. The proposed methods are evaluated on the IFN/ENIT Arabic handwriting database. Philippe Dreuw, David Rybach, Christian Gollan, Hermann Ney |
ICDAR | 4 |
| 2009 | Investigating the use of morphological decomposition and diacritization for improving Arabic LVCSRabstractOne of the challenges related to large vocabulary Arabic speech recognition is the rich morphology nature of Arabic language which leads to both high out-of-vocabulary (OOV) rates and high language model (LM) perplexities.Another challenge is the absence of the short vowels (diacritics) from the Arabic written transcripts which causes a large difference between spoken and written language and thus a weaker connection between the acoustic and language models.In this work, we try to address these two important challenges by introducing both morphological decomposition and diacritization in Arabic language modeling.Finally, we are able to obtain about 3.7% relative reduction in word error rate (WER) with respect to a comparable non-diacritized full-words system running on our test set. Amr El-Desoky Mousa, Christian Gollan, David Rybach, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 5 |
| 2009 | Optimizing CRFs for SLU tasks in various languages using modified training criteriaabstractIn this paper, we present improvements of our state-of-the-art concept tagger based on conditional random fields.Statistical models have been optimized for three tasks of varying complexity in three languages (French, Italian, and Polish).Modified training criteria have been investigated leading to small improvements.The respective corpora as well as parameter optimization results for all models are presented in detail.A comparison of the selected features between languages as well as a close look at the tuning of the regularization parameter is given.The experimental results show in what level the optimizations of the single systems are portable between languages. Stefan Hahn, Patrick Lehnen, Georg Heigold, Hermann Ney |
INTERSPEECH | 4 |
| 2009 | Investigations on convex optimization using log-linear HMMs for digit string recognitionabstractDiscriminative methods are an important technique to refine the acoustic model in speech recognition.Conventional discriminative training is initialized with some baseline model and the parameters are re-estimated in a separate step.This approach has proven to be successful, but it includes many heuristics, approximations, and parameters to be tuned.This tuning involves much engineering and makes it difficult to reproduce and compare experiments.In contrast to the conventional training, convex optimization techniques provide a sound approach to estimate all model parameters from scratch.Such a straight approach hopefully dispense with additional heuristics, e.g.scaling of posteriors.This paper addresses the question how well this concept using log-linear models carries over to practice.Experimental results are reported for a digit string recognition task, which allows for the investigation of this issue without approximations. Georg Heigold, David Rybach, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 4 |
| 2009 | Log-linear model combination with word-dependent scaling factorsabstractLog-linear model combination is the standard approach in LVCSR to combine several knowledge sources, usually an acoustic and a language model.Instead of using a single scaling factor per knowledge source, we make the scaling factor wordand pronunciation-dependent.In this work, we combine three acoustic models, a pronunciation model, and a language model for a Mandarin BN/BC task.The achieved error rate reduction of 2% relative is small but consistent for two test sets.An analysis of the results shows that the major contribution comes from the improved interdependency of language and acoustic model. Björn Hoffmeister, Ruoying Liang, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 4 |
| 2009 | Bayes risk approximations using time overlap with an application to system combinationabstractThe computation of the Minimum Bayes Risk (MBR) decoding rule for word lattices needs approximations. We investigate a class of approximations where the Levenshtein alignment is approximated under the condition that competing lattice arcs overlap in time. The approximations have their origins in MBR decoding and in discriminative training. We develop modified versions and propose a new, conceptually extremely simple confusion network algorithm. The MBR decoding rule is extended to scope with several lattices, which enables us to apply all the investigated approximations to system combination. All approximations are tested on a Mandarin and on an English LVCSR task for a single system and for system combination. The new methods are competitive in error rate and show some advantages over the standard approaches to MBR decoding. Index Terms: speech recognition, minimum bayes risk, confusion network, system combination, discriminative training Björn Hoffmeister, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 3 |
| 2009 | Large-scale Polish SLUabstractIn this paper, we present state-of-the art concept tagging results on a new corpus for Polish SLU.For this language, it is the first large-scale corpus (~200 different concepts) which has been semantically annotated and will be made publicly available.Conditional Random Fields have proven to lead to best results for string-to-string translation problems.Using this approach, we achieve a concept error rate of 22.6% on an evaluation corpus.To additionally extract attribute values, a combination of a statistical and a rule-based approach is used leading to a CER of 30.2%. Patrick Lehnen, Stefan Hahn, Hermann Ney, Agnieszka Mykowiecka 0001 |
INTERSPEECH | 3 |
| 2009 | Cross-language bootstrapping for unsupervised acoustic model training: rapid development of a Polish speech recognition systemabstractThis paper describes the rapid development of a Polish language speech recognition system. The system development was performed without access to any transcribed acoustic training data. This was achieved through the combined use of cross-language bootstrapping and confidence based unsupervised acoustic model training. A Spanish acoustic model was ported to Polish, through the use of a manually constructed phoneme mapping. This initial model was refined through iterative recognition and retraining of the untranscribed audio data. The system was trained and evaluated on recordings from the European Parliament, and included several state-of-the-art speech recognition techniques in addition to the use of unsupervised model training. Confidence based speaker adaptive training using features space transform adaptation, as well as vocal tract length normalization and maximum likelihood linear regression, was used to refine the acoustic model. Through the combination of the different techniques, good performance was achieved on the domain of parliamentary speeches. Jonas Lööf, Christian Gollan, Hermann Ney |
INTERSPEECH | 3 |
| 2009 | Development of the GALE 2008 Mandarin LVCSR systemabstractThis paper describes the current improvements of the RWTH Mandarin LVCSR system.We introduce vocal tract length normalization for the Gammatone features and present comparable results for Gammatone based feature extraction and classical feature extraction.In order to benefit from the huge amount of data of 1600h available in the GALE project we have trained the acoustic models up to 8M Gaussians.We present detailed character error rates for the different number of Gaussians.Different kinds of systems are developed and a two stage decoding framework is applied, which uses cross-adaptation and a subsequent lattice-based system combination.In addition to various acoustic front-ends, these systems use different kinds of neural network toneme posterior features.We present detailed recognition results of the development cycle and the different acoustic front-ends of the systems.Finally, we compare the ultimate evaluation system to our last years system and can report a 10% relative improvement. Christian Plahl, Björn Hoffmeister, Georg Heigold, Jonas Lööf, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 6 |
| 2009 | The RWTH aachen university open source speech recognition systemabstractWe announce the public availability of the RWTH Aachen University speech recognition toolkit.The toolkit includes state of the art speech recognition technology for acoustic model training and decoding.Speaker adaptation, speaker adaptive training, unsupervised training, a finite state automata library, and an efficient tree search decoder are notable components.Comprehensive documentation, example setups for training and recognition, and a tutorial are provided to support newcomers. David Rybach, Christian Gollan, Georg Heigold, Björn Hoffmeister, Jonas Lööf, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 7 |
| 2009 | Statistical Approaches to Computer-Assisted TranslationabstractCurrent machine translation (MT) systems are still not perfect. In practice, the output from these systems needs to be edited to correct errors. A way of increasing the productivity of the whole translation process (MT plus human work) is to incorporate the human correction activities within the translation process itself, thereby shifting the MT paradigm to that of computer-assisted translation. This model entails an iterative process in which the human translator activity is included in the loop: In each iteration, a prefix of the translation is validated (accepted or amended) by the human and the system computes its best (or n-best) translation suffix hypothesis to complete this prefix. A successful framework for MT is the so-called statistical (or pattern recognition) framework. Interestingly, within this framework, the adaptation of MT systems to the interactive scenario affects mainly the search process, allowing a great reuse of successful techniques and models. In this article, alignment templates, phrase-based models, and stochastic finite-state transducers are used to develop computer-assisted translation systems. These systems were assessed in a European project (TransType2) in two real tasks: The translation of printer manuals; manuals and the translation of the Bulletin of the European Union. In each task, the following three pairs of languages were involved (in both translation directions): English-Spanish, English-German, and English-French. Sergio Barrachina 0001, Oliver Bender, Francisco Casacuberta, Jorge Civera, Elsa Cubel, Shahram Khadivi, Antonio L. Lagarda, Hermann Ney, Jesús Tomás, Enrique Vidal 0001, Juan Miguel Vilar |
Comput. Linguistics | 8 |
| 2009 | Edit distances with block movements and error rate confidence estimates
Gregor Leusch, Hermann Ney |
Mach. Transl. | 2 |
| 2009 | Applications of Statistical Machine Translation Approaches to Spoken Language UnderstandingabstractIn this paper, we investigate two statistical methods for spoken language understanding based on statistical machine translation. The first approach employs the source-channel paradigm, whereas the other uses the maximum entropy framework. Starting with an annotated corpus, we describe the problem of natural language understanding as a translation from a source sentence to a formal language target sentence. We analyze the quality of different alignment models and feature functions and show that the direct maximum entropy approach outperforms the source channel-based method. Furthermore, we investigate how both methods perform if the input sentences contain speech recognition errors. Finally, we investigate a new approach to combine speech recognition and spoken language understanding. For this purpose, we employ minimum error rate training which directly optimizes the final evaluation criterion. By combining all knowledge sources in a log-linear way, we show that we can decrease both the word error rate and the slot error rate. Experiments were carried out on two German inhouse corpora for spoken dialogue systems. Klaus Macherey, Oliver Bender, Hermann Ney |
IEEE Trans. Speech Audio Process. | 3 |
| 2008 | Bayesian Semi-Supervised Chinese Word Segmentation for Statistical Machine Translation
Jia Xu 0004, Jianfeng Gao 0001, Kristina Toutanova, Hermann Ney |
COLING | 4 |
| 2008 | Pan, zoom, scan - Time-coherent, trained automatic video croppingabstractWe present a method to fully automatically fit videos in 16:9 format on 4:3 screens and vice versa. It can be applied to arbitrary aspect ratios and can be used to make videos suitable for mobile viewing devices with small and possibly uncommonly sized displays. The cropping sequence is optimised over time to create smooth transitions and thus leads to an excellent viewing experience. Current televisions have simple and often disturbing methods which either show the centre region of the image, distort the image, or pad it with black borders. The technique presented here can fully automatically find the ldquorightrdquo viewing area for each image in a video sequence. It works in real-time with only very little time-shift. We employ different low-level features and a log-linear model to learn how to find the right area. The method is able to automatically decide whether padding with black borders is necessary or whether all relevant image areas fit on screen by cropping the image. Evaluation is done on ten videos from five different types of content and the baseline methods are clearly outperformed. Thomas Deselaers, Philippe Dreuw, Hermann Ney |
CVPR | 3 |
| 2008 | Triplet Lexicon Models for Statistical Machine Translation
Sasa Hasan, Juri Ganitkevitch, Hermann Ney, Jesús Andrés-Ferrer |
EMNLP | 3 |
| 2008 | Complexity of Finding the BLEU-optimal Hypothesis in a Confusion Network
Gregor Leusch, Evgeny Matusov, Hermann Ney |
EMNLP | 3 |
| 2008 | Efficient approximations to model-based joint tracking and recognition of continuous sign languageabstractWe propose several tracking adaptation approaches to recover from early tracking errors in sign language recognition by optimizing the obtained tracking paths w.r.t. to the hypothesized word sequences of an automatic sign language recognition system. Hand or head tracking is usually only optimized according to a tracking criterion. As a consequence, methods which depend on accurate detection and tracking of body parts lead to recognition errors in gesture and sign language processing. We analyze an integrated tracking and recognition approach addressing these problems and propose approximation approaches over multiple hand hypotheses to ease the time complexity of the integrated approach. Most state-of-the-art systems consider tracking as a preprocessing feature extraction part. Experiments on a publicly available benchmark database show that the proposed methods strongly improve the recognition accuracy of the system. Philippe Dreuw, Jens Forster, Thomas Deselaers, Hermann Ney |
FG | 4 |
| 2008 | A GIS-like training algorithm for log-linear models with hidden variablesabstractConditional random fields (CRFs) are often estimated using an entropy based criterion in combination with generalized iterative scaling (GIS). GIS offers, upon others, the immediate advantages that it is locally convergent, completely parameter free, and guarantees an improvement of the criterion in each step. GIS, however, is limited in two aspects. GIS cannot be applied when the model incorporates hidden variables, and it can only be applied to optimize the maximum mutual information criterion (MMI). Here, we extend the GIS algorithm to resolve these two limitations. The new approach allows for training log-linear models with hidden variables and optimizes discriminative training criteria different from maximum mutual information (MMI), including minimum phone error (MPE). The proposed GIS-like method shares the above-mentioned theoretical properties of GIS. The framework is tested for optical character recognition on the USPS task, and for speech recognition on the Sietill task for continuous digit string recognition. Georg Heigold, Thomas Deselaers, Ralf Schlüter, Hermann Ney |
ICASSP | 4 |
| 2008 | Modified MMI/MPE: a direct evaluation of the margin in speech recognitionabstractIn this paper we show how common speech recognition training criteria such as the Minimum Phone Error criterion or the Maximum Mutual Information criterion can be extended to incorporate a margin term. Different margin-based training algorithms have been proposed to refine existing training algorithms for general machine learning problems. However, for speech recognition, some special problems have to be addressed and all approaches proposed either lack practical applicability or the inclusion of a margin term enforces significant changes to the underlying model, e.g. the optimization algorithm, the loss function, or the parameterization of the model. In our approach, the conventional training criteria are modified to incorporate a margin term. This allows us to do large-margin training in speech recognition using the same efficient algorithms for accumulation and optimization and to use the same software as for conventional discriminative training. We show that the proposed criteria are equivalent to Support Vector Machines with suitable smooth loss functions, approximating the non-smooth hinge loss function or the hard error (e.g. phone error). Experimental results are given for two different tasks: the rather simple digit string recognition task Sietill which severely suffers from overfitting and the large vocabulary European Parliament Plenary Sessions English task which is supposed to be dominated by the risk and the generalization does not seem to be such an issue. Georg Heigold, Thomas Deselaers, Ralf Schlüter, Hermann Ney |
ICML | 4 |
| 2008 | SVMs, Gaussian mixtures, and their generative/discriminative fusionabstractWe present a new technique that employs support vector machines and Gaussian mixture densities to create a generative/discriminative joint classifier. In the past, several approaches to fuse the advantages of generative and discriminative approaches were presented, often leading to improved robustness and recognition accuracy. The presented method directly fuses both approaches, effectively allowing to fully exploit the advantages of both. The fusion of SVMs and GMDs is done by representing SVMs in the framework of GMDs without changing the training and without changing the decision boundary. The new classifier is evaluated on four tasks from the UCI machine learning repository. It is shown that for the relatively rare cases where SVMs have problems, the combined method outperforms both individual ones. Thomas Deselaers, Georg Heigold, Hermann Ney |
ICPR | 3 |
| 2008 | Bag-of-visual-words models for adult image classification and filteringabstractWe present a method to classify images into different categories of pornographic content to create a system for filtering pornographic images from network traffic. Although different systems for this application were presented in the past, most of these systems are based on simple skin colour features and have rather poor performance. Recent advances in the image recognition field in particular for the classification of objects have shown that bag-of-visual-words-approaches are a good method for many image classification problems. The system we present here, is based on this approach, uses a task-specific visual vocabulary and is trained and evaluated on an image database of 8500 images from different categories. It is shown that it clearly outperforms earlier systems on this dataset and further evaluation on two novel web-traffic collections shows the good performance of the proposed system. Thomas Deselaers, Lexi Pimenidis, Hermann Ney |
ICPR | 3 |
| 2008 | Learning weighted distances for relevance feedback in image retrievalabstractWe present a new method for relevance feedback in image retrieval and a scheme to learn weighted distances which can be used in combination with different relevance feedback methods. User feedback is a crucial step in image retrieval to maximise retrieval performance as was shown in recent image retrieval evaluations. Machine learning is expected to be able to learn how to rank images according to users needs. Most image retrieval systems incorporate user feedback using rather heuristic means and only few groups have formally investigated how to maximise the benefit from it using machine learning techniques. We incorporate our distance-learning method into our new relevance feedback scheme and into two different approaches from the literature. The methods are compared on two publicly available databases, one which is purely content-based and one which uses additional textual information. It is shown that the new relevance feedback scheme outperforms the other methods and that all methods benefit from weighted distance learning. Thomas Deselaers, Roberto Paredes, Enrique Vidal 0001, Hermann Ney |
ICPR | 4 |
| 2008 | White-space models for offline Arabic handwriting recognitionabstractWe propose to explicitly model white-spaces for Arabic handwriting recognition within different writing variants. Position-dependent character shapes in Arabic handwriting allow for large white-spaces between characters even within words. Here, a separate character model for white-spaces in combination with a lexicon using different writing variants and character model length adaptation is proposed. Current handwriting recognition systems model the white-spaces implicitly within the character models leading to possibly degraded models, or try to explicitly segment the Arabic words into pieces of Arabic words being prone to segmentation errors. Several white-space modeling approaches are analyzed on the well known IFN/ENIT database and outperform the best reported error rates. Philippe Dreuw, Stephan M. Jonas, Hermann Ney |
ICPR | 3 |
| 2008 | Language model adaptation for a speech to sign language translation system using web frequencies and a MAP frameworkabstractThis paper presents a successful technique for creating a new language model (LM) that adapts the original target LM used by a machine translation (MT) system.This technique is especially useful for situations where there are very scarce resources for training the target side (Spanish Sign Language (LSE) in our case) in order to properly estimate the target LM, the Sign Language Model (SLM), used by the MT system.The technique uses information from the source language, Spanish in our task, and from the phrase-based translation matrix in order to create a new LM, estimated using web frequencies, which adapts the counts of the SLM through the Maximum A Posteriori method (MAP).The corpus consists of common used sentences spoken by an officer when assisting people in applying for, or renewing, the National Identification Document.The proposed technique allows relative reductions of 15.5% on perplexity and 2.7% on WER for translation, which are close to half the maximum performance obtainable when only the LM is optimized. Luis Fernando D'Haro, Rubén San-Segundo-Hernández, Ricardo de Córdoba, Jan Bungeroth, Daniel Stein, Hermann Ney |
INTERSPEECH | 6 |
| 2008 | Towards automatic learning in LVCSR: rapid development of a Persian broadcast transcription systemabstractWe present a new method for automatic learning and refining of pronunciations for large vocabulary continuous speech recognition which starts from a small amount of transcribed data and uses automatic transcription techniques for additional untranscribed speech data.The recognition performance of speech recognition systems usually depends on the available amount and quality of the transcribed training data.The creation of such data is a costly and tedious process and the approach presented here allows training with small amounts of annotated data.The model parameters of a statistical joint-multigram grapheme-to-phoneme converter are iteratively estimated using small amounts of manual and relatively larger amounts of automatic transcriptions and thus the system improves itself in an unsupervised manner.Using the new approach, we create a Persian broadcast transcription system from less than five hours of transcribed speech and 52 hours of untranscribed audio data. Christian Gollan, Hermann Ney |
INTERSPEECH | 2 |
| 2008 | System combination for spoken language understandingabstractOne of the first steps in an SLU system usually is the extraction of flat concepts. Within this paper, we present five methods for concept tagging and give experimental results on the state-of-the-art MEDIA corpus for both, manual transcriptions (REF) and ASR input (ASR). Compared to previous publications, some single systems could be improved and the ASR results are presented for the first time. We could improve the tagging performance of the best known result on this task by approx. 7 % relatively from 16.2 % to 15.0 % CER for REF using light-weight system combination (ROVER). For the ASR task, we achieve improvements by approx. 3 % relatively from 29.8 % to 28.9 % CER. An analysis of the differences in performance on both tasks is also given. Index Terms: Spoken dialogue systems, system combination 1. Stefan Hahn, Patrick Lehnen, Hermann Ney |
INTERSPEECH | 3 |
| 2008 | On the equivalence of Gaussian and log-linear HMMsabstractThe acoustic models of conventional state-of-the-art speech recognition systems use generative Gaussian HMMs. In the past few years, discriminative models like for example Conditional Random Fields (CRFs) have been proposed to refine the acoustic models. CRFs directly model the class posteriors, the quantities of interest in recognition. CRFs are undirected models, and CRFs do not assume local normalization constraints as HMMs do. This paper addresses the issue to what extent such less restricted models add flexiblity to the model compared with the generative counterpart. This work extends our previous work in that it provides the technical details used for showing the equivalence of Gaussian and log-linear HMMs. The correctness of the proposed equivalence transformation for conditional probabilities is demonstrated on a simple concept tagging task. Georg Heigold, Patrick Lehnen, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 4 |
| 2008 | iCNC and iROVER: the limits of improving system combination with classification?abstractWe show how ROVER and confusion network combination (CNC) can be improved with classification. The general idea of improving combination with classification is that each word is assigned to a certain location and at each location a classifier decides which of the provided alternatives is most likely correct. We investigate four variations of this idea and three different classifiers, which are trained on various features derived from ASR lattices. For our experiments, we use highly optimized ROVER and CNC systems as baseline, which already give a relative reduction in WER of more than 20 % for the TC-STAR 2007 English task. With our methods we can further improve the result of the corresponding standard combination method. Index Terms: speech recognition, system combination 1. Björn Hoffmeister, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 3 |
| 2008 | Speaker adaptive training using shift-MLLRabstractIn this paper a novel method for speaker adaptive training (SAT), based on Gaussian mean offset adaptation, so called Shift-MLLR, is presented.The method differs from previous SAT methods, where linear transformations of Gaussian means or features are utilized, in that only an offset vector is used for adaptation, but instead the number of regression classes is increased.This is shown to allow an efficient implementation.Furthermore, the use of word posterior confidence measures for Shift-MLLR is investigated, also in combination with the proposed SAT method.The presented methods are integrated into a state of the art speech recognition system, and performance is contrasted with Shift-MLLR without SAT, as well as with MLLR.Large and consistent improvements in word error rate are observed from the new SAT method, as well as from confidence based Shift-MLLR.The combination of the new speaker adaptive training method with confidence based estimation show consistent improvements. Jonas Lööf, Christian Gollan, Hermann Ney |
INTERSPEECH | 3 |
| 2008 | Spoken language translation systems ************ ASR word lattice translation with exhaustive reordering is possible
Evgeny Matusov, Björn Hoffmeister, Hermann Ney |
INTERSPEECH | 3 |
| 2008 | Recent improvements of the RWTH GALE Mandarin LVCSR systemabstractThis paper describes the current improvements of the RWTH Mandarin LVCSR system. We introduce a new reduced toneme set developed at RWTH. We are using different toneme sets and pronunciation lexica. For the purpose of discriminative training we will show a fast way to transform word lattices between systems using different toneme sets and pronunciation lexica. In addition to various acoustic front-ends, the current systems use different kinds of neural network toneme posterior features. While different kinds of systems are developed, a two stage decoding framework for combining these systems is applied. We show detailed recognition results of the development cycle of the systems. Finally, two methods to integrate tonal features are compared. Christian Plahl, Björn Hoffmeister, Mei-Yuh Hwang, Danju Lu, Georg Heigold, Jonas Lööf, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 8 |
| 2008 | The ATIS Sign Language Corpus
Jan Bungeroth, Daniel Stein, Philippe Dreuw, Hermann Ney, Sara Morrissey, Andy Way, Lynette van Zijl |
LREC | 4 |
| 2008 | Benchmark Databases for Video-Based Automatic Sign Language Recognition
Philippe Dreuw, Carol Neidle, Vassilis Athitsos, Stan Sclaroff, Hermann Ney |
LREC | 5 |
| 2008 | A Comparison of Various Methods for Concept Tagging for Spoken Language Understanding
Stefan Hahn, Patrick Lehnen, Christian Raymond, Hermann Ney |
LREC | 4 |
| 2008 | A Multi-Genre SMT System for Arabic to French
Sasa Hasan, Hermann Ney |
LREC | 2 |
| 2008 | Automatic Evaluation Measures for Statistical Machine Translation System Optimization
Arne Mauser, Sasa Hasan, Hermann Ney |
LREC | 3 |
| 2008 | Features for image retrieval: an experimental comparison
Thomas Deselaers, Daniel Keysers, Hermann Ney |
Inf. Retr. | 3 |
| 2008 | Deformations, patches, and discriminative models for automatic annotation of medical radiographs
Thomas Deselaers, Hermann Ney |
Pattern Recognit. Lett. | 2 |
| 2008 | Joint-sequence models for grapheme-to-phoneme conversion
Maximilian Bisani, Hermann Ney |
Speech Commun. | 2 |
| 2008 | Integration of Speech Recognition and Machine Translation in Computer-Assisted TranslationabstractParallel integration of automatic speech recognition (ASR) models and statistical machine translation (MT) models is an unexplored research area in comparison to the large amount of works done on integrating them in series, i.e., speech-to-speech translation. Parallel integration of these models is possible when we have access to the speech of a target language text and to its corresponding source language text, like a computer-assisted translation system. To our knowledge, only a few methods for integrating ASR models with MT models in parallel have been studied. In this paper, we systematically study a number of different translation models in the context of theN-best list rescoring. As an alternative to theN-best list rescoring, we use ASR word graphs in order to arrive at a tighter integration of ASR and MT models. The experiments are carried out on two tasks: English-to-German with an ASR vocabulary size of 17 K words, and Spanish-to-English with an ASR vocabulary of 58 K words. For the best method, the MT models reduce the ASR word error rate by a relative of 18% and 29% on the 17 K and the 58 K tasks, respectively. Shahram Khadivi, Hermann Ney |
IEEE Trans. Speech Audio Process. | 2 |
| 2008 | System Combination for Machine Translation of Spoken and Written LanguageabstractThis paper describes an approach for computing a consensus translation from the outputs of multiple machine translation (MT) systems. The consensus translation is computed by weighted majority voting on a confusion network, similarly to the well-established ROVER approach of Fiscus for combining speech recognition hypotheses. To create the confusion network, pairwise word alignments of the original MT hypotheses are learned using an enhanced statistical alignment algorithm that explicitly models word reordering. The context of a whole corpus of automatic translations rather than a single sentence is taken into account in order to achieve high alignment quality. The confusion network is rescored with a special language model, and the consensus translation is extracted as the best path. The proposed system combination approach was evaluated in the framework of the TC-STAR speech translation project. Up to six state-of-the-art statistical phrase-based translation systems from different project partners were combined in the experiments. Significant improvements in translation quality from Spanish to English and from English to Spanish in comparison with the best of the individual MT systems were achieved under official evaluation conditions. Evgeny Matusov, Gregor Leusch, Rafael E. Banchs, Nicola Bertoldi, Daniel Déchelotte, Marcello Federico, Muntsin Kolss, Young-Suk Lee 0001, José B. Mariño, Matthias Paulik, Salim Roukos, Holger Schwenk, Hermann Ney |
IEEE Trans. Speech Audio Process. | 13 |
| 2007 | Minimum Bayes Risk Decoding for BLEU
Nicola Ehling, Richard Zens, Hermann Ney |
ACL | 3 |
| 2007 | The RWTH Arabic-to-English spoken language translation systemabstractWe present the RWTH phrase-based statistical machine translation system designed for the translation of Arabic speech into English text. This system was used in the Global Autonomous Language Exploitation (GALE) Go/No-Go Translation Evaluation 2007. Using a two-pass approach, we first generate n-best translation candidates and then rerank these candidates using additional models. We give a short review of the decoder as well as of the models used in both passes. We stress the difficulties of spoken language translation, i.e. how to combine the recognition and translation systems and how to compensate for missing punctuation. In addition, we cover our work on domain adaptation for the applied language models. We present translation results for the official GALE 2006 evaluation set and the GALE 2007 development set. Oliver Bender, Evgeny Matusov, Stefan Hahn, Sasa Hasan, Shahram Khadivi, Hermann Ney |
ASRU | 6 |
| 2007 | Development of the 2007 RWTH Mandarin LVCSR systemabstractThis paper describes the development of the RWTH Mandarin LVCSR system. Different acoustic front-ends together with multiple system cross-adaptation are used in a two stage decoding framework. We describe the system in detail and present systematic recognition results. Especially, we compare a variety of approaches for cross-adapting to multiple systems. During the development we did a comparative study on different methods for integrating tone and phoneme posterior features. Furthermore, we apply lattice based consensus decoding and system combination methods. In these methods, the effect of minimizing character instead of word errors is compared. The final system obtains a character error rate of 17.7% on the GALE 2006 evaluation data. Björn Hoffmeister, Christian Plahl, Peter Fritz, Georg Heigold, Jonas Lööf, Ralf Schlüter, Hermann Ney |
ASRU | 7 |
| 2007 | Advances in Arabic broadcast news transcription at RWTHabstractThis paper describes the RWTH speech recognition system for Arabic. Several design aspects of the system, including cross-adaptation, multiple system design and combination, are analyzed. We summarize the semi-automatic lexicon generation for Arabic using a statistical approach to grapheme-to-phoneme conversion and pronunciation statistics. Furthermore, a novel ASR-based audio segmentation algorithm is presented. Finally, we discuss practical approaches for parallelized acoustic training and memory efficient lattice rescoring. Systematic results are reported on recent GALE evaluation corpora. David Rybach, Stefan Hahn, Christian Gollan, Ralf Schlüter, Hermann Ney |
ASRU | 5 |
| 2007 | A Systematic Comparison of Training Criteria for Statistical Machine Translation
Richard Zens, Sasa Hasan, Hermann Ney |
EMNLP-CoNLL | 3 |
| 2007 | Cross-Site and Intra-Site ASR System Combination: Comparisons on Lattice and 1-Best MethodsabstractWe evaluate system combination techniques for automatic speech recognition using systems from multiple sites who participated in the TC-STAR 2006 evaluation. Both lattice and 1-best combination techniques are tested for cross-site and intra-site tasks. For pairwise combinations the lattice based approaches can outperform 1-best ROVER with confidence scores, but 1-best ROVER results are equal (or even better) when combining three or four systems. Björn Hoffmeister, Dustin Hillard, Stefan Hahn, Ralf Schlüter, Mari Ostendorf, Hermann Ney |
ICASSP (4) | 6 |
| 2007 | Gammatone Features and Feature Combination for Large Vocabulary Speech RecognitionabstractIn this work, an acoustic feature set based on a gammatone filterbank is introduced for large vocabulary speech recognition. The gammatone features presented here lead to competitive results on the EPPS English task, and considerable improvements were obtained by subsequent combination to a number of standard acoustic features, i.e. MFCC, PLP, MF-PLP, and VTLN plus voicedness. Best results were obtained when combining gammatone features to all other features using weighted ROVER, resulting in a relative improvement of about 12% in word error rate compared to the best single feature system. We also found that ROVER gives better results for feature combination than both log-linear model combination and LDA. Ralf Schlüter, Ilja Bezrukov, Hermann Wagner, Hermann Ney |
ICASSP (4) | 4 |
| 2007 | Speech recognition with state-based nearest neighbour classifiersabstractWe present a system that uses nearest neighbour classification on the state level of the hidden Markov model. Common speech recognition systems nowadays use Gaussian mixtures with a very high number of densities. We propose to carry this idea to the extreme, such that each observation is a prototype of its own. This approach is well-known and widely used in other areas of pattern recognition and has some immediate advantages over other classification approaches, but has never been applied to speech recognition. We evaluate the proposed method on the SieTill corpus of continuous digit strings and on the large vocabulary EPPS English task. It is shown that nearest neighbour outperforms conventional systems when training data is sparse. Index Terms: automatic speech recognition, nearest neighbour classification, kernel densities Thomas Deselaers, Georg Heigold, Hermann Ney |
INTERSPEECH | 3 |
| 2007 | Speech recognition techniques for a sign language recognition systemabstractOne of the most significant differences between automatic sign language recognition (ASLR) and automatic speech recognition (ASR) is due to the computer vision problems, whereas the corresponding problems in speech signal processing have been solved due to intensive research in the last 30 years.We present our approach where we start from a large vocabulary speech recognition system to profit from the insights that have been obtained in ASR research.The system developed is able to recognize sentences of continuous sign language independent of the speaker.The features used are obtained from standard video cameras without any special data acquisition devices.In particular, we focus on feature and model combination techniques applied in ASR, and the usage of pronunciation and language models (LM) in sign language.These techniques can be used for all kind of sign language recognition systems, and for many video analysis problems where the temporal context is important, e.g. for action or gesture recognition.On a publicly available benchmark database consisting of 201 sentences and 3 signers, we can achieve a 17% WER. Philippe Dreuw, David Rybach, Thomas Deselaers, Morteza Zahedi, Hermann Ney |
INTERSPEECH | 5 |
| 2007 | An improved method for unsupervised training of LVCSR systemsabstractIn this paper, we introduce an improved method for unsupervised training where the data selection or filtering process is done on state level. We describe in detail the setup of the experiments and introduce the state confidence scores on word and allophone state level for performing the data selection for mixture training on state level. Although we are using a relatively small amount of 180 hours of untranscribed recordings in addition to the available carefully manually transcribed transcriptions of 100 hours, we are able to significantly improve our final speaker adaptive acoustic model. Furthermore, we present promising results by doing system combination using the acoustic models trained on different confidence thresholds. These methods are evaluated on the EPPS corpus starting from the RWTH European English parliamentary speech transcription system. A significant improvement of 7 % relative is achieved using less data for unsupervised training than conventional systems require. Christian Gollan, Stefan Hahn, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 4 |
| 2007 | On the equivalence of Gaussian HMM and Gaussian HMM-like hidden conditional random fieldsabstractIn this work we show that Gaussian HMMs (GHMMs) are equivalent to GHMM-like Hidden Conditional Random Fields (HCRFs). Hence, improvements of HCRFs over GHMMs found in literature are not due to a refined acoustic modeling but rather come from the more robust formulation of the underlying optimization problem or spurious local optima. Conventional GHMMs are usually estimated with a criterion on segment level whereas hybrid approaches are based on a formulation of the criterion on frame level. In contrast to CRFs, these approaches do not provide scores or do not support more than two classes in a natural way. In this work we analyze these two classes of criteria and propose a refined frame based criterion, which is shown to be an approximation of the associated criterion on segment level. Experimental results concerning these issues are reported for the German digit string recognition task Sietill and the large vocabulary English European Parliament Plenary Sessions (EPPS) task. Index Terms: speech recognition, parameter estimation, maximum entropy methods Georg Heigold, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 3 |
| 2007 | The RWTH 2007 TC-STAR evaluation system for european English and SpanishabstractIn this work, the RWTH automatic speech recognition systems developed for the third TC-STAR evaluation campaign 2007 are presented.The RWTH systems make systematic use of internal system combination, combining systems with differences in feature extraction, adaptation methods, and training data used.To take advantage of this, novel feature extraction methods were employed; this year saw the introduction of Gammatone features and MLP based phone posterior features.Further improvements were achieved using unsupervised training, and it is notable that these improvements were achieved using a fairly low amount of automatically transcribed data.Also contributing to the improvements over last year was the switch to MPE training, and the introduction of projecting SAT transforms. Jonas Lööf, Christian Gollan, Stefan Hahn, Georg Heigold, Björn Hoffmeister, Christian Plahl, David Rybach, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 9 |
| 2007 | Efficient estimation of speaker-specific projecting feature transformsabstractThis paper introduces a new, efficient approach for estimating projecting feature transforms for speech recognition.It is based on the MMI criterion, a likelihood ratio criterion motivated by a simplification of the MMI criterion, and is shown to be closely related to HLDA.In comparison to current methods, the new method is faster, making it more suitable for speaker adaptive training, where the number of speakers, and therefore the number of transforms are substantial.The proposed method was integrated into the RWTH parliamentary speeches transcription system.Experimental results are presented using speaker specific projecting transforms, both when used in recognition only and when used for speaker adaptive training, showing consistent improvements.Furthermore, the observed improvements are shown to be additive to the improvement of MLLR.Comparisons to DLT are presented, and results are presented for a new projecting DLT method. Jonas Lööf, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 3 |
| 2007 | Improving speech translation with automatic boundary predictionabstractThis paper investigates the influence of automatic sentence boundary and sub-sentence punctuation prediction on machine translation (MT) of automatically recognized speech.We use prosodic and lexical cues to determine sentence boundaries, and successfully combine two complementary approaches to sentence boundary prediction.We also introduce a new feature for segmentation prediction that directly considers the assumptions of the phrase translation model.In addition, we show how automatically predicted commas can be used to constrain reordering in MT search.We evaluate the presented methods using a state-of-the-art phrase-based statistical MT system on two large vocabulary tasks.We find that careful optimization of the segmentation parameters directly for translation quality improves the translation results in comparison to independent optimization for segmentation quality of the predicted source language sentence boundaries. Evgeny Matusov, Dustin Hillard, Mathew Magimai-Doss, Dilek Hakkani-Tür, Mari Ostendorf, Hermann Ney |
INTERSPEECH | 6 |
| 2007 | Domain dependent statistical machine translation
Jia Xu 0004, Yonggang Deng, Hermann Ney |
MTSummit | 4 |
| 2007 | Combining data-driven MT systems for improved sign language translation
Sara Morrissey, Andy Way, Daniel Stein, Jan Bungeroth, Hermann Ney |
MTSummit | 5 |
| 2007 | Efficient Phrase-Table Representation for Machine Translation with Applications to Online MT and Speech Translation
Richard Zens, Hermann Ney |
HLT-NAACL | 2 |
| 2007 | Word-Level Confidence Estimation for Machine TranslationabstractThis article introduces and evaluates several different word-level confidence measures for machine translation. These measures provide a method for labeling each word in an automatically generated translation as correct or incorrect. All approaches to confidence estimation presented here are based on word posterior probabilities. Different concepts of word posterior probabilities as well as different ways of calculating them will be introduced and compared. They can be divided into two categories: System-based methods that explore knowledge provided by the translation system that generated the translations, and direct methods that are independent of the translation system. The system-based techniques make use of system output, such as word graphs or N-best lists. The word posterior probability is determined by summing the probabilities of the sentences in the translation hypothesis space that contains the target word. The direct confidence measures take other knowledge sources, such as word or phrase lexica, into account. They can be applied to output from nonstatistical machine translation systems as well. Experimental assessment of the different confidence measures on various translation tasks and in several language pairs will be presented. Moreover,the application of confidence measures for rescoring of translation hypotheses will be investigated. Nicola Ueffing, Hermann Ney |
Comput. Linguistics | 2 |
| 2007 | The CLEF 2005 Automatic Medical Image Annotation Task
Thomas Deselaers, Henning Müller, Paul D. Clough, Hermann Ney, Thomas M. Deserno |
Int. J. Comput. Vis. | 4 |
| 2007 | Deformation Models for Image RecognitionabstractWe present the application of different nonlinear image deformation models to the task of image recognition. The deformation models are especially suited for local changes as they often occur in the presence of image object variability. We show that, among the discussed models, there is one approach that combines simplicity of implementation, low-computational complexity, and highly competitive performance across various real-world image recognition tasks. We show experimentally that the model performs very well for four different handwritten digit recognition tasks and for the classification of medical images, thus showing high generalization capacity. In particular, an error rate of 0.54 percent on the MNIST benchmark is achieved, as well as the lowest reported error rate, specifically 12.6 percent, in the 2005 international ImageCLEF evaluation of medical image categorization. Daniel Keysers, Thomas Deselaers, Christian Gollan, Hermann Ney |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2007 | Using multiple acoustic feature sets for speech recognition
András Zolnay, Daniil Kocharov, Ralf Schlüter, Hermann Ney |
Speech Commun. | 4 |
| 2006 | Integration of Speech to Computer-Assisted Translation Using Finite-State Automata
Shahram Khadivi, Richard Zens, Hermann Ney |
ACL | 3 |
| 2006 | Patch-based Object Recognition Using Discriminatively Trained Gaussian MixturesabstractWe present an approach using Gaussian mixture models for part-based object recognition where spatial relationships of the parts are explicitly modeled and parameters of the generative model are tuned discriminatively. These extensions lead to great improvements of the classification accuracy. Furthermore we evaluate several improvements over our baseline system which incrementally improve the obtained results which compare favorable well to other published results for the three Caltech tasks and the PASCAL evaluation 05 tasks. 1 Andre Hegerath, Thomas Deselaers, Hermann Ney |
BMVC | 3 |
| 2006 | Geometric Features for Improving Continuous Appearance-based Sign Language RecognitionabstractIn this paper we present an appearance-based sign language recognition system which uses a weighted combination of different features in the statistical framework of a large vocabulary speech recognition system. The performance of the approach is systematically evaluated and it is shown that a significant improvement can be gained over a baseline system when appropriate features are suitably combined. In particular, the word error rate is improved from 50 % for the baseline system to 30 % for the optimized system. 1 Morteza Zahedi, Philippe Dreuw, David Rybach, Thomas Deselaers, Hermann Ney |
BMVC | 5 |
| 2006 | CDER: Efficient MT Evaluation Using Block Movements
Gregor Leusch, Nicola Ueffing, Hermann Ney |
EACL | 3 |
| 2006 | Computing Consensus Translation for Multiple Machine Translation Systems Using Enhanced Hypothesis Alignment
Evgeny Matusov, Nicola Ueffing, Hermann Ney |
EACL | 3 |
| 2006 | A Flexible Architecture for CAT Applications
Sasa Hasan, Shahram Khadivi, Richard Zens, Hermann Ney |
EAMT | 4 |
| 2006 | Morpho-Syntax Based Statistical Methods for Automatic Sign Language Translation
Daniel Stein, Jan Bungeroth, Hermann Ney |
EAMT | 3 |
| 2006 | Vtln Warping Factor Estimation Using Accumulation of Sufficient StatisticsabstractIn this paper we present an efficient and flexible approach to VTLN warping factor estimation. Due to the equivalence of frequency warping and linear transformation of cepstral coefficients, warping factors can be efficiently estimated by accumulating the sufficient statistics for linear transformation estimation, and searching the constrained space of transformations given by the explicit mapping between warping factors and linear transformation matrices. We show that the positive effect of using a properly normalized optimization criterion for warping factor estimation, which has been previously demonstrated for a signal analysis front-end without a filterbank, carries over to a MFCC front-end, resulting in a net improvement in word error rate Jonas Lööf, Hermann Ney, Srinivasan Umesh |
ICASSP (1) | 2 |
| 2006 | Integrating Speech Recognition and Machine Translation: Where do We Stand?abstractThis paper describes state-of-the-art interfaces between speech recognition and machine translation. We modify two different machine translation systems to effectively process dense speech recognition lattices. In addition, we describe how to fully integrate speech translation with machine translation based on weighted finite-state transducers. With a thorough set of experiments, we show that both the acoustic model scores and the source language model positively and significantly affect the translation quality. We have found consistent improvements on three different corpora compared with translations of single best recognition results Evgeny Matusov, Stephan Kanthak, Hermann Ney |
ICASSP (5) | 3 |
| 2006 | Text-Independent Voice Conversion Based on Unit SelectionabstractSo far, most of the voice conversion training procedures are text-dependent, i.e., they are based on parallel training utterances of source and large speaker. Since several applications (e.g. speech-to-speech translation or dubbing) require text-independent training, over the last two years, training techniques that use non-parallel data were proposed In this paper, we present a new approach that applies unit selection to find corresponding time frames in source and target speech. By means of a subjective experiment it is shown that this technique achieves the same performance as the conventional text-dependent training David Suendermann-Oeft, Harald Höge, Antonio Bonafonte, Hermann Ney, Alan W. Black, Shri Narayanan |
ICASSP (1) | 4 |
| 2006 | Frame based system combination and a comparison with weighted ROVER and CNCabstractIn this paper we present a novel ASR system combination technique able to combine systems producing word graphs of different structure and with different segmentations. The new method is based on the definition of a time frame-wise word error cost function in a minimum Bayes risk framework. In contrast to confusion network combination (CNC), it preserves both the word graph structure and the word boundaries. First experimental results are presented on the European Parliament Plenary Sessions (EPPS) task for European Spanish and British English. The new approach to system combination is compared to both ROVER and CNC. In addition, we also apply datadriven weighting schemes for all system combination approaches addressed in this work. For the experiments presented, a variety of internal systems as well as an additional external system were combined. Index Terms: speech recognition, system combination, word posteriors. 1. Björn Hoffmeister, Tobias Klein, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 4 |
| 2006 | The 2006 RWTH parliamentary speeches transcription systemabstractIn this work, investigations in the course of the developement of RWTH automatic speech recognition systems developed for the second TC-STAR evaluation campaign 2006 are presented.The systems were designed to transcribe parliamentary speeches taken from the European Parliament Plenary Sessions (EPPS) in European English and Spanish, as well as speeches from the Spanish Parliament.The RWTH systems apply a two pass search strategy with a fourgram one-pass decoder including a fast vocal tract length normalization variant as first pass.The systems further include several adaptation and normalization methods, minimum classification error trained models, and bayes risk minimization.For all relevant individual components contrastive results are presented on the EPPS Spanish and English data, including investigations which did not yet enter the evaluation systems. Jonas Lööf, Maximilian Bisani, Christian Gollan, Georg Heigold, Björn Hoffmeister, Christian Plahl, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 8 |
| 2006 | Feature combination using linear discriminant analysis and its pitfallsabstractIn this paper, Linear Discriminant Analysis (LDA) is investigated with respect to the combination of different acoustic features for automatic speech recognition. It is shown that the combination of acoustic features using LDA does not consistently lead to improvements in word error rate. A detailed analysis of the recognition results on the Verbmobil (VM II) and on the English portion of the European Parliament Plenary Sessions (EPPS) corpus is given. This includes an independent analysis of the effect of the dimension of the input to LDA, the effect of strongly correlated input features, as well as a detailed numerical analysis of the generalized eigenvalue problem underlying LDA. Relative improvements in word error rate of up to 5 % were observed for LDA-based combination of multiple acoustic features. 1. Ralf Schlüter, András Zolnay, Hermann Ney |
INTERSPEECH | 3 |
| 2006 | Text-independent cross-language voice conversionabstractSo far, cross-language voice conversion requires at least one bilingual speaker and parallel speech data to perform the training. This paper shows how these obstacles can be overcome by means of a recently presented text-independent training method based on unit selection. The new method is evaluated in the framework of the European speech-to-speech translation project TC-Star and achieves a performance similar to that of text-dependent intralingual voice conversion. David Suendermann-Oeft, Harald Höge, Antonio Bonafonte, Hermann Ney, Julia Hirschberg |
INTERSPEECH | 4 |
| 2006 | A German Sign Language Corpus of the Domain Weather Report
Jan Bungeroth, Daniel Stein, Philippe Dreuw, Morteza Zahedi, Hermann Ney |
LREC | 5 |
| 2006 | Creating a Large-Scale Arabic to French Statistical MachineTranslation System
Sasa Hasan, Anas El Isbihani, Hermann Ney |
LREC | 3 |
| 2006 | Training a Statistical Machine Translation System without GIZA++
Arne Mauser, Evgeny Matusov, Hermann Ney |
LREC | 3 |
| 2006 | POS-based Word Reorderings for Statistical Machine Translation
Maja Popovic, Hermann Ney |
LREC | 2 |
| 2006 | Error Analysis of Statistical Machine Translation Output
David Vilar, Jia Xu 0004, Luis Fernando D'Haro, Hermann Ney |
LREC | 4 |
| 2006 | Modeling spontaneous speech variability in professional dictation
Hauke Schramm, Xavier L. Aubert, Bart Bakker, Carsten Meyer, Hermann Ney |
Speech Commun. | 5 |
| 2006 | Quantile based histogram equalization for noise robust large vocabulary speech recognitionabstractThe noise robustness of automatic speech recognition systems can be improved by reducing an eventual mismatch between the training and test data distributions during feature extraction. Based on the quantiles of these distributions the parameters of transformation functions can be reliably estimated with small amounts of data. This paper will give a detailed review of quantile equalization applied to the Mel scaled filter bank, including considerations about the application in online systems and improvements through a second transformation step that combines neighboring filter channels. The recognition tests have shown that previous experimental observations on small vocabulary recognition tasks can be confirmed on the larger vocabulary Aurora 4 noisy Wall Street Journal database. The word error rate could be reduced from 45.7% to 25.5% (clean training) and from 19.5% to 17.0% (multicondition training). Florian Hilger, Hermann Ney |
IEEE Trans. Speech Audio Process. | 2 |
| 2005 | Discriminative Training for Object Recognition Using Image PatchesabstractWe present a method for automatically learning discriminative image patches for the recognition of given object classes. The approach applies discriminative training of log-linear models to image patch histograms. We show that it works well on three tasks and performs significantly better than other methods using the same features. For example, the method decides that patches containing an eye are most important for distinguishing face from background images. The recognition performance is very competitive with error rates presented in other publications. In particular, a new best error rate for the Caltech motorbikes data of 1.5% is achieved. Thomas Deselaers, Daniel Keysers, Hermann Ney |
CVPR (2) | 3 |
| 2005 | Comparison of generation strategies for interactive machine translation
Oliver Bender, Sasa Hasan, David Vilar, Richard Zens, Hermann Ney |
EAMT | 5 |
| 2005 | Clustered language models based on regular expressions for SMT
Sasa Hasan, Hermann Ney |
EAMT | 2 |
| 2005 | Efficient statistical machine translation with constrained reordering
Evgeny Matusov, Stephan Kanthak, Hermann Ney |
EAMT | 3 |
| 2005 | Exploiting phrasal lexica and additional morpho-syntactic language resources for statistical machine translation with scarce training data
Maja Popovic, Hermann Ney |
EAMT | 2 |
| 2005 | Application of word-level confidence measures in interactive statistical machine translation
Nicola Ueffing, Hermann Ney |
EAMT | 2 |
| 2005 | Sentence segmentation using IBM word alignment model 1
Jia Xu 0004, Richard Zens, Hermann Ney |
EAMT | 3 |
| 2005 | Cross Domain Automatic Transcription on the TC-STAR EPPS CorpusabstractThis paper describes the ongoing development of the British English European Parliament Plenary Session corpus. This corpus will be part of the speech-to-speech translation evaluation infrastructure of the European TC-STAR project. Furthermore, we present first recognition results on the English speech recordings. The transcription system has been derived from an older speech recognition system built for the North-American broadcast news task. We report on the measures taken for rapid cross-domain porting and present encouraging results. Christian Gollan, Maximilian Bisani, Stephan Kanthak, Ralf Schlüter, Hermann Ney |
ICASSP (1) | 5 |
| 2005 | A Study on Residual Prediction Techniques for Voice ConversionabstractSeveral well-studied voice conversion techniques use line spectral frequencies as features to represent the spectral envelopes of the processed speech frames. In order to return to the time domain, these features are converted to linear predictive coefficients that serve as coefficients of a filter applied to an unknown residual signal. We compare several residual prediction approaches that have already been proposed in the literature dealing with voice conversion. We also present a novel technique that outperforms the others in terms of voice conversion performance and sound quality. David Suendermann-Oeft, Antonio Bonafonte, Hermann Ney, Harald Höge |
ICASSP (1) | 3 |
| 2005 | Acoustic Feature Combination for Robust Speech RecognitionabstractIn this paper, we consider the use of multiple acoustic features of the speech signal for robust speech recognition. We investigate the combination of various auditory based (mel frequency cepstrum coefficients, perceptual linear prediction, etc.) and articulatory based (voicedness) features. Features are combined by linear discriminant analysis and log-linear model combination based techniques. We describe the two feature combination techniques and compare the experimental results. Experiments performed on the large-vocabulary task VerbMobil II (German conversational speech) show that the accuracy of automatic speech recognition systems can be improved by the combination of different acoustic features. András Zolnay, Ralf Schlüter, Hermann Ney |
ICASSP (1) | 3 |
| 2005 | Open vocabulary speech recognition with flat hybrid modelsabstractToday's speech recognition systems are able to recognize arbitrary sentences over a large but finite vocabulary.However, many important speech recognition tasks feature an open, constantly changing vocabulary.(E.g.broadcast news transcription, translation of political debates, etc. Ideally, a system designed for such open vocabulary tasks would be able to recognize arbitrary, even previously unseen words.To some extent this can be achieved by using sub-lexical language models.We demonstrate that, by using a simple flat hybrid model, we can significantly improve a well-optimized state-ofthe-art speech recognition system over a wide range of out-of-vocabulary rates. Maximilian Bisani, Hermann Ney |
INTERSPEECH | 2 |
| 2005 | Automatic text dictation in computer-assisted translationabstractIn this paper, we study the incorporation of statistical machine translation models to automatic speech recognition models in the framework of computer-assisted translation. The system is given a source language text to be translated and it shows the source text to the human translator to translate it orally. The system captures the user speech which is the dictation of the target language sentence. Since the system has simultaneous access to the source language text and the speech signal of the target language text, it is possible to improve the speech recognition accuracy by incorporating the statistical machine translation models. We show that statistical translation models have a high impact on improving the speech recognition results. Using these models, we achieve a relative word error rate reduction of 17%. 1. Shahram Khadivi, András Zolnay, Hermann Ney |
INTERSPEECH | 3 |
| 2005 | Articulatory motivated acoustic features for speech recognitionabstractIn this paper, we consider the use of multiple acoustic features of the speech signal for continuous speech recognition. A novel articulatory motivated acoustic feature is introduced, namely the spectrum derivative feature. The new feature is tested in combination with the standard Mel Frequency Cepstral Coefficients (MFCC) and the voicedness features. Linear Discriminant Analysis is applied to find the optimal combination of different acoustic features. Experiments have been performed on small and large vocabulary tasks. Significant improvements in word error rate have been obtained by combining the MFCC feature with the articulatory motivated voicedness and spectrum derivative features: improvements of up to 25 % on the small-vocabulary task and improvements of up to 4 % on the large-vocabulary task relative to using MFCC alone with the same overall number of parameters in the system. 1. Daniil Kocharov, András Zolnay, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 4 |
| 2005 | Investigations on error minimizing training criteria for discriminative training in automatic speech recognitionabstractDiscriminative training criteria have been shown to consistently outperform maximum likelihood trained speech recognition systems. In this paper we employ the Minimum Classification Error (MCE) criterion to optimize the parameters of the acoustic model of a large scale speech recognition system. The statistics for both the correct and the competing model are solely collected on word lattices without the use of N-best lists. Thus, particularly for long utterances, the number of sentence alternatives taken into account is significantly larger compared to N-best lists. The MCE criterion is embedded in an extended unifying approach for a class of discriminative training criteria which allows for direct comparison of the performance gain obtained with the improvements of other commonly used criteria such as Maximum Mutual Information (MMI) and Minimum Word Error (MWE). Experiments conducted on large vocabulary tasks show a consistent performance gain for MCE over MMI. Moreover, the improvements obtained with MCE turn out to be in the same order of magnitude as the performance gains obtained with the MWE criterion. 1. Wolfgang Macherey, Lars Haferkamp, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 4 |
| 2005 | On the integration of speech recognition and statistical machine translationabstractThis paper focuses on the interface between speech recognition and machine translation in a speech translation system. Based on a thorough theoretical framework, we exploit word lattices of automatic speech recognition hypotheses as input to our translation system which is based on weighted finite-state transducers. We show that acoustic recognition scores of the recognized words in the lattices positively and significantly affect the translation quality. In experiments, we have found consistent improvements on three different corpora compared with translations of single best recognized results. In addition we build and evaluate a fully integrated speech translation model. 1. Evgeny Matusov, Stephan Kanthak, Hermann Ney |
INTERSPEECH | 3 |
| 2005 | Bayes risk minimization using metric loss functionsabstractIn this work, fundamental properties of Bayes decision rule using general loss functions are derived analytically and are verified experimentally for automatic speech recognition. It is shown that, for maximum posterior probabilities larger than 1/2, Bayes decision rule with a metric loss function always decides on the posterior maximizing class independent of the specific choice of (metric) loss function. Also for maximum posterior probabilities less than 1/2, a condition is derived under which the Bayes risk using a general metric loss function is still minimized by the posterior maximizing class. For a speech recognition task with low initial word error rate, it is shown that nearly 2/3 of the test utterances fulfil these conditions and need not be considered for Bayes risk minimization with Levenshtein loss, which reduces the computational complexity of Bayes risk minimization. In addition, bounds for the difference between the Bayes risk for the posterior maximizing class and minimum Bayes risk are derived, which can serve as cost estimates for Bayes risk minimization approaches. Ralf Schlüter, T. Scharrenbach, Volker Steinbiss, Hermann Ney |
INTERSPEECH | 4 |
| 2005 | Evaluation of VTLN-based voice conversion for embedded speech synthesisabstractRecently, we demonstrated that vocal tract length normalization (VTLN) can be applied to voice conversion tasks.In particular, when the conversion algorithm is performed in time domain, this technique is very resource-efficient and, consequently, suitable for embedded applications.In this paper, we use VTLNbased voice conversion as a novel feature of a small footprint speech synthesizer running on mobile devices.The characteristics of this feature are investigated by means of extensive subjective tests. David Suendermann-Oeft, Guntram Strecha, Antonio Bonafonte, Harald Höge, Hermann Ney |
INTERSPEECH | 5 |
| 2005 | Implementing frequency-warping and VTLN through linear transformation of conventional MFCCabstractIn this paper, we show that frequency-warping (including VTLN) can be implemented through linear transformation of conventional MFCC.Unlike the Pitz-Ney [1] continuous domain approach, we directly determine the relation between frequency-warping and the linear-transformation in the discrete-domain.The advantage of such an approach is that it can be applied to any frequency-warping and is not limited to cases where an analytical closed-form solution can be found.The proposed method exploits the bandlimited interpolation idea (in the frequency-domain) to do the necessary frequency-warping and yields exact results as long as the cepstral coefficients are que-frency limited.This idea of quefrencylimitedness shows the importance of the filter-bank smoothing of the spectra which has been ignored in [1,2].Furthermore, unlike [1], since we operate in the discrete domain, we can also apply the usual discrete-cosine transform (i.e.DCT-II) on the logarithm of the filter-bank output to get conventional MFCC features.Therefore, using our proposed method, we can linearly transform conventional MFCC cepstra to do VTLN and we do not require any recomputation of the warped-features.We provide experimental results in support of this approach. Srinivasan Umesh, András Zolnay, Hermann Ney |
INTERSPEECH | 3 |
| 2005 | Statistical Machine Translation of European Parliamentary SpeechesabstractIn this paper we present the ongoing work at RWTH Aachen University for building a speech-to-speech translation system within the TC-Star project. The corpus we work on consists of parliamentary speeches held in the European Plenary Sessions. To our knowledge, this is the first project that focuses on speech-to-speech translation applied to a real-life task. We describe the statistical approach used in the development of our system and analyze its performance under different conditions: dealing with syntactically correct input, dealing with the exact transcription of speech and dealing with the (noisy) output of an automatic speech recognition system. Experimental results show that our system is able to perform adequately in each of these conditions. David Vilar, Evgeny Matusov, Sasa Hasan, Richard Zens, Hermann Ney |
MTSummit | 5 |
| 2005 | Automatic Filtering of Bilingual Corpora for Statistical Machine Translation
Shahram Khadivi, Hermann Ney |
NLDB | 2 |
| 2005 | Vocal Tract Normalization Equals Linear Transformation in Cepstral SpaceabstractVocal tract normalization (VTN) is a widely used speaker normalization technique which reduces the effect of different lengths of the human vocal tract and results in an improved recognition accuracy of automatic speech recognition systems. We show that VTN results in a linear transformation in the cepstral domain, which so far have been considered as independent approaches of speaker normalization. We are now able to compute the Jacobian determinant of the transformation matrix, which allows the normalization of the probability distributions used in speaker-normalization for automatic speech recognition. We show that VTN can be viewed as a special case of Maximum Likelihood Linear Regression (MLLR). Consequently, we can explain previous experimental results that improvements obtained by VTN and subsequent MLLR are not additive in some cases. For three typical warping functions the transformation matrix is calculated analytically and we show that the matrices are diagonal dominant and thus can be approximated by quindiagonal matrices. Michael Pitz, Hermann Ney |
IEEE Trans. Speech Audio Process. | 2 |
| 2005 | Unsupervised training of acoustic models for large vocabulary continuous speech recognitionabstractFor large vocabulary continuous speech recognition systems, the amount of acoustic training data is of crucial importance. In the past, large amounts of speech were thus recorded from various sources and had to be transcribed manually. It is thus desirable to train a recognizer with as little manually transcribed acoustic data as possible. Since untranscribed speech is available in various forms nowadays, the unsupervised training of a speech recognizer on recognized transcriptions is studied in this paper. A low-cost recognizer trained with between one and six h of manually transcribed speech is used to recognize 72 h of untranscribed acoustic data. These transcriptions are then used in combination with a confidence measure to train an improved recognizer. The effect of the confidence measure which is used to detect possible recognition errors is studied systematically. Finally, the unsupervised training is applied iteratively. Starting with only one h of transcribed acoustic data, a recognition system is trained fully automatically. With this iterative training procedure, the word error rates are reduced from 71.3% to 38.3% on the Broadcast News'96 evaluation test set and from 65.6% to 29.3% on the Broadcast News'98 evaluation test set. In comparison with an optimized system trained with the manually generated transcriptions of the complete 72 h training corpus, the word error rates increase by 14.3% relative and 18.6% relative, respectively. Frank Wessel, Hermann Ney |
IEEE Trans. Speech Audio Process. | 2 |
| 2004 | FSA: An Efficient and Flexible C++ Toolkit for Finite State Automata Using On-Demand ComputationabstractIn this paper we present the RWTH FSA toolkit --- an efficient implementation of algorithms for creating and manipulating weighted finite-state automata. The toolkit has been designed using the principle of on-demand computation and offers a large range of widely used algorithms. To prove the superior efficiency of the toolkit, we compare the implementation to that of other publically available toolkits. We also show that on-demand computations help to reduce memory requirements significantly without any loss in speed. To increase its flexibility, the RWTH FSA toolkit supports high-level interfaces to the programming language Python as well as a command-line tool for interactive manipulation of FSAs. Furthermore, we show how to utilize the toolkit to rapidly build a fast and accurate statistical machine translation system. Future extensibility of the toolkit is ensured as it will be publically available as open source software. Stephan Kanthak, Hermann Ney |
ACL | 2 |
| 2004 | Symmetric Word Alignments for Statistical Machine Translation
Evgeny Matusov, Richard Zens, Hermann Ney |
COLING | 3 |
| 2004 | Improving Word Alignment Quality using Morpho-syntactic Information
Hermann Ney, Maja Popovic |
COLING | 1 |
| 2004 | Improved Word Alignment Using a Symmetric Lexicon Model
Richard Zens, Evgeny Matusov, Hermann Ney |
COLING | 3 |
| 2004 | Reordering Constraints for Phrase-Based Statistical Machine Translation
Richard Zens, Hermann Ney, Taro Watanabe, Eiichiro Sumita |
COLING | 2 |
| 2004 | Error Measures and Bayes Decision Rules Revisited with Applications to POS Tagging
Hermann Ney, Maja Popovic, David Suendermann-Oeft |
EMNLP | 1 |
| 2004 | Bootstrap estimates for confidence intervals in ASR performance evaluationabstractThe field of speech recognition has clearly benefited from precisely defined testing conditions and objective performance measures such as word error rate. In the development and evaluation of new methods, the question arises whether the empirically observed difference in performance is due to a genuine advantage of one system over the other, or just an effect of chance. However, many publications still do not concern themselves with the statistical significance of the results reported. We present a bootstrap method for significance analysis which is, at the same time, intuitive, precise and and easy to use. Unlike some methods, we make no (possibly ill-founded) approximations and the results are immediately interpretable in terms of word error rate. Maximilian Bisani, Hermann Ney |
ICASSP (1) | 2 |
| 2004 | Discriminative training with tied covariance matricesabstractDiscriminative training techniques have proved to be a powerful method for improving large vocabulary speech recognition systems based on Gaussian mixture hidden Markov models.Typically, the optimization of discriminative objective functions is done using the extended Baum algorithm.Since for continuous distributions no proof of fast and stable convergence is known up to now, parameter re-estimation depends on setting the iteration constants in the update rules heuristically, ensuring that the new variances are positive definite.In case of density specific variances this leads to a system of quadratic inequalities.However, if tied variances are used, the inequalities become more complicated and often the resulting constants are too large to be appropriate for discriminative training.In this paper we present an alternative approach to setting the iteration constants to alleviate this problem.First experimental results show that the new method leads to improved convergence speed and test set performance. Wolfgang Macherey, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 3 |
| 2004 | Statistical machine translation and its challengesabstractIn addition to speech recognition and syntactic parsing, during the last 10 years, the statistical approach has found widespread use in machine translation of both written language and spoken language. In many comparative evaluations, the statistical approach was found to be competitive or superior to the existing conventional approaches. Since the first statistical approach was proposed at the end of the 80s, many attempts have been made to improve the state of the art. Like other natural language processing tasks, machine translation requires four major components: a decision rule, a set of probability models, a training criterion and an efficient generation of the target sentence. We will consider each of these four components in more detail and point out promising research directions. Hermann Ney |
INTERSPEECH | 1 |
| 2004 | A first step towards text-independent voice conversionabstractSo far, all conventional voice conversion approaches are text-dependent, i.e., they need equivalent training utterances of source and target speaker. Since several recently proposed applications call for renouncing this requirement, in this paper, we present an algorithm which finds corresponding time frames within text-independent training data. The performance of this algorithm is tested by means of a voice conversion framework based on linear transformation of the spectral envelope. Experimental results are reported on a Spanish cross-gender corpus utilizing several objective error measures. Hermann Ney, David Suendermann-Oeft, Antonio Bonafonte, Harald Höge |
INTERSPEECH | 1 |
| 2004 | Towards the Use of Word Stems and Suffixes for Statistical Machine Translation
Maja Popovic, Hermann Ney |
LREC | 2 |
| 2004 | Improvements in Phrase-Based Statistical Machine Translation
Richard Zens, Hermann Ney |
HLT-NAACL | 2 |
| 2004 | Statistical Machine Translation with Scarce Resources Using Morpho-syntactic InformationabstractIn statistical machine translation, correspondences between the words in the source and the target language are learned from parallel corpora, and often little or no linguistic knowledge is used to structure the underlying models. In particular, existing statistical systems for machine translation often treat different inflected forms of the same lemma as if they were independent of one another. The bilingual training data can be better exploited by explicitly taking into account the interdependencies of related inflected forms. We propose the construction of hierarchical lexicon models on the basis of equivalence classes of words. In addition, we introduce sentence-level restructuring transformations which aim at the assimilation of word order in related sentences. We have systematically investigated the amount of bilingual training data required to maintain an acceptable quality of machine translation. The combination of the suggested methods for improving translation quality in frameworks with scarce resources has been successfully tested: We were able to reduce the amount of bilingual training data to less than 10% of the original corpus, while losing only 1.6% in translation quality. The improvement of the translation results is demonstrated on two German-English corpora taken from the Verbmobil task and the Nespole! task. Sonja Nießen, Hermann Ney |
Comput. Linguistics | 2 |
| 2004 | The Alignment Template Approach to Statistical Machine TranslationabstractA phrase-based statistical machine translation approach — the alignment template approach — is described. This translation approach allows for general many-to-many relations between words. Thereby, the context of words is taken into account in the translation model, and local changes in word order from source to target language can be learned explicitly. The model is described using a log-linear modeling approach, which is a generalization of the often used source-channel approach. Thereby, the model is easier to extend than classical statistical machine translation systems. We describe in detail the process for learning phrasal translations, the feature functions used, and the search algorithm. The evaluation of this approach is performed on three different tasks. For the German-English speech Verbmobil task, we analyze the effect of various system components. On the French-English Canadian Hansards task, the alignment template system obtains significantly better results than a single-word-based translation model. In the Chinese-English 2002 National Institute of Standards and Technology (NIST) machine translation evaluation it yields statistically significantly better NIST scores than all competing research and commercial translation systems. Franz Josef Och, Hermann Ney |
Comput. Linguistics | 2 |
| 2004 | Some approaches to statistical and finite-state speech-to-speech translation
Francisco Casacuberta, Hermann Ney, Franz Josef Och, Enrique Vidal 0001, Juan Miguel Vilar, Sergio Barrachina 0001, Ismael García-Varea, David Llorens, Carlos D. Martínez-Hinarejos, Sirko Molau |
Comput. Speech Lang. | 2 |