Simon Wiesler

dblp:89/8760 · DBLP profile ↗
← Back
28ranked-venue papers
12as first author
6since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 27 · 11 first-author · 6 since 2021Artificial intelligence and machine learning · 15 · 6 first-author · 5 since 2021
YearPublicationVenuePosition
2023 Hierarchical Attention-Based Contextual Biasing For Personalized Speech Recognition Using Neural Transducers
abstract
Although end-to-end (E2E) automatic speech recognition (ASR) systems excel in general tasks, they frequently struggle with accurately recognizing personal rare words. Leveraging contextual information to bias the internal states of E2E ASR model has proven to be an effective solution. However most existing work focuses on biasing for a single domain and it is still challenging to expand such contextualization mechanisms to many domains. To address this limitation, in this work we propose a hierarchical attention architecture to scale contextual biasing to a wide range of domains simultaneously. Given multiple catalogs of contextual information, the high-level attention determines which source of catalog to focus on and the low-level attention learns to attend to the most relevant entity within the focused catalog. Experiments on diverse domains demonstrate the proposed architecture results in $35 \%$ to $60 \%$ relative WER improvements on personal rare words and outperforms existing approaches.
Sibo Tong, Philip Harding, Simon Wiesler
ASRU3
2023 Slot-Triggered Contextual Biasing For Personalized Speech Recognition Using Neural Transducers
abstract
End-to-end (E2E) automatic speech recognition (ASR) models have been found to perform well on general transcription tasks but often fail to correctly recognize words that occur infrequently in the training data. Personalization is important for a variety of tasks, including virtual assistants where recall of infrequently observed words such as contact names, song titles and place names is critical. In these cases contextual information is often available which can be used to bias the E2E ASR model. Contextual biasing (CB) has been shown to be effective for this task, however most existing work focuses on biasing for a single domain and so in this work we focus on the application of biasing to multiple domains. We propose a method whereby the E2E ASR model is trained to emit opening and closing tags around slot content which are used to both selectively enable biasing and decide which catalog to use for biasing. Our method is shown to not only efficiently scale to multiple slots, but also further improves accuracy on slot content.
Sibo Tong, Philip Harding, Simon Wiesler
ICASSP3
2023 Multi-View Frequency-Attention Alternative to CNN Frontends for Automatic Speech Recognition
Belen Alastruey, Lukas Drude, Jahn Heymann, Simon Wiesler
INTERSPEECH4
2023 Selective Biasing with Trie-based Contextual Adapters for Personalised Speech Recognition using Neural Transducers
Philip Harding, Sibo Tong, Simon Wiesler
INTERSPEECH3
2023 Model-Internal Slot-triggered Biasing for Domain Expansion in Neural Transducer ASR Models
Yiting Lu, Philip Harding, Kanthashree Mysore Sathyendra, Sibo Tong, Xuandi Fu, Feng-Ju Chang, Simon Wiesler, Grant P. Strimel
INTERSPEECH8
2021 Improving RNN-T ASR Accuracy Using Context Audio
abstract
We present a training scheme for streaming automatic speech recognition (ASR) based on recurrent neural network transducers (RNN-T) which allows the encoder network to learn to exploit context audio from a stream, using segmented or partially labeled sequences of the stream during training. We show that the use of context audio during training and inference can lead to word error rate reductions of more than 6% in a realistic production setting for a voice assistant ASR system. We investigate the effect of the proposed training approach on acoustically challenging data containing background speech and present data points which indicate that this approach helps the network learn both speaker and environment adaptation. To gain further insight into the ability of a long short-term memory (LSTM) based ASR encoder to exploit long-term context, we also visualize RNN-T loss gradients with respect to the input.
Andreas Schwarz, Ilya Sklyar, Simon Wiesler
Interspeech3
2020 Subword Regularization: An Analysis of Scalability and Generalization for End-to-End Automatic Speech Recognition
abstract
Subwords are the most widely used output units in end-to-end speech recognition. They combine the best of two worlds by modeling the majority of frequent words directly and at the same time allow open vocabulary speech recognition by backing off to shorter units or characters to construct words unseen during training. However, mapping text to subwords is ambiguous and often multiple segmentation variants are possible. Yet, many systems are trained using only the most likely segmentation. Recent research suggests that sampling subword segmentations during training acts as a regularizer for neural machine translation and speech recognition models, leading to performance improvements. In this work, we conduct a principled investigation on the regularizing effect of the subword segmentation sampling method for a streaming end-to-end speech recognition task. In particular, we evaluate the subword regularization contribution depending on the size of the training dataset. Our results suggest that subword regularization provides a consistent improvement of (2-8%) relative word-error-rate reduction, even in a large-scale setting with datasets up to a size of 20k hours. Further, we analyze the effect of subword regularization on recognition of unseen words and its implications on beam diversity.
Egor Lakomkin, Jahn Heymann, Ilya Sklyar, Simon Wiesler
INTERSPEECH4
2020 Improving Speech Recognition of Compound-Rich Languages
Prabhat Pandey, Volker Leutnant, Simon Wiesler, Jahn Heymann, Daniel Willett
INTERSPEECH3
2015 Investigation of mixture splitting concept for training linear bottlenecks of deep neural network acoustic models
abstract
A Gaussian or log-linear mixture model trained by maximum likelihood may be trained further using discriminative training. It is desirable that the mixture splitting is also done during the discriminative training, to achieve better mixture density distribution. In previous work such a discriminative splitting approach was presented. Similarly, the resolution of a deep neural network may also be increased by splitting. In this paper, discriminative splitting is applied as a way of initializing a linear bottleneck between two layers of a DNN. Experiments for a single hidden layer and six hidden layer cases show the potential of this approach as an alternative method of pre-training for linear bottlenecks for MLP hidden layers.
Muhammad Ali Tahir, Simon Wiesler, Ralf Schlüter, Hermann Ney
ICASSP2
2015 Sequence-discriminative training of recurrent neural networks
abstract
We investigate sequence-discriminative training of long shortterm memory recurrent neural networks using the maximum mutual information criterion. We show that although recurrent neural networks already make use of the whole observation sequence and are able to incorporate more contextual information than feed forward networks, their performance can be improved with sequence-discriminative training. Experiments are performed on two publicly available handwriting recognition tasks containing English and French handwriting. On the English corpus, we obtain a relative improvement in WER of over 11% with maximum mutual information (MMI) training compared to cross-entropy training. On the French corpus, we observed that it is necessary to interpolate the MMI objective function with cross-entropy.
Paul Voigtlaender, Patrick Doetsch, Simon Wiesler, Ralf Schlüter, Hermann Ney
ICASSP3
2015 Investigations on sequence training of neural networks
abstract
In this paper we present an investigation of sequence-discriminative training of deep neural networks for automatic speech recognition. We evaluate different sequence-discriminative training criteria (MMI and MPE) and optimization algorithms (including SGD and Rprop) using the RASR toolkit. Further, we compare the training of the whole network with that of the output layer only. Technical details necessary for a robust training are studied, since there is no consensus yet on the ultimate training recipe. The investigation extends our previous work on training linear bottleneck networks from scratch showing the consistently positive effect of sequence training.
Simon Wiesler, Pavel Golik, Ralf Schlüter, Hermann Ney
ICASSP1
2014 The RWTH English lecture recognition system
abstract
In this paper, we describe the RWTH speech recognition system for English lectures developed within the Translectures project. A difficulty in the development of an English lectures recognition system, is the high ratio of non-native speakers. We address this problem by using very effective deep bottleneck features trained on multilingual data. The acoustic model is trained on large amounts of data from different domains and with different dialects. Large improvements are obtained from unsupervised acoustic adaptation. Another challenge is the frequent use of technical terms and the wide range of topics. In our recognition system, slides, which are attached to most lectures, are used for improving lexical coverage and language model adaptation.
Simon Wiesler, Kazuki Irie, Zoltán Tüske, Ralf Schlüter, Hermann Ney
ICASSP1
2014 RASR/NN: The RWTH neural network toolkit for speech recognition
abstract
This paper describes the new release of RASR — the open source version of the well-proven speech recognition toolkit developed and used at RWTH Aachen University. The focus is put on the implementation of the NN module for training neural network acoustic models. We describe code design, configuration, and features of the NN module. The key feature is a high flexibility regarding the network topology, choice of activation functions, training criteria, and optimization algorithm, as well as a built-in support for efficient GPU computing. The evaluation of run-time performance and recognition accuracy is performed exemplary with a deep neural network as acoustic model in a hybrid NN/HMM system. The results show that RASR achieves a state-of-the-art performance on a real-world large vocabulary task, while offering a complete pipeline for building and applying large scale speech recognition systems.
Simon Wiesler, Alexander Richard, Pavel Golik, Ralf Schlüter, Hermann Ney
ICASSP1
2014 Mean-normalized stochastic gradient for large-scale deep learning
abstract
Deep neural networks are typically optimized with stochastic gradient descent (SGD). In this work, we propose a novel second-order stochastic optimization algorithm. The algorithm is based on analytic results showing that a non-zero mean of features is harmful for the optimization. We prove convergence of our algorithm in a convex setting. In our experiments we show that our proposed algorithm converges faster than SGD. Further, in contrast to earlier work, our algorithm allows for training models with a factorized structure from scratch. We found this structure to be very useful not only because it accelerates training and decoding, but also because it is a very effective means against overfitting. Combining our proposed optimization algorithm with this model structure, model size can be reduced by a factor of eight and still improvements in recognition error rate are obtained. Additional gains are obtained by improving the Newbob learning rate strategy.
Simon Wiesler, Alexander Richard, Ralf Schlüter, Hermann Ney
ICASSP1
2013 A critical evaluation of stochastic algorithms for convex optimization
abstract
Log-linear models find a wide range of applications in pattern recognition. The training of log-linear models is a convex optimization problem. In this work, we compare the performance of stochastic and batch optimization algorithms. Stochastic algorithms are fast on large data sets but can not be parallelized well. In our experiments on a broadcast conversations recognition task, stochastic methods yield competitive results after only a short training period, but when spending enough computational resources for parallelization, batch algorithms are competitive with stochastic algorithms. We obtained slight improvements by using a stochastic second order algorithm. Our best log-linear model outperforms the maximum likelihood trained Gaussian mixture model baseline although being ten times smaller.
Simon Wiesler, Alexander Richard, Ralf Schlüter, Hermann Ney
ICASSP1
2013 Improving LVCSR with hidden conditional random fields for grapheme-to-phoneme conversion
abstract
In virtually every state-of-the-art large vocabulary continuous speech recognition (LVCSR) system, grapheme-to-phoneme (G2P) conversion is applied to generalize beyond a fixed set of words given by a background lexicon. The overall performance of the G2P system has a strong effect on the recognition qual-ity. Typically, generative models based on joint-n-grams are used, although some discriminative models have a competitive performance but the training time may be quite large. In this work, the effect of using discriminative G2P modeling based on hidden conditional random fields (HCRFs) is ana-lyzed. Besides measuring and comparing the G2P qualities on a textual level, one focus is the performance of LVCSR systems. Although the HCRF model does not outperform the generative one on text data, we could improve our English QUAERO ASR system by 1-3 % relative on a couple of test corpora over a strong baseline by only replacing the G2P strategy. Index Terms: grapheme-to-phoneme conversion, G2P, LVCSR, HCRF, hidden conditional random fields
Stefan Hahn, Patrick Lehnen, Simon Wiesler, Ralf Schlüter, Hermann Ney
INTERSPEECH3
2013 Investigations on hessian-free optimization for cross-entropy training of deep neural networks
abstract
Context-dependent deep neural network HMMs have been shown to achieve recognition accuracy superior to Gaussian mixture models in a number of recent works.Typically, neural networks are optimized with stochastic gradient descent.On large datasets, stochastic gradient descent improves quickly during the beginning of the optimization.But since it does not make use of second order information, its asymptotic convergence behavior is slow.In regions with pathological curvature, stochastic gradient descent may almost stagnate and thereby falsely indicate convergence.Another drawback of stochastic gradient descent is that it can only be parallelized within minibatches.The Hessian-free algorithm is a second order batch optimization algorithm that does not suffer from these problems.In a recent work, Hessian-free optimization has been applied to a training of deep neural networks according to a sequence criterion.In that work, improvements in accuracy and training time have been reported.In this paper, we analyze the properties of the Hessian-free optimization algorithm and investigate whether it is suited for cross-entropy training of deep neural networks as well.
Simon Wiesler, Jinyu Li 0001
INTERSPEECH1
2012 Basis vector orthogonalization for an improved kernel gradient matching pursuit method
abstract
With the aim of achieving a computationally efficient optimization of kernel-based probabilistic models for various problems, such as sequential pattern recognition, we have already developed the kernel gradient matching pursuit method as an approximation technique for kernel-based classification. The conventional kernel gradient matching pursuit method approximates the optimal parameter vector by using a linear combination of a small number of basis vectors. In this paper, we propose an improved kernel gradient matching pursuit method that introduces orthogonality constraints to the obtained basis vector set. We verified the efficiency of the proposed method by conducting recognition experiments based on handwritten image datasets and speech datasets. We realized a scalable kernel optimization that incorporated various models, handled very high-dimensional features (>;100 K features), and enabled the use of large scale datasets (>; 10 M samples).
Yotaro Kubo, Shinji Watanabe 0001, Atsushi Nakamura, Simon Wiesler, Ralf Schlüter, Hermann Ney
ICASSP4
2012 Accelerated Batch Learning of Convex Log-linear Models for LVCSR
abstract
This paper describes a log-linear modeling framework suitable for large-scale speech recognition tasks. We introduce modifications to our training procedure that are required for extending our previous work on log-linear models to larger tasks. We give a detailed description of the training procedure with a focus on aspects that impact computational efficiency. The performance of our approach is evaluated on the English Quaero corpus, a challenging broadcast conversations task. The log-linear model consistently outperforms the maximum likelihood baseline system. Comparable performance to a system with minimum-phone-error training is achieved. Index Terms: acoustic modeling, discriminative models 1.
Simon Wiesler, Ralf Schlüter, Hermann Ney
INTERSPEECH1
2011 A convergence analysis of log-linear training and its application to speech recognition
abstract
Log-linear models are a promising approach for speech recognition. Typically, log-linear models are trained according to a strictly convex criterion. Optimization algorithms are guaranteed to converge to the unique global optimum of the objective function from any initialization. For large-scale applications, considerations in the limit of infinite iterations are not sufficient. We show that log-linear training can be a highly ill-conditioned optimization problem, resulting in extremely slow convergence. Conversely, the optimization problem can be preconditioned by feature transformations. Making use of our convergence analysis, we improve our log-linear speech recognition system and achieve a strong reduction of its training time. In addition, we validate our analysis on a continuous handwriting recognition task.
Simon Wiesler, Ralf Schlüter, Hermann Ney
ASRU1
2011 Subspace pursuit method for kernel-log-linear models
abstract
This paper presents a novel method for reducing the dimensionality of kernel spaces. Recently, to maintain the convexity of training, log linear models without mixtures have been used as emission probability density functions in hidden Markov models for automatic speech recognition. In that framework, nonlinearly-transformed high-dimensional features are used to achieve the nonlinear classification of the original observation vectors without using mixtures. In this paper, with the goal of using high-dimensional features in kernel spaces, the cutting plane subspace pursuit method proposed for support vector machines is generalized and applied to log-linear models. The experimental results show that the proposed method achieved an efficient approximation of the feature space by using a limited number of basis vectors.
Yotaro Kubo, Simon Wiesler, Ralf Schlüter, Hermann Ney, Shinji Watanabe 0001, Atsushi Nakamura, Tetsunori Kobayashi
ICASSP2
2011 The RWTH 2010 Quaero ASR evaluation system for English, French, and German
abstract
Recognizing Broadcast Conversational (BC) speech data is a difficult task, which can be regarded as one of the major challenges beyond the recognition of Broadcast News (BN).
Martin Sundermeyer, Markus Nußbaum-Thom, Simon Wiesler, Christian Plahl, Amr El-Desoky Mousa, Stefan Hahn, David Nolden, Ralf Schlüter, Hermann Ney
ICASSP3
2011 Feature selection for log-linear acoustic models
abstract
Log-linear acoustic models have been shown to be competitive with Gaussian mixture models in speech recognition. Their high training time can be reduced by feature selection. We compare a simple univariate feature selection algorithm with ReliefF - an efficient multivariate algorithm. An alternative to feature selection is ℓ1-regularized training, which leads to sparse models. We observe that this gives no speedup when sparse features are used, hence feature selection methods are preferable. For dense features, ℓ1-regularization can reduce training and recognition time. We generalize the well known Rprop algorithm for the optimization of ℓ1-regularized functions. Experiments on the Wall Street Journal corpus showed that a large number of sparse features could be discarded without loss of performance. A strong regularization led to slight performance degradations, but can be useful on large tasks, where training the full model is not tractable.
Simon Wiesler, Alexander Richard, Yotaro Kubo, Ralf Schlüter, Hermann Ney
ICASSP1
2011 A Convergence Analysis of Log-Linear Training
abstract
Log-linear models are widely used probability models for statistical pattern recognition. Typically, log-linear models are trained according to a convex criterion. In recent years, the interest in log-linear models has greatly increased. The optimization of log-linear model parameters is costly and therefore an important topic, in particular for large-scale applications. Different optimization algorithms have been evaluated empirically in many papers. In this work, we analyze the optimization problem analytically and show that the training of log-linear models can be highly ill-conditioned. We verify our findings on two handwriting tasks. By making use of our convergence analysis, we obtain good results on a large-scale continuous handwriting recognition task with a simple and generic approach.
Simon Wiesler, Hermann Ney
NIPS1
2010 Discriminative HMMS, log-linear models, and CRFS: What is the difference?
abstract
Recently, there have been many papers studying discriminative acoustic modeling techniques like conditional random fields or discriminative training of conventional Gaussian HMMs. This paper will give an overview of the recent work and progress. We will strictly distinguish between the type of acoustic models on the one hand and the training criterion on the other hand. We will address two issues in more detail: the relation between conventional Gaussian HMMs and conditional random fields and the advantages of formulating the training criterion as a convex optimization problem. Experimental results for various speech tasks will be presented to carefully evaluate the different concepts and approaches, including both a digit string and large vocabulary continuous speech recognition tasks.
Georg Heigold, Simon Wiesler, Markus Nußbaum-Thom, Patrick Lehnen, Ralf Schlüter, Hermann Ney
ICASSP2
2010 The RWTH 2009 quaero ASR evaluation system for English and German
abstract
In this work, the RWTH automatic speech recognition systems for English and German for the second Quaero evaluation campaign 2009 are presented. The systems are designed to transcribe web data, European parliament plenary sessions and broadcast news data. Another challenge in the 2009 evaluation is that almost no in-domain training data is provided and the test data contains a large variety of speech types. The RWTH participates for the English and German languages with the best results for German and competitive results for the English. Contributing to the enhancements are the systematic use of hierarchical neural network based posterior features, system combination, speaker adaptation, cross speaker adaptation, domain dependent modeling and the usage of additional training data.
Markus Nußbaum-Thom, Simon Wiesler, Martin Sundermeyer, Christian Plahl, Stefan Hahn, Ralf Schlüter, Hermann Ney
INTERSPEECH2
2010 A discriminative splitting criterion for phonetic decision trees
abstract
Phonetic decision trees are a key concept in acoustic modeling for large vocabulary continuous speech recognition.Although discriminative training has become a major line of research in speech recognition and all state-of-the-art acoustic models are trained discriminatively, the conventional phonetic decision tree approach still relies on the maximum likelihood principle.In this paper we develop a splitting criterion based on the minimization of the classification error.An improvement of more than 10% relative over a discriminatively trained baseline system on the Wall Street Journal corpus suggests that the proposed approach is promising.
Simon Wiesler, Georg Heigold, Markus Nußbaum-Thom, Ralf Schlüter, Hermann Ney
INTERSPEECH1
2009 Investigations on features for log-linear acoustic models in continuous speech recognition
abstract
Hidden Markov Models with Gaussian Mixture Models as emission probabilities (GHMMs) are the underlying structure of all state-of-the-art speech recognition systems. Using Gaussian mixture distributions follows the generative approach where the class-conditional probability is modeled, although for classification only the posterior probability is needed. Though being very successful in related tasks like Natural Language Processing (NLP), in speech recognition direct modeling of posterior probabilities with log-linear models has rarely been used and has not been applied successfully to continuous speech recognition. In this paper we report competitive results for a speech recognizer with a log-linear acoustic model on the Wall Street Journal corpus, a Large Vocabulary Continuous Speech Recognition (LVCSR) task. We trained this model from scratch, i.e. without relying on an existing GHMM system. Previously the use of data dependent sparse features for log-linear models has been proposed. We compare them with polynomial features and show that the combination of polynomial and data dependent sparse features leads to better results.
Simon Wiesler, Markus Nußbaum-Thom, Georg Heigold, Ralf Schlüter, Hermann Ney
ASRU1