Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Georg Heigold

dblp:46/2236 · DBLP profile ↗
← Back
54ranked-venue papers
19as first author
5since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 38 · 10 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 38 · 13 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
10 papers
Video understanding and tracking · 23% Image recognition and object detection · 22% Representation and self-supervised learning · 19%
Computer graphics and multimedia
3 papers
Audio and music processing · 100%

Topics — the 26 heaviest of 28, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Representation and self-supervised learning › representation learning
object-centric representation learning
1.022022
Conditional Object-Centric Learning from Video · ICLR 2022
Object-Centric Learning with Slot Attention · NeurIPS 2020
Machine learning › Deep learning architectures and training › transformer
vision transformer
1.022021
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale · ICLR 2021
ViViT: A Video Vision Transformer · ICCV 2021
Computer vision › Image recognition and object detection
object localization
0.712023
Video OWL-ViT: Temporally-consistent open-world localization in video · ICCV 2023
Computer vision › Video understanding and tracking
object tracking
0.712023
Video OWL-ViT: Temporally-consistent open-world localization in video · ICCV 2023
Computer vision › Video understanding and tracking › video representation learning
object-centric video learning
0.612022
Conditional Object-Centric Learning from Video · ICLR 2022
Computer vision › Video understanding and tracking
video classification
0.512021
ViViT: A Video Vision Transformer · ICCV 2021
Computer vision › Image recognition and object detection
object discovery
0.412020
Object-Centric Learning with Slot Attention · NeurIPS 2020
Machine learning › Representation and self-supervised learning › representation learning › object-centric representation learning
slot attention
0.412020
Object-Centric Learning with Slot Attention · NeurIPS 2020
Computer vision › Image recognition and object detection › object discovery
unsupervised object discovery
0.412020
Object-Centric Learning with Slot Attention · NeurIPS 2020
Audio and music processing
speech recognition
0.332013
Investigations on an EM-Style Optimization Algorithm for Discriminative Training of HMMs · IEEE ACM Trans. Audio Speech Lang. Process. 2013
Equivalence of Generative and Log-Linear Models · IEEE Trans. Speech Audio Process. 2011
Optimization Algorithms and Applications for Speech and Language Processing · IEEE Trans. Speech Audio Process. 2013
Natural language and speech › Information extraction and text analysis › morphological analysis
morphological tagging
0.312017
Cross-lingual Character-Level Neural Morphological Tagging · EMNLP 2017
Audio and music processing › speech recognition
discriminative training
0.212013
Investigations on an EM-Style Optimization Algorithm for Discriminative Training of HMMs · IEEE ACM Trans. Audio Speech Lang. Process. 2013
Machine learning › Deep learning architectures and training
transformer
0.112021
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale · ICLR 2021
Natural language and speech › Speech recognition and synthesis
automatic speech recognition
0.112012
WFST Enabled Solutions to ASR Problems: Beyond HMM Decoding · IEEE Trans. Speech Audio Process. 2012
Natural language and speech › Speech recognition and synthesis › acoustic modeling
discriminative acoustic model training
0.112012
WFST Enabled Solutions to ASR Problems: Beyond HMM Decoding · IEEE Trans. Speech Audio Process. 2012
Computer vision › Image recognition and object detection › character recognition
handwritten digit recognition
0.112012
Latent Log-Linear Models for Handwritten Digit Classification · IEEE Trans. Pattern Anal. Mach. Intell. 2012
Machine learning › Probabilistic and Bayesian machine learning › structured models
latent variable model
0.112012
Latent Log-Linear Models for Handwritten Digit Classification · IEEE Trans. Pattern Anal. Mach. Intell. 2012
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › exponential family
maximum entropy models
0.112012
Latent Log-Linear Models for Handwritten Digit Classification · IEEE Trans. Pattern Anal. Mach. Intell. 2012
Natural language and speech › Machine translation
system combination
0.112012
WFST Enabled Solutions to ASR Problems: Beyond HMM Decoding · IEEE Trans. Speech Audio Process. 2012
Natural language and speech › Speech recognition and synthesis › search and decoding
weighted finite-state transducer decoding
0.112012
WFST Enabled Solutions to ASR Problems: Beyond HMM Decoding · IEEE Trans. Speech Audio Process. 2012
Audio and music processing › speech recognition
acoustic modeling
0.112011
Equivalence of Generative and Log-Linear Models · IEEE Trans. Speech Audio Process. 2011
Machine learning › Transfer learning and domain adaptation
cross-lingual transfer
0.112017
Cross-lingual Character-Level Neural Morphological Tagging · EMNLP 2017
Natural language and speech › Speech recognition and synthesis
acoustic model training
0.112008
Modified MMI/MPE: a direct evaluation of the margin in speech recognition · ICML 2008
Natural language and speech › Speech recognition and synthesis › acoustic model training
discriminative training
0.112008
Modified MMI/MPE: a direct evaluation of the margin in speech recognition · ICML 2008
Natural language and speech › Speech recognition and synthesis › acoustic model training › discriminative training
large margin training
0.112008
Modified MMI/MPE: a direct evaluation of the margin in speech recognition · ICML 2008
Natural language and speech › Language models and text generation › decoding
minimum bayes risk decoding
0.012012
WFST Enabled Solutions to ASR Problems: Beyond HMM Decoding · IEEE Trans. Speech Audio Process. 2012

Methods — techniques the papers use, named apart from their topics

transformer decoder · 0.7open-vocabulary detection · 0.7image-text pretraining · 0.7slot attention · 0.6contrastive learning · 0.6spatiotemporal tokenization · 0.5self-attention · 0.5regularization · 0.5pre-trained image model · 0.5large-scale pretraining · 0.5extended baum-welch · 0.3sparse optimization · 0.2rprop · 0.2minimum phone error · 0.2maximum mutual information · 0.2expectation-maximization · 0.2baum-welch · 0.2GIS · 0.2
YearPublicationVenuePosition
2025 Massive Sound Embedding Benchmark (MSEB)
abstract
Audio is a critical component of multimodal perception, and any truly intelligent system must demonstrate a wide range of auditory capabilities. These capabilities include transcription, classification, retrieval, reasoning, segmentation, clustering, reranking, and reconstruction. Fundamentally, each task involves transforming a raw audio signal into a meaningful 'embedding'—be it a single vector, a sequence of continuous or discrete representations, or another structured form—which then serves as the basis for generating the task's final response. To accelerate progress towards robust machine auditory intelligence, we present the Massive Sound Embedding Benchmark (MSEB): an extensible framework designed to evaluate the auditory components of any multimodal system. In its first release, MSEB offers a comprehensive suite of eight core tasks, with more planned for the future, supported by diverse datasets, including the new, large-scale Simple Voice Questions (SVQ) dataset. Our initial experiments establish clear performance headrooms, highlighting the significant opportunity to improve real-world multimodal experiences where audio is a core signal. We encourage the research community to use MSEB to assess their algorithms and contribute to its growth. The library is publicly hosted at https://github.com/google-research/mseb.
Georg Heigold, Ehsan Variani, Tom Bagby, Cyril Allauzen, Ji Ma 0004, Shankar Kumar, Michael Riley 0001
NeurIPS1
2023 Video OWL-ViT: Temporally-consistent open-world localization in video
abstract
We present an architecture and a training recipe that adapts pretrained open-world image models to localization in videos. Understanding the open visual world (without being constrained by fixed label spaces) is crucial for many real-world vision tasks. Contrastive pre-training on large image-text datasets has recently led to significant improvements for image-level tasks. For more structured tasks involving object localization applying pre-trained models is more challenging. This is particularly true for video tasks, where task-specific data is limited. We show successful transfer of open-world models by building on the OWL-ViT open-vocabulary detection model and adapting it to video by adding a transformer decoder. The decoder propagates object representations recurrently through time by using the output tokens for one frame as the object queries for the next. Our model is end-to-end trainable on video data and enjoys improved temporal consistency compared to tracking-by-detection baselines, while retaining the open-world capabilities of the backbone detector. We evaluate our model on the challenging TAO-OW benchmark and demonstrate that open-world capabilities, learned from large-scale image-text pretraining, can be transferred successfully to open-world localization across diverse videos.
Georg Heigold, Daniel Keysers, Matthias Minderer, Mario Lucic, Alexey A. Gritsenko, Fisher Yu 0001, Alex Bewley, Thomas Kipf
ICCV1
2022 Conditional Object-Centric Learning from Video
Thomas Kipf, Gamaleldin F. Elsayed, Aravindh Mahendran, Austin Stone, Sara Sabour, Georg Heigold, Rico Jonschkowski, Alexey Dosovitskiy, Klaus Greff
ICLR6
2021 ViViT: A Video Vision Transformer
abstract
We present pure-transformer based models for video classification, drawing upon the recent success of such models in image classification. Our model extracts spatiotemporal tokens from the input video, which are then encoded by a series of transformer layers. In order to handle the long sequences of tokens encountered in video, we propose several, efficient variants of our model which factorise the spatial- and temporal-dimensions of the input. Although transformer-based models are known to only be effective when large training datasets are available, we show how we can effectively regularise the model during training and leverage pretrained image models to be able to train on comparatively small datasets. We conduct thorough ablation studies, and achieve state-of-the-art results on multiple video classification benchmarks including Kinetics 400 and 600, Epic Kitchens, Something-Something v2 and Moments in Time, outperforming prior methods based on deep 3D convolutional networks.
Anurag Arnab, Mostafa Dehghani 0001, Georg Heigold, Chen Sun 0002, Mario Lucic, Cordelia Schmid
ICCV3
2021 An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov 0003, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani 0001, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, Neil Houlsby
ICLR9
2020 Object-Centric Learning with Slot Attention
abstract
Learning object-centric representations of complex scenes is a promising step towards enabling efficient abstract reasoning from low-level perceptual features. Yet, most deep learning approaches learn distributed representations that do not capture the compositional properties of natural scenes. In this paper, we present the Slot Attention module, an architectural component that interfaces with perceptual representations such as the output of a convolutional neural network and produces a set of task-dependent abstract representations which we call slots. These slots are exchangeable and can bind to any object in the input by specializing through a competitive procedure over multiple rounds of attention. We empirically demonstrate that Slot Attention can extract object-centric representations that enable generalization to unseen compositions when trained on unsupervised object discovery and supervised property prediction tasks.
Francesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran, Georg Heigold, Jakob Uszkoreit, Alexey Dosovitskiy, Thomas Kipf
NeurIPS5
2017 An Extensive Empirical Evaluation of Character-Based Morphological Tagging for 14 Languages
abstract
This paper investigates neural characterbased morphological tagging for languages with complex morphology and large tag sets.Character-based approaches are attractive as they can handle rarelyand unseen words gracefully.We evaluate on 14 languages and observe consistent gains over a state-of-the-art morphological tagger across all languages except for English and French, where we match the state-of-the-art.We compare two architectures for computing characterbased word vectors using recurrent (RNN) and convolutional (CNN) nets.We show that the CNN based approach performs slightly worse and less consistently than the RNN based approach.Small but systematic gains are observed when combining the two architectures by ensembling.
Georg Heigold, Günter Neumann, Josef van Genabith
EACL (1)1
2017 Cross-lingual Character-Level Neural Morphological Tagging
abstract
Even for common NLP tasks, sufficient supervision is not available in many languages-morphological tagging is no exception.In the work presented here, we explore a transfer learning scheme, whereby we train character-level recurrent neural taggers to predict morphological taggings for high-resource languages and low-resource languages together.Learning joint character representations among multiple related languages successfully enables knowledge transfer from the high-resource languages to the low-resource ones, improving accuracy by up to 30%.
Ryan Cotterell, Georg Heigold
EMNLP2
2016 Scaling character-based morphological tagging to fourteen languages
abstract
This paper investigates neural character-based morphological tagging for languages with complex morphology and large tag sets. Character-based approaches are attractive as they can handle rarely- and unseen words gracefully. More specifically, beside a rich morphology, non-canonical language, change of language or other linguistic variability can heavily degrade the accuracy of natural language processing of web and CMC data. We evaluate on 14 languages and observe consistent gains over a state-of-the-art morphological tagger across all languages except for English and French, where we match the state-of-the-art. The gains are clearly correlated with the amount of training data. We present supplementary experiments to explore whether and to what extent unsupervised data through pre-trained word vectors can compensate for limited amounts of supervised data. Moreover, we show preliminary results to study the effect of noisy input data by flipping characters at random.
Georg Heigold, Josef van Genabith, Günter Neumann
IEEE BigData1
2016 End-to-end text-dependent speaker verification
abstract
In this paper we present a data-driven, integrated approach to speaker verification, which maps a test utterance and a few reference utterances directly to a single score for verification and jointly optimizes the system's components using the same evaluation protocol and metric as at test time. Such an approach will result in simple and efficient systems, requiring little domain-specific knowledge and making few model assumptions. We implement the idea by formulating the problem as a single neural network architecture, including the estimation of a speaker model on only a few utterances, and evaluate it on our internal "Ok Google" benchmark for text-dependent speaker verification. The proposed approach appears to be very effective for big data applications Like ours that require highly accurate, easy-to-maintain systems with a small footprint.
Georg Heigold, Ignacio Moreno, Samy Bengio, Noam Shazeer
ICASSP1
2015 A Gaussian Mixture Model layer jointly optimized with discriminative features within a Deep Neural Network architecture
abstract
This article proposes and evaluates a Gaussian Mixture Model (GMM) represented as the last layer of a Deep Neural Network (DNN) architecture and jointly optimized with all previous layers using Asynchronous Stochastic Gradient Descent (ASGD). The resulting “Deep GMM” architecture was investigated with special attention to the following issues: (1) The extent to which joint optimization improves over separate optimization of the DNN-based feature extraction layers and the GMM layer; (2) The extent to which depth (measured in number of layers, for a matched total number of parameters) helps a deep generative model based on the GMM layer, compared to a vanilla DNN model; (3) Head-to-head performance of Deep GMM architectures vs. equivalent DNN architectures of comparable depth, using the same optimization criterion (frame-level Cross Entropy (CE)) and optimization method (ASGD); (4) Expanded possibilities for modeling offered by the Deep GMM generative model. The proposed Deep GMMs were found to yield Word Error Rates (WERs) competitive with state-of-the-art DNN systems, at the cost of pre-training using standard DNNs to initialize the Deep GMM feature extraction layers. An extension to Deep Subspace GMMs is described, resulting in additional gains.
Ehsan Variani, Erik McDermott, Georg Heigold
ICASSP3
2014 Small-footprint keyword spotting using deep neural networks
abstract
Our application requires a keyword spotting system with a small memory footprint, low computational cost, and high precision. To meet these requirements, we propose a simple approach based on deep neural networks. A deep neural network is trained to directly predict the keyword(s) or subword units of the keyword(s) followed by a posterior handling method producing a final confidence score. Keyword recognition results achieve 45% relative improvement with respect to a competitive Hidden Markov Model-based system, while performance in the presence of babble noise shows 39% relative improvement.
Guoguo Chen, Carolina Parada, Georg Heigold
ICASSP3
2014 Asynchronous stochastic optimization for sequence training of deep neural networks
abstract
This paper explores asynchronous stochastic optimization for sequence training of deep neural networks. Sequence training requires more computation than frame-level training using pre-computed frame data. This leads to several complications for stochastic optimization, arising from significant asynchrony in model updates under massive parallelization, and limited data shuffling due to utterance-chunked processing. We analyze the impact of these two issues on the efficiency and performance of sequence training. In particular, we suggest a framework to formalize the reasoning about the asynchrony and present experimental results on both small and large scale Voice Search tasks to validate the effectiveness and efficiency of asynchronous stochastic optimization.
Georg Heigold, Erik McDermott, Vincent Vanhoucke, Andrew W. Senior, Michiel Bacchiani
ICASSP1
2014 GMM-free DNN acoustic model training
abstract
While deep neural networks (DNNs) have become the dominant acoustic model (AM) for speech recognition systems, they are still dependent on Gaussian mixture models (GMMs) for alignments both for supervised training and for context dependent (CD) tree building. Here we explore bootstrapping DNN AM training without GMM AMs and show that CD trees can be built with DNN alignments which are better matched to the DNN model and its features. We show that these trees and alignments result in better models than from the GMM alignments and trees. By removing the GMM acoustic model altogether we simplify the system required to train a DNN from scratch.
Andrew W. Senior, Georg Heigold, Michiel Bacchiani, Hank Liao
ICASSP2
2014 Asynchronous, online, GMM-free training of a context dependent acoustic model for speech recognition
abstract
We propose an algorithm that allows online training of a con-text dependent DNN model. It designs a state inventory based on DNN features and jointly optimizes the DNN parameters and alignment of the training data. The process allows flat starting a model from scratch and avoids any dependency on a GMM/HMM model to bootstrap the training process. A 15k state model trained with the proposed algorithm reduced the er-ror rate on a mobile speech task with 24 % compared to a system bootstrapped from a CI HMM/GMM and with 16 % compared to a system bootstrapped from a CD HMM/GMM system. Index Terms: Deep Neural Networks, online training 1.
Michiel Bacchiani, Andrew W. Senior, Georg Heigold
INTERSPEECH3
2014 Word embeddings for speech recognition
abstract
Speech recognition systems have used the concept of states as a way to decompose words into sub-word units for decades. As the number of such states now reaches the number of words used to train acoustic models, it is interesting to consider ap-proaches that relax the assumption that words are made of states. We present here an alternative construction, where words are projected into a continuous embedding space where words that sound alike are nearby in the Euclidean sense. We show how embeddings can still allow to score words that were not in the training dictionary. Initial experiments using a lattice rescor-ing approach and model combination on a large realistic dataset show improvements in word error rate. Index Terms: embeddings, deep learning, speech recognition. 1.
Samy Bengio, Georg Heigold
INTERSPEECH2
2014 Asynchronous stochastic optimization for sequence training of deep neural networks: towards big data
abstract
Previous work presented a proof of concept for sequence training of deep neural networks (DNNs) using asynchronous stochastic optimization, mainly focusing on a small-scale task. The approach offers the potential to leverage both the efficiency of stochastic gradient descent and the scalability of parallel computation. This study presents results for four different voice search tasks to confirm the effectiveness and efficiency of the proposed framework across different conditions: amount of data (from 60 hours to 20,000 hours), type of speech (read speech vs. spontaneous speech), quality of data (supervised vs. unsupervised data), and language. Significant gains over baselines (DNNs trained at the frame level) are found to hold across these conditions. The experimental results are analyzed, and additional practical details for the approach are provided. Furthermore, different sequence training criteria are compared.
Erik McDermott, Georg Heigold, Pedro J. Moreno 0001, Andrew W. Senior, Michiel Bacchiani
INTERSPEECH2
2014 Sequence discriminative distributed training of long short-term memory recurrent neural networks
Hasim Sak, Oriol Vinyals, Georg Heigold, Andrew W. Senior, Erik McDermott, Rajat Monga, Mark Z. Mao
INTERSPEECH3
2013 Multilingual acoustic models using distributed deep neural networks
abstract
Today's speech recognition technology is mature enough to be useful for many practical applications. In this context, it is of paramount importance to train accurate acoustic models for many languages within given resource constraints such as data, processing power, and time. Multilingual training has the potential to solve the data issue and close the performance gap between resource-rich and resource-scarce languages. Neural networks lend themselves naturally to parameter sharing across languages, and distributed implementations have made it feasible to train large networks. In this paper, we present experimental results for cross- and multi-lingual network training of eleven Romance languages on 10k hours of data in total. The average relative gains over the monolingual baselines are 4%/2% (data-scarce/data-rich languages) for cross- and 7%/2% for multi-lingual training. However, the additional gain from jointly training the languages on all data comes at an increased training time of roughly four weeks, compared to two weeks (monolingual) and one week (crosslingual).
Georg Heigold, Vincent Vanhoucke, Andrew W. Senior, Patrick Nguyen, Marc'Aurelio Ranzato, Matthieu Devin, Jeffrey Dean
ICASSP1
2013 Deep neural networks with auxiliary Gaussian mixture models for real-time speech recognition
abstract
We present a framework that improves real-time speech recognition performance using deep neural networks (DNNs) with auxiliary Gaussian mixture models (GMMs). The DNNs and the auxiliary GMMs share the same hidden Markov model (HMM) state inventory. First, online incremental feature-space adaptation is performed using the GMM acoustic model. The speaker-adapted features are used to improve the recognition performance of both GMM and DNN models. Second, the acoustic scores from GMMs and DNN are combined at the state-level during decoding. Experiments on a large vocabulary speech recognition task show that both approaches improve recognition performance consistently and that the gains are mostly additive, resulting in about 5% relative improvement over the competitive DNN baseline in both Portuguese and English systems.
Georg Heigold
ICASSP3
2013 An empirical study of learning rates in deep neural networks for speech recognition
abstract
Recent deep neural network systems for large vocabulary speech recognition are trained with minibatch stochastic gradient descent but use a variety of learning rate scheduling schemes. We investigate several of these schemes, particularly AdaGrad. Based on our analysis of its limitations, we propose a new variant `AdaDec' that decouples long-term learning-rate scheduling from per-parameter learning rate variation. AdaDec was found to result in higher frame accuracies than other methods. Overall, careful choice of learning rate schemes leads to faster convergence and lower word error rates.
Andrew W. Senior, Georg Heigold, Marc'Aurelio Ranzato
ICASSP2
2013 Multiframe deep neural networks for acoustic modeling
abstract
Deep neural networks have been shown to perform very well as acoustic models for automatic speech recognition. Compared to Gaussian mixtures however, they tend to be very expensive computationally, making them challenging to use in real-time applications. One key advantage of such neural networks is their ability to learn from very long observation windows going up to 400 ms. Given this very long temporal context, it is tempting to wonder whether one can run neural networks at a lower frame rate than the typical 10 ms, and whether there might be computational benefits to doing so. This paper describes a method of tying the neural network parameters over time which achieves comparable performance to the typical frame-synchronous model, while achieving up to a 4X reduction in the computational cost of the neural network activations.
Vincent Vanhoucke, Matthieu Devin, Georg Heigold
ICASSP3
2013 Investigations on an EM-Style Optimization Algorithm for Discriminative Training of HMMs
abstract
Today's speech recognition systems are based on hidden Markov models (HMMs) with Gaussian mixture models whose parameters are estimated using a discriminative training criterion such as Maximum Mutual Information (MMI) or Minimum Phone Error (MPE). Currently, the optimization is almost always done with (empirical variants of) Extended Baum-Welch (EBW). This type of optimization requires sophisticated update schemes for the step sizes and a considerable amount of parameter tuning, and only little is known about its convergence behavior. In this paper, we derive an EM-style algorithm for discriminative training of HMMs. Like Expectation-Maximization (EM) for the generative training of HMMs, the proposed algorithm improves the training criterion on each iteration, converges to a local optimum, and is completely parameter-free. We investigate the feasibility of the proposed EM-style algorithm for discriminative training of two tasks, namely grapheme-to-phoneme conversion and spoken digit string recognition.
Georg Heigold, Hermann Ney, Ralf Schlüter
IEEE ACM Trans. Audio Speech Lang. Process.1
2013 Optimization Algorithms and Applications for Speech and Language Processing
abstract
Optimization techniques have been used for many years in the formulation and solution of computational problems arising in speech and language processing. Such techniques are found in the Baum-Welch, extended Baum-Welch (EBW), Rprop, and GIS algorithms, for example. Additionally, the use of regularization terms has been seen in other applications of sparse optimization. This paper outlines a range of problems in which optimization formulations and algorithms play a role, giving some additional details on certain application problems in machine translation, speaker/language recognition, and automatic speech recognition. Several approaches developed in the speech and language processing communities are described in a way that makes them more recognizable as optimization procedures. Our survey is not exhaustive and is complemented by other papers in this volume.
Stephen J. Wright 0001, Dimitri Kanevsky, Li Deng 0001, Xiaodong He 0001, Georg Heigold, Haizhou Li 0001
IEEE Trans. Speech Audio Process.5
2012 Investigations on exemplar-based features for speech recognition towards thousands of hours of unsupervised, noisy data
abstract
The acoustic models in state-of-the-art speech recognition systems are based on phones in context that are represented by hidden Markov models. This modeling approach may be limited in that it is hard to incorporate long-span acoustic context. Exemplar-based approaches are an attractive alter-native, in particular if massive data and computational power are available. Yet, most of the data at Google are unsupervised and noisy. This paper investigates an exemplar-based approach under this yet not well understood data regime. A log-linear rescoring framework is used to combine the exemplar-based features on the word level with the first-pass model. This approach guarantees at least baseline performance and focuses on the refined modeling of words with sufficient data. Experimental results for the Voice Search and the YouTube tasks are presented.
Georg Heigold, Patrick Nguyen, Mitch Weintraub, Vincent Vanhoucke
ICASSP1
2012 Overview of large scale optimization for discriminative training in speech recognition
abstract
Over the past few decades, a variety of specialized approaches have been proposed to solve large problems in speech recognition. Conventional optimization techniques have not been widely applied, because the problems do not readily admit an objective for evaluating a given set of parameters and because of the large number of parameters. This situation is changing, due to recent developments in algorithmic optimization. In this paper, we review the specialized algorithms, including methods derived from the extended Baum-Welch (EBW) approach, Rprop, and GIS. We discuss optimization frameworks that could also potentially be applied, and outline some connections between the optimization methods and existing specialized methods.
Dimitri Kanevsky, Georg Heigold, Stephen J. Wright 0001, Hermann Ney
ICASSP2
2012 Posterior-Scaled MPE: Novel Discriminative Training Criteria
abstract
We recently discovered novel discriminative training criteria following a principled approach. In this approach training criteria are developed from error bounds on the global error for pattern classification tasks that depend on non-trivial loss functions. Automatic speech recognition (ASR) is a prominent example for such a task depending on the non-trivial Levenshtein loss. In this context, the posterior-scaled Minimum Phoneme Error (MPE) training criterion, which is the state-of-the-art discriminative training criterion in ASR, was shown to be an approximation to one of the novel criteria. Here, we describe the implementation of the posterior-scaled MPE criterion in a transducer-based framework, and compare this criterion to other discriminative training criteria on an ASR task. This comparison indicates that the posterior-scaled MPE criterion performs better than other discriminative criteria including MPE. Index Terms: error bounds, discriminative training criteria, margin, MPE
Markus Nußbaum-Thom, Zoltán Tüske, Georg Heigold, Ralf Schlüter, Hermann Ney
INTERSPEECH3
2012 Latent Log-Linear Models for Handwritten Digit Classification
abstract
We present latent log-linear models, an extension of log-linear models incorporating latent variables, and we propose two applications thereof: log-linear mixture models and image deformation-aware log-linear models. The resulting models are fully discriminative, can be trained efficiently, and the model complexity can be controlled. Log-linear mixture models offer additional flexibility within the log-linear modeling framework. Unlike previous approaches, the image deformation-aware model directly considers image deformations and allows for a discriminative training of the deformation parameters. Both are trained using alternating optimization. For certain variants, convergence to a stationary point is guaranteed and, in practice, even variants without this guarantee converge and find models that perform well. We tune the methods on the USPS data set and evaluate on the MNIST data set, demonstrating the generalization capabilities of our proposed models. Our models, although using significantly fewer parameters, are able to obtain competitive results with models proposed in the literature.
Thomas Deselaers, Tobias Gass, Georg Heigold, Hermann Ney
IEEE Trans. Pattern Anal. Mach. Intell.3
2012 WFST Enabled Solutions to ASR Problems: Beyond HMM Decoding
abstract
During the last decade, weighted finite-state transducers (WFSTs) have become popular in speech recognition. While their main field of application remains hidden Markov model (HMM) decoding, the WFST framework is now also seen as a brick in solutions to many other central problems in automatic speech recognition (ASR). These solutions are less known, and this work aims at giving an overview of the applications of WFSTs in large-vocabulary continuous speech recognition (LVCSR) besides HMM decoding: discriminative acoustic model training, Bayes risk decoding, and system combination. The application of the WFST framework has a big practical impact: we show how the framework helps to structure problems, to develop generic solutions, and to delegate complex computations to WFST toolkits. In this paper, we review the literature, discuss existing approaches, and provide new insights into WFST enabled solutions. We also present a novel, purely WFST-based algorithm for computing the exact Bayes risk hypothesis from a lattice with the Levenshtein distance as loss function. We present the problems and their solutions in a unified framework and discuss the advantages and limits of using WFSTs. We do not provide new experimental results, but refer to the existing literature. Our work helps to identify where and how the transducer framework can contribute to a compact and generic solution to LVCSR problems.
Björn Hoffmeister, Georg Heigold, David Rybach, Ralf Schlüter, Hermann Ney
IEEE Trans. Speech Audio Process.2
2011 EM-style optimization of hidden conditional random fields for grapheme-to-phoneme conversion
abstract
We have recently proposed an EM-style algorithm to optimize log-linear models with hidden variables. In this paper, we use this algorithm to optimize a hidden conditional random field, i.e., a conditional random field with hidden variables. Similar to hidden Markov models, the alignments are the hidden variables in the examples considered. Here, EM-style algorithms are iterative optimization algorithms which are guaranteed to improve the training criterion in each iteration without the need for tuning step sizes, sophisticated update schemes or numerical line optimization (with hardly predictable complexity). This is a rather strong property which conventional gradient-based optimization algorithms do not have. We present experimental results for a grapheme-to-phoneme conversion task and compare the convergence behavior of the EM-style algorithm with L-BFGS and Rprop.
Georg Heigold, Stefan Hahn, Patrick Lehnen, Hermann Ney
ICASSP1
2011 Confidence- and margin-based MMI/MPE discriminative training for off-line handwriting recognition
Philippe Dreuw, Georg Heigold, Hermann Ney
Int. J. Document Anal. Recognit.2
2011 Equivalence of Generative and Log-Linear Models
abstract
Conventional speech recognition systems are based on hidden Markov models (HMMs) with Gaussian mixture models (GHMMs). Discriminative log-linear models are an alternative modeling approach and have been investigated recently in speech recognition. GHMMs are directed models with constraints, e.g., positivity of variances and normalization of conditional probabilities, while log-linear models do not use such constraints. This paper compares the posterior form of typical generative models related to speech recognition with their log-linear model counterparts. The key result will be the derivation of the equivalence of these two different approaches under weak assumptions. In particular, we study Gaussian mixture models, part-of-speech bigram tagging models, and eventually, the GHMMs. This result unifies two important but fundamentally different modeling paradigms in speech recognition on the functional level. Furthermore, this paper will present comparative experimental results for various speech tasks of different complexity, including a digit string and large-vocabulary continuous speech recognition tasks.
Georg Heigold, Hermann Ney, Patrick Lehnen, Tobias Gass, Ralf Schlüter
IEEE Trans. Speech Audio Process.1
2010 Discriminative HMMS, log-linear models, and CRFS: What is the difference?
abstract
Recently, there have been many papers studying discriminative acoustic modeling techniques like conditional random fields or discriminative training of conventional Gaussian HMMs. This paper will give an overview of the recent work and progress. We will strictly distinguish between the type of acoustic models on the one hand and the training criterion on the other hand. We will address two issues in more detail: the relation between conventional Gaussian HMMs and conditional random fields and the advantages of formulating the training criterion as a convex optimization problem. Experimental results for various speech tasks will be presented to carefully evaluate the different concepts and approaches, including both a digit string and large vocabulary continuous speech recognition tasks.
Georg Heigold, Simon Wiesler, Markus Nußbaum-Thom, Patrick Lehnen, Ralf Schlüter, Hermann Ney
ICASSP1
2010 A discriminative splitting criterion for phonetic decision trees
abstract
Phonetic decision trees are a key concept in acoustic modeling for large vocabulary continuous speech recognition.Although discriminative training has become a major line of research in speech recognition and all state-of-the-art acoustic models are trained discriminatively, the conventional phonetic decision tree approach still relies on the maximum likelihood principle.In this paper we develop a splitting criterion based on the minimization of the classification error.An improvement of more than 10% relative over a discriminatively trained baseline system on the Wall Street Journal corpus suggests that the proposed approach is promising.
Simon Wiesler, Georg Heigold, Markus Nußbaum-Thom, Ralf Schlüter, Hermann Ney
INTERSPEECH2
2010 Object classification by fusing SVMs and Gaussian mixtures
Thomas Deselaers, Georg Heigold, Hermann Ney
Pattern Recognit.2
2009 Generalized likelihood ratio discriminant analysis
abstract
Linear Discriminant Analysis (LDA) has been established as an important means for dimension reduction and decorrelation in speech recognition. The major points of criticism of LDA are that it uses an ad hoc and non-discriminative training criterion, and that the estimation is performed in a separate preprocessing step. This paper presents a new discriminative training method for the estimation of (projecting) linear feature transforms. More precisely, the problem is formulated in the loglinear framework, resulting in a convex optimization problem. Experimental results are provided for a digit string recognition task to compare the performance and robustness of the proposed approach (in combination with ML or MMI optimized acoustic models) with conventional LDA. Also, first experiments for a large vocabulary task are presented.
Muhammad Ali Tahir, Georg Heigold, Christian Plahl, Ralf Schlüter, Hermann Ney
ASRU2
2009 Investigations on features for log-linear acoustic models in continuous speech recognition
abstract
Hidden Markov Models with Gaussian Mixture Models as emission probabilities (GHMMs) are the underlying structure of all state-of-the-art speech recognition systems. Using Gaussian mixture distributions follows the generative approach where the class-conditional probability is modeled, although for classification only the posterior probability is needed. Though being very successful in related tasks like Natural Language Processing (NLP), in speech recognition direct modeling of posterior probabilities with log-linear models has rarely been used and has not been applied successfully to continuous speech recognition. In this paper we report competitive results for a speech recognizer with a log-linear acoustic model on the Wall Street Journal corpus, a Large Vocabulary Continuous Speech Recognition (LVCSR) task. We trained this model from scratch, i.e. without relying on an existing GHMM system. Previously the use of data dependent sparse features for log-linear models has been proposed. We compare them with polynomial features and show that the combination of polynomial and data dependent sparse features leads to better results.
Simon Wiesler, Markus Nußbaum-Thom, Georg Heigold, Ralf Schlüter, Hermann Ney
ASRU3
2009 Modified MPE/MMI in a transducer-based framework
abstract
In this paper we show how common training criteria like for example MPE or MMI can be extended to incorporate a margin term. In addition, a transducer-based training implementation is presented, which covers a large variety of discriminative training criteria for ASR, including the standard MMI, MPE, and MCE criteria, as well as the modifications to these criteria presented here. The modified criteria are directly related with the conventional large margin formulation of SVMs. In the proposed approach, we can take advantage of the generalization guarantees of large margin classifiers while keeping the existing framework for the discriminative training, including the efficient algorithms for conventional MPE or MMI. On the conceptual side, this allows for a direct evaluation of the margin term. Finally, experimental results are presented for different large vocabulary continuous speech recognition tasks (one of which is trained on a very large amount of training data) using these modified criteria.
Georg Heigold, Ralf Schlüter, Hermann Ney
ICASSP1
2009 A flat direct model for speech recognition
abstract
We introduce a direct model for speech recognition that assumes an unstructured, i.e., flat text output. The flat model allows us to model arbitrary attributes and dependences of the output. This is different from the HMMs typically used for speech recognition. This conventional modeling approach is based on sequential data and makes rigid assumptions on the dependences. HMMs have proven to be convenient and appropriate for large vocabulary continuous speech recognition. Our task under consideration, however, is the Windows Live Search for Mobile (WLS4M) task. This is a cellphone application that allows users to interact with web-based information portals. In particular, the set of valid outputs can be considered discrete and finite (although probably large, i.e., unseen events are an issue). Hence, a flat direct model lends itself to this task, making the adding of different knowledge sources and dependences straightforward and cheap. Using e.g. HMM posterior, m-gram, and spotter features, significant improvements over the conventional HMM system were observed.
Georg Heigold, Geoffrey Zweig, Xiao Li 0006, Patrick Nguyen
ICASSP1
2009 Confidence-Based Discriminative Training for Model Adaptation in Offline Arabic Handwriting Recognition
abstract
We present a novel confidence-based discriminative training for model adaptation approach for an HMM based Arabic handwriting recognition system to handle different handwriting styles and their variations.Most current approaches are maximum-likelihood trained HMM systems and try to adapt their models to different writing styles using writer adaptive training, unsupervised clustering, or additional writer specific data.Discriminative training based on the maximum mutual information criterion is used to train writer independent handwriting models. For model adaptation during decoding, an unsupervised confidence-based discriminative training on a word and frame level within a two-pass decoding process is proposed. Additionally, the training criterion is extended to incorporate a margin term.The proposed methods are evaluated on the IFN/ENIT Arabic handwriting database, where the proposed novel adaptation approach can decrease the word-error-rate by 33% relative.
Philippe Dreuw, Georg Heigold, Hermann Ney
ICDAR2
2009 Optimizing CRFs for SLU tasks in various languages using modified training criteria
abstract
In this paper, we present improvements of our state-of-the-art concept tagger based on conditional random fields.Statistical models have been optimized for three tasks of varying complexity in three languages (French, Italian, and Polish).Modified training criteria have been investigated leading to small improvements.The respective corpora as well as parameter optimization results for all models are presented in detail.A comparison of the selected features between languages as well as a close look at the tuning of the regularization parameter is given.The experimental results show in what level the optimizations of the single systems are portable between languages.
Stefan Hahn, Patrick Lehnen, Georg Heigold, Hermann Ney
INTERSPEECH3
2009 Investigations on convex optimization using log-linear HMMs for digit string recognition
abstract
Discriminative methods are an important technique to refine the acoustic model in speech recognition.Conventional discriminative training is initialized with some baseline model and the parameters are re-estimated in a separate step.This approach has proven to be successful, but it includes many heuristics, approximations, and parameters to be tuned.This tuning involves much engineering and makes it difficult to reproduce and compare experiments.In contrast to the conventional training, convex optimization techniques provide a sound approach to estimate all model parameters from scratch.Such a straight approach hopefully dispense with additional heuristics, e.g.scaling of posteriors.This paper addresses the question how well this concept using log-linear models carries over to practice.Experimental results are reported for a digit string recognition task, which allows for the investigation of this issue without approximations.
Georg Heigold, David Rybach, Ralf Schlüter, Hermann Ney
INTERSPEECH1
2009 Development of the GALE 2008 Mandarin LVCSR system
abstract
This paper describes the current improvements of the RWTH Mandarin LVCSR system.We introduce vocal tract length normalization for the Gammatone features and present comparable results for Gammatone based feature extraction and classical feature extraction.In order to benefit from the huge amount of data of 1600h available in the GALE project we have trained the acoustic models up to 8M Gaussians.We present detailed character error rates for the different number of Gaussians.Different kinds of systems are developed and a two stage decoding framework is applied, which uses cross-adaptation and a subsequent lattice-based system combination.In addition to various acoustic front-ends, these systems use different kinds of neural network toneme posterior features.We present detailed recognition results of the development cycle and the different acoustic front-ends of the systems.Finally, we compare the ultimate evaluation system to our last years system and can report a 10% relative improvement.
Christian Plahl, Björn Hoffmeister, Georg Heigold, Jonas Lööf, Ralf Schlüter, Hermann Ney
INTERSPEECH3
2009 The RWTH aachen university open source speech recognition system
abstract
We announce the public availability of the RWTH Aachen University speech recognition toolkit.The toolkit includes state of the art speech recognition technology for acoustic model training and decoding.Speaker adaptation, speaker adaptive training, unsupervised training, a finite state automata library, and an efficient tree search decoder are notable components.Comprehensive documentation, example setups for training and recognition, and a tutorial are provided to support newcomers.
David Rybach, Christian Gollan, Georg Heigold, Björn Hoffmeister, Jonas Lööf, Ralf Schlüter, Hermann Ney
INTERSPEECH3
2008 A GIS-like training algorithm for log-linear models with hidden variables
abstract
Conditional random fields (CRFs) are often estimated using an entropy based criterion in combination with generalized iterative scaling (GIS). GIS offers, upon others, the immediate advantages that it is locally convergent, completely parameter free, and guarantees an improvement of the criterion in each step. GIS, however, is limited in two aspects. GIS cannot be applied when the model incorporates hidden variables, and it can only be applied to optimize the maximum mutual information criterion (MMI). Here, we extend the GIS algorithm to resolve these two limitations. The new approach allows for training log-linear models with hidden variables and optimizes discriminative training criteria different from maximum mutual information (MMI), including minimum phone error (MPE). The proposed GIS-like method shares the above-mentioned theoretical properties of GIS. The framework is tested for optical character recognition on the USPS task, and for speech recognition on the Sietill task for continuous digit string recognition.
Georg Heigold, Thomas Deselaers, Ralf Schlüter, Hermann Ney
ICASSP1
2008 Modified MMI/MPE: a direct evaluation of the margin in speech recognition
abstract
In this paper we show how common speech recognition training criteria such as the Minimum Phone Error criterion or the Maximum Mutual Information criterion can be extended to incorporate a margin term. Different margin-based training algorithms have been proposed to refine existing training algorithms for general machine learning problems. However, for speech recognition, some special problems have to be addressed and all approaches proposed either lack practical applicability or the inclusion of a margin term enforces significant changes to the underlying model, e.g. the optimization algorithm, the loss function, or the parameterization of the model. In our approach, the conventional training criteria are modified to incorporate a margin term. This allows us to do large-margin training in speech recognition using the same efficient algorithms for accumulation and optimization and to use the same software as for conventional discriminative training. We show that the proposed criteria are equivalent to Support Vector Machines with suitable smooth loss functions, approximating the non-smooth hinge loss function or the hard error (e.g. phone error). Experimental results are given for two different tasks: the rather simple digit string recognition task Sietill which severely suffers from overfitting and the large vocabulary European Parliament Plenary Sessions English task which is supposed to be dominated by the risk and the generalization does not seem to be such an issue.
Georg Heigold, Thomas Deselaers, Ralf Schlüter, Hermann Ney
ICML1
2008 SVMs, Gaussian mixtures, and their generative/discriminative fusion
abstract
We present a new technique that employs support vector machines and Gaussian mixture densities to create a generative/discriminative joint classifier. In the past, several approaches to fuse the advantages of generative and discriminative approaches were presented, often leading to improved robustness and recognition accuracy. The presented method directly fuses both approaches, effectively allowing to fully exploit the advantages of both. The fusion of SVMs and GMDs is done by representing SVMs in the framework of GMDs without changing the training and without changing the decision boundary. The new classifier is evaluated on four tasks from the UCI machine learning repository. It is shown that for the relatively rare cases where SVMs have problems, the combined method outperforms both individual ones.
Thomas Deselaers, Georg Heigold, Hermann Ney
ICPR2
2008 On the equivalence of Gaussian and log-linear HMMs
abstract
The acoustic models of conventional state-of-the-art speech recognition systems use generative Gaussian HMMs. In the past few years, discriminative models like for example Conditional Random Fields (CRFs) have been proposed to refine the acoustic models. CRFs directly model the class posteriors, the quantities of interest in recognition. CRFs are undirected models, and CRFs do not assume local normalization constraints as HMMs do. This paper addresses the issue to what extent such less restricted models add flexiblity to the model compared with the generative counterpart. This work extends our previous work in that it provides the technical details used for showing the equivalence of Gaussian and log-linear HMMs. The correctness of the proposed equivalence transformation for conditional probabilities is demonstrated on a simple concept tagging task.
Georg Heigold, Patrick Lehnen, Ralf Schlüter, Hermann Ney
INTERSPEECH1
2008 Recent improvements of the RWTH GALE Mandarin LVCSR system
abstract
This paper describes the current improvements of the RWTH Mandarin LVCSR system. We introduce a new reduced toneme set developed at RWTH. We are using different toneme sets and pronunciation lexica. For the purpose of discriminative training we will show a fast way to transform word lattices between systems using different toneme sets and pronunciation lexica. In addition to various acoustic front-ends, the current systems use different kinds of neural network toneme posterior features. While different kinds of systems are developed, a two stage decoding framework for combining these systems is applied. We show detailed recognition results of the development cycle of the systems. Finally, two methods to integrate tonal features are compared.
Christian Plahl, Björn Hoffmeister, Mei-Yuh Hwang, Danju Lu, Georg Heigold, Jonas Lööf, Ralf Schlüter, Hermann Ney
INTERSPEECH5
2007 Development of the 2007 RWTH Mandarin LVCSR system
abstract
This paper describes the development of the RWTH Mandarin LVCSR system. Different acoustic front-ends together with multiple system cross-adaptation are used in a two stage decoding framework. We describe the system in detail and present systematic recognition results. Especially, we compare a variety of approaches for cross-adapting to multiple systems. During the development we did a comparative study on different methods for integrating tone and phoneme posterior features. Furthermore, we apply lattice based consensus decoding and system combination methods. In these methods, the effect of minimizing character instead of word errors is compared. The final system obtains a character error rate of 17.7% on the GALE 2006 evaluation data.
Björn Hoffmeister, Christian Plahl, Peter Fritz, Georg Heigold, Jonas Lööf, Ralf Schlüter, Hermann Ney
ASRU4
2007 Speech recognition with state-based nearest neighbour classifiers
abstract
We present a system that uses nearest neighbour classification on the state level of the hidden Markov model. Common speech recognition systems nowadays use Gaussian mixtures with a very high number of densities. We propose to carry this idea to the extreme, such that each observation is a prototype of its own. This approach is well-known and widely used in other areas of pattern recognition and has some immediate advantages over other classification approaches, but has never been applied to speech recognition. We evaluate the proposed method on the SieTill corpus of continuous digit strings and on the large vocabulary EPPS English task. It is shown that nearest neighbour outperforms conventional systems when training data is sparse. Index Terms: automatic speech recognition, nearest neighbour classification, kernel densities
Thomas Deselaers, Georg Heigold, Hermann Ney
INTERSPEECH2
2007 On the equivalence of Gaussian HMM and Gaussian HMM-like hidden conditional random fields
abstract
In this work we show that Gaussian HMMs (GHMMs) are equivalent to GHMM-like Hidden Conditional Random Fields (HCRFs). Hence, improvements of HCRFs over GHMMs found in literature are not due to a refined acoustic modeling but rather come from the more robust formulation of the underlying optimization problem or spurious local optima. Conventional GHMMs are usually estimated with a criterion on segment level whereas hybrid approaches are based on a formulation of the criterion on frame level. In contrast to CRFs, these approaches do not provide scores or do not support more than two classes in a natural way. In this work we analyze these two classes of criteria and propose a refined frame based criterion, which is shown to be an approximation of the associated criterion on segment level. Experimental results concerning these issues are reported for the German digit string recognition task Sietill and the large vocabulary English European Parliament Plenary Sessions (EPPS) task. Index Terms: speech recognition, parameter estimation, maximum entropy methods
Georg Heigold, Ralf Schlüter, Hermann Ney
INTERSPEECH1
2007 The RWTH 2007 TC-STAR evaluation system for european English and Spanish
abstract
In this work, the RWTH automatic speech recognition systems developed for the third TC-STAR evaluation campaign 2007 are presented.The RWTH systems make systematic use of internal system combination, combining systems with differences in feature extraction, adaptation methods, and training data used.To take advantage of this, novel feature extraction methods were employed; this year saw the introduction of Gammatone features and MLP based phone posterior features.Further improvements were achieved using unsupervised training, and it is notable that these improvements were achieved using a fairly low amount of automatically transcribed data.Also contributing to the improvements over last year was the switch to MPE training, and the introduction of projecting SAT transforms.
Jonas Lööf, Christian Gollan, Stefan Hahn, Georg Heigold, Björn Hoffmeister, Christian Plahl, David Rybach, Ralf Schlüter, Hermann Ney
INTERSPEECH4
2006 The 2006 RWTH parliamentary speeches transcription system
abstract
In this work, investigations in the course of the developement of RWTH automatic speech recognition systems developed for the second TC-STAR evaluation campaign 2006 are presented.The systems were designed to transcribe parliamentary speeches taken from the European Parliament Plenary Sessions (EPPS) in European English and Spanish, as well as speeches from the Spanish Parliament.The RWTH systems apply a two pass search strategy with a fourgram one-pass decoder including a fast vocal tract length normalization variant as first pass.The systems further include several adaptation and normalization methods, minimum classification error trained models, and bayes risk minimization.For all relevant individual components contrastive results are presented on the EPPS Spanish and English data, including investigations which did not yet enter the evaluation systems.
Jonas Lööf, Maximilian Bisani, Christian Gollan, Georg Heigold, Björn Hoffmeister, Christian Plahl, Ralf Schlüter, Hermann Ney
INTERSPEECH4