EDBT 2026 Demo / reviewers in the wild / expert
Georg Heigold
dblp:46/2236
· DBLP profile ↗
54ranked-venue papers
19as first author
5since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 38 · 10 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 38 · 13 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
10 papers |
Video understanding and tracking · 23% Image recognition and object detection · 22% Representation and self-supervised learning · 19% | |
| Computer graphics and multimedia
3 papers |
Audio and music processing · 100% |
Topics — the 26 heaviest of 28, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Representation and self-supervised learning › representation learning
object-centric representation learning |
1.0 | 2 | 2022 | Conditional Object-Centric Learning from Video · ICLR 2022 Object-Centric Learning with Slot Attention · NeurIPS 2020 |
Machine learning › Deep learning architectures and training › transformer
vision transformer |
1.0 | 2 | 2021 | An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale · ICLR 2021 ViViT: A Video Vision Transformer · ICCV 2021 |
Computer vision › Image recognition and object detection
object localization |
0.7 | 1 | 2023 | Video OWL-ViT: Temporally-consistent open-world localization in video · ICCV 2023 |
Computer vision › Video understanding and tracking
object tracking |
0.7 | 1 | 2023 | Video OWL-ViT: Temporally-consistent open-world localization in video · ICCV 2023 |
Computer vision › Video understanding and tracking › video representation learning
object-centric video learning |
0.6 | 1 | 2022 | Conditional Object-Centric Learning from Video · ICLR 2022 |
Computer vision › Video understanding and tracking
video classification |
0.5 | 1 | 2021 | ViViT: A Video Vision Transformer · ICCV 2021 |
Computer vision › Image recognition and object detection
object discovery |
0.4 | 1 | 2020 | Object-Centric Learning with Slot Attention · NeurIPS 2020 |
Machine learning › Representation and self-supervised learning › representation learning › object-centric representation learning
slot attention |
0.4 | 1 | 2020 | Object-Centric Learning with Slot Attention · NeurIPS 2020 |
Computer vision › Image recognition and object detection › object discovery
unsupervised object discovery |
0.4 | 1 | 2020 | Object-Centric Learning with Slot Attention · NeurIPS 2020 |
Audio and music processing
speech recognition |
0.3 | 3 | 2013 | Investigations on an EM-Style Optimization Algorithm for Discriminative Training of HMMs · IEEE ACM Trans. Audio Speech Lang. Process. 2013 Equivalence of Generative and Log-Linear Models · IEEE Trans. Speech Audio Process. 2011 Optimization Algorithms and Applications for Speech and Language Processing · IEEE Trans. Speech Audio Process. 2013 |
Natural language and speech › Information extraction and text analysis › morphological analysis
morphological tagging |
0.3 | 1 | 2017 | Cross-lingual Character-Level Neural Morphological Tagging · EMNLP 2017 |
Audio and music processing › speech recognition
discriminative training |
0.2 | 1 | 2013 | Investigations on an EM-Style Optimization Algorithm for Discriminative Training of HMMs · IEEE ACM Trans. Audio Speech Lang. Process. 2013 |
Machine learning › Deep learning architectures and training
transformer |
0.1 | 1 | 2021 | An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale · ICLR 2021 |
Natural language and speech › Speech recognition and synthesis
automatic speech recognition |
0.1 | 1 | 2012 | WFST Enabled Solutions to ASR Problems: Beyond HMM Decoding · IEEE Trans. Speech Audio Process. 2012 |
Natural language and speech › Speech recognition and synthesis › acoustic modeling
discriminative acoustic model training |
0.1 | 1 | 2012 | WFST Enabled Solutions to ASR Problems: Beyond HMM Decoding · IEEE Trans. Speech Audio Process. 2012 |
Computer vision › Image recognition and object detection › character recognition
handwritten digit recognition |
0.1 | 1 | 2012 | Latent Log-Linear Models for Handwritten Digit Classification · IEEE Trans. Pattern Anal. Mach. Intell. 2012 |
Machine learning › Probabilistic and Bayesian machine learning › structured models
latent variable model |
0.1 | 1 | 2012 | Latent Log-Linear Models for Handwritten Digit Classification · IEEE Trans. Pattern Anal. Mach. Intell. 2012 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › exponential family
maximum entropy models |
0.1 | 1 | 2012 | Latent Log-Linear Models for Handwritten Digit Classification · IEEE Trans. Pattern Anal. Mach. Intell. 2012 |
Natural language and speech › Machine translation
system combination |
0.1 | 1 | 2012 | WFST Enabled Solutions to ASR Problems: Beyond HMM Decoding · IEEE Trans. Speech Audio Process. 2012 |
Natural language and speech › Speech recognition and synthesis › search and decoding
weighted finite-state transducer decoding |
0.1 | 1 | 2012 | WFST Enabled Solutions to ASR Problems: Beyond HMM Decoding · IEEE Trans. Speech Audio Process. 2012 |
Audio and music processing › speech recognition
acoustic modeling |
0.1 | 1 | 2011 | Equivalence of Generative and Log-Linear Models · IEEE Trans. Speech Audio Process. 2011 |
Machine learning › Transfer learning and domain adaptation
cross-lingual transfer |
0.1 | 1 | 2017 | Cross-lingual Character-Level Neural Morphological Tagging · EMNLP 2017 |
Natural language and speech › Speech recognition and synthesis
acoustic model training |
0.1 | 1 | 2008 | Modified MMI/MPE: a direct evaluation of the margin in speech recognition · ICML 2008 |
Natural language and speech › Speech recognition and synthesis › acoustic model training
discriminative training |
0.1 | 1 | 2008 | Modified MMI/MPE: a direct evaluation of the margin in speech recognition · ICML 2008 |
Natural language and speech › Speech recognition and synthesis › acoustic model training › discriminative training
large margin training |
0.1 | 1 | 2008 | Modified MMI/MPE: a direct evaluation of the margin in speech recognition · ICML 2008 |
Natural language and speech › Language models and text generation › decoding
minimum bayes risk decoding |
0.0 | 1 | 2012 | WFST Enabled Solutions to ASR Problems: Beyond HMM Decoding · IEEE Trans. Speech Audio Process. 2012 |
Methods — techniques the papers use, named apart from their topics
transformer decoder · 0.7open-vocabulary detection · 0.7image-text pretraining · 0.7slot attention · 0.6contrastive learning · 0.6spatiotemporal tokenization · 0.5self-attention · 0.5regularization · 0.5pre-trained image model · 0.5large-scale pretraining · 0.5extended baum-welch · 0.3sparse optimization · 0.2rprop · 0.2minimum phone error · 0.2maximum mutual information · 0.2expectation-maximization · 0.2baum-welch · 0.2GIS · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Massive Sound Embedding Benchmark (MSEB)abstractAudio is a critical component of multimodal perception, and any truly intelligent system must demonstrate a wide range of auditory capabilities. These capabilities include transcription, classification, retrieval, reasoning, segmentation, clustering, reranking, and reconstruction. Fundamentally, each task involves transforming a raw audio signal into a meaningful 'embedding'—be it a single vector, a sequence of continuous or discrete representations, or another structured form—which then serves as the basis for generating the task's final response. To accelerate progress towards robust machine auditory intelligence, we present the Massive Sound Embedding Benchmark (MSEB): an extensible framework designed to evaluate the auditory components of any multimodal system. In its first release, MSEB offers a comprehensive suite of eight core tasks, with more planned for the future, supported by diverse datasets, including the new, large-scale Simple Voice Questions (SVQ) dataset. Our initial experiments establish clear performance headrooms, highlighting the significant opportunity to improve real-world multimodal experiences where audio is a core signal. We encourage the research community to use MSEB to assess their algorithms and contribute to its growth. The library is publicly hosted at https://github.com/google-research/mseb. Georg Heigold, Ehsan Variani, Tom Bagby, Cyril Allauzen, Ji Ma 0004, Shankar Kumar, Michael Riley 0001 |
NeurIPS | 1 |
| 2023 | Video OWL-ViT: Temporally-consistent open-world localization in videoabstractWe present an architecture and a training recipe that adapts pretrained open-world image models to localization in videos. Understanding the open visual world (without being constrained by fixed label spaces) is crucial for many real-world vision tasks. Contrastive pre-training on large image-text datasets has recently led to significant improvements for image-level tasks. For more structured tasks involving object localization applying pre-trained models is more challenging. This is particularly true for video tasks, where task-specific data is limited. We show successful transfer of open-world models by building on the OWL-ViT open-vocabulary detection model and adapting it to video by adding a transformer decoder. The decoder propagates object representations recurrently through time by using the output tokens for one frame as the object queries for the next. Our model is end-to-end trainable on video data and enjoys improved temporal consistency compared to tracking-by-detection baselines, while retaining the open-world capabilities of the backbone detector. We evaluate our model on the challenging TAO-OW benchmark and demonstrate that open-world capabilities, learned from large-scale image-text pretraining, can be transferred successfully to open-world localization across diverse videos. Georg Heigold, Daniel Keysers, Matthias Minderer, Mario Lucic, Alexey A. Gritsenko, Fisher Yu 0001, Alex Bewley, Thomas Kipf |
ICCV | 1 |
| 2022 | Conditional Object-Centric Learning from Video
Thomas Kipf, Gamaleldin F. Elsayed, Aravindh Mahendran, Austin Stone, Sara Sabour, Georg Heigold, Rico Jonschkowski, Alexey Dosovitskiy, Klaus Greff |
ICLR | 6 |
| 2021 | ViViT: A Video Vision TransformerabstractWe present pure-transformer based models for video classification, drawing upon the recent success of such models in image classification. Our model extracts spatiotemporal tokens from the input video, which are then encoded by a series of transformer layers. In order to handle the long sequences of tokens encountered in video, we propose several, efficient variants of our model which factorise the spatial- and temporal-dimensions of the input. Although transformer-based models are known to only be effective when large training datasets are available, we show how we can effectively regularise the model during training and leverage pretrained image models to be able to train on comparatively small datasets. We conduct thorough ablation studies, and achieve state-of-the-art results on multiple video classification benchmarks including Kinetics 400 and 600, Epic Kitchens, Something-Something v2 and Moments in Time, outperforming prior methods based on deep 3D convolutional networks. Anurag Arnab, Mostafa Dehghani 0001, Georg Heigold, Chen Sun 0002, Mario Lucic, Cordelia Schmid |
ICCV | 3 |
| 2021 | An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov 0003, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani 0001, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, Neil Houlsby |
ICLR | 9 |
| 2020 | Object-Centric Learning with Slot AttentionabstractLearning object-centric representations of complex scenes is a promising step towards enabling efficient abstract reasoning from low-level perceptual features. Yet, most deep learning approaches learn distributed representations that do not capture the compositional properties of natural scenes. In this paper, we present the Slot Attention module, an architectural component that interfaces with perceptual representations such as the output of a convolutional neural network and produces a set of task-dependent abstract representations which we call slots. These slots are exchangeable and can bind to any object in the input by specializing through a competitive procedure over multiple rounds of attention. We empirically demonstrate that Slot Attention can extract object-centric representations that enable generalization to unseen compositions when trained on unsupervised object discovery and supervised property prediction tasks. Francesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran, Georg Heigold, Jakob Uszkoreit, Alexey Dosovitskiy, Thomas Kipf |
NeurIPS | 5 |
| 2017 | An Extensive Empirical Evaluation of Character-Based Morphological Tagging for 14 LanguagesabstractThis paper investigates neural characterbased morphological tagging for languages with complex morphology and large tag sets.Character-based approaches are attractive as they can handle rarelyand unseen words gracefully.We evaluate on 14 languages and observe consistent gains over a state-of-the-art morphological tagger across all languages except for English and French, where we match the state-of-the-art.We compare two architectures for computing characterbased word vectors using recurrent (RNN) and convolutional (CNN) nets.We show that the CNN based approach performs slightly worse and less consistently than the RNN based approach.Small but systematic gains are observed when combining the two architectures by ensembling. Georg Heigold, Günter Neumann, Josef van Genabith |
EACL (1) | 1 |
| 2017 | Cross-lingual Character-Level Neural Morphological TaggingabstractEven for common NLP tasks, sufficient supervision is not available in many languages-morphological tagging is no exception.In the work presented here, we explore a transfer learning scheme, whereby we train character-level recurrent neural taggers to predict morphological taggings for high-resource languages and low-resource languages together.Learning joint character representations among multiple related languages successfully enables knowledge transfer from the high-resource languages to the low-resource ones, improving accuracy by up to 30%. Ryan Cotterell, Georg Heigold |
EMNLP | 2 |
| 2016 | Scaling character-based morphological tagging to fourteen languagesabstractThis paper investigates neural character-based morphological tagging for languages with complex morphology and large tag sets. Character-based approaches are attractive as they can handle rarely- and unseen words gracefully. More specifically, beside a rich morphology, non-canonical language, change of language or other linguistic variability can heavily degrade the accuracy of natural language processing of web and CMC data. We evaluate on 14 languages and observe consistent gains over a state-of-the-art morphological tagger across all languages except for English and French, where we match the state-of-the-art. The gains are clearly correlated with the amount of training data. We present supplementary experiments to explore whether and to what extent unsupervised data through pre-trained word vectors can compensate for limited amounts of supervised data. Moreover, we show preliminary results to study the effect of noisy input data by flipping characters at random. Georg Heigold, Josef van Genabith, Günter Neumann |
IEEE BigData | 1 |
| 2016 | End-to-end text-dependent speaker verificationabstractIn this paper we present a data-driven, integrated approach to speaker verification, which maps a test utterance and a few reference utterances directly to a single score for verification and jointly optimizes the system's components using the same evaluation protocol and metric as at test time. Such an approach will result in simple and efficient systems, requiring little domain-specific knowledge and making few model assumptions. We implement the idea by formulating the problem as a single neural network architecture, including the estimation of a speaker model on only a few utterances, and evaluate it on our internal "Ok Google" benchmark for text-dependent speaker verification. The proposed approach appears to be very effective for big data applications Like ours that require highly accurate, easy-to-maintain systems with a small footprint. Georg Heigold, Ignacio Moreno, Samy Bengio, Noam Shazeer |
ICASSP | 1 |
| 2015 | A Gaussian Mixture Model layer jointly optimized with discriminative features within a Deep Neural Network architectureabstractThis article proposes and evaluates a Gaussian Mixture Model (GMM) represented as the last layer of a Deep Neural Network (DNN) architecture and jointly optimized with all previous layers using Asynchronous Stochastic Gradient Descent (ASGD). The resulting “Deep GMM” architecture was investigated with special attention to the following issues: (1) The extent to which joint optimization improves over separate optimization of the DNN-based feature extraction layers and the GMM layer; (2) The extent to which depth (measured in number of layers, for a matched total number of parameters) helps a deep generative model based on the GMM layer, compared to a vanilla DNN model; (3) Head-to-head performance of Deep GMM architectures vs. equivalent DNN architectures of comparable depth, using the same optimization criterion (frame-level Cross Entropy (CE)) and optimization method (ASGD); (4) Expanded possibilities for modeling offered by the Deep GMM generative model. The proposed Deep GMMs were found to yield Word Error Rates (WERs) competitive with state-of-the-art DNN systems, at the cost of pre-training using standard DNNs to initialize the Deep GMM feature extraction layers. An extension to Deep Subspace GMMs is described, resulting in additional gains. Ehsan Variani, Erik McDermott, Georg Heigold |
ICASSP | 3 |
| 2014 | Small-footprint keyword spotting using deep neural networksabstractOur application requires a keyword spotting system with a small memory footprint, low computational cost, and high precision. To meet these requirements, we propose a simple approach based on deep neural networks. A deep neural network is trained to directly predict the keyword(s) or subword units of the keyword(s) followed by a posterior handling method producing a final confidence score. Keyword recognition results achieve 45% relative improvement with respect to a competitive Hidden Markov Model-based system, while performance in the presence of babble noise shows 39% relative improvement. Guoguo Chen, Carolina Parada, Georg Heigold |
ICASSP | 3 |
| 2014 | Asynchronous stochastic optimization for sequence training of deep neural networksabstractThis paper explores asynchronous stochastic optimization for sequence training of deep neural networks. Sequence training requires more computation than frame-level training using pre-computed frame data. This leads to several complications for stochastic optimization, arising from significant asynchrony in model updates under massive parallelization, and limited data shuffling due to utterance-chunked processing. We analyze the impact of these two issues on the efficiency and performance of sequence training. In particular, we suggest a framework to formalize the reasoning about the asynchrony and present experimental results on both small and large scale Voice Search tasks to validate the effectiveness and efficiency of asynchronous stochastic optimization. Georg Heigold, Erik McDermott, Vincent Vanhoucke, Andrew W. Senior, Michiel Bacchiani |
ICASSP | 1 |
| 2014 | GMM-free DNN acoustic model trainingabstractWhile deep neural networks (DNNs) have become the dominant acoustic model (AM) for speech recognition systems, they are still dependent on Gaussian mixture models (GMMs) for alignments both for supervised training and for context dependent (CD) tree building. Here we explore bootstrapping DNN AM training without GMM AMs and show that CD trees can be built with DNN alignments which are better matched to the DNN model and its features. We show that these trees and alignments result in better models than from the GMM alignments and trees. By removing the GMM acoustic model altogether we simplify the system required to train a DNN from scratch. Andrew W. Senior, Georg Heigold, Michiel Bacchiani, Hank Liao |
ICASSP | 2 |
| 2014 | Asynchronous, online, GMM-free training of a context dependent acoustic model for speech recognitionabstractWe propose an algorithm that allows online training of a con-text dependent DNN model. It designs a state inventory based on DNN features and jointly optimizes the DNN parameters and alignment of the training data. The process allows flat starting a model from scratch and avoids any dependency on a GMM/HMM model to bootstrap the training process. A 15k state model trained with the proposed algorithm reduced the er-ror rate on a mobile speech task with 24 % compared to a system bootstrapped from a CI HMM/GMM and with 16 % compared to a system bootstrapped from a CD HMM/GMM system. Index Terms: Deep Neural Networks, online training 1. Michiel Bacchiani, Andrew W. Senior, Georg Heigold |
INTERSPEECH | 3 |
| 2014 | Word embeddings for speech recognitionabstractSpeech recognition systems have used the concept of states as a way to decompose words into sub-word units for decades. As the number of such states now reaches the number of words used to train acoustic models, it is interesting to consider ap-proaches that relax the assumption that words are made of states. We present here an alternative construction, where words are projected into a continuous embedding space where words that sound alike are nearby in the Euclidean sense. We show how embeddings can still allow to score words that were not in the training dictionary. Initial experiments using a lattice rescor-ing approach and model combination on a large realistic dataset show improvements in word error rate. Index Terms: embeddings, deep learning, speech recognition. 1. Samy Bengio, Georg Heigold |
INTERSPEECH | 2 |
| 2014 | Asynchronous stochastic optimization for sequence training of deep neural networks: towards big dataabstractPrevious work presented a proof of concept for sequence training of deep neural networks (DNNs) using asynchronous stochastic optimization, mainly focusing on a small-scale task. The approach offers the potential to leverage both the efficiency of stochastic gradient descent and the scalability of parallel computation. This study presents results for four different voice search tasks to confirm the effectiveness and efficiency of the proposed framework across different conditions: amount of data (from 60 hours to 20,000 hours), type of speech (read speech vs. spontaneous speech), quality of data (supervised vs. unsupervised data), and language. Significant gains over baselines (DNNs trained at the frame level) are found to hold across these conditions. The experimental results are analyzed, and additional practical details for the approach are provided. Furthermore, different sequence training criteria are compared. Erik McDermott, Georg Heigold, Pedro J. Moreno 0001, Andrew W. Senior, Michiel Bacchiani |
INTERSPEECH | 2 |
| 2014 | Sequence discriminative distributed training of long short-term memory recurrent neural networks
Hasim Sak, Oriol Vinyals, Georg Heigold, Andrew W. Senior, Erik McDermott, Rajat Monga, Mark Z. Mao |
INTERSPEECH | 3 |
| 2013 | Multilingual acoustic models using distributed deep neural networksabstractToday's speech recognition technology is mature enough to be useful for many practical applications. In this context, it is of paramount importance to train accurate acoustic models for many languages within given resource constraints such as data, processing power, and time. Multilingual training has the potential to solve the data issue and close the performance gap between resource-rich and resource-scarce languages. Neural networks lend themselves naturally to parameter sharing across languages, and distributed implementations have made it feasible to train large networks. In this paper, we present experimental results for cross- and multi-lingual network training of eleven Romance languages on 10k hours of data in total. The average relative gains over the monolingual baselines are 4%/2% (data-scarce/data-rich languages) for cross- and 7%/2% for multi-lingual training. However, the additional gain from jointly training the languages on all data comes at an increased training time of roughly four weeks, compared to two weeks (monolingual) and one week (crosslingual). Georg Heigold, Vincent Vanhoucke, Andrew W. Senior, Patrick Nguyen, Marc'Aurelio Ranzato, Matthieu Devin, Jeffrey Dean |
ICASSP | 1 |
| 2013 | Deep neural networks with auxiliary Gaussian mixture models for real-time speech recognitionabstractWe present a framework that improves real-time speech recognition performance using deep neural networks (DNNs) with auxiliary Gaussian mixture models (GMMs). The DNNs and the auxiliary GMMs share the same hidden Markov model (HMM) state inventory. First, online incremental feature-space adaptation is performed using the GMM acoustic model. The speaker-adapted features are used to improve the recognition performance of both GMM and DNN models. Second, the acoustic scores from GMMs and DNN are combined at the state-level during decoding. Experiments on a large vocabulary speech recognition task show that both approaches improve recognition performance consistently and that the gains are mostly additive, resulting in about 5% relative improvement over the competitive DNN baseline in both Portuguese and English systems. Georg Heigold |
ICASSP | 3 |
| 2013 | An empirical study of learning rates in deep neural networks for speech recognitionabstractRecent deep neural network systems for large vocabulary speech recognition are trained with minibatch stochastic gradient descent but use a variety of learning rate scheduling schemes. We investigate several of these schemes, particularly AdaGrad. Based on our analysis of its limitations, we propose a new variant `AdaDec' that decouples long-term learning-rate scheduling from per-parameter learning rate variation. AdaDec was found to result in higher frame accuracies than other methods. Overall, careful choice of learning rate schemes leads to faster convergence and lower word error rates. Andrew W. Senior, Georg Heigold, Marc'Aurelio Ranzato |
ICASSP | 2 |
| 2013 | Multiframe deep neural networks for acoustic modelingabstractDeep neural networks have been shown to perform very well as acoustic models for automatic speech recognition. Compared to Gaussian mixtures however, they tend to be very expensive computationally, making them challenging to use in real-time applications. One key advantage of such neural networks is their ability to learn from very long observation windows going up to 400 ms. Given this very long temporal context, it is tempting to wonder whether one can run neural networks at a lower frame rate than the typical 10 ms, and whether there might be computational benefits to doing so. This paper describes a method of tying the neural network parameters over time which achieves comparable performance to the typical frame-synchronous model, while achieving up to a 4X reduction in the computational cost of the neural network activations. Vincent Vanhoucke, Matthieu Devin, Georg Heigold |
ICASSP | 3 |
| 2013 | Investigations on an EM-Style Optimization Algorithm for Discriminative Training of HMMsabstractToday's speech recognition systems are based on hidden Markov models (HMMs) with Gaussian mixture models whose parameters are estimated using a discriminative training criterion such as Maximum Mutual Information (MMI) or Minimum Phone Error (MPE). Currently, the optimization is almost always done with (empirical variants of) Extended Baum-Welch (EBW). This type of optimization requires sophisticated update schemes for the step sizes and a considerable amount of parameter tuning, and only little is known about its convergence behavior. In this paper, we derive an EM-style algorithm for discriminative training of HMMs. Like Expectation-Maximization (EM) for the generative training of HMMs, the proposed algorithm improves the training criterion on each iteration, converges to a local optimum, and is completely parameter-free. We investigate the feasibility of the proposed EM-style algorithm for discriminative training of two tasks, namely grapheme-to-phoneme conversion and spoken digit string recognition. Georg Heigold, Hermann Ney, Ralf Schlüter |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2013 | Optimization Algorithms and Applications for Speech and Language ProcessingabstractOptimization techniques have been used for many years in the formulation and solution of computational problems arising in speech and language processing. Such techniques are found in the Baum-Welch, extended Baum-Welch (EBW), Rprop, and GIS algorithms, for example. Additionally, the use of regularization terms has been seen in other applications of sparse optimization. This paper outlines a range of problems in which optimization formulations and algorithms play a role, giving some additional details on certain application problems in machine translation, speaker/language recognition, and automatic speech recognition. Several approaches developed in the speech and language processing communities are described in a way that makes them more recognizable as optimization procedures. Our survey is not exhaustive and is complemented by other papers in this volume. Stephen J. Wright 0001, Dimitri Kanevsky, Li Deng 0001, Xiaodong He 0001, Georg Heigold, Haizhou Li 0001 |
IEEE Trans. Speech Audio Process. | 5 |
| 2012 | Investigations on exemplar-based features for speech recognition towards thousands of hours of unsupervised, noisy dataabstractThe acoustic models in state-of-the-art speech recognition systems are based on phones in context that are represented by hidden Markov models. This modeling approach may be limited in that it is hard to incorporate long-span acoustic context. Exemplar-based approaches are an attractive alter-native, in particular if massive data and computational power are available. Yet, most of the data at Google are unsupervised and noisy. This paper investigates an exemplar-based approach under this yet not well understood data regime. A log-linear rescoring framework is used to combine the exemplar-based features on the word level with the first-pass model. This approach guarantees at least baseline performance and focuses on the refined modeling of words with sufficient data. Experimental results for the Voice Search and the YouTube tasks are presented. Georg Heigold, Patrick Nguyen, Mitch Weintraub, Vincent Vanhoucke |
ICASSP | 1 |
| 2012 | Overview of large scale optimization for discriminative training in speech recognitionabstractOver the past few decades, a variety of specialized approaches have been proposed to solve large problems in speech recognition. Conventional optimization techniques have not been widely applied, because the problems do not readily admit an objective for evaluating a given set of parameters and because of the large number of parameters. This situation is changing, due to recent developments in algorithmic optimization. In this paper, we review the specialized algorithms, including methods derived from the extended Baum-Welch (EBW) approach, Rprop, and GIS. We discuss optimization frameworks that could also potentially be applied, and outline some connections between the optimization methods and existing specialized methods. Dimitri Kanevsky, Georg Heigold, Stephen J. Wright 0001, Hermann Ney |
ICASSP | 2 |
| 2012 | Posterior-Scaled MPE: Novel Discriminative Training CriteriaabstractWe recently discovered novel discriminative training criteria following a principled approach. In this approach training criteria are developed from error bounds on the global error for pattern classification tasks that depend on non-trivial loss functions. Automatic speech recognition (ASR) is a prominent example for such a task depending on the non-trivial Levenshtein loss. In this context, the posterior-scaled Minimum Phoneme Error (MPE) training criterion, which is the state-of-the-art discriminative training criterion in ASR, was shown to be an approximation to one of the novel criteria. Here, we describe the implementation of the posterior-scaled MPE criterion in a transducer-based framework, and compare this criterion to other discriminative training criteria on an ASR task. This comparison indicates that the posterior-scaled MPE criterion performs better than other discriminative criteria including MPE. Index Terms: error bounds, discriminative training criteria, margin, MPE Markus Nußbaum-Thom, Zoltán Tüske, Georg Heigold, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 3 |
| 2012 | Latent Log-Linear Models for Handwritten Digit ClassificationabstractWe present latent log-linear models, an extension of log-linear models incorporating latent variables, and we propose two applications thereof: log-linear mixture models and image deformation-aware log-linear models. The resulting models are fully discriminative, can be trained efficiently, and the model complexity can be controlled. Log-linear mixture models offer additional flexibility within the log-linear modeling framework. Unlike previous approaches, the image deformation-aware model directly considers image deformations and allows for a discriminative training of the deformation parameters. Both are trained using alternating optimization. For certain variants, convergence to a stationary point is guaranteed and, in practice, even variants without this guarantee converge and find models that perform well. We tune the methods on the USPS data set and evaluate on the MNIST data set, demonstrating the generalization capabilities of our proposed models. Our models, although using significantly fewer parameters, are able to obtain competitive results with models proposed in the literature. Thomas Deselaers, Tobias Gass, Georg Heigold, Hermann Ney |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2012 | WFST Enabled Solutions to ASR Problems: Beyond HMM DecodingabstractDuring the last decade, weighted finite-state transducers (WFSTs) have become popular in speech recognition. While their main field of application remains hidden Markov model (HMM) decoding, the WFST framework is now also seen as a brick in solutions to many other central problems in automatic speech recognition (ASR). These solutions are less known, and this work aims at giving an overview of the applications of WFSTs in large-vocabulary continuous speech recognition (LVCSR) besides HMM decoding: discriminative acoustic model training, Bayes risk decoding, and system combination. The application of the WFST framework has a big practical impact: we show how the framework helps to structure problems, to develop generic solutions, and to delegate complex computations to WFST toolkits. In this paper, we review the literature, discuss existing approaches, and provide new insights into WFST enabled solutions. We also present a novel, purely WFST-based algorithm for computing the exact Bayes risk hypothesis from a lattice with the Levenshtein distance as loss function. We present the problems and their solutions in a unified framework and discuss the advantages and limits of using WFSTs. We do not provide new experimental results, but refer to the existing literature. Our work helps to identify where and how the transducer framework can contribute to a compact and generic solution to LVCSR problems. Björn Hoffmeister, Georg Heigold, David Rybach, Ralf Schlüter, Hermann Ney |
IEEE Trans. Speech Audio Process. | 2 |
| 2011 | EM-style optimization of hidden conditional random fields for grapheme-to-phoneme conversionabstractWe have recently proposed an EM-style algorithm to optimize log-linear models with hidden variables. In this paper, we use this algorithm to optimize a hidden conditional random field, i.e., a conditional random field with hidden variables. Similar to hidden Markov models, the alignments are the hidden variables in the examples considered. Here, EM-style algorithms are iterative optimization algorithms which are guaranteed to improve the training criterion in each iteration without the need for tuning step sizes, sophisticated update schemes or numerical line optimization (with hardly predictable complexity). This is a rather strong property which conventional gradient-based optimization algorithms do not have. We present experimental results for a grapheme-to-phoneme conversion task and compare the convergence behavior of the EM-style algorithm with L-BFGS and Rprop. Georg Heigold, Stefan Hahn, Patrick Lehnen, Hermann Ney |
ICASSP | 1 |
| 2011 | Confidence- and margin-based MMI/MPE discriminative training for off-line handwriting recognition
Philippe Dreuw, Georg Heigold, Hermann Ney |
Int. J. Document Anal. Recognit. | 2 |
| 2011 | Equivalence of Generative and Log-Linear ModelsabstractConventional speech recognition systems are based on hidden Markov models (HMMs) with Gaussian mixture models (GHMMs). Discriminative log-linear models are an alternative modeling approach and have been investigated recently in speech recognition. GHMMs are directed models with constraints, e.g., positivity of variances and normalization of conditional probabilities, while log-linear models do not use such constraints. This paper compares the posterior form of typical generative models related to speech recognition with their log-linear model counterparts. The key result will be the derivation of the equivalence of these two different approaches under weak assumptions. In particular, we study Gaussian mixture models, part-of-speech bigram tagging models, and eventually, the GHMMs. This result unifies two important but fundamentally different modeling paradigms in speech recognition on the functional level. Furthermore, this paper will present comparative experimental results for various speech tasks of different complexity, including a digit string and large-vocabulary continuous speech recognition tasks. Georg Heigold, Hermann Ney, Patrick Lehnen, Tobias Gass, Ralf Schlüter |
IEEE Trans. Speech Audio Process. | 1 |
| 2010 | Discriminative HMMS, log-linear models, and CRFS: What is the difference?abstractRecently, there have been many papers studying discriminative acoustic modeling techniques like conditional random fields or discriminative training of conventional Gaussian HMMs. This paper will give an overview of the recent work and progress. We will strictly distinguish between the type of acoustic models on the one hand and the training criterion on the other hand. We will address two issues in more detail: the relation between conventional Gaussian HMMs and conditional random fields and the advantages of formulating the training criterion as a convex optimization problem. Experimental results for various speech tasks will be presented to carefully evaluate the different concepts and approaches, including both a digit string and large vocabulary continuous speech recognition tasks. Georg Heigold, Simon Wiesler, Markus Nußbaum-Thom, Patrick Lehnen, Ralf Schlüter, Hermann Ney |
ICASSP | 1 |
| 2010 | A discriminative splitting criterion for phonetic decision treesabstractPhonetic decision trees are a key concept in acoustic modeling for large vocabulary continuous speech recognition.Although discriminative training has become a major line of research in speech recognition and all state-of-the-art acoustic models are trained discriminatively, the conventional phonetic decision tree approach still relies on the maximum likelihood principle.In this paper we develop a splitting criterion based on the minimization of the classification error.An improvement of more than 10% relative over a discriminatively trained baseline system on the Wall Street Journal corpus suggests that the proposed approach is promising. Simon Wiesler, Georg Heigold, Markus Nußbaum-Thom, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 2 |
| 2010 | Object classification by fusing SVMs and Gaussian mixtures
Thomas Deselaers, Georg Heigold, Hermann Ney |
Pattern Recognit. | 2 |
| 2009 | Generalized likelihood ratio discriminant analysisabstractLinear Discriminant Analysis (LDA) has been established as an important means for dimension reduction and decorrelation in speech recognition. The major points of criticism of LDA are that it uses an ad hoc and non-discriminative training criterion, and that the estimation is performed in a separate preprocessing step. This paper presents a new discriminative training method for the estimation of (projecting) linear feature transforms. More precisely, the problem is formulated in the loglinear framework, resulting in a convex optimization problem. Experimental results are provided for a digit string recognition task to compare the performance and robustness of the proposed approach (in combination with ML or MMI optimized acoustic models) with conventional LDA. Also, first experiments for a large vocabulary task are presented. Muhammad Ali Tahir, Georg Heigold, Christian Plahl, Ralf Schlüter, Hermann Ney |
ASRU | 2 |
| 2009 | Investigations on features for log-linear acoustic models in continuous speech recognitionabstractHidden Markov Models with Gaussian Mixture Models as emission probabilities (GHMMs) are the underlying structure of all state-of-the-art speech recognition systems. Using Gaussian mixture distributions follows the generative approach where the class-conditional probability is modeled, although for classification only the posterior probability is needed. Though being very successful in related tasks like Natural Language Processing (NLP), in speech recognition direct modeling of posterior probabilities with log-linear models has rarely been used and has not been applied successfully to continuous speech recognition. In this paper we report competitive results for a speech recognizer with a log-linear acoustic model on the Wall Street Journal corpus, a Large Vocabulary Continuous Speech Recognition (LVCSR) task. We trained this model from scratch, i.e. without relying on an existing GHMM system. Previously the use of data dependent sparse features for log-linear models has been proposed. We compare them with polynomial features and show that the combination of polynomial and data dependent sparse features leads to better results. Simon Wiesler, Markus Nußbaum-Thom, Georg Heigold, Ralf Schlüter, Hermann Ney |
ASRU | 3 |
| 2009 | Modified MPE/MMI in a transducer-based frameworkabstractIn this paper we show how common training criteria like for example MPE or MMI can be extended to incorporate a margin term. In addition, a transducer-based training implementation is presented, which covers a large variety of discriminative training criteria for ASR, including the standard MMI, MPE, and MCE criteria, as well as the modifications to these criteria presented here. The modified criteria are directly related with the conventional large margin formulation of SVMs. In the proposed approach, we can take advantage of the generalization guarantees of large margin classifiers while keeping the existing framework for the discriminative training, including the efficient algorithms for conventional MPE or MMI. On the conceptual side, this allows for a direct evaluation of the margin term. Finally, experimental results are presented for different large vocabulary continuous speech recognition tasks (one of which is trained on a very large amount of training data) using these modified criteria. Georg Heigold, Ralf Schlüter, Hermann Ney |
ICASSP | 1 |
| 2009 | A flat direct model for speech recognitionabstractWe introduce a direct model for speech recognition that assumes an unstructured, i.e., flat text output. The flat model allows us to model arbitrary attributes and dependences of the output. This is different from the HMMs typically used for speech recognition. This conventional modeling approach is based on sequential data and makes rigid assumptions on the dependences. HMMs have proven to be convenient and appropriate for large vocabulary continuous speech recognition. Our task under consideration, however, is the Windows Live Search for Mobile (WLS4M) task. This is a cellphone application that allows users to interact with web-based information portals. In particular, the set of valid outputs can be considered discrete and finite (although probably large, i.e., unseen events are an issue). Hence, a flat direct model lends itself to this task, making the adding of different knowledge sources and dependences straightforward and cheap. Using e.g. HMM posterior, m-gram, and spotter features, significant improvements over the conventional HMM system were observed. Georg Heigold, Geoffrey Zweig, Xiao Li 0006, Patrick Nguyen |
ICASSP | 1 |
| 2009 | Confidence-Based Discriminative Training for Model Adaptation in Offline Arabic Handwriting RecognitionabstractWe present a novel confidence-based discriminative training for model adaptation approach for an HMM based Arabic handwriting recognition system to handle different handwriting styles and their variations.Most current approaches are maximum-likelihood trained HMM systems and try to adapt their models to different writing styles using writer adaptive training, unsupervised clustering, or additional writer specific data.Discriminative training based on the maximum mutual information criterion is used to train writer independent handwriting models. For model adaptation during decoding, an unsupervised confidence-based discriminative training on a word and frame level within a two-pass decoding process is proposed. Additionally, the training criterion is extended to incorporate a margin term.The proposed methods are evaluated on the IFN/ENIT Arabic handwriting database, where the proposed novel adaptation approach can decrease the word-error-rate by 33% relative. Philippe Dreuw, Georg Heigold, Hermann Ney |
ICDAR | 2 |
| 2009 | Optimizing CRFs for SLU tasks in various languages using modified training criteriaabstractIn this paper, we present improvements of our state-of-the-art concept tagger based on conditional random fields.Statistical models have been optimized for three tasks of varying complexity in three languages (French, Italian, and Polish).Modified training criteria have been investigated leading to small improvements.The respective corpora as well as parameter optimization results for all models are presented in detail.A comparison of the selected features between languages as well as a close look at the tuning of the regularization parameter is given.The experimental results show in what level the optimizations of the single systems are portable between languages. Stefan Hahn, Patrick Lehnen, Georg Heigold, Hermann Ney |
INTERSPEECH | 3 |
| 2009 | Investigations on convex optimization using log-linear HMMs for digit string recognitionabstractDiscriminative methods are an important technique to refine the acoustic model in speech recognition.Conventional discriminative training is initialized with some baseline model and the parameters are re-estimated in a separate step.This approach has proven to be successful, but it includes many heuristics, approximations, and parameters to be tuned.This tuning involves much engineering and makes it difficult to reproduce and compare experiments.In contrast to the conventional training, convex optimization techniques provide a sound approach to estimate all model parameters from scratch.Such a straight approach hopefully dispense with additional heuristics, e.g.scaling of posteriors.This paper addresses the question how well this concept using log-linear models carries over to practice.Experimental results are reported for a digit string recognition task, which allows for the investigation of this issue without approximations. Georg Heigold, David Rybach, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 1 |
| 2009 | Development of the GALE 2008 Mandarin LVCSR systemabstractThis paper describes the current improvements of the RWTH Mandarin LVCSR system.We introduce vocal tract length normalization for the Gammatone features and present comparable results for Gammatone based feature extraction and classical feature extraction.In order to benefit from the huge amount of data of 1600h available in the GALE project we have trained the acoustic models up to 8M Gaussians.We present detailed character error rates for the different number of Gaussians.Different kinds of systems are developed and a two stage decoding framework is applied, which uses cross-adaptation and a subsequent lattice-based system combination.In addition to various acoustic front-ends, these systems use different kinds of neural network toneme posterior features.We present detailed recognition results of the development cycle and the different acoustic front-ends of the systems.Finally, we compare the ultimate evaluation system to our last years system and can report a 10% relative improvement. Christian Plahl, Björn Hoffmeister, Georg Heigold, Jonas Lööf, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 3 |
| 2009 | The RWTH aachen university open source speech recognition systemabstractWe announce the public availability of the RWTH Aachen University speech recognition toolkit.The toolkit includes state of the art speech recognition technology for acoustic model training and decoding.Speaker adaptation, speaker adaptive training, unsupervised training, a finite state automata library, and an efficient tree search decoder are notable components.Comprehensive documentation, example setups for training and recognition, and a tutorial are provided to support newcomers. David Rybach, Christian Gollan, Georg Heigold, Björn Hoffmeister, Jonas Lööf, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 3 |
| 2008 | A GIS-like training algorithm for log-linear models with hidden variablesabstractConditional random fields (CRFs) are often estimated using an entropy based criterion in combination with generalized iterative scaling (GIS). GIS offers, upon others, the immediate advantages that it is locally convergent, completely parameter free, and guarantees an improvement of the criterion in each step. GIS, however, is limited in two aspects. GIS cannot be applied when the model incorporates hidden variables, and it can only be applied to optimize the maximum mutual information criterion (MMI). Here, we extend the GIS algorithm to resolve these two limitations. The new approach allows for training log-linear models with hidden variables and optimizes discriminative training criteria different from maximum mutual information (MMI), including minimum phone error (MPE). The proposed GIS-like method shares the above-mentioned theoretical properties of GIS. The framework is tested for optical character recognition on the USPS task, and for speech recognition on the Sietill task for continuous digit string recognition. Georg Heigold, Thomas Deselaers, Ralf Schlüter, Hermann Ney |
ICASSP | 1 |
| 2008 | Modified MMI/MPE: a direct evaluation of the margin in speech recognitionabstractIn this paper we show how common speech recognition training criteria such as the Minimum Phone Error criterion or the Maximum Mutual Information criterion can be extended to incorporate a margin term. Different margin-based training algorithms have been proposed to refine existing training algorithms for general machine learning problems. However, for speech recognition, some special problems have to be addressed and all approaches proposed either lack practical applicability or the inclusion of a margin term enforces significant changes to the underlying model, e.g. the optimization algorithm, the loss function, or the parameterization of the model. In our approach, the conventional training criteria are modified to incorporate a margin term. This allows us to do large-margin training in speech recognition using the same efficient algorithms for accumulation and optimization and to use the same software as for conventional discriminative training. We show that the proposed criteria are equivalent to Support Vector Machines with suitable smooth loss functions, approximating the non-smooth hinge loss function or the hard error (e.g. phone error). Experimental results are given for two different tasks: the rather simple digit string recognition task Sietill which severely suffers from overfitting and the large vocabulary European Parliament Plenary Sessions English task which is supposed to be dominated by the risk and the generalization does not seem to be such an issue. Georg Heigold, Thomas Deselaers, Ralf Schlüter, Hermann Ney |
ICML | 1 |
| 2008 | SVMs, Gaussian mixtures, and their generative/discriminative fusionabstractWe present a new technique that employs support vector machines and Gaussian mixture densities to create a generative/discriminative joint classifier. In the past, several approaches to fuse the advantages of generative and discriminative approaches were presented, often leading to improved robustness and recognition accuracy. The presented method directly fuses both approaches, effectively allowing to fully exploit the advantages of both. The fusion of SVMs and GMDs is done by representing SVMs in the framework of GMDs without changing the training and without changing the decision boundary. The new classifier is evaluated on four tasks from the UCI machine learning repository. It is shown that for the relatively rare cases where SVMs have problems, the combined method outperforms both individual ones. Thomas Deselaers, Georg Heigold, Hermann Ney |
ICPR | 2 |
| 2008 | On the equivalence of Gaussian and log-linear HMMsabstractThe acoustic models of conventional state-of-the-art speech recognition systems use generative Gaussian HMMs. In the past few years, discriminative models like for example Conditional Random Fields (CRFs) have been proposed to refine the acoustic models. CRFs directly model the class posteriors, the quantities of interest in recognition. CRFs are undirected models, and CRFs do not assume local normalization constraints as HMMs do. This paper addresses the issue to what extent such less restricted models add flexiblity to the model compared with the generative counterpart. This work extends our previous work in that it provides the technical details used for showing the equivalence of Gaussian and log-linear HMMs. The correctness of the proposed equivalence transformation for conditional probabilities is demonstrated on a simple concept tagging task. Georg Heigold, Patrick Lehnen, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 1 |
| 2008 | Recent improvements of the RWTH GALE Mandarin LVCSR systemabstractThis paper describes the current improvements of the RWTH Mandarin LVCSR system. We introduce a new reduced toneme set developed at RWTH. We are using different toneme sets and pronunciation lexica. For the purpose of discriminative training we will show a fast way to transform word lattices between systems using different toneme sets and pronunciation lexica. In addition to various acoustic front-ends, the current systems use different kinds of neural network toneme posterior features. While different kinds of systems are developed, a two stage decoding framework for combining these systems is applied. We show detailed recognition results of the development cycle of the systems. Finally, two methods to integrate tonal features are compared. Christian Plahl, Björn Hoffmeister, Mei-Yuh Hwang, Danju Lu, Georg Heigold, Jonas Lööf, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 5 |
| 2007 | Development of the 2007 RWTH Mandarin LVCSR systemabstractThis paper describes the development of the RWTH Mandarin LVCSR system. Different acoustic front-ends together with multiple system cross-adaptation are used in a two stage decoding framework. We describe the system in detail and present systematic recognition results. Especially, we compare a variety of approaches for cross-adapting to multiple systems. During the development we did a comparative study on different methods for integrating tone and phoneme posterior features. Furthermore, we apply lattice based consensus decoding and system combination methods. In these methods, the effect of minimizing character instead of word errors is compared. The final system obtains a character error rate of 17.7% on the GALE 2006 evaluation data. Björn Hoffmeister, Christian Plahl, Peter Fritz, Georg Heigold, Jonas Lööf, Ralf Schlüter, Hermann Ney |
ASRU | 4 |
| 2007 | Speech recognition with state-based nearest neighbour classifiersabstractWe present a system that uses nearest neighbour classification on the state level of the hidden Markov model. Common speech recognition systems nowadays use Gaussian mixtures with a very high number of densities. We propose to carry this idea to the extreme, such that each observation is a prototype of its own. This approach is well-known and widely used in other areas of pattern recognition and has some immediate advantages over other classification approaches, but has never been applied to speech recognition. We evaluate the proposed method on the SieTill corpus of continuous digit strings and on the large vocabulary EPPS English task. It is shown that nearest neighbour outperforms conventional systems when training data is sparse. Index Terms: automatic speech recognition, nearest neighbour classification, kernel densities Thomas Deselaers, Georg Heigold, Hermann Ney |
INTERSPEECH | 2 |
| 2007 | On the equivalence of Gaussian HMM and Gaussian HMM-like hidden conditional random fieldsabstractIn this work we show that Gaussian HMMs (GHMMs) are equivalent to GHMM-like Hidden Conditional Random Fields (HCRFs). Hence, improvements of HCRFs over GHMMs found in literature are not due to a refined acoustic modeling but rather come from the more robust formulation of the underlying optimization problem or spurious local optima. Conventional GHMMs are usually estimated with a criterion on segment level whereas hybrid approaches are based on a formulation of the criterion on frame level. In contrast to CRFs, these approaches do not provide scores or do not support more than two classes in a natural way. In this work we analyze these two classes of criteria and propose a refined frame based criterion, which is shown to be an approximation of the associated criterion on segment level. Experimental results concerning these issues are reported for the German digit string recognition task Sietill and the large vocabulary English European Parliament Plenary Sessions (EPPS) task. Index Terms: speech recognition, parameter estimation, maximum entropy methods Georg Heigold, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 1 |
| 2007 | The RWTH 2007 TC-STAR evaluation system for european English and SpanishabstractIn this work, the RWTH automatic speech recognition systems developed for the third TC-STAR evaluation campaign 2007 are presented.The RWTH systems make systematic use of internal system combination, combining systems with differences in feature extraction, adaptation methods, and training data used.To take advantage of this, novel feature extraction methods were employed; this year saw the introduction of Gammatone features and MLP based phone posterior features.Further improvements were achieved using unsupervised training, and it is notable that these improvements were achieved using a fairly low amount of automatically transcribed data.Also contributing to the improvements over last year was the switch to MPE training, and the introduction of projecting SAT transforms. Jonas Lööf, Christian Gollan, Stefan Hahn, Georg Heigold, Björn Hoffmeister, Christian Plahl, David Rybach, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 4 |
| 2006 | The 2006 RWTH parliamentary speeches transcription systemabstractIn this work, investigations in the course of the developement of RWTH automatic speech recognition systems developed for the second TC-STAR evaluation campaign 2006 are presented.The systems were designed to transcribe parliamentary speeches taken from the European Parliament Plenary Sessions (EPPS) in European English and Spanish, as well as speeches from the Spanish Parliament.The RWTH systems apply a two pass search strategy with a fourgram one-pass decoder including a fast vocal tract length normalization variant as first pass.The systems further include several adaptation and normalization methods, minimum classification error trained models, and bayes risk minimization.For all relevant individual components contrastive results are presented on the EPPS Spanish and English data, including investigations which did not yet enter the evaluation systems. Jonas Lööf, Maximilian Bisani, Christian Gollan, Georg Heigold, Björn Hoffmeister, Christian Plahl, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 4 |