Tomás Mikolov

dblp:45/8055 · DBLP profile ↗
← Back
35ranked-venue papers
11as first author
4since 2021 · last 2023
0000-0002-6938-5426ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 28 · 8 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 7 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
10 papers
Representation and self-supervised learning · 24% Deep learning architectures and training · 18% Machine translation · 10%
Databases, data mining, and information retrieval
1 paper
Data integration and cleaning · 50% Knowledge graphs · 50%

Topics — the 20 heaviest of 22, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training
recurrent neural network
0.522017
Variable Computation in Recurrent Neural Networks · ICLR (Poster) 2017
Inferring Algorithmic Patterns with Stack-Augmented Recurrent Nets · NIPS 2015
Knowledge graphs › knowledge graph embedding
entity embedding
0.412019
Place Deduplication with Embeddings · WWW 2019
Data integration and cleaning
entity resolution
0.412019
Place Deduplication with Embeddings · WWW 2019
Machine learning › Representation and self-supervised learning › word representation › word embedding
bilingual word embedding
0.312018
Loss in Translation: Learning Bilingual Word Mapping with a Retrieval Criterion · EMNLP 2018
Machine learning › Learning paradigms › supervised learning
multimodal classification
0.312018
Efficient Large-Scale Multi-Modal Classification · AAAI 2018
Natural language and speech › Machine translation
word translation
0.312018
Loss in Translation: Learning Bilingual Word Mapping with a Retrieval Criterion · EMNLP 2018
Robotics › Motion planning and robot control › robot control
neural network controller
0.212016
Learning Simple Algorithms from Examples · ICML 2016
Machine learning › Reinforcement learning › value-based reinforcement learning
q-learning
0.212016
Learning Simple Algorithms from Examples · ICML 2016
Machine learning › Representation and self-supervised learning › word representation
distributed representation
0.222014
Distributed Representations of Sentences and Documents · ICML 2014
Distributed Representations of Words and Phrases and their Compositionality · NIPS 2013
Machine learning › Learning theory › online learning
sequence prediction
0.212015
Inferring Algorithmic Patterns with Stack-Augmented Recurrent Nets · NIPS 2015
Machine learning › Representation and self-supervised learning › text embedding
document embedding
0.212014
Distributed Representations of Sentences and Documents · ICML 2014
Natural language and speech › Language models and text generation
text representation
0.212014
Distributed Representations of Sentences and Documents · ICML 2014
Machine learning › Deep learning architectures and training › recurrent neural network
recurrent neural network training
0.212013
On the difficulty of training recurrent neural networks · ICML (3) 2013
Machine learning › Representation and self-supervised learning › word representation
skip-gram
0.212013
Distributed Representations of Words and Phrases and their Compositionality · NIPS 2013
Machine learning › Deep learning architectures and training › training dynamics
vanishing and exploding gradients
0.212013
On the difficulty of training recurrent neural networks · ICML (3) 2013
Computer vision › Vision and language › cross-modal alignment
visual-semantic embedding
0.212013
DeViSE: A Deep Visual-Semantic Embedding Model · NIPS 2013
Machine learning › Representation and self-supervised learning › word representation
word embedding
0.212013
Distributed Representations of Words and Phrases and their Compositionality · NIPS 2013
Machine learning › Transfer learning and domain adaptation › zero-shot learning
zero-shot classification
0.212013
DeViSE: A Deep Visual-Semantic Embedding Model · NIPS 2013
Natural language and speech › Language models and text generation
language modeling
0.112011
A Fast Re-scoring Strategy to Capture Long-Distance Dependencies · EMNLP 2011
Natural language and speech › Machine translation
statistical machine translation
0.112011
A Fast Re-scoring Strategy to Capture Long-Distance Dependencies · EMNLP 2011

Methods — techniques the papers use, named apart from their topics

graph embedding · 0.4data-driven pipeline · 0.4retrieval criterion · 0.3orthogonal matrix alignment · 0.3multimodal fusion · 0.3feature discretization · 0.3convolutional neural network · 0.3adaptive computation · 0.3q-learning · 0.2neural network · 0.2trainable memory · 0.2stack augmentation · 0.2
YearPublicationVenuePosition
2023 Preserving Semantics in Textual Adversarial Attacks
abstract
The growth of hateful online content, or hate speech, has been associated with a global increase in violent crimes against minorities [23]. Harmful online content can be produced easily, automatically and anonymously. Even though, some form of auto-detection is already achieved through text classifiers in NLP, they can be fooled by adversarial attacks. To strengthen existing systems and stay ahead of attackers, we need better adversarial attacks. In this paper, we show that up to 70% of adversarial examples generated by adversarial attacks should be discarded because they do not preserve semantics. We address this core weakness and propose a new, fully supervised sentence embedding technique called Semantics-Preserving-Encoder (SPE). Our method outperforms existing sentence encoders used in adversarial attacks by achieving 1.2× ∼ 5.1× better real attack success rate. We release our code as a plugin that can be used in any existing adversarial attack to improve its quality and speed up its execution. (The code, datasets and test examples are available at https://github.com/DavidHerel/semantics-preserving-encoder.)
David Herel, Hugo Cisneros, Tomás Mikolov
ECAI3
2022 Classification of Discrete Dynamical Systems Based on Transients
abstract
In order to develop systems capable of artificial evolution, we need to identify which systems can produce complex behavior. We present a novel classification method applicable to any class of deterministic discrete space and time dynamical systems. The method is based on classifying the asymptotic behavior of the average computation time in a given system before entering a loop. We were able to identify a critical region of behavior that corresponds to a phase transition from ordered behavior to chaos across various classes of dynamical systems. To show that our approach can be applied to many different computational systems, we demonstrate the results of classifying cellular automata, Turing machines, and random Boolean networks. Further, we use this method to classify 2D cellular automata to automatically find those with interesting, complex dynamics. We believe that our work can be used to design systems in which complex structures emerge. Also, it can be used to compare various versions of existing attempts to model open-ended evolution (Channon, 2006; Ofria & Wilke, 2004; Ray, 1991).
Barbora Hudcová, Tomás Mikolov
Artif. Life2
2022 Emergence of Self-Reproducing Metabolisms as Recursive Algorithms in an Artificial Chemistry
abstract
One of the main goals of Artificial Life is to research the conditions for the emergence of life, not necessarily as it is, but as it could be. Artificial chemistries are one of the most important tools for this purpose because they provide us with a basic framework to investigate under which conditions metabolisms capable of reproducing themselves, and ultimately, of evolving, can emerge. While there have been successful attempts at producing examples of emergent self-reproducing metabolisms, the set of rules involved remain too complex to shed much light on the underlying principles at work. In this article, we hypothesize that the key property needed for self-reproducing metabolisms to emerge is the existence of an autocatalyzed subset of Turing-complete reactions. We validate this hypothesis with a minimalistic artificial chemistry with conservation laws, which is based on a Turing-complete rewriting system called combinatory logic. Our experiments show that a single run of this chemistry, starting from a tabula rasa state, discovers-with no external intervention-a wide range of emergent structures including ones that self-reproduce in each cycle. All of these structures take the form of recursive algorithms that acquire basic constituents from the environment and decompose them in a process that is remarkably similar to biological metabolisms.
Germán Kruszewski, Tomás Mikolov
Artif. Life2
2021 Language Modeling and Artificial Intelligence
Tomás Mikolov
Interspeech1
2019 Place Deduplication with Embeddings
abstract
Thanks to the advancing mobile location services, people nowadays can post about places to share visiting experience on-the-go. A large place graph not only helps users explore interesting destinations, but also provides opportunities for understanding and modeling the real world. To improve coverage and flexibility of the place graph, many platforms import places data from multiple sources, which unfortunately leads to the emergence of numerous duplicated places that severely hinder subsequent location-related services. In this work, we take the anonymous place graph from Facebook as an example to systematically study the problem of place deduplication: We carefully formulate the problem, study its connections to various related tasks that lead to several promising basic models, and arrive at a systematic two-step data-driven pipeline based on place embedding with multiple novel techniques that works significantly better than the state-of-the-art.
Carl Yang 0001, Do Huy Hoang, Tomás Mikolov, Jiawei Han 0001
WWW3
2018 Efficient Large-Scale Multi-Modal Classification
abstract
While the incipient internet was largely text-based, the modern digital world is becoming increasingly multi-modal. Here, we examine multi-modal classification where one modality is discrete, e.g. text, and the other is continuous, e.g. visual representations transferred from a convolutional neural network. In particular, we focus on scenarios where we have to be able to classify large quantities of data quickly. We investigate various methods for performing multi-modal fusion and analyze their trade-offs in terms of classification accuracy and computational efficiency. Our findings indicate that the inclusion of continuous information improves performance over text-only on a range of multi-modal classification tasks, even with simple fusion methods. In addition, we experiment with discretizing the continuous features in order to speed up and simplify the fusion process even further. Our results show that fusion with discretized features outperforms text-only classification, at a fraction of the computational cost of full multi-modal fusion, with the additional benefit of improved interpretability.
Douwe Kiela, Edouard Grave, Armand Joulin, Tomás Mikolov
AAAI4
2018 Loss in Translation: Learning Bilingual Word Mapping with a Retrieval Criterion
abstract
Continuous word representations learned separately on distinct languages can be aligned so that their words become comparable in a common space.Existing works typically solve a quadratic problem to learn a orthogonal matrix aligning a bilingual lexicon, and use a retrieval criterion for inference.In this paper, we propose an unified formulation that directly optimizes a retrieval criterion in an end-to-end fashion.Our experiments on standard benchmarks show that our approach outperforms the state of the art on word translation, with the biggest improvements observed for distant language pairs such as English-Chinese.
Armand Joulin, Piotr Bojanowski, Tomás Mikolov, Hervé Jégou, Edouard Grave
EMNLP3
2018 Learning Word Vectors for 157 Languages
Edouard Grave, Piotr Bojanowski, Prakhar Gupta, Armand Joulin, Tomás Mikolov
LREC5
2018 Advances in Pre-Training Distributed Word Representations
Tomás Mikolov, Edouard Grave, Piotr Bojanowski, Christian Puhrsch, Armand Joulin
LREC1
2017 Variable Computation in Recurrent Neural Networks
Yacine Jernite, Edouard Grave, Armand Joulin, Tomás Mikolov
ICLR (Poster)4
2017 Learning Simpler Language Models with the Differential State Framework
abstract
Learning useful information across long time lags is a critical and difficult problem for temporal neural models in tasks such as language modeling. Existing architectures that address the issue are often complex and costly to train. The differential state framework (DSF) is a simple and high-performing design that unifies previously introduced gated neural models. DSF models maintain longer-term memory by learning to interpolate between a fast-changing data-driven representation and a slowly changing, implicitly stable state. Within the DSF framework, a new architecture is presented, the delta-RNN. This model requires hardly any more parameters than a classical, simple recurrent network. In language modeling at the word and character levels, the delta-RNN outperforms popular complex architectures, such as the long short-term memory (LSTM) and the gated recurrent unit (GRU), and, when regularized, performs comparably to several state-of-the-art baselines. At the subword level, the delta-RNN's performance is comparable to that of complex gated architectures.
Alexander Ororbia, Tomás Mikolov, David Reitter
Neural Comput.2
2017 Enriching Word Vectors with Subword Information
abstract
Continuous word representations, trained on large unlabeled corpora are useful for many natural language processing tasks. Popular models that learn such representations ignore the morphology of words, by assigning a distinct vector to each word. This is a limitation, especially for languages with large vocabularies and many rare words. In this paper, we propose a new approach based on the skipgram model, where each word is represented as a bag of character n-grams. A vector representation is associated to each character n-gram; words being represented as the sum of these representations. Our method is fast, allowing to train models on large corpora quickly and allows us to compute word representations for words that did not appear in the training data. We evaluate our word representations on nine different languages, both on word similarity and analogy tasks. By comparing to recently proposed morphological word representations, we show that our vectors achieve state-of-the-art performance on these tasks.
Piotr Bojanowski, Edouard Grave, Armand Joulin, Tomás Mikolov
Trans. Assoc. Comput. Linguistics4
2016 A Roadmap Towards Machine Intelligence
Tomás Mikolov, Armand Joulin, Marco Baroni
CICLing (1)1
2016 Learning Simple Algorithms from Examples
abstract
We present an approach for learning simple algorithms such as copying, multi-digit addition and single digit multiplication directly from examples. Our framework consists of a set of interfaces, accessed by a controller. Typical interfaces are 1-D tapes or 2-D grids that hold the input and output data. For the controller, we explore a range of neural network-based models which vary in their ability to abstract the underlying algorithm from training instances and generalize to test examples with many thousands of digits. The controller is trained using Q-learning with several enhancements and we show that the bottleneck is in the capabilities of the controller rather than in the search incurred by Q-learning.
Wojciech Zaremba, Tomás Mikolov, Armand Joulin, Rob Fergus
ICML2
2015 Inferring Algorithmic Patterns with Stack-Augmented Recurrent Nets
abstract
Despite the recent achievements in machine learning, we are still very far from achieving real artificial intelligence. In this paper, we discuss the limitations of standard deep learning approaches and show that some of these limitations can be overcome by learning how to grow the complexity of a model in a structured way. Specifically, we study the simplest sequence prediction problems that are beyond the scope of what is learnable with standard recurrent networks, algorithmically generated sequences which can only be learned by models which have the capacity to count and to memorize sequences. We show that some basic algorithms can be learned from sequential data using a recurrent network associated with a trainable memory.
Armand Joulin, Tomás Mikolov
NIPS2
2014 Distributed Representations of Sentences and Documents
abstract
Many machine learning algorithms require the input to be represented as a fixed length feature vector. When it comes to texts, one of the most common representations is bag-of-words. Despite their popularity, bag-of-words models have two major weaknesses: they lose the ordering of the words and they also ignore semantics of the words. For example, "powerful," "strong" and "Paris" are equally distant. In this paper, we propose an unsupervised algorithm that learns vector representations of sentences and text documents. This algorithm represents each document by a dense vector which is trained to predict words in the document. Its construction gives our algorithm the potential to overcome the weaknesses of bag-of-words models. Empirical results show that our technique outperforms bag-of-words models as well as other techniques for text representations. Finally, we achieve new state-of-the-art results on several text classification and sentiment analysis tasks.
Quoc V. Le, Tomás Mikolov
ICML2
2014 One billion word benchmark for measuring progress in statistical language modeling
abstract
We propose a new benchmark corpus to be used for measuring progress in statistical language modeling. With almost one billion words of training data, we hope this benchmark will be useful to quickly evaluate novel language modeling techniques, and to compare their contribution when combined with other advanced techniques. We show performance of several well-known types of language models, with the best results achieved with a recurrent neural network based language model. The baseline unpruned Kneser-Ney 5-gram model achieves perplexity 67.6; a combination of techniques leads to 35% reduction in perplexity, or 10% reduction in cross-entropy (bits), over that baseline. The benchmark is available as a code.google.com project; besides the scripts needed to rebuild the training/held-out data, it also makes available log-probability values for each word in each of ten held-out data sets, for each of the baseline n-gram models.
Ciprian Chelba, Tomás Mikolov, Mike Schuster, Qi Ge, Thorsten Brants, Phillipp Koehn, Tony Robinson
INTERSPEECH2
2013 On the difficulty of training recurrent neural networks
abstract
There are two widely known issues with properly training recurrent neural networks, the vanishing and the exploding gradient problems detailed in Bengio et al. (1994). In this paper we attempt to improve the understanding of the underlying issues by exploring these problems from an analytical, a geometric and a dynamical systems perspective. Our analysis is used to justify a simple yet effective solution. We propose a gradient norm clipping strategy to deal with exploding gradients and a soft constraint for the vanishing gradients problem. We validate empirically our hypothesis and proposed solutions in the experimental section.
Razvan Pascanu, Tomás Mikolov, Yoshua Bengio
ICML (3)2
2013 Linguistic Regularities in Continuous Space Word Representations
Tomás Mikolov, Scott Yih, Geoffrey Zweig
HLT-NAACL1
2013 Combining Heterogeneous Models for Measuring Relational Similarity
Alisa Zhila, Scott Yih, Christopher Meek, Geoffrey Zweig, Tomás Mikolov
HLT-NAACL5
2013 DeViSE: A Deep Visual-Semantic Embedding Model
abstract
Modern visual recognition systems are often limited in their ability to scale to large numbers of object categories. This limitation is in part due to the increasing difficulty of acquiring sufficient training data in the form of labeled images as the number of object categories grows. One remedy is to leverage data from other sources -- such as text data -- both to train visual models and to constrain their predictions. In this paper we present a new deep visual-semantic embedding model trained to identify visual objects using both labeled image data as well as semantic information gleaned from unannotated text. We demonstrate that this model matches state-of-the-art performance on the 1000-class ImageNet object recognition challenge while making more semantically reasonable errors, and also show that the semantic information can be exploited to make predictions about tens of thousands of image labels not observed during training. Semantic knowledge improves such zero-shot predictions by up to 65%, achieving hit rates of up to 10% across thousands of novel labels never seen by the visual model.
Andrea Frome, Gregory S. Corrado, Jonathon Shlens, Samy Bengio, Jeffrey Dean, Marc'Aurelio Ranzato, Tomás Mikolov
NIPS7
2013 Distributed Representations of Words and Phrases and their Compositionality
abstract
The recently introduced continuous Skip-gram model is an efficient method for learning high-quality distributed vector representations that capture a large number of precise syntactic and semantic word relationships. In this paper we present several improvements that make the Skip-gram model more expressive and enable it to learn higher quality vectors more rapidly. We show that by subsampling frequent words we obtain significant speedup, and also learn higher quality representations as measured by our tasks. We also introduce Negative Sampling, a simplified variant of Noise Contrastive Estimation (NCE) that learns more accurate vectors for frequent words compared to the hierarchical softmax. An inherent limitation of word representations is their indifference to word order and their inability to represent idiomatic phrases. For example, the meanings of Canada'' and "Air'' cannot be easily combined to obtain "Air Canada''. Motivated by this example, we present a simple and efficient method for finding phrases, and show that their vector representations can be accurately learned by the Skip-gram model. "
Tomás Mikolov, Ilya Sutskever, Kai Chen 0010, Gregory S. Corrado, Jeffrey Dean
NIPS1
2013 Approximate inference: A sampling based modeling technique to capture complex dependencies in a language model
Anoop Deoras, Tomás Mikolov, Stefan Kombrink, Kenneth Church 0001
Speech Commun.2
2012 Improving language models for ASR using translated in-domain data
abstract
Acquisition of in-domain training data to build speech recognition systems for under-resourced languages can be a costly, time-demanding and tedious process. In this work, we propose the use of machine translation to translate English transcripts of telephone speech into Czech language in order to improve a Czech CTS speech recognition system. The translated transcripts are used as additional language model training data in a scenario where the baseline language model is trained on off- and close-domain data only. We report perplexities, OOV and word error rates and examine different data sets and translators on their suitability for the described task.
Stefan Kombrink, Tomás Mikolov, Martin Karafiát, Lukás Burget
ICASSP2
2012 Context dependent recurrent neural network language model
abstract
Recurrent neural network language models (RNNLMs) have recently demonstrated state-of-the-art performance across a variety of tasks. In this paper, we improve their performance by providing a contextual real-valued input vector in association with each word. This vector is used to convey contextual information about the sentence being modeled. By performing Latent Dirichlet Allocation using a block of preceding text, we achieve a topic-conditioned RNNLM. This approach has the key advantage of avoiding the data fragmentation associated with building multiple topic models on different data subsets. We report perplexity results on the Penn Treebank data, where we achieve a new state-of-the-art. We further apply the model to the Wall Street Journal speech recognition task, where we observe improvements in word-error-rate.
Tomás Mikolov, Geoffrey Zweig
SLT1
2011 Strategies for training large scale neural network language models
abstract
We describe how to effectively train neural network based language models on large data sets. Fast convergence during training and better overall performance is observed when the training data are sorted by their relevance. We introduce hash-based implementation of a maximum entropy model, that can be trained as a part of the neural network model. This leads to significant reduction of computational complexity. We achieved around 10% relative reduction of word error rate on English Broadcast News speech recognition task, against large 4-gram model trained on 400M tokens.
Tomás Mikolov, Anoop Deoras, Daniel Povey, Lukás Burget, Jan Cernocký
ASRU1
2011 A Fast Re-scoring Strategy to Capture Long-Distance Dependencies
Anoop Deoras, Tomás Mikolov, Kenneth Church 0001
EMNLP2
2011 Variational approximation of long-span language models for lvcsr
abstract
Long-span language models that capture syntax and semantics are seldom used in the first pass of large vocabulary continuous speech recognition systems due to the prohibitive search-space of sentence-hypotheses. Instead, an N-best list of hypotheses is created using tractable n-gram models, and rescored using the long-span models. It is shown in this paper that computationally tractable variational approximations of the long-span models are a better choice than standard n-gram models for first pass decoding. They not only result in a better first pass output, but also produce a lattice with a lower oracle word error rate, and rescoring the N-best list from such lattices with the long-span models requires a smaller N to attain the same accuracy. Empirical results on the WSJ, MIT Lectures, NIST 2007 Meeting Recognition and NIST 2001 Conversational Telephone Recognition data sets are presented to support these claims.
Anoop Deoras, Tomás Mikolov, Stefan Kombrink, Martin Karafiát, Sanjeev Khudanpur
ICASSP2
2011 Extensions of recurrent neural network language model
abstract
We present several modifications of the original recurrent neural network language model (RNN LM).While this model has been shown to significantly outperform many competitive language modeling techniques in terms of accuracy, the remaining problem is the computational complexity. In this work, we show approaches that lead to more than 15 times speedup for both training and testing phases. Next, we show importance of using a backpropagation through time algorithm. An empirical comparison with feedforward networks is also provided. In the end, we discuss possibilities how to reduce the amount of parameters in the model. The resulting RNN model can thus be smaller, faster both during training and testing, and more accurate than the basic one.
Tomás Mikolov, Stefan Kombrink, Lukás Burget, Jan Cernocký, Sanjeev Khudanpur
ICASSP1
2011 Recurrent Neural Network Based Language Modeling in Meeting Recognition
abstract
We use recurrent neural network (RNN) based language models to improve the BUT English meeting recognizer. On the baseline setup using the original language models we decrease word error rate (WER) more than 1% absolute by n-best list rescoring and language model adaptation. When n-gram language models are trained on the same moderately sized data set as the RNN models, improvements are higher yielding a system which performs comparable to the baseline. A noticeable improvement was observed with unsupervised adaptation of RNN models. Furthermore, we examine the influence of word history on WER and show how to speed-up rescoring by caching common prefix strings. Index Terms: automatic speech recognition, language modeling, recurrent neural networks, rescoring, adaptation
Stefan Kombrink, Tomás Mikolov, Martin Karafiát, Lukás Burget
INTERSPEECH2
2011 Empirical Evaluation and Combination of Advanced Language Modeling Techniques
abstract
We present results obtained with several advanced language modeling techniques, including class based model, cache model, maximum entropy model, structured language model, random forest language model and several types of neural network based language models. We show results obtained after combining all these models by using linear interpolation. We conclude that for both small and moderately sized tasks, we obtain new state of the art results with combination of models, that is significantly better than performance of any individual model. Obtained perplexity reductions against Good-Turing trigram baseline are over 50% and against modified Kneser-Ney smoothed 5-gram over 40%. Index Terms: language modeling, neural networks, model combination, speech recognition
Tomás Mikolov, Anoop Deoras, Stefan Kombrink, Lukás Burget, Jan Cernocký
INTERSPEECH1
2010 Recurrent neural network based language model
abstract
A new recurrent neural network based language model (RNN LM) with applications to speech recognition is presented. Results indicate that it is possible to obtain around 50% reduction of perplexity by using mixture of several RNN LMs, compared to a state of the art backoff language model. Speech recognition experiments show around 18% reduction of word error rate on the Wall Street Journal task when comparing models trained on the same amount of data, and around 5% on the much harder NIST RT05 task, even when the backoff model is trained on much more data than the RNN LM. We provide ample empirical evidence to suggest that connectionist language models are superior to standard n-gram techniques, except their high computational (training) complexity. Index Terms: language modeling, recurrent neural networks, speech recognition
Tomás Mikolov, Martin Karafiát, Lukás Burget, Jan Cernocký, Sanjeev Khudanpur
INTERSPEECH1
2009 Neural network based language models for highly inflective languages
abstract
Speech recognition of inflectional and morphologically rich languages like Czech is currently quite a challenging task, because simple n-gram techniques are unable to capture important regularities in the data. Several possible solutions were proposed, namely class based models, factored models, decision trees and neural networks. This paper describes improvements obtained in recognition of spoken Czech lectures using language models based on neural networks. Relative reductions in word error rate are more than 15% over baseline obtained with adapted 4-gram backoff language model using modified Kneser-Ney smoothing.
Tomás Mikolov, Jirí Kopecký, Lukás Burget, Ondrej Glembek, Jan Cernocký
ICASSP1
2008 Advances in phonotactic language recognition
abstract
This paper summarizes recent advances in PRLM language recognition within the context of the NIST 2007 LR evaluations (LRE). We present a comparison of binary decision tree (BT) vs. N -gram models when adaptation from a universal (background) model (UBM) is used, we introduce multi-models— anchor-model-like approach to scoring, and we adopt the framework of intersession variation using factor analysis.
Ondrej Glembek, Pavel Matejka, Lukás Burget, Tomás Mikolov
INTERSPEECH4
2008 BUT language recognition system for NIST 2007 evaluations
abstract
This paper describes Brno University of Technology (BUT) system for 2007 NIST Language recognition (LRE) evaluation. The system is a fusion of 4 acoustic and 9 phonotactic subsystems. We have investigated several new topics such as discriminatively trained language models in phonotactic systems, and eigen-channel adaptation in model and feature domain in acoustic systems. We also point out the importance of calibration and fusion. All results are presented on NIST 2007 LRE data.
Pavel Matejka, Lukás Burget, Ondrej Glembek, Petr Schwarz, Valiantsina Hubeika, Michal Fapso, Tomás Mikolov, Oldrich Plchot, Jan Cernocký
INTERSPEECH7