Michael Riley 0001

dblp:73/924 · also Michael D. Riley · DBLP profile ↗
← Back
77ranked-venue papers
7as first author
10since 2021 · last 2025
0009-0002-9449-025XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 50 · 7 first-author · 5 since 2021Artificial intelligence and machine learning · 48 · 4 first-author · 6 since 2021Theory of computation · 8Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Massive Sound Embedding Benchmark (MSEB)
abstract
Audio is a critical component of multimodal perception, and any truly intelligent system must demonstrate a wide range of auditory capabilities. These capabilities include transcription, classification, retrieval, reasoning, segmentation, clustering, reranking, and reconstruction. Fundamentally, each task involves transforming a raw audio signal into a meaningful 'embedding'—be it a single vector, a sequence of continuous or discrete representations, or another structured form—which then serves as the basis for generating the task's final response. To accelerate progress towards robust machine auditory intelligence, we present the Massive Sound Embedding Benchmark (MSEB): an extensible framework designed to evaluate the auditory components of any multimodal system. In its first release, MSEB offers a comprehensive suite of eight core tasks, with more planned for the future, supported by diverse datasets, including the new, large-scale Simple Voice Questions (SVQ) dataset. Our initial experiments establish clear performance headrooms, highlighting the significant opportunity to improve real-world multimodal experiences where audio is a core signal. We encourage the research community to use MSEB to assess their algorithms and contribute to its growth. The library is publicly hosted at https://github.com/google-research/mseb.
Georg Heigold, Ehsan Variani, Tom Bagby, Cyril Allauzen, Ji Ma 0004, Shankar Kumar, Michael Riley 0001
NeurIPS7
2024 Accelerating Blockwise Parallel Language Models with Draft Refinement
abstract
Autoregressive language models have achieved remarkable advancements, yet their potential is often limited by the slow inference speeds associated with sequential token generation. Blockwise parallel decoding (BPD) was proposed by Stern et al. [42] as a method to improve inference speed of language models by simultaneously predicting multiple future tokens, termed block drafts, which are subsequently verified by the autoregressive model. This paper advances the understanding and improvement of block drafts in two ways. First, we analyze token distributions generated across multiple prediction heads. Second, leveraging these insights, we propose algorithms to improve BPD inference speed by refining the block drafts using task-independent \ngram and neural language models as lightweight rescorers. Experiments demonstrate that by refining block drafts of open-sourced Vicuna and Medusa LLMs, the mean accepted token length are increased by 5-25% relative. This results in over a 3x speedup in wall clock time compared to standard autoregressive decoding in open-source 7B and 13B LLMs.
Taehyeon Kim 0001, Ananda Theertha Suresh, Kishore Papineni, Michael Riley 0001, Sanjiv Kumar, Adrian Benton
NeurIPS4
2023 Large-Scale Language Model Rescoring on Long-Form Data
abstract
In this work, we study the impact of Large-scale Language Models (LLM) on Automated Speech Recognition (ASR) of YouTube videos, which we use as a source for long-form ASR. We demonstrate up to 8% relative reduction in Word Error Eate (WER) on US English (en-us) and code-switched Indian English (en-in) long-form ASR test sets and a reduction of up to 30% relative on Salient Term Error Rate (STER) over a strong first-pass baseline that uses a maximum-entropy based language model. Improved lattice processing that results in a lattice with a proper (non-tree) digraph topology and carrying context from the 1-best hypothesis of the previous segment(s) results in significant wins in rescoring with LLMs. We also find that the gains in performance from the combination of LLMs trained on vast quantities of available data (such as C4 [1]) and conventional neural LMs is additive and significantly outperforms a strong first-pass baseline with a maximum entropy LM.
Tongzhou Chen, Cyril Allauzen, Daniel S. Park, David Rybach, W. Ronny Huang, Rodrigo Cabrera, Kartik Audhkhasi, Bhuvana Ramabhadran, Pedro J. Moreno 0001, Michael Riley 0001
ICASSP11
2023 Alignment Entropy Regularization
abstract
Existing training criteria in automatic speech recognition (ASR) permit the model to freely explore more than one time alignments between the feature and label sequences. In this paper, we use entropy to measure a model’s uncertainty, i.e. how it chooses to distribute the probability mass over the set of allowed alignments. Furthermore, we evaluate the effect of entropy regularization in encouraging the model to distribute the probability mass only on a smaller subset of allowed alignments. Experiments show that entropy regularization enables a much simpler decoding method without sacrificing word error rate, and provides better time alignment quality.
Ehsan Variani, David Rybach, Cyril Allauzen, Michael Riley 0001
ICASSP5
2023 Last: Scalable Lattice-Based Speech Modelling in Jax
abstract
We introduce LAST, a LAttice-based Speech Transducer library in JAX. With an emphasis on flexibility, ease-of-use, and scalability, LAST implements differentiable weighted finite state automaton (WFSA) algorithms needed for training & inference that scale to a large WFSA such as a recognition lattice over the entire utterance. Despite these WFSA algorithms being well-known in the literature, new challenges arise from performance characteristics of modern architectures, and from nuances in automatic differentiation. We describe a suite of generally applicable techniques employed in LAST to address these challenges, and demonstrate their effectiveness with benchmarks on TPUv3 and V100 GPU.
Ehsan Variani, Tom Bagby, Michael Riley 0001
ICASSP4
2022 On Adaptive Weight Interpolation of the Hybrid Autoregressive Transducer
Ehsan Variani, Michael Riley 0001, David Rybach, Cyril Allauzen, Tongzhou Chen, Bhuvana Ramabhadran
INTERSPEECH2
2022 Global Normalization for Streaming Speech Recognition in a Modular Framework
abstract
We introduce the Globally Normalized Autoregressive Transducer (GNAT) for addressing the label bias problem in streaming speech recognition. Our solution admits a tractable exact computation of the denominator for the sequence-level normalization. Through theoretical and empirical results, we demonstrate that by switching to a globally normalized model, the word error rate gap between streaming and non-streaming speech-recognition models can be greatly reduced (by more than 50% on the Librispeech dataset). This model is developed in a modular framework which encompasses all the common neural speech recognition models. The modularity of this framework enables controlled comparison of modelling choices and creation of new models. A JAX implementation of our models has been open sourced.
Ehsan Variani, Michael Riley 0001, David Rybach, Matt Shannon, Cyril Allauzen
NeurIPS3
2022 Spatial Model Personalization in Gboard
abstract
We introduce a framework for adapting a virtual keyboard to individual user behavior by modifying a Gaussian spatial model to use personalized key center offset means and, optionally, learned covariances. Through numerous real-world studies, we determine the importance of training data quantity and weights, as well as the number of clusters into which to group keys to avoid overfitting. While past research has shown potential of this technique using artificially-simple virtual keyboards and games or fixed typing prompts, we demonstrate effectiveness using the highly-tuned Gboard app with a representative set of users and their real typing behaviors. Across a variety of top languages, we achieve small-but-significant improvements in both typing speed and decoder accuracy.
Gary Sivek, Michael Riley 0001
Proc. ACM Hum. Comput. Interact.2
2021 A Hybrid Seq-2-Seq ASR Design for On-Device and Server Applications
Cyril Allauzen, Ehsan Variani, Michael Riley 0001, David Rybach, Hao Zhang 0010
Interspeech3
2021 Approximating Probabilistic Models as Weighted Finite Automata
abstract
Abstract Weighted finite automata (WFAs) are often used to represent probabilistic models, such as ngram language models, because among other things, they are efficient for recognition tasks in time and space. The probabilistic source to be represented as a WFA, however, may come in many forms. Given a generic probabilistic model over sequences, we propose an algorithm to approximate it as a WFA such that the Kullback-Leibler divergence between the source model and the WFA target model is minimized. The proposed algorithm involves a counting step and a difference of convex optimization step, both of which can be performed efficiently.We demonstrate the usefulness of our approach on various tasks, including distilling n-gram models from neural models, building compact language models, and building open-vocabulary character models. The algorithms used for these experiments are available in an open-source software library.
Ananda Theertha Suresh, Brian Roark, Michael Riley 0001, Vlad Schogol
Comput. Linguistics3
2020 Hybrid Autoregressive Transducer (HAT)
abstract
This paper proposes and evaluates the hybrid autoregressive transducer (HAT) model, a time-synchronous encoder-decoder model that preserves the modularity of conventional automatic speech recognition systems. The HAT model provides a way to measure the quality of the internal language model that can be used to decide whether inference with an external language model is beneficial or not. We evaluate our proposed model on a large-scale voice search task. Our experiments show significant improvements in WER compared to the state-of-the-art approaches1.
Ehsan Variani, David Rybach, Cyril Allauzen, Michael Riley 0001
ICASSP4
2020 Learning discrete distributions: user vs item-level privacy
abstract
Much of the literature on differential privacy focuses on item-level privacy, where loosely speaking, the goal is to provide privacy per item or training example. However, recently many practical applications such as federated learning require preserving privacy for all items of a single user, which is much harder to achieve. Therefore understanding the theoretical limit of user-level privacy becomes crucial. We study the fundamental problem of learning discrete distributions over $k$ symbols with user-level differential privacy. If each user has $m$ samples, we show that straightforward applications of Laplace or Gaussian mechanisms require the number of users to be $\mathcal{O}(k/(m\alpha^2) + k/\epsilon\alpha)$ to achieve an $\ell_1$ distance of $\alpha$ between the true and estimated distributions, with the privacy-induced penalty $k/\epsilon\alpha$ independent of the number of samples per user $m$. Moreover, we show that any mechanism that only operates on the final aggregate should require a user complexity of the same order. We then propose a mechanism such that the number of users scales as $\tilde{\mathcal{O}}(k/(m\alpha^2) + k/\sqrt{m}\epsilon\alpha)$ and further show that it is nearly-optimal under certain regimes. Thus the privacy penalty is $\tilde{\Theta}(\sqrt{m})$ times smaller compared to the standard mechanisms. We also propose general techniques for obtaining lower bounds on restricted differentially private estimators and a lower bound on the total variation between binomial distributions, both of which might be of independent interest.
Yuhan Liu 0007, Ananda Theertha Suresh, Felix X. Yu, Sanjiv Kumar, Michael Riley 0001
NeurIPS5
2019 Federated Learning of N-Gram Language Models
abstract
Mingqing Chen, Ananda Theertha Suresh, Rajiv Mathews, Adeline Wong, Cyril Allauzen, Françoise Beaufays, Michael Riley. Proceedings of the 23rd Conference on Computational Natural Language Learning (CoNLL). 2019.
Mingqing Chen, Ananda Theertha Suresh, Rajiv Mathews, Adeline Wong, Cyril Allauzen, Françoise Beaufays, Michael Riley 0001
CoNLL7
2018 Semantic Lattice Processing in Contextual Automatic Speech Recognition for Google Assistant
Leonid Velikovich, Justin Scheiner, Petar S. Aleksic, Pedro J. Moreno 0001, Michael Riley 0001
INTERSPEECH6
2018 Algorithms for Weighted Finite Automata with Failure Transitions
Cyril Allauzen, Michael Riley 0001
CIAA2
2017 On lattice generation for large vocabulary speech recognition
abstract
Lattice generation is an essential feature of the decoder for many speech recognition applications. In this paper, we first review lattice generation methods for WFST-based decoding and describe in a uniform formalism two established approaches for state-of-the-art speech recognition systems: the phone pair and the N-best histories approaches. We then present a novel optimization method, pruned determinization followed by minimization, that produces a deterministic minimal lattice that retains all paths within specified weight and lattice size thresholds. Experimentally, we show that before optimization, the phone-pair and the N-best histories approaches each have conditions where they perform better when evaluated on video transcription and mixed voice search and dictation tasks. However, once this lattice optimization procedure is applied, the phone pair approach has the lowest oracle WER for a given lattice density by a significant margin. We further show that the pruned determinization presented here is efficient to use during decoding unlike classical weighted determinization from which it is derived. Finally, we consider on-the-fly lattice rescoring in which the lattice generation and combination with the secondary LM are done in one step. We compare the phone pair and N-best histories approaches for this scenario and find the former superior in our experiments.
David Rybach, Michael Riley 0001, Johan Schalkwyk
ASRU2
2017 A disambiguation algorithm for weighted automata
Mehryar Mohri, Michael Riley 0001
Theor. Comput. Sci.2
2016 Contextual Prediction Models for Speech Recognition
Yoni Halpern, Keith B. Hall, Vlad Schogol, Michael Riley 0001, Brian Roark, Gleb Skobeltsyn, Martin Bäuml
INTERSPEECH4
2016 Learning N-Gram Language Models from Uncertain Data
Vitaly Kuznetsov, Hank Liao, Mehryar Mohri, Michael Riley 0001, Brian Roark
INTERSPEECH4
2015 Rapid vocabulary addition to context-dependent decoder graphs
Cyril Allauzen, Michael Riley 0001
INTERSPEECH2
2015 Composition-based on-the-fly rescoring for salient n-gram biasing
abstract
We introduce a technique for dynamically applying contextually-derived language models to a state-of-the-art speech recognition system. These generally small-footprint models can be seen as a generalization of cache-based models [1], whereby contextually salient n-grams are derived from relevant sources (not just user generated language) to produce a model intended for combination with the baseline language model. The derived models are applied during first-pass decoding as a form of on-the-fly composition between the decoder search graph and the set of weighted contextual n-grams. We present a construction algorithm which takes a trie representing the contextual n-grams and produces a weighted finite state automaton which is more compact than a standard n-gram machine. Finally, we present a set of empirical results on the recognition of spoken search queries where a contextual model encoding recent trending queries is applied using the proposed technique.
Keith B. Hall, Eunjoon Cho, Cyril Allauzen, Françoise Beaufays, Noah Coccaro, Kaisuke Nakajima, Michael Riley 0001, Brian Roark, David Rybach, Linda Zhang 0004
INTERSPEECH7
2015 Automata and graph compression
abstract
We present a theoretical framework for the compression of automata, which are widely used representations in speech processing, natural language processing and many other tasks. As a corollary, our framework further covers graph compression. We introduce a probabilistic process of graph and automata generation that is similar to stationary ergodic processes and that covers real-world phenomena. We also introduce a universal compression scheme LZA for this probabilistic model and show that LZA significantly outperforms other compression techniques such as gzip and the UNIX compress command for several synthetic and real data sets.
Mehryar Mohri, Michael Riley 0001, Ananda Theertha Suresh
ISIT2
2015 On the Disambiguation of Weighted Automata
Mehryar Mohri, Michael Riley 0001
CIAA2
2014 Encoding linear models as weighted finite-state transducers
abstract
We present algorithms, implemented as an extension to the OpenFst library, that yield a class of transducers that encode linear models for structured inference tasks like segmentation and tagging.This allows the use of general finite-state operations with such models.For instance, finite-state composition can be used to apply the model to lattice input (or other more general automata) and then the result automaton can be passed to subsequent processing such as general shortest path algorithms.We demonstrate the use of the library extension on graphemeto-phoneme conversion, encoding multiple varieties of linear models for that task, and achieve solid PER/WER gains over previous best reported results on g2p conversion of a publicly available dataset (CMU).
Cyril Allauzen, Keith B. Hall, Michael Riley 0001, Brian Roark
INTERSPEECH4
2014 Pushdown Automata in Statistical Machine Translation
abstract
This article describes the use of pushdown automata (PDA) in the context of statistical machine translation and alignment under a synchronous context-free grammar. We use PDAs to compactly represent the space of candidate translations generated by the grammar when applied to an input sentence. General-purpose PDA algorithms for replacement, composition, shortest path, and expansion are presented. We describe HiPDT, a hierarchical phrase-based decoder using the PDA representation and these algorithms. We contrast the complexity of this decoder with a decoder based on a finite state automata representation, showing that PDAs provide a more suitable framework to achieve exact decoding for larger synchronous context-free grammars and smaller language models. We assess this experimentally on a large-scale Chinese-to-English alignment and translation task. In translation, we propose a two-pass decoding strategy involving a weaker language model in the first-pass to address the results of PDA complexity analysis. We study in depth the experimental conditions and tradeoffs in which HiPDT can achieve state-of-the-art performance for large-scale SMT.
Cyril Allauzen, William J. Byrne, Adrià de Gispert, Gonzalo Iglesias, Michael Riley 0001
Comput. Linguistics5
2014 Direct construction of compact context-dependency transducers from data
David Rybach, Michael Riley 0001, Christopher Alberti
Comput. Speech Lang.2
2013 Smoothed marginal distribution constraints for language modeling
Brian Roark, Cyril Allauzen, Michael Riley 0001
ACL (1)3
2013 Pre-initialized composition for large-vocabulary speech recognition
abstract
This paper describes a modified composition algorithm that is used for combining two finite-state transducers, representing the context-dependent lexicon and the language model respec-tively, in large vocabulary speech recogntion. This algorithm is a hybrid between the static and dynamic expansion of the re-sultant transducer, which maps from context-dependent phones to words and is searched during decoding. The approach is to pre-compute part of the recognition transducer and leave the balance to be expanded during decoding. This method allows for a fine-grained trade-off between space and time in recogni-tion. For example, the time overhead of purely dynamic expan-sion can be reduced by over six-fold with only a 20 % increase in memory in a collection of large-vocabulary recognition tasks available on the Google Android platform.
Cyril Allauzen, Michael Riley 0001
INTERSPEECH2
2012 Mobile music modeling, analysis and recognition
abstract
We present an analysis of music modeling and recognition techniques in the context of mobile music matching, substantially improving on the techniques presented in [1]. We accomplish this by adapting the features specifically to this task, and by introducing new modeling techniques that enable using a corpus of noisy and channel-distorted data to improve mobile music recognition quality. We report the results of an extensive empirical investigation of the system's robustness under realistic channel effects and distortions. We show an improvement of recognition accuracy by explicit duration modeling of music phonemes and by integrating the expected noise environment into the training process. Finally, we propose the use of frame-to-phoneme alignment for high-level structure analysis of polyphonic music.
Pavel Golik, Boulos Harb, Ananya Misra, Michael Riley 0001, Alex Rudnick, Eugene Weinstein
ICASSP4
2012 Voice Query Refinement
abstract
We describe a system for the refinement of spoken search queries. Given an initial query (Northern Italian restaurants in New York), instead of requiring a fully-specified followup query (Korean restaurants in New York), a more natural, abbreviated update query (Korean instead) may be spoken. The system consists of a parsing step to identify the type and arguments of the refinement, a candidate generation step to enumerate the possible refinements, and a model classification step to select the best refinement. We present results on test query refinements given both to this system and to human judges that show the automated system outperforms the human judges on that data set. Index terms: spoken dialog systems, voice search, query refinement 1.
Cyril Allauzen, Edward Benson, Ciprian Chelba, Michael Riley 0001, Johan Schalkwyk
INTERSPEECH4
2012 A Pushdown Transducer Extension for the OpenFst Library
Cyril Allauzen, Michael Riley 0001
CIAA2
2011 Hierarchical Phrase-based Translation Representations
Gonzalo Iglesias, Cyril Allauzen, William J. Byrne, Adrià de Gispert, Michael Riley 0001
EMNLP5
2011 Bayesian Language Model Interpolation for Mobile Speech Input
abstract
This paper explores various static interpolation methods for approximating a single dynamically-interpolated language model used for a variety of recognition tasks on the Google Android platform. The goal is to find the statically-interpolated firstpass LM that best reduces search errors in a two-pass system or that even allows eliminating the more complex dynamic second pass entirely. Static interpolation weights that are uniform, prior-weighted, and the maximum likelihood, maximum a posteriori, and Bayesian solutions are considered. Analysis argues and recognition experiments on Android test data show that a Bayesian interpolation approach performs best.
Cyril Allauzen, Michael Riley 0001
INTERSPEECH2
2010 Direct construction of compact context-dependency transducers from data
abstract
This paper describes a new method for building compact context-dependency transducers for finite-state transducer-based ASR decoders.Instead of the conventional phonetic decisiontree growing followed by FST compilation, this approach incorporates the phonetic context splitting directly into the transducer construction.The objective function of the split optimization is augmented with a regularization term that measures the number of transducer states introduced by a split.We give results on a large spoken-query task for various n-phone orders and other phonetic features that show this method can greatly reduce the size of the resulting context-dependency transducer with no significant impact on recognition accuracy.This permits using context sizes and features that might otherwise be unmanageable.
David Rybach, Michael Riley 0001
INTERSPEECH2
2010 Expected Sequence Similarity Maximization
Cyril Allauzen, Shankar Kumar, Wolfgang Macherey, Mehryar Mohri, Michael Riley 0001
HLT-NAACL5
2010 Filters for Efficient Composition of Weighted Finite-State Transducers
Cyril Allauzen, Michael Riley 0001, Johan Schalkwyk
CIAA2
2009 WEB-derived pronunciations
abstract
Pronunciation information is available in large quantities on the Web, in the form of IPA and ad-hoc transcriptions. We describe techniques for extracting candidate pronunciations from Web pages and associating them with orthographic words, filtering out poorly extracted pronunciations, normalizing IPA pronunciations to better conform to a common transcription standard, and generating phonemic from ad-hoc transcriptions. We show improvements on a letter-to-phoneme task when using web-derived vs. Pronlex pronunciations.
Arnab Ghoshal, Martin Jansche, Sanjeev Khudanpur, Michael Riley 0001, Morgan Ulinski
ICASSP4
2009 A generalized composition algorithm for weighted finite-state transducers
abstract
This paper describes a weighted finite-state transducer composition algorithm that generalizes the concept of the composition filter and presents filters that remove useless epsilon paths and push forward labels and weights along epsilon paths. This filtering permits the compostion of large speech recognition contextdependent lexicons and language models much more efficiently in time and space than previously possible. We present experiments on Broadcast News and a spoken query task that demonstrate an ∼5 % to 10 % overhead for dynamic, runtime composition compared to a static, offline composition of the recognition transducer. To our knowledge, this is the first such system with so little overhead.
Cyril Allauzen, Michael Riley 0001, Johan Schalkwyk
INTERSPEECH2
2009 Web derived pronunciations for spoken term detection
abstract
Indexing and retrieval of speech content in various forms such as broadcast news, customer care data and on-line media has gained a lot of interest for a wide range of applications, from customer analytics to on-line media search. For most retrieval applications, the speech content is typically first converted to a lexical or phonetic representation using automatic speech recognition (ASR). The first step in searching through indexes built on these representations is the generation of pronunciations for named entities and foreign language query terms. This paper summarizes the results of the work conducted during the 2008 JHU Summer Workshop by the Multilingual Spoken Term Detection team, on mining the web for pronunciations and analyzing their impact on spoken term detection. We will first present methods to use the vast amount of pronunciation information available on the Web, in the form of IPA and ad-hoc transcriptions. We describe techniques for extracting candidate pronunciations from Web pages and associating them with orthographic words, filtering out poorly extracted pronunciations, normalizing IPA pronunciations to better conform to a common transcription standard, and generating phonemic representations from ad-hoc transcriptions. We then present an analysis of the effectiveness of using these pronunciations to represent Out-Of-Vocabulary (OOV) query terms on the performance of a spoken term detection (STD) system. We will provide comparisons of Web pronunciations against automated techniques for pronunciation generation as well as pronunciations generated by human experts. Our results cover a range of speech indexes based on lattices, confusion networks and one-best transcriptions at both word and word fragments levels.
Dogan Can, Erica Cooper, Arnab Ghoshal, Martin Jansche, Sanjeev Khudanpur, Bhuvana Ramabhadran, Michael Riley 0001, Murat Saraclar, Abhinav Sethy, Morgan Ulinski, Christopher M. White
SIGIR7
2008 Sample Selection Bias Correction Theory
Corinna Cortes, Mehryar Mohri, Michael Riley 0001, Afshin Rostamizadeh
ALT3
2007 OpenFst: A General and Efficient Weighted Finite-State Transducer Library
Cyril Allauzen, Michael Riley 0001, Johan Schalkwyk, Wojciech Skut, Mehryar Mohri
CIAA2
2006 Efficient Computation of the Relative Entropy of Probabilistic Automata
Corinna Cortes, Mehryar Mohri, Ashish Rastogi, Michael Riley 0001
LATIN4
2006 MAP adaptation of stochastic grammars
Michiel Bacchiani, Michael Riley 0001, Brian Roark, Richard Sproat
Comput. Speech Lang.2
2004 Statistical Modeling for Unit Selection in Speech Synthesis
abstract
Traditional concatenative speech synthesis systems use a number of heuristics to define the target and concatenation costs, essential for the design of the unit selection component. In contrast to these approaches, we introduce a general statistical modeling framework for unit selection inspired by automatic speech recognition. Given appropriate data, techniques based on that framework can result in a more accurate unit selection, thereby improving the general quality of a speech synthesizer. They can also lead to a more modular and a substantially more efficient system.We present a new unit selection system based on statistical modeling. To overcome the original absence of data, we use an existing high-quality unit selection system to generate a corpus of unit sequences. We show that the concatenation cost can be accurately estimated from this corpus using a statistical n-gram language model over units. We used weighted automata and transducers for the representation of the components of the system and designed a new and more efficient composition algorithm making use of string potentials for their combination. The resulting statistical unit selection is shown to be about 2.6 times faster than the last release of the AT&T Natural Voices Product while preserving the same quality, and offers much flexibility for the use and integration of new and more complex components.
Mehryar Mohri, Cyril Allauzen, Michael Riley 0001
ACL3
2004 A generalized construction of integrated speech recognition transducers
abstract
We showed in previous work that weighted finite-state transducers provide a common representation for many components of a speech recognition system and described general algorithms for combining these representations to build a single optimized and compact transducer integrating all these components, directly mapping from HMM states to words. This approach works well for certain well-controlled input transducers, but presents some problems related to the efficiency of composition and the applicability of determinization and weight-pushing with more general transducers. We generalize our prior construction of the integrated speech recognition transducer to work with an arbitrary number of component transducers and, to a large extent, release the constraints imposed on the type of input transducers by providing more general solutions to these problems. This generalization allowed us to deal with cases where our prior optimization did not apply. Our experiments in the AT&T HMIHY 0300 task and an AT&T VoiceTone task show the efficiency of our generalized optimization technique. We report a 1.6 recognition speed-up in the HMIHY 0300 task, 1.8 speed-up in a VoiceTone task using a word-based language model, and 1.7 using a class-based model.
Cyril Allauzen, Mehryar Mohri, Michael Riley 0001, Brian Roark
ICASSP (1)3
2004 Methods for task adaptation of acoustic models with limited transcribed in-domain data
abstract
Application specific acoustic models provide the best recognition accuracy, but they are expensive, because they require the transcription of tens or hundreds of hours of in-domain speech for training. Therefore, this paper focuses on the acoustic model estimation given limited in-domain transcribed speech data, and large amounts of (typically available) transcribed out-of-domain data. First, we evaluate several combinations of known methods to optimize the adaptation/training of acoustic models on the limited in-domain speech data. Then, we propose to use Gaussian sharing to combine in-domain models with out-of-domain models, and a data generation process to simulate the presence of more speakers in the in-domain data. In a spoken language dialog application, we contrast our methods against an upper accuracy bound of 69.1 % (model trained on many in-domain data) and a lower bound of 60.8 % (no in-domain data). Using only 2 hours of in-domain speech for model estimation, we improve the accuracy by 5.1 % (to 65.9%) over the lower bound; data generation and Gaussian sharing contibute 2.2% to this improvement. With 9 hours of in-domain speech, the improvement of accuracy is 6.5%, to 67.3%. 1.
Enrico Bocchieri, Michael Riley 0001, Murat Saraclar
INTERSPEECH2
2002 A comparison of two LVR search optimization techniques
abstract
This paper presents a detailed comparison between two search optimization techniques for large vocabulary speech recognition -- one based on word-conditioned tree search (WCTS) and one based on weighted finite-state transducers (WFSTs). Existing North American Business News systems from RWTH and AT&T representing each of the two approaches, were modified to remove variations in model data and acoustic likelihood computation. An experimental comparison showed that the WFST-based system explored fewer search states and had less runtime overhead than the WCTS-based system for a given word error rate. This is attributed to differences in the pre-compilation, degree of non-determinism, and path weight distribution in the respective search graphs.
Stephan Kanthak, Hermann Ney, Michael Riley 0001, Mehryar Mohri
INTERSPEECH3
2002 An efficient algorithm for the n-best-strings problem
abstract
problem in a weighted automaton. This problem arises commonly in speech recognition applications when a ranked list of unique recognizer hypotheses is desired. We believe this is the first n-best algorithm to remove redundant hypotheses before rather than after the n-best determination. We give a detailed description of the algorithm and demonstrate its correctness. We report experimental results showing its efficiency and practicality even for large n in a 40; 000-word vocabulary North American Business News (NAB) task. In particular, we show that 1000-best generation in this task requires negligible added time over recognizer lattice generation.
Mehryar Mohri, Michael Riley 0001
INTERSPEECH2
2002 Towards automatic closed captioning : low latency real time broadcast news transcription
abstract
In this paper, we present a low latency real-time Broadcast News recognition system capable of transcribing live television newscasts with reasonable accuracy. We describe our recent modeling and efficiency improvements that yield a 22 % word error rate on the Hub4e98 test set while running faster than real-time. These include the discriminative training of a feature transform and the acoustic model, and the optimization of the likelihood computation. We give experimental results that show the accuracy of the system at different speeds. We also explain how we achieved low latency, presenting measurements that show the typical system latency is less than 1 second. 1.
Murat Saraclar, Michael Riley 0001, Enrico Bocchieri, Vincent Goffin
INTERSPEECH2
2002 Weighted finite-state transducers in speech recognition
Mehryar Mohri, Fernando Pereira 0003, Michael Riley 0001
Comput. Speech Lang.3
2001 A weight pushing algorithm for large vocabulary speech recognition
abstract
Weighted finite-state transducers provide a general framework for the representation of the components of speech recognition systems; language models, pronunciation dictionaries, contextdependent models, HMM-level acoustic models, and the output word or phone lattices can all be represented by weighted automata and transducers. In general, a representation is not unique and there may be different weighted transducers realizing the same mapping. In particular, even when they have exactly the same topology with the same input and output labels, two equivalent transducers may differ by the way the weights are distributed along each path. We present
Mehryar Mohri, Michael Riley 0001
INTERSPEECH2
2000 The Design Principles of a Weighted Finite-State Transducer Library
Mehryar Mohri, Fernando Pereira 0003, Michael Riley 0001
Theor. Comput. Sci.3
1999 Rapid unit selection from a large speech corpus for concatenative speech synthesis
abstract
Concatenative Text-to-Speech (TTS) systems such as those described by Hunt and Black [6] can select at synthesis time from a very large number of recorded units. The selected units are chosen to minimize a combination of target and join costs for a given sentence. However, the join costs, in particular, can be quite expensive to compute, even when this computation has been optimized. If possible, we would avoid this computation by precomputing and caching all the possible join costs, but their number is prohibitive. Although the search space of possible joins is large, we have found that only a small fraction are selected in practice. By synthesizing a large quantity of text and logging the units actually selected, we were able to gather usage statistics and construct a practical and efficient cache of concatenation costs. Use of this cache dramatically decreased the runtime of the AT&T Next-Generation TTS system [1] with negligible effect on speech quality. Experiments show that by ca...
Marc C. Beutnagel, Mehryar Mohri, Michael Riley 0001
EUROSPEECH3
1999 Efficient general lattice generation and rescoring
Andrej Ljolje, Fernando Pereira 0003, Michael Riley 0001
EUROSPEECH3
1999 The AT&t large vocabulary conversational speech recognition system
abstract
In the frame of the INCO-Copernicus program of European Commission we have started to develop an audio-visual pronunciation teaching and training method and software system for hearing and speech-handicapped persons to help them to control their speech production. A teaching method is drawn up for progression from the individual sound preparation to practice of the sounds in sentences. The main aim is to develop an audio-visual articulation training and teaching system for all participant languages, these being English, Swedish, Slovenian and Hungarian. The basic part is a general language-independent measuring system and database editor. This database editor makes it possible to construct modules for all participant languages and for different speech disabilities. Two modules are under development for its construction in all languages, one of them being for teaching and training vowels for hearing-impaired children, while the other one is for correction of misarticulated fricative sounds.
Andrej Ljolje, Michael Riley 0001, Donald Hindle
EUROSPEECH2
1999 Integrated context-dependent networks in very large vocabulary speech recognition
Mehryar Mohri, Michael Riley 0001
EUROSPEECH2
1999 Network optimizations for large-vocabulary speech recognition
Mehryar Mohri, Michael Riley 0001
Speech Commun.2
1999 Stochastic pronunciation modelling from hand-labelled phonetic corpora
Michael Riley 0001, William J. Byrne, Michael Finke, Sanjeev Khudanpur, Andrej Ljolje, John W. McDonough, Harriet J. Nock, Murat Saraclar, Charles Wooters, George Zavaliagkos
Speech Commun.1
1998 Pronunciation modelling using a hand-labelled corpus for conversational speech recognition
abstract
Accurately modelling pronunciation variability in conversational speech is an important component of an automatic speech recognition system. We describe some of the projects undertaken in this direction during and after WS97, the Fifth LVCSR Summer Workshop, held at Johns Hopkins University, Baltimore, in July-August, 1997. We first illustrate a use of hand-labelled phonetic transcriptions of a portion of the Switchboard corpus, in conjunction with statistical techniques, to learn alternatives to canonical pronunciations of words. We then describe the use of these alternate pronunciations in an automatic speech recognition system. We demonstrate that the improvement in recognition performance from pronunciation modelling persists as the system is enhanced with better acoustic and language models.
William J. Byrne, Michael Finke, Sanjeev Khudanpur, John W. McDonough, Harriet J. Nock, Michael Riley 0001, Murat Saraclar, Charles Wooters, George Zavaliagkos
ICASSP6
1998 Full expansion of context-dependent networks in large vocabulary speech recognition
abstract
We combine our earlier approach to context-dependent network representation with our algorithm for determining weighted networks to build optimized networks for large-vocabulary speech recognition combining an n-gram language model, a pronunciation dictionary and context-dependency modeling. While fully-expanded networks have been used before in restrictive settings (medium vocabulary or no cross-word contexts), we demonstrate that our network determination method makes it practical to use fully-expanded networks also in large-vocabulary recognition with full cross-word context modeling. For the DARPA North American Business News task (NAB), we give network sizes and recognition speeds and accuracies using bigram and trigram grammars with vocabulary sizes ranging from 10000 to 160000 words. With our construction, the fully-expanded NAB context-dependent networks contain only about twice as many arcs as the corresponding language models. Interestingly, we also find that, with these networks, real-time word accuracy is improved by increasing the vocabulary size and n-gram order.
Mehryar Mohri, Michael Riley 0001, Donald Hindle, Andrej Ljolje, Fernando Pereira 0003
ICASSP2
1997 A spoken language system for automated call routing
abstract
We are interested in the problem of understanding fluently spoken language. In particular, we consider people's responses to the open-ended prompt of "How may I help you?". We then further restrict the problem to classifying and automatically routing such a call, based on the meaning of the user's response. Thus, we aim at extracting a relatively small number of semantic actions from the utterances of a very large set of users who are not trained to the system's capabilities and limitations. In this paper, we describe the main components of our speech understanding system: the large vocabulary recognizer and the language understanding module performing the call-type classification. In particular, we propose automatic algorithms for selecting phrases from a training corpus in order to enhance the prediction power of the standard word n-gram. The phrase language models are integrated into stochastic finite state machines which outperform standard word n-gram language models. From the speech recognizer output we recognize and exploit automatically-acquired salient phrase fragments to make a call-type classification. This system is evaluated on a database of 10 K fluently spoken utterances collected from interactions between users and human agents.
Giuseppe Riccardi, Allen L. Gorin, Andrej Ljolje, Michael Riley 0001
ICASSP4
1997 The Watson speech recognition engine
abstract
In 1995, AT&T Research (then within Bell Labs) began work on a software-only automated speech recognition system named Watson(TM). The goal was ambitious; Watson was to serve as a single code base supporting applications ranging from PC-desktop command and control through to scaleable telephony interactive voice services. Furthermore, the software was to be the new code base for the research group, allowing fast deployment of new algorithmic advances from the lab into the field. A set of C++ objects has been developed which support these objectives. This paper gives an overview of the Watson automatic speech recognizer software architecture, describes the algorithms employed, and provides performance numbers for some sample tasks.
R. Douglas Sharp, Enrico Bocchieri, Cecilia Castillo, Sarangarajan Parthasarathy, Christopher Rath, Michael Riley 0001, James Rowland
ICASSP6
1997 Weighted determinization and minimization for large vocabulary speech recognition
abstract
Speech recognition requires solving many space and time problems that can have a critical effect on the overall system performance. We describe the use of two general new algorithms [5] that transform recognition networks into equivalent ones that require much less time and space in large-vocabulary speech recognition. The new algorithms generalize classical automata determinization and minimization to deal properly with the probabilities of alternative hypotheses and with the relationships between units (distributions, phones, words) at different levels in the recognition system. 1. INTRODUCTION The networks used in the search stage of speech recognition systems are often highly redundant. Many paths correspond to the same word contents (word lattices and language models), or to the same phonemes (pronunciation dictionaries) for instance, with distinct weights or probabilities. More generally, at a given state of a network there might be several thousand alternative outgoing arcs, ma...
Mehryar Mohri, Michael Riley 0001
EUROSPEECH2
1997 Transducer composition for context-dependent network expansion
abstract
Context-dependent models for language units are essential in high-accuracy speech recognition. However, standard speech recognition frameworks are based on the substitution of lowerlevel models for higher-level units. Since substitution cannot express context-dependency constraints, actual recognizers use restrictive model-structure assumptions and specialized code for context-dependent models, leading to decreased flexibility and lost opportunities for automatic model optimization. Instead, we propose a recognition framework that builds in the possibility of context dependency from the start by using weighted finite-state transduction rather than substitution. The framework is implemented with a general demand-driven transducer composition algorithm that allows great flexibility in model structure, form of context dependency and network expansion method, while achieving competitive recognition performance. 1. INTRODUCTION 1.1. The Substitution Architecture In the standard architectur...
Michael Riley 0001, Fernando Pereira 0003, Mehryar Mohri
EUROSPEECH1
1996 Compilation of Weighted Finite-State Transducers from Decision Trees
abstract
We report on a method for compiling decision trees into weighted finite-state transducers. The key assumptions are that the tree predictions specify how to rewrite symbols from an input string, and the decision at each tree node is stateable in terms of regular expressions on the input string. Each leaf node can then be treated as a separate rule where the left and right contexts are constructable from the decisions made traversing the tree from the root to the leaf. These rules are compiled into transducers using the weighted rewite-rule rule-compilation algorithm described in (Mohri and Sproat, 1996).
Richard Sproat, Michael Riley 0001
ACL2
1995 The AT&t 60,000 word speech-to-text system
Michael Riley 0001, Andrej Ljolje, Donald Hindle, Fernando Pereira 0003
EUROSPEECH1
1994 Prediction of word confusabilities for speech recognition
David B. Roe, Michael Riley 0001
ICSLP2
1993 Automatic segmentation of speech for TTS
Andrej Ljolje, Michael Riley 0001
EUROSPEECH2
1992 Efficient grammar processing for a spoken language translation system
abstract
A problem with many speech understanding systems is that grammars that are more suitable for representing the relation between sentences and their meanings, such as context free grammars (CFGs) and augmented phrase structure grammars (APSGs), are computationally very demanding. On the other hand, finite state grammars are efficient, but cannot represent directly the sentence-meaning relation. The authors describe how speech recognition and language analysis can be tightly coupled by developing an APSG for the analysis component and deriving automatically from it a finite-state approximation that is used as the recognition language model. Using this technique, the authors have built an efficient translation system that is fast compared to others with comparably sized language models.>
David B. Roe, Fernando Pereira 0003, Richard Sproat, Michael Riley 0001, Pedro J. Moreno 0001, Alejandro Macarrón Larumbe
ICASSP4
1992 Optimal speech recognition using phone recognition and lexical access
Andrej Ljolje, Michael Riley 0001
ICSLP2
1992 Recognizing phonemes vs. recognizing phones: a comparison
Michael Riley 0001, Andrej Ljolje
ICSLP1
1992 A spoken language translator for restricted-domain context-free languages
David B. Roe, Pedro J. Moreno 0001, Richard Sproat, Fernando Pereira 0003, Michael Riley 0001, Alejandro Macarrón Larumbe
Speech Commun.5
1991 Automatic segmentation and labeling of speech
abstract
The authors investigate an automatic approach to segmentation of labeled speech and labeling and segmentation of speech when only the orthographic transcription of speech is available. The technique is based on a phone recognition system based on a trigram phonotactic model, gamma distribution phone duration models, and a spectral model based on five different structures for phone models of varying contextual dependencies. The alignment of speech with a given phone sequence is performed as a very constrained phone recognition task with the phonotactic model based only on the given phone sequence. When only orthographic transcription is provided, a classification-tree-based prediction of most likely phone realizations is used as an input network for the phone recognizer. The maximum likelihood phone sequence is then treated as the true phone sequence and its segment boundaries are compared with the reference boundaries.>
Andrej Ljolje, Michael Riley 0001
ICASSP2
1991 A statistical model for generating pronunciation networks
abstract
Methods to predict detailed phonetic pronunciations from a coarse phonemic transcription are described. The phonemic base forms, obtainable from orthographic text by dictionary lookup and other means, do not specify fine phonetic detail such as flapping, glottal stop insertion, or the formation of syllabic nasals and liquids. These phenomena depend on the phonetic context (often spanning word boundaries), stress environment, speaking rate, and dialect. A procedure is presented that builds decision trees, trained on the TIMIT database, using some of these features to predict pronunciation alternatives. The resulting phonetic network predicts the correct pronunciation of a phoneme on test data from the same corpus approximately 83% of the time and the correct phone was in the top five guesses 99% of the time.>
Michael Riley 0001
ICASSP1
1991 Lexical access with a statistically-derived phonetic network
abstract
A probabilistic approach to lexical access from a recognized phone sequence is presented. Lexical access is seen as finding the word sequence that maximizes the lexical likelihood of a sequence of phones and durations as recognized by a phone recognizer. This is theoretically correct for minimum error rate recognition within the model presented and is intuitively pleasing since it means that the confusion matrix of the phone recognizer will be learned and its regularities exploited. The lexical likelihoods are estimated from training data provided by the phone recognizer using statistical decision trees. Classification trees are used to estimate the phone realiziation distributions and regression trees are used to estimate the phone duration distributions. We find they can capture effectively allophonic variation, alternative pronunciation, word co-articulation and segmental durations. We describe a simpified, but efficient implementation of these models to lexical access in the DARPA resource management recognitiion task.
Michael Riley 0001, Andrej Ljolje
EUROSPEECH1
1991 Toward a spoken language translator for restricted-domain context-free languages
David B. Roe, Fernando Pereira 0003, Richard Sproat, Michael Riley 0001, Pedro J. Moreno 0001, Alejandro Macarrón Larumbe
EUROSPEECH4
1987 Beyond quasi-stationarity: Designing time-frequency representations for speech signals
abstract
This work addresses two related questions. The first is what joint time-frequency energy representations are most appropriate for speech signals, in particular, for the analysis of formant structure. Quasi-stationarity is not assumed, since it neglects dynamic regions. A set of desired properties is proposed, and a subclass of the quadratic transforms that best meets these criteria is derived, which consists of two-dimensionally smoothed Wigner distributions with gaussian kernels. The second question addressed is how to obtain suitable symbolic descriptions of the phonetically relevant features in these time-frequency surfaces. We propose time-frequency ridges in these surfaces, the 2-D analog of spectral peaks, which can be found by examining the derivatives of the time-frequency surface produced above.
Michael Riley 0001
ICASSP1