EDBT 2026 Demo / reviewers in the wild / expert
Cyril Allauzen
dblp:85/2685
· DBLP profile ↗
55ranked-venue papers
27as first author
12since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 31 · 10 first-author · 9 since 2021Artificial intelligence and machine learning · 29 · 10 first-author · 7 since 2021Theory of computation · 12 · 12 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Massive Sound Embedding Benchmark (MSEB)abstractAudio is a critical component of multimodal perception, and any truly intelligent system must demonstrate a wide range of auditory capabilities. These capabilities include transcription, classification, retrieval, reasoning, segmentation, clustering, reranking, and reconstruction. Fundamentally, each task involves transforming a raw audio signal into a meaningful 'embedding'—be it a single vector, a sequence of continuous or discrete representations, or another structured form—which then serves as the basis for generating the task's final response. To accelerate progress towards robust machine auditory intelligence, we present the Massive Sound Embedding Benchmark (MSEB): an extensible framework designed to evaluate the auditory components of any multimodal system. In its first release, MSEB offers a comprehensive suite of eight core tasks, with more planned for the future, supported by diverse datasets, including the new, large-scale Simple Voice Questions (SVQ) dataset. Our initial experiments establish clear performance headrooms, highlighting the significant opportunity to improve real-world multimodal experiences where audio is a core signal. We encourage the research community to use MSEB to assess their algorithms and contribute to its growth. The library is publicly hosted at https://github.com/google-research/mseb. Georg Heigold, Ehsan Variani, Tom Bagby, Cyril Allauzen, Ji Ma 0004, Shankar Kumar, Michael Riley 0001 |
NeurIPS | 4 |
| 2024 | A* shortest string decoding for non-idempotent semiringsabstractThe single shortest path algorithm is undefined for weighted finite-state automata over nonidempotent semirings because such semirings do not guarantee the existence of a shortest path.However, in non-idempotent semirings admitting an order satisfying a monotonicity condition (such as the plus-times or log semirings), the shortest string is well-defined.We describe an algorithm which finds the shortest string for a weighted non-deterministic automaton over such semirings using the backwards shortest distance of an equivalent deterministic automaton (DFA) as a heuristic for A* search performed over a companion idempotent semiring, This algorithm is proven to return the shortest string.There may be exponentially more states in the equivalent DFA, but the proposed algorithm needs to visit only a small fraction of them if determinization is performed "on the fly". Kyle Gorman, Cyril Allauzen |
EACL (1) | 2 |
| 2024 | Multilingual and Fully Non-Autoregressive ASR with Large Language Model Fusion: A Comprehensive StudyabstractIn the era of large models, the autoregressive nature of decoding often results in latency serving as a significant bottleneck. We propose a non-autoregressive LM-fused ASR system that effectively leverages the parallelization capabilities of accelerator hardware. Our approach combines the Universal Speech Model (USM) and the PaLM 2 language model in per-segment scoring mode, achieving an average relative WER improvement across all languages of 10.8% on FLEURS and 3.6% on YouTube captioning. Furthermore, our comprehensive ablation study analyzes key parameters such as LLM size, context length, vocabulary size, fusion methodology. For instance, we explore the impact of LLM size ranging from 128M to 340B parameters on ASR performance. This study provides valuable insights into the factors influencing the effectiveness of practical large-scale LM-fused speech recognition systems. W. Ronny Huang, Cyril Allauzen, Tongzhou Chen, Kilol Gupta, James Qin, Yu Zhang 0033, Yongqiang Wang 0011, Shuo-Yiin Chang, Tara N. Sainath |
ICASSP | 2 |
| 2023 | Large-Scale Language Model Rescoring on Long-Form DataabstractIn this work, we study the impact of Large-scale Language Models (LLM) on Automated Speech Recognition (ASR) of YouTube videos, which we use as a source for long-form ASR. We demonstrate up to 8% relative reduction in Word Error Eate (WER) on US English (en-us) and code-switched Indian English (en-in) long-form ASR test sets and a reduction of up to 30% relative on Salient Term Error Rate (STER) over a strong first-pass baseline that uses a maximum-entropy based language model. Improved lattice processing that results in a lattice with a proper (non-tree) digraph topology and carrying context from the 1-best hypothesis of the previous segment(s) results in significant wins in rescoring with LLMs. We also find that the gains in performance from the combination of LLMs trained on vast quantities of available data (such as C4 [1]) and conventional neural LMs is additive and significantly outperforms a strong first-pass baseline with a maximum entropy LM. Tongzhou Chen, Cyril Allauzen, Daniel S. Park, David Rybach, W. Ronny Huang, Rodrigo Cabrera, Kartik Audhkhasi, Bhuvana Ramabhadran, Pedro J. Moreno 0001, Michael Riley 0001 |
ICASSP | 2 |
| 2023 | E2E Segmentation in a Two-Pass Cascaded Encoder ASR ModelabstractWe explore unifying a neural segmenter with two-pass cascaded encoder ASR into a single model. A key challenge is allowing the segmenter (which runs in real-time, synchronously with the decoder) to finalize the non-causal 2nd pass (which runs 900 ms behind real-time) without introducing user-perceived latency or deletion errors during inference. We propose a design where the neural segmenter is integrated with the causal 1st pass decoder to emit a end-of-segment (EOS) signal in real-time. The EOS signal is then used to finalize the non-causal 2nd pass. We experiment with different ways to finalize the 2nd pass, and find that a dummy frame injection strategy allows for simultaneous high quality 2nd pass results and low finalization latency. On a real-world long-form captioning task (YouTube), we achieve 2.4% relative WER and 140 ms EOS latency gains over a baseline VAD-based segmenter with the same cascaded encoder. W. Ronny Huang, Shuo-Yiin Chang, Tara N. Sainath, Yanzhang He, David Rybach, Robert David 0002, Rohit Prabhavalkar, Cyril Allauzen, Cal Peyser, Trevor Strohman |
ICASSP | 8 |
| 2023 | Improving Contextual Biasing with Text InjectionabstractIn this work, we present a model-based approach to improving contextual biasing that improves quality without drastically increasing model computation during inference. Specifically, we look at injecting text data during training which is representative of contextually-relevant context that will be seen at inference, using a modality-matching text injection method known as JOIST. As JOIST injects text data directly into the E2E model, there is no additional model computation during inference, which is a big difference compared to most model-based biasing techniques. We find that our proposed approach, when combined with an FST-based context model, improves recognition of contacts between 5–15% relative. Tara N. Sainath, Rohit Prabhavalkar, Diamantino Caseiro, Pat Rondon, Cyril Allauzen |
ICASSP | 5 |
| 2023 | Alignment Entropy RegularizationabstractExisting training criteria in automatic speech recognition (ASR) permit the model to freely explore more than one time alignments between the feature and label sequences. In this paper, we use entropy to measure a model’s uncertainty, i.e. how it chooses to distribute the probability mass over the set of allowed alignments. Furthermore, we evaluate the effect of entropy regularization in encouraging the model to distribute the probability mass only on a smaller subset of allowed alignments. Experiments show that entropy regularization enables a much simpler decoding method without sacrificing word error rate, and provides better time alignment quality. Ehsan Variani, David Rybach, Cyril Allauzen, Michael Riley 0001 |
ICASSP | 4 |
| 2022 | E2E Segmenter: Joint Segmenting and Decoding for Long-Form ASRabstractImproving the performance of end-to-end ASR models on long utterances ranging from minutes to hours in length is an ongoing challenge in speech recognition. A common solution is to segment the audio in advance using a separate voice activity detector (VAD) that decides segment boundary locations based purely on acoustic speech/non-speech information. VAD segmenters, however, may be sub-optimal for real-world speech where, e.g., a complete sentence that should be taken as a whole may contain hesitations in the middle ("set an alarm for... 5 o'clock"). We propose to replace the VAD with an end-to-end ASR model capable of predicting segment boundaries in a streaming fashion, allowing the segmentation decision to be conditioned not only on better acoustic features but also on semantic features from the decoded text with negligible extra computation. In experiments on real world long-form audio (YouTube) with lengths of up to 30 minutes, we demonstrate 8.5% relative WER improvement and 250 ms reduction in median end-of-segment latency compared to the VAD segmenter baseline on a state-of-the-art Conformer RNN-T model. W. Ronny Huang, Shuo-Yiin Chang, David Rybach, Tara N. Sainath, Rohit Prabhavalkar, Cal Peyser, Zhiyun Lu, Cyril Allauzen |
INTERSPEECH | 8 |
| 2022 | On Adaptive Weight Interpolation of the Hybrid Autoregressive Transducer
Ehsan Variani, Michael Riley 0001, David Rybach, Cyril Allauzen, Tongzhou Chen, Bhuvana Ramabhadran |
INTERSPEECH | 4 |
| 2022 | Global Normalization for Streaming Speech Recognition in a Modular FrameworkabstractWe introduce the Globally Normalized Autoregressive Transducer (GNAT) for addressing the label bias problem in streaming speech recognition. Our solution admits a tractable exact computation of the denominator for the sequence-level normalization. Through theoretical and empirical results, we demonstrate that by switching to a globally normalized model, the word error rate gap between streaming and non-streaming speech-recognition models can be greatly reduced (by more than 50% on the Librispeech dataset). This model is developed in a modular framework which encompasses all the common neural speech recognition models. The modularity of this framework enables controlled comparison of modelling choices and creation of new models. A JAX implementation of our models has been open sourced. Ehsan Variani, Michael Riley 0001, David Rybach, Matt Shannon, Cyril Allauzen |
NeurIPS | 6 |
| 2021 | A Hybrid Seq-2-Seq ASR Design for On-Device and Server Applications
Cyril Allauzen, Ehsan Variani, Michael Riley 0001, David Rybach, Hao Zhang 0010 |
Interspeech | 1 |
| 2021 | An Efficient Streaming Non-Recurrent On-Device End-to-End Model with Improvements to Rare-Word Modeling
Tara N. Sainath, Yanzhang He, Arun Narayanan, Rami Botros, Ruoming Pang, David Rybach, Cyril Allauzen, Ehsan Variani, James Qin, Quoc-Nam Le-The, Shuo-Yiin Chang, Bo Li 0028, Anmol Gulati, Chung-Cheng Chiu, Diamantino Caseiro, Wei Li 0133, Qiao Liang 0001, Pat Rondon |
Interspeech | 7 |
| 2020 | Hybrid Autoregressive Transducer (HAT)abstractThis paper proposes and evaluates the hybrid autoregressive transducer (HAT) model, a time-synchronous encoder-decoder model that preserves the modularity of conventional automatic speech recognition systems. The HAT model provides a way to measure the quality of the internal language model that can be used to decide whether inference with an external language model is beneficial or not. We evaluate our proposed model on a large-scale voice search task. Our experiments show significant improvements in WER compared to the state-of-the-art approaches1. Ehsan Variani, David Rybach, Cyril Allauzen, Michael Riley 0001 |
ICASSP | 3 |
| 2019 | Federated Learning of N-Gram Language ModelsabstractMingqing Chen, Ananda Theertha Suresh, Rajiv Mathews, Adeline Wong, Cyril Allauzen, Françoise Beaufays, Michael Riley. Proceedings of the 23rd Conference on Computational Natural Language Learning (CoNLL). 2019. Mingqing Chen, Ananda Theertha Suresh, Rajiv Mathews, Adeline Wong, Cyril Allauzen, Françoise Beaufays, Michael Riley 0001 |
CoNLL | 5 |
| 2019 | Contextual Recovery of Out-of-Lattice Named Entities in Automatic Speech Recognition
Jack Serrino, Leonid Velikovich, Petar S. Aleksic, Cyril Allauzen |
INTERSPEECH | 4 |
| 2018 | Algorithms for Weighted Finite Automata with Failure Transitions
Cyril Allauzen, Michael Riley 0001 |
CIAA | 1 |
| 2015 | Improved recognition of contact names in voice commandsabstractThe recognition of contact names in mobile-device voice commands is a challenging problem. Some of the difficulties include potentially infinite vocabularies, low probability of contact tokens in the language model (LM), increased false triggering of contact voice commands when none are spoken, and very large and noisy contact name lists. In this paper we suggest solutions for each of these difficulties. We address low prior probability and out-of-vocabulary contact name problems by using class-based language models, and creating on-the-fly user dependent small language models containing only relevant names. These models are compiled dynamically based on analysis of the mobile device state. Since these solutions can increase biasing towards contact names during recognition, it is crucial to monitor false triggering. To properly balance this bias we introduce the concept of a contacts insertion reward. This reward is tuned using both positive and negative test sets. We show significant recognition performance improvements on data sets in three languages, without negatively impacting the overall system performance. The improvements are obtained in both offline evaluations as well as on live traffic experiments. Petar S. Aleksic, Cyril Allauzen, David Elson, Aleksandar Kracun, Diego Melendo Casado, Pedro J. Moreno 0001 |
ICASSP | 2 |
| 2015 | Bringing contextual information to google speech recognitionabstractIn automatic speech recognition on mobile devices, very often what a user says strongly depends on the particular context he or she is in. The n-grams relevant to the context are often not known in advance. The context can depend on, for example, particular dialog state, options presented to the user, conversation topic, location, etc. Speech recognition of sentences that include these n-grams can be challenging, as they are often not well represented in a language model (LM) or even include out-of-vocabulary (OOV) words. In this paper, we propose a solution for using contextual information to improve speech recognition accuracy. We utilize an on-the-fly rescoring mechanism to adjust the LM weights of a small set of n-grams relevant to the particular context during speech decoding. Our solution handles out of vocabulary words. It also addresses efficient combination of multiple sources of context and it even allows biasing class based language models. We show significant speech recognition accuracy improvements on several datasets, using various types of contexts, without negatively impacting the overall system. The improvements are obtained in both offline and live experiments. Petar S. Aleksic, Mohammadreza Ghodsi, Assaf Hurwitz Michaely, Cyril Allauzen, Keith B. Hall, Brian Roark, David Rybach, Pedro J. Moreno 0001 |
INTERSPEECH | 4 |
| 2015 | Rapid vocabulary addition to context-dependent decoder graphs
Cyril Allauzen, Michael Riley 0001 |
INTERSPEECH | 1 |
| 2015 | Composition-based on-the-fly rescoring for salient n-gram biasingabstractWe introduce a technique for dynamically applying contextually-derived language models to a state-of-the-art speech recognition system. These generally small-footprint models can be seen as a generalization of cache-based models [1], whereby contextually salient n-grams are derived from relevant sources (not just user generated language) to produce a model intended for combination with the baseline language model. The derived models are applied during first-pass decoding as a form of on-the-fly composition between the decoder search graph and the set of weighted contextual n-grams. We present a construction algorithm which takes a trie representing the contextual n-grams and produces a weighted finite state automaton which is more compact than a standard n-gram machine. Finally, we present a set of empirical results on the recognition of spoken search queries where a contextual model encoding recent trending queries is applied using the proposed technique. Keith B. Hall, Eunjoon Cho, Cyril Allauzen, Françoise Beaufays, Noah Coccaro, Kaisuke Nakajima, Michael Riley 0001, Brian Roark, David Rybach, Linda Zhang 0004 |
INTERSPEECH | 3 |
| 2014 | Encoding linear models as weighted finite-state transducersabstractWe present algorithms, implemented as an extension to the OpenFst library, that yield a class of transducers that encode linear models for structured inference tasks like segmentation and tagging.This allows the use of general finite-state operations with such models.For instance, finite-state composition can be used to apply the model to lattice input (or other more general automata) and then the result automaton can be passed to subsequent processing such as general shortest path algorithms.We demonstrate the use of the library extension on graphemeto-phoneme conversion, encoding multiple varieties of linear models for that task, and achieve solid PER/WER gains over previous best reported results on g2p conversion of a publicly available dataset (CMU). Cyril Allauzen, Keith B. Hall, Michael Riley 0001, Brian Roark |
INTERSPEECH | 2 |
| 2014 | Pushdown Automata in Statistical Machine TranslationabstractThis article describes the use of pushdown automata (PDA) in the context of statistical machine translation and alignment under a synchronous context-free grammar. We use PDAs to compactly represent the space of candidate translations generated by the grammar when applied to an input sentence. General-purpose PDA algorithms for replacement, composition, shortest path, and expansion are presented. We describe HiPDT, a hierarchical phrase-based decoder using the PDA representation and these algorithms. We contrast the complexity of this decoder with a decoder based on a finite state automata representation, showing that PDAs provide a more suitable framework to achieve exact decoding for larger synchronous context-free grammars and smaller language models. We assess this experimentally on a large-scale Chinese-to-English alignment and translation task. In translation, we propose a two-pass decoding strategy involving a weaker language model in the first-pass to address the results of PDA complexity analysis. We study in depth the experimental conditions and tradeoffs in which HiPDT can achieve state-of-the-art performance for large-scale SMT. Cyril Allauzen, William J. Byrne, Adrià de Gispert, Gonzalo Iglesias, Michael Riley 0001 |
Comput. Linguistics | 1 |
| 2013 | Smoothed marginal distribution constraints for language modeling
Brian Roark, Cyril Allauzen, Michael Riley 0001 |
ACL (1) | 2 |
| 2013 | Mixture of mixture n-gram language modelsabstractThis paper presents a language model adaptation technique to build a single static language model from a set of language models each trained on a separate text corpus while aiming to maximize the likelihood of an adaptation data set given as a development set of sentences. The proposed model can be considered as a mixture of mixture language models. The mixture model at the top level is a sentence-level mixture model where each sentence is assumed to be drawn from one of a discrete set of topic or task clusters. After selecting a cluster, each n-gram is assumed to be drawn from one of the given n-gram language models. We estimate cluster mixture weights and n-gram language model mixture weights for each cluster using the expectation-maximization (EM) algorithm to seek the parameter estimates maximizing the likelihood of the development sentences. This mixture of mixture models can be represented efficiently as a static n-gram language model using the previously proposed Bayesian language model interpolation technique. We show a significant improvement with this technique (both perplexity and WER) compared to the standard one level interpolation scheme. Hasim Sak, Cyril Allauzen, Kaisuke Nakajima, Françoise Beaufays |
ASRU | 2 |
| 2013 | Language model verbalization for automatic speech recognitionabstractTranscribing speech in properly formatted written language presents some challenges for automatic speech recognition systems. The difficulty arises from the conversion ambiguity between verbal and written language in both directions. Non-lexical vocabulary items such as numeric entities, dates, times, abbreviations and acronyms are particularly ambiguous. This paper describes a finite-state transducer based approach that improves proper transcription of these entities. The approach involves training a language model in the written language domain, and integrating verbal expansions of vocabulary items as a finite-state model into the decoding graph construction. We build an inverted finite-state transducer to map written vocabulary items to alternate verbal expansions using rewrite rules. Then, this verbalizer transducer is composed with the n-gram language model to obtain a verbalized language model, whose input labels are in the verbal language domain while output labels are in the written language domain. We show that the proposed approach is very effective in improving the recognition accuracy of numeric entities. Hasim Sak, Françoise Beaufays, Kaisuke Nakajima, Cyril Allauzen |
ICASSP | 4 |
| 2013 | Pre-initialized composition for large-vocabulary speech recognitionabstractThis paper describes a modified composition algorithm that is used for combining two finite-state transducers, representing the context-dependent lexicon and the language model respec-tively, in large vocabulary speech recogntion. This algorithm is a hybrid between the static and dynamic expansion of the re-sultant transducer, which maps from context-dependent phones to words and is searched during decoding. The approach is to pre-compute part of the recognition transducer and leave the balance to be expanded during decoding. This method allows for a fine-grained trade-off between space and time in recogni-tion. For example, the time overhead of purely dynamic expan-sion can be reduced by over six-fold with only a 20 % increase in memory in a collection of large-vocabulary recognition tasks available on the Google Android platform. Cyril Allauzen, Michael Riley 0001 |
INTERSPEECH | 1 |
| 2013 | Written-domain language modeling for automatic speech recognitionabstractLanguage modeling for automatic speech recognition (ASR) systems has been traditionally in the verbal domain. In this paper, we present finite-state modeling techniques that we developed for language modeling in the written domain. The first technique we describe is for the verbalization of written-domain vocabulary items, which include lexical and non-lexical entities. The second technique is the decomposition–recomposition approach to address the out-of-vocabulary (OOV) and the data sparsity problems with non-lexical entities such as URLs, email addresses, phone numbers, and dollar amounts. We evaluate the proposed written-domain language modeling approaches on a very large vocabulary speech recognition system for English. We show that the written-domain language modeling improves the speech recognition and the ASR transcript rendering accuracy in the written domain over a baseline system using a verbal-domain language model. In addition, the writtendomain system is much simpler since it does not require complex and error-prone text normalization and denormalization rules, which are generally required for verbal-domain language modeling. Hasim Sak, Yun-Hsuan Sung, Françoise Beaufays, Cyril Allauzen |
INTERSPEECH | 4 |
| 2012 | Voice Query RefinementabstractWe describe a system for the refinement of spoken search queries. Given an initial query (Northern Italian restaurants in New York), instead of requiring a fully-specified followup query (Korean restaurants in New York), a more natural, abbreviated update query (Korean instead) may be spoken. The system consists of a parsing step to identify the type and arguments of the refinement, a candidate generation step to enumerate the possible refinements, and a model classification step to select the best refinement. We present results on test query refinements given both to this system and to human judges that show the automated system outperforms the human judges on that data set. Index terms: spoken dialog systems, voice search, query refinement 1. Cyril Allauzen, Edward Benson, Ciprian Chelba, Michael Riley 0001, Johan Schalkwyk |
INTERSPEECH | 1 |
| 2012 | A Pushdown Transducer Extension for the OpenFst Library
Cyril Allauzen, Michael Riley 0001 |
CIAA | 1 |
| 2011 | Hierarchical Phrase-based Translation Representations
Gonzalo Iglesias, Cyril Allauzen, William J. Byrne, Adrià de Gispert, Michael Riley 0001 |
EMNLP | 2 |
| 2011 | Bayesian Language Model Interpolation for Mobile Speech InputabstractThis paper explores various static interpolation methods for approximating a single dynamically-interpolated language model used for a variety of recognition tasks on the Google Android platform. The goal is to find the statically-interpolated firstpass LM that best reduces search errors in a two-pass system or that even allows eliminating the more complex dynamic second pass entirely. Static interpolation weights that are uniform, prior-weighted, and the maximum likelihood, maximum a posteriori, and Bayesian solutions are considered. Analysis argues and recognition experiments on Android test data show that a Bayesian interpolation approach performs best. Cyril Allauzen, Michael Riley 0001 |
INTERSPEECH | 1 |
| 2011 | Unary Data Structures for Language ModelsabstractLanguage models are important components of speech recognition and machine translation systems. Trained on billions of words, and consisting of billions of parameters, language models often are the single largest components of these systems. There have been many proposed techniques to reduce the storage requirements for language models. A technique based upon pointer-free compact storage of ordinal trees shows compression competitive with the best proposed systems, while retaining the full finite state structure, and without using computationally expensive block compression schemes or lossy quantization techniques. Index Terms: n-gram language models, unary data structures 1. Jeffrey S. Sorensen, Cyril Allauzen |
INTERSPEECH | 2 |
| 2010 | On-demand language model interpolation for mobile speech inputabstractGoogle offers several speech features on the Android mobile operating system: search by voice, voice input to any text field, and an API for application developers. As a result, our speech recognition service must support a wide range of usage scenarios and speaking styles: relatively short search queries, addresses, business names, dictated SMS and e-mail messages, and a long tail of spoken input to any of the applications users may install. We present a method of on-demand language model interpolation in which contextual information about each utterance determines interpolation weights among a number of n-gram language models. On-demand interpolation results in an 11.2% relative reduction in WER compared to using a single language model to handle all traffic. Index Terms: language modeling, interpolation, mobile Brandon Ballinger, Cyril Allauzen, Alexander Gruenstein, Johan Schalkwyk |
INTERSPEECH | 2 |
| 2010 | Expected Sequence Similarity Maximization
Cyril Allauzen, Shankar Kumar, Wolfgang Macherey, Mehryar Mohri, Michael Riley 0001 |
HLT-NAACL | 1 |
| 2010 | Large-Scale Training of SVMs with Automata Kernels
Cyril Allauzen, Corinna Cortes, Mehryar Mohri |
CIAA | 1 |
| 2010 | Filters for Efficient Composition of Weighted Finite-State Transducers
Cyril Allauzen, Michael Riley 0001, Johan Schalkwyk |
CIAA | 1 |
| 2009 | A generalized composition algorithm for weighted finite-state transducersabstractThis paper describes a weighted finite-state transducer composition algorithm that generalizes the concept of the composition filter and presents filters that remove useless epsilon paths and push forward labels and weights along epsilon paths. This filtering permits the compostion of large speech recognition contextdependent lexicons and language models much more efficiently in time and space than previously possible. We present experiments on Broadcast News and a spoken query task that demonstrate an ∼5 % to 10 % overhead for dynamic, runtime composition compared to a static, offline composition of the recognition transducer. To our knowledge, this is the first such system with so little overhead. Cyril Allauzen, Michael Riley 0001, Johan Schalkwyk |
INTERSPEECH | 1 |
| 2008 | General Algorithms for Testing the Ambiguity of Finite Automata
Cyril Allauzen, Mehryar Mohri, Ashish Rastogi |
Developments in Language Theory | 1 |
| 2008 | Sequence kernels for predicting protein essentialityabstractThe problem of identifying the minimal gene set required to sustain life is of crucial importance in understanding cellular mechanisms and designing therapeutic drugs. This work describes several kernel-based solutions for predicting essential genes that outperform existing models while using less training data. Our first solution is based on a semi-manually designed kernel derived from the Pfam database, which includes several Pfam domains. We then present novel and general domain-based sequence kernels that capture sequence similarity with respect to several domains made of large sets of protein sequences. We show how to deal with the large size of the problem -- several thousands of domains with individual domains sometimes containing thousands of sequences -- by representing and efficiently computing these kernels using automata. We report results of extensive experiments demonstrating that they compare favorably with the Pfam kernel in predicting protein essentiality, while requiring no manual tuning. Cyril Allauzen, Mehryar Mohri, Ameet Talwalkar |
ICML | 1 |
| 2008 | 3-Way Composition of Weighted Finite-State Transducers
Cyril Allauzen, Mehryar Mohri |
CIAA | 1 |
| 2007 | OpenFst: A General and Efficient Weighted Finite-State Transducer Library
Cyril Allauzen, Michael Riley 0001, Johan Schalkwyk, Wojciech Skut, Mehryar Mohri |
CIAA | 1 |
| 2006 | A Unified Construction of the Glushkov, Follow, and Antimirov Automata
Cyril Allauzen, Mehryar Mohri |
MFCS | 1 |
| 2005 | The AT&T WATSON Speech RecognizerabstractThis paper describes the AT&T WATSON real-time speech recognizer, the product of several decades of research at AT&T. The recognizer handles a wide range of vocabulary sizes and is based on continuous-density hidden Markov models for acoustic modeling and finite state networks for language modeling. The recognition network is optimized for efficient search. We identify the algorithms used for high-accuracy, real-time and low-latency recognition. We present results for small and large vocabulary tasks taken from the AT&T VoiceTone/sup /spl reg// service, showing word accuracy improvement of about 5% absolute and real-time processing speed-up by a factor between 2 and 3. Vincent Goffin, Cyril Allauzen, Enrico Bocchieri, Dilek Hakkani-Tür, Andrej Ljolje, Sarangarajan Parthasarathy, Mazin G. Rahim, Giuseppe Riccardi, Murat Saraclar |
ICASSP (1) | 2 |
| 2005 | Robust access to large structured data using voice form-fillingabstractA method for accurate and scalable form-filling by voice is presented. A form consists of a number of fields. Accurate speech recognition is achieved by applying task-specific inter-field constraints. The task constraints are specified typically by providing a database of valid form-entries, such as an employee directory containing the name, location, and telephone number. Scalability to very large vocabularies, number of fields, and the ability to accept a variety of user responses, is achieved by a two-pass recognition scheme. An index-based retrieval method is used in the first-pass to produce a shortlist of form-entries. These are rescored in the second-pass to obtain the final result. Experiments on a simple corporate directory access application are presented to demonstrate that the new approach compares favorably, in terms of computing needs, with a traditional one-pass speech recognition system. Experiments on a national street address recognition application are presented to demonstrate that the new approach scales very well to large tasks Cyril Allauzen, R. Munkong |
INTERSPEECH | 2 |
| 2004 | Statistical Modeling for Unit Selection in Speech SynthesisabstractTraditional concatenative speech synthesis systems use a number of heuristics to define the target and concatenation costs, essential for the design of the unit selection component. In contrast to these approaches, we introduce a general statistical modeling framework for unit selection inspired by automatic speech recognition. Given appropriate data, techniques based on that framework can result in a more accurate unit selection, thereby improving the general quality of a speech synthesizer. They can also lead to a more modular and a substantially more efficient system.We present a new unit selection system based on statistical modeling. To overcome the original absence of data, we use an existing high-quality unit selection system to generate a corpus of unit sequences. We show that the concatenation cost can be accurately estimated from this corpus using a statistical n-gram language model over units. We used weighted automata and transducers for the representation of the components of the system and designed a new and more efficient composition algorithm making use of string potentials for their combination. The resulting statistical unit selection is shown to be about 2.6 times faster than the last release of the AT&T Natural Voices Product while preserving the same quality, and offers much flexibility for the use and integration of new and more complex components. Mehryar Mohri, Cyril Allauzen, Michael Riley 0001 |
ACL | 2 |
| 2004 | A generalized construction of integrated speech recognition transducersabstractWe showed in previous work that weighted finite-state transducers provide a common representation for many components of a speech recognition system and described general algorithms for combining these representations to build a single optimized and compact transducer integrating all these components, directly mapping from HMM states to words. This approach works well for certain well-controlled input transducers, but presents some problems related to the efficiency of composition and the applicability of determinization and weight-pushing with more general transducers. We generalize our prior construction of the integrated speech recognition transducer to work with an arbitrary number of component transducers and, to a large extent, release the constraints imposed on the type of input transducers by providing more general solutions to these problems. This generalization allowed us to deal with cases where our prior optimization did not apply. Our experiments in the AT&T HMIHY 0300 task and an AT&T VoiceTone task show the efficiency of our generalized optimization technique. We report a 1.6 recognition speed-up in the HMIHY 0300 task, 1.8 speed-up in a VoiceTone task using a word-based language model, and 1.7 using a class-based model. Cyril Allauzen, Mehryar Mohri, Michael Riley 0001, Brian Roark |
ICASSP (1) | 1 |
| 2004 | A General Weighted Grammar Library
Cyril Allauzen, Mehryar Mohri, Brian Roark |
CIAA | 1 |
| 2004 | An optimal pre-determinization algorithm for weighted transducers
Cyril Allauzen, Mehryar Mohri |
Theor. Comput. Sci. | 1 |
| 2003 | Generalized Algorithms for Constructing Statistical Language ModelsabstractRecent text and speech processing applications such as speech mining raise new and more general problems related to the construction of language models. We present and describe in detail several new and efficient algorithms to address these more general problems and report experimental results demonstrating their usefulness. We give an algorithm for computing efficiently the expected counts of any sequence in a word lattice output by a speech recognizer or any arbitrary weighted automaton; describe a new technique for creating exact representations of n-gram language models by weighted automata whose size is practical for offline use even for a vocabulary size of about 500,000 words and an n-gram order n = 6; and present a simple and more general technique for constructing class-based language models that allows each class to represent an arbitrary weighted automaton. An efficient implementation of our algorithms and techniques has been incorporated in a general software library for language modeling, the GRM Library, that includes many other text and grammar processing functionalities. Cyril Allauzen, Mehryar Mohri, Brian Roark |
ACL | 1 |
| 2003 | Generalized optimization algorithm for speech recognition transducersabstractWeighted transducers provide a common representation for the components of a speech recognition system. In previous work, we showed that these components can be combined off-line into a single compact recognition transducer that maps directly HMM state sequences to word sequences. The construction of that recognition transducer and its efficiency of use critically depend on the use of a general optimization algorithm, determinization. However, not all weighted automata and transducers used in large-vocabulary speech recognition are determinizable. We present a general algorithm that can make an arbitrary weighted transducer determinizable and generalize our previous optimization technique for building an integrated recognition transducer to deal with arbitrary weighted transducers used in speech recognition. We report experimental results in a large- vocabulary speech recognition task, How May I Help You (HMIHY), showing that our generalized technique leads to a recognition transducer that performs as well as our original solution in the case of classical n-gram models while inserting less special symbols, and that it leads to a substantial improvement of the recognition speed, factor of 2.6, in the same task when using a class-based language model. Cyril Allauzen, Mehryar Mohri |
ICASSP (1) | 1 |
| 2003 | An Efficient Pre-determinization Algorithm
Cyril Allauzen, Mehryar Mohri |
CIAA | 1 |
| 2002 | p-Subsequentiable Transducers
Cyril Allauzen, Mehryar Mohri |
CIAA | 1 |
| 2001 | Efficient Experimental String Matching by Weak Factor Recognition
Cyril Allauzen, Maxime Crochemore, Mathieu Raffinot |
CPM | 1 |
| 2000 | Simple Optimal String Matching Algorithm
Cyril Allauzen, Mathieu Raffinot |
CPM | 1 |
| 1999 | Factor Oracle: A New Structure for Pattern Matching
Cyril Allauzen, Maxime Crochemore, Mathieu Raffinot |
SOFSEM | 1 |