VLDB 2026 Research / reviewers in the wild / expert
Brian Roark
dblp:09/246
· DBLP profile ↗
82ranked-venue papers
19as first author
8since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 64 · 17 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 31 · 2 first-authorHuman-computer interaction and ubiquitous computing · 1Theory of computation · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
17 papers |
Information extraction and text analysis · 35% Speech recognition and synthesis · 34% Language models and text generation · 12% | |
| Computer graphics and multimedia
1 paper |
Audio and music processing · 100% |
Topics — the 30 heaviest of 35, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Speech recognition and synthesis › speech analysis
language identification |
0.9 | 1 | 2025 | Improving Informally Romanized Language Identification · EMNLP 2025 |
Machine learning › Trustworthy machine learning › language model interpretability
linguistic representation analysis |
0.4 | 1 | 2019 | Meaning to Form: Measuring Systematicity as Information · ACL (1) 2019 |
Natural language and speech › Information extraction and text analysis › multilingual NLP
multilingual language modeling |
0.4 | 1 | 2019 | What Kind of Language Is Hard to Language-Model? · ACL (1) 2019 |
Machine learning › Representation and self-supervised learning
mutual information |
0.4 | 1 | 2019 | Meaning to Form: Measuring Systematicity as Information · ACL (1) 2019 |
Natural language and speech › Information extraction and text analysis
syntactic parsing |
0.4 | 5 | 2011 | Beam-Width Prediction for Efficient Context-Free Parsing · ACL 2011 Deriving lexical and syntactic expectation-based measures for psycholinguistic modeling via incremental top-down parsing · EMNLP 2009 PCFGs with Syntactic and Prosodic Indicators of Speech Repairs · ACL 2006 |
Natural language and speech › Language models and text generation
language modeling |
0.3 | 3 | 2013 | Smoothed marginal distribution constraints for language modeling · ACL (1) 2013 Discriminative Syntactic Language Modeling for Speech Recognition · ACL 2005 Generalized Algorithms for Constructing Statistical Language Models · ACL 2003 |
Natural language and speech › Information extraction and text analysis › error detection
grammatical error detection |
0.2 | 1 | 2014 | Data Driven Grammatical Error Detection in Transcripts of Children's Speech · EMNLP 2014 |
Natural language and speech › Speech recognition and synthesis
automatic speech recognition |
0.2 | 4 | 2005 | Discriminative Syntactic Language Modeling for Speech Recognition · ACL 2005 Discriminative Language Modeling with Conditional Random Fields and the Perceptron Algorithm · ACL 2004 Generalized Algorithms for Constructing Statistical Language Models · ACL 2003 |
Natural language and speech › Speech recognition and synthesis
pronunciation modeling |
0.2 | 1 | 2013 | Pair Language Models for Deriving Alternative Pronunciations and Spellings from Pronunciation Dictionaries · EMNLP 2013 |
Audio and music processing › speech recognition
language modeling |
0.1 | 1 | 2012 | Discriminative Language Modeling With Linguistic and Statistically Derived Features · IEEE Trans. Speech Audio Process. 2012 |
Natural language and speech › Language models and text generation › decoding › decoding strategy
beam search |
0.1 | 1 | 2011 | Beam-Width Prediction for Efficient Context-Free Parsing · ACL 2011 |
Natural language and speech › Information extraction and text analysis › syntactic parsing
context-free parsing |
0.1 | 1 | 2011 | Beam-Width Prediction for Efficient Context-Free Parsing · ACL 2011 |
Natural language and speech › Speech recognition and synthesis › acoustic model training
discriminative training |
0.1 | 1 | 2011 | Minimum Imputed-Risk: Unsupervised Discriminative Training for Machine Translation · EMNLP 2011 |
Natural language and speech › Machine translation
statistical machine translation |
0.1 | 1 | 2011 | Minimum Imputed-Risk: Unsupervised Discriminative Training for Machine Translation · EMNLP 2011 |
Medical and health informatics › clinical diagnosis › neurodegenerative disease diagnosis
mild cognitive impairment detection |
0.1 | 1 | 2011 | Spoken Language Derived Measures for Detecting Mild Cognitive Impairment · IEEE Trans. Speech Audio Process. 2011 |
Natural language and speech › Language models and text generation
evaluation of language models |
0.1 | 1 | 2019 | What Kind of Language Is Hard to Language-Model? · ACL (1) 2019 |
Natural language and speech › Information extraction and text analysis
psycholinguistic modeling |
0.1 | 1 | 2009 | Deriving lexical and syntactic expectation-based measures for psycholinguistic modeling via incremental top-down parsing · EMNLP 2009 |
Natural language and speech › Information extraction and text analysis › text segmentation
discourse segmentation |
0.1 | 1 | 2007 | The utility of parse-derived features for automatic discourse segmentation · ACL 2007 |
Natural language and speech › Speech recognition and synthesis › spoken language understanding
speech parsing |
0.1 | 1 | 2006 | PCFGs with Syntactic and Prosodic Indicators of Speech Repairs · ACL 2006 |
Natural language and speech › Speech recognition and synthesis › spontaneous speech processing
speech repair detection |
0.1 | 1 | 2006 | PCFGs with Syntactic and Prosodic Indicators of Speech Repairs · ACL 2006 |
Compilers and program optimization
parsing |
0.1 | 2 | 2002 | Markov Parsing: Lattice Rescoring with a Statistical Parser · ACL 2002 Efficient probabilistic top-down and left-corner parsing · ACL 1999 |
Natural language and speech › Speech recognition and synthesis › spoken language understanding
spoken language processing |
0.1 | 1 | 2014 | Data Driven Grammatical Error Detection in Transcripts of Children's Speech · EMNLP 2014 |
Natural language and speech › Language models and text generation › language modeling
discriminative language modeling |
0.0 | 1 | 2004 | Discriminative Language Modeling with Conditional Random Fields and the Perceptron Algorithm · ACL 2004 |
Natural language and speech › Information extraction and text analysis › syntactic parsing › statistical parsing
discriminative parsing |
0.0 | 1 | 2004 | Incremental Parsing with the Perceptron Algorithm · ACL 2004 |
Natural language and speech › Information extraction and text analysis › syntactic parsing
incremental parsing |
0.0 | 1 | 2004 | Incremental Parsing with the Perceptron Algorithm · ACL 2004 |
Audio and music processing
speech recognition |
0.0 | 1 | 2012 | Discriminative Language Modeling With Linguistic and Statistically Derived Features · IEEE Trans. Speech Audio Process. 2012 |
Natural language and speech › Machine translation
unsupervised machine translation |
0.0 | 1 | 2011 | Minimum Imputed-Risk: Unsupervised Discriminative Training for Machine Translation · EMNLP 2011 |
Natural language and speech › Speech recognition and synthesis › automatic speech recognition › language modeling for speech recognition
lattice rescoring |
0.0 | 1 | 2002 | Markov Parsing: Lattice Rescoring with a Statistical Parser · ACL 2002 |
Natural language and speech › Information extraction and text analysis
lexical semantics |
0.0 | 1 | 2009 | Deriving lexical and syntactic expectation-based measures for psycholinguistic modeling via incremental top-down parsing · EMNLP 2009 |
Natural language and speech › Information extraction and text analysis › lexical semantics › verb semantics
selectional preference |
0.0 | 1 | 2009 | Deriving lexical and syntactic expectation-based measures for psycholinguistic modeling via incremental top-down parsing · EMNLP 2009 |
Methods — techniques the papers use, named apart from their topics
synthetic training data · 0.9linear classifier · 0.9surprisal-based difficulty estimation · 0.4recurrent neural network · 0.4mutual information · 0.4mixed-effects modeling · 0.4dependency parsing · 0.2data-driven error labeling · 0.2smoothing · 0.2finite state transducer · 0.2topic-sensitive features · 0.1feature-based discriminative training · 0.1forced alignment · 0.1automatic parsing · 0.1ROC analysis · 0.1incremental statistical parsing · 0.0a-star search · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mining Naturally Romanized Seed Corpora without Romanizations
Adrian Benton, Alexander Gutkin, Christo Kirov, Brian Roark |
LREC | 4 |
| 2025 | Improving Informally Romanized Language IdentificationabstractThe Latin script is often used to informally write languages with non-Latin native scripts.In many cases (e.g., most languages in India), the lack of conventional spelling in the Latin script results in high spelling variability.Such romanization renders languages that are normally easily distinguished due to being written in different scripts -Hindi and Urdu, for example -highly confusable.In this work, we increase language identification (LID) accuracy for romanized text by improving the methods used to synthesize training sets.We find that training on synthetic samples which incorporate natural spelling variation yields higher LID system accuracy than including available naturally occurring examples in the training set, or even training higher capacity models.We demonstrate new state-of-the-art LID performance on romanized text from 20 Indic languages in the Bhasha-Abhijnaanam evaluation set (Madhani et al., 2023a), improving test F1 from the reported 74.7% (using a pretrained neural model) to 85.4% using a linear classifier trained solely on synthetic data and 88.2% when also training on available harvested text. Adrian Benton, Alexander Gutkin, Christo Kirov, Brian Roark |
EMNLP | 4 |
| 2024 | Context-aware Transliteration of Romanized South Asian LanguagesabstractAbstract While most transliteration research is focused on single tokens such as named entities—for example, transliteration of from the Gujarati script to the Latin script “Ahmedabad” footnoteThe most populous city in the Indian state of Gujarat. the informal romanization prevalent in South Asia and elsewhere often requires transliteration of full sentences. The lack of large parallel text collections of full sentence (as opposed to single word) transliterations necessitates incorporation of contextual information into transliteration via non-parallel resources, such as via mono-script text collections. In this article, we present a number of methods for improving transliteration in context for such a use scenario. Some of these methods in fact improve performance without making use of sentential context, allowing for better quantification of the degree to which contextual information in particular is responsible for system improvements. Our final systems, which ultimately rely upon ensembles including large pretrained language models fine-tuned on simulated parallel data, yield substantial improvements over the best previously reported results for full sentence transliteration from Latin to native script on all 12 languages in the Dakshina dataset (Roark et al. 2020), with an overall 3.3% absolute (18.6% relative) mean word-error rate reduction. Christo Kirov, Cibu Johny, Anna Katanova, Alexander Gutkin, Brian Roark |
Comput. Linguistics | 5 |
| 2022 | Criteria for Useful Automatic Romanization in South Asian LanguagesabstractThis paper presents a number of possible criteria for systems that transliterate South Asian languages from their native scripts into the Latin script, a process known as romanization. These criteria are related to either fidelity to human linguistic behavior (pronunciation transparency, naturalness and conventionality) or processing utility for people (ease of input) as well as under-the-hood in systems (invertibility and stability across languages and scripts). When addressing these differing criteria several linguistic considerations, such as modeling of prominent phonological processes and their relation to orthography, need to be taken into account. We discuss these key linguistic details in the context of Brahmic scripts and languages that use them, such as Hindi and Malayalam. We then present the core features of several romanization algorithms, implemented in a finite state transducer (FST) formalism, that address differing criteria. Implementations of these algorithms have been released as part of the Nisaba finite-state script processing library. Isin Demirsahin, Cibu Johny, Alexander Gutkin, Brian Roark |
LREC | 4 |
| 2022 | Extensions to Brahmic script processing within the Nisaba library: new scripts, languages and utilitiesabstractThe Brahmic family of scripts is used to record some of the most spoken languages in the world and is arguably the most diverse family of writing systems. In this work, we present several substantial extensions to Brahmic script functionality within the open-source Nisaba library of finite-state script normalization and processing utilities (Johny et al., 2021). First, we extend coverage from the original ten scripts to an additional ten scripts of South Asia and beyond, including some used to record endangered languages such as Dogri. Second, we augment the language layer so that scripts used by multiple languages in distinct ways can be processed correctly for more languages, such as the Bengali script when used for the low-resource language Santali. We document key changes to the finite-state engine required to support these new languages and scripts. Finally, we add new script processing utilities, including lightweight script-level reading normalization that (unlike existing visual normalization) does not preserve visual invariance, and a fixed-input transliteration mechanism specifically tailored to Brahmic text entry with ASCII characters. Alexander Gutkin, Cibu Johny, Raiomond Doctor, Lawrence Wolf-Sonkin, Brian Roark |
LREC | 5 |
| 2021 | Disambiguatory Signals are Stronger in Word-initial PositionsabstractPsycholinguistic studies of human word processing and lexical access provide ample evidence of the preferred nature of word-initial versus word-final segments, e.g., in terms of attention paid by listeners (greater) or the likelihood of reduction by speakers (lower).This has led to the conjecture-as in Wedel et al. (2019b), but common elsewhere-that languages have evolved to provide more information earlier in words than later.Informationtheoretic methods to establish such tendencies in lexicons have suffered from several methodological shortcomings that leave open the question of whether this high word-initial informativeness is actually a property of the lexicon or simply an artefact of the incremental nature of recognition.In this paper, we point out the confounds in existing methods for comparing the informativeness of segments early in the word versus later in the word, and present several new measures that avoid these confounds.When controlling for these confounds, we still find evidence across hundreds of languages that indeed there is a cross-linguistic tendency to front-load information in words. 1 Tiago Pimentel, Ryan Cotterell, Brian Roark |
EACL | 3 |
| 2021 | Finding Concept-specific Biases in Form-Meaning AssociationsabstractTiago Pimentel, Brian Roark, Søren Wichmann, Ryan Cotterell, Damián Blasi. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Tiago Pimentel, Brian Roark, Søren Wichmann, Ryan Cotterell, Damián E. Blasi |
NAACL-HLT | 2 |
| 2021 | Approximating Probabilistic Models as Weighted Finite AutomataabstractAbstract Weighted finite automata (WFAs) are often used to represent probabilistic models, such as ngram language models, because among other things, they are efficient for recognition tasks in time and space. The probabilistic source to be represented as a WFA, however, may come in many forms. Given a generic probabilistic model over sequences, we propose an algorithm to approximate it as a WFA such that the Kullback-Leibler divergence between the source model and the WFA target model is minimized. The proposed algorithm involves a counting step and a difference of convex optimization step, both of which can be performed efficiently.We demonstrate the usefulness of our approach on various tasks, including distilling n-gram models from neural models, building compact language models, and building open-vocabulary character models. The algorithms used for these experiments are available in an open-source software library. Ananda Theertha Suresh, Brian Roark, Michael Riley 0001, Vlad Schogol |
Comput. Linguistics | 2 |
| 2020 | Language-Agnostic Multilingual ModelingabstractMultilingual Automated Speech Recognition (ASR) systems allow for the joint training of data-rich and data-scarce languages in a single model. This enables data and parameter sharing across languages, which is especially beneficial for the data-scarce languages. However, most state-of-the-art multilingual models require the encoding of language information and therefore are not as flexible or scalable when expanding to newer languages. Language-independent multilingual models help to address this issue, and are also better suited for multicultural societies where several languages are frequently used together (but often rendered with different writing systems). In this paper, we propose a new approach to building a language-agnostic multilingual ASR system which transforms all languages to one writing system through a many-to-one transliteration transducer. Thus, similar sounding acoustics are mapped to a single, canonical target sequence of graphemes, effectively separating the modeling and rendering problems. We show with four Indic languages, namely, Hindi, Bengali, Tamil and Kannada, that the language-agnostic multilingual model achieves up to 10% relative reduction in Word Error Rate (WER) over a language-dependent multilingual model. Arindrima Datta, Bhuvana Ramabhadran, Jesse Emond, Anjuli Kannan, Brian Roark |
ICASSP | 5 |
| 2020 | Processing South Asian Languages Written in the Latin Script: the Dakshina DatasetabstractThis paper describes the Dakshina dataset, a new resource consisting of text in both the Latin and native scripts for 12 South Asian languages. The dataset includes, for each language: 1) native script Wikipedia text; 2) a romanization lexicon; and 3) full sentence parallel data in both a native script of the language and the basic Latin alphabet. We document the methods used for preparation and selection of the Wikipedia text in each language; collection of attested romanizations for sampled lexicons; and manual romanization of held-out sentences from the native script collections. We additionally provide baseline results on several tasks made possible by the dataset, including single word transliteration, full sentence transliteration, and language modeling of native script and romanized text. Brian Roark, Lawrence Wolf-Sonkin, Christo Kirov, Sabrina J. Mielke, Cibu Johny, Isin Demirsahin, Keith B. Hall |
LREC | 1 |
| 2020 | Phonotactic Complexity and its Trade-offsabstractWe present methods for calculating a measure of phonotactic complexity—bits per phoneme— that permits a straightforward cross-linguistic comparison. When given a word, represented as a sequence of phonemic segments such as symbols in the international phonetic alphabet, and a statistical model trained on a sample of word types from the language, we can approximately measure bits per phoneme using the negative log-probability of that word under the model. This simple measure allows us to compare the entropy across languages, giving insight into how complex a language’s phonotactics is. Using a collection of 1016 basic concept words across 106 languages, we demonstrate a very strong negative correlation of − 0.74 between bits per phoneme and the average length of words. Tiago Pimentel, Brian Roark, Ryan Cotterell |
Trans. Assoc. Comput. Linguistics | 2 |
| 2019 | What Kind of Language Is Hard to Language-Model?abstractHow language-agnostic are current state-ofthe-art NLP tools?Are there some types of language that are easier to model with current methods?In prior work (Cotterell et al., 2018) we attempted to address this question for language modeling, and observed that recurrent neural network language models do not perform equally well over all the highresource European languages found in the Europarl corpus.We speculated that inflectional morphology may be the primary culprit for the discrepancy.In this paper, we extend these earlier experiments to cover 69 languages from 13 language families using a multilingual Bible corpus.Methodologically, we introduce a new paired-sample multiplicative mixed-effects model to obtain language difficulty coefficients from at-least-pairwise parallel corpora.In other words, the model is aware of inter-sentence variation and can handle missing data.Exploiting this model, we show that "translationese" is not any easier to model than natively written language in a fair comparison.Trying to answer the question of what features difficult languages have in common, we try and fail to reproduce our earlier (Cotterell et al., 2018) observation about morphological complexity and instead reveal far simpler statistics of the data that seem to drive complexity in a much larger sample. Difficulty estimation from sentence surprisal Sabrina J. Mielke, Ryan Cotterell, Kyle Gorman, Brian Roark, Jason Eisner |
ACL (1) | 4 |
| 2019 | Meaning to Form: Measuring Systematicity as InformationabstractA longstanding debate in semiotics centers on the relationship between linguistic signs and their corresponding semantics: is there an arbitrary relationship between a word form and its meaning, or does some systematic phenomenon pervade?For instance, does the character bigram gl have any systematic relationship to the meaning of words like glisten, gleam and glow?In this work, we offer a holistic quantification of the systematicity of the sign using mutual information and recurrent neural networks.We employ these in a data-driven and massively multilingual approach to the question, examining 106 languages.We find a statistically significant reduction in entropy when modeling a word form conditioned on its semantic representation.Encouragingly, we also recover wellattested English examples of systematic affixes.We conclude with the meta-point: Our approximate effect size (measured in bits) is quite small-despite some amount of systematicity between form and meaning, an arbitrary relationship and its resulting benefits dominate human language. Tiago Pimentel, Arya McCarthy, Damián E. Blasi, Brian Roark, Ryan Cotterell |
ACL (1) | 4 |
| 2019 | Neural Models of Text Normalization for Speech ApplicationsabstractMachine learning, including neural network techniques, have been applied to virtually every domain in natural language processing. One problem that has been somewhat resistant to effective machine learning solutions is text normalization for speech applications such as text-to-speech synthesis (TTS). In this application, one must decide, for example, that 123 is verbalized as one hundred twenty three in 123 pages but as one twenty three in 123 King Ave. For this task, state-of-the-art industrial systems depend heavily on hand-written language-specific grammars. We propose neural network models that treat text normalization for TTS as a sequence-to-sequence problem, in which the input is a text token in context, and the output is the verbalization of that token. We find that the most effective model, in accuracy and efficiency, is one where the sentential context is computed once and the results of that computation are combined with the computation of each token in sequence to compute the verbalization. This model allows for a great deal of flexibility in terms of representing the context, and also allows us to integrate tagging and segmentation into the process. These models perform very well overall, but occasionally they will predict wildly inappropriate verbalizations, such as reading 3 cm as three kilometers. Although rare, such verbalizations are a major issue for TTS applications. We thus use finite-state covering grammars to guide the neural models, either during training and decoding, or just during decoding, away from such “unrecoverable” errors. Such grammars can largely be learned from data. Hao Zhang 0010, Richard Sproat, Axel H. Ng, Felix Stahlberg, Xiaochang Peng, Kyle Gorman, Brian Roark |
Comput. Linguistics | 7 |
| 2018 | Transliteration Based Approaches to Improve Code-Switched Speech Recognition PerformanceabstractCode-switching is a commonly occurring phenomenon in many multilingual communities, wherein a speaker switches between languages within a single utterance. Conventional Word Error Rate (WER) is not sufficient for measuring the performance of code-mixed languages due to ambiguities in transcription, misspellings and borrowing of words from two different writing systems. These rendering errors artificially inflate the WER of an Automated Speech Recognition (ASR) system and complicate its evaluation. Furthermore, these errors make it harder to accurately evaluate modeling errors originating from code-switched language and acoustic models. In this work, we propose the use of a new metric, transliteration-optimized Word Error Rate (toWER) that smoothes out many of these irregularities by mapping all text to one writing system and demonstrate a correlation with the amount of code-switching present in a language. We also present a novel approach to acoustic and language modeling for bilingual code-switched Indic languages using the same transliteration approach to normalize the data for three types of language models, namely, a conventional n-gram language model, a maximum entropy based language model and a Long Short Term Memory (LSTM) language model, and a state-of-the-art Connectionist Temporal Classification (CTC) acoustic model. We demonstrate the robustness of the proposed approach on several Indic languages from Google Voice Search traffic with significant gains in ASR performance up to 10% relative over the state-of-the-art baseline. Jesse Emond, Bhuvana Ramabhadran, Brian Roark, Pedro J. Moreno 0001 |
SLT | 3 |
| 2016 | Contextual Prediction Models for Speech Recognition
Yoni Halpern, Keith B. Hall, Vlad Schogol, Michael Riley 0001, Brian Roark, Gleb Skobeltsyn, Martin Bäuml |
INTERSPEECH | 5 |
| 2016 | Learning N-Gram Language Models from Uncertain Data
Vitaly Kuznetsov, Hank Liao, Mehryar Mohri, Michael Riley 0001, Brian Roark |
INTERSPEECH | 5 |
| 2015 | Bringing contextual information to google speech recognitionabstractIn automatic speech recognition on mobile devices, very often what a user says strongly depends on the particular context he or she is in. The n-grams relevant to the context are often not known in advance. The context can depend on, for example, particular dialog state, options presented to the user, conversation topic, location, etc. Speech recognition of sentences that include these n-grams can be challenging, as they are often not well represented in a language model (LM) or even include out-of-vocabulary (OOV) words. In this paper, we propose a solution for using contextual information to improve speech recognition accuracy. We utilize an on-the-fly rescoring mechanism to adjust the LM weights of a small set of n-grams relevant to the particular context during speech decoding. Our solution handles out of vocabulary words. It also addresses efficient combination of multiple sources of context and it even allows biasing class based language models. We show significant speech recognition accuracy improvements on several datasets, using various types of contexts, without negatively impacting the overall system. The improvements are obtained in both offline and live experiments. Petar S. Aleksic, Mohammadreza Ghodsi, Assaf Hurwitz Michaely, Cyril Allauzen, Keith B. Hall, Brian Roark, David Rybach, Pedro J. Moreno 0001 |
INTERSPEECH | 6 |
| 2015 | Composition-based on-the-fly rescoring for salient n-gram biasingabstractWe introduce a technique for dynamically applying contextually-derived language models to a state-of-the-art speech recognition system. These generally small-footprint models can be seen as a generalization of cache-based models [1], whereby contextually salient n-grams are derived from relevant sources (not just user generated language) to produce a model intended for combination with the baseline language model. The derived models are applied during first-pass decoding as a form of on-the-fly composition between the decoder search graph and the set of weighted contextual n-grams. We present a construction algorithm which takes a trie representing the contextual n-grams and produces a weighted finite state automaton which is more compact than a standard n-gram machine. Finally, we present a set of empirical results on the recognition of spoken search queries where a contextual model encoding recent trending queries is applied using the proposed technique. Keith B. Hall, Eunjoon Cho, Cyril Allauzen, Françoise Beaufays, Noah Coccaro, Kaisuke Nakajima, Michael Riley 0001, Brian Roark, David Rybach, Linda Zhang 0004 |
INTERSPEECH | 8 |
| 2015 | Graph-Based Word Alignment for Clinical Language EvaluationabstractAmong the more recent applications for natural language processing algorithms has been the analysis of spoken language data for diagnostic and remedial purposes, fueled by the demand for simple, objective, and unobtrusive screening tools for neurological disorders such as dementia. The automated analysis of narrative retellings in particular shows potential as a component of such a screening tool since the ability to produce accurate and meaningful narratives is noticeably impaired in individuals with dementia and its frequent precursor, mild cognitive impairment, as well as other neurodegenerative and neurodevelopmental disorders. In this article, we present a method for extracting narrative recall scores automatically and highly accurately from a word-level alignment between a retelling and the source narrative. We propose improvements to existing machine translation-based systems for word alignment, including a novel method of word alignment relying on random walks on a graph that achieves alignment accuracy superior to that of standard expectation maximization-based techniques for word alignment in a fraction of the time required for expectation maximization. In addition, the narrative recall score features extracted from these high-quality word alignments yield diagnostic classification accuracy comparable to that achieved using manually assigned scores and significantly higher than that achieved with summary-level text similarity metrics used in other areas of NLP. These methods can be trivially adapted to spontaneous language samples elicited with non-linguistic stimuli, thereby demonstrating the flexibility and generalizability of these methods. Emily Tucker Prud'hommeaux, Brian Roark |
Comput. Linguistics | 2 |
| 2014 | Data Driven Grammatical Error Detection in Transcripts of Children's SpeechabstractWe investigate grammatical error detection in spoken language, and present a data-driven method to train a dependency parser to automatically identify and label grammatical errors.This method is agnostic to the label set used, and the only manual annotations needed for training are grammatical error labels.We find that the proposed system is robust to disfluencies, so that a separate stage to elide disfluencies is not required.The proposed system outperforms two baseline systems on two different corpora that use different sets of error tags.It is able to identify utterances with grammatical errors with an F1-score as high as 0.623, as compared to a baseline F1 of 0.350 on the same data. Eric Morley, Anna Eva Hallin, Brian Roark |
EMNLP | 3 |
| 2014 | Backoff inspired features for maximum entropy language modelsabstractMaximum Entropy (MaxEnt) language models [1, 2] are linear models that are typically regularized via well-known L1 or L2 terms in the likelihood objective, hence avoiding the need for the kinds of backoff or mixture weights used in smoothed n-gram language models using Katz backoff [3] and similar tech-niques. Even though backoff cost is not required to regularize the model, we investigate the use of backoff features in Max-Ent models, as well as some backoff-inspired variants. These features are shown to improve model quality substantially, as shown in perplexity and word-error rate reductions, even in very large scale training scenarios of tens or hundreds of billions of words and hundreds of millions of features. Index Terms: maximum entropy modeling, language model-ing, n-gram models, linear models Fadi Biadsy, Keith B. Hall, Pedro J. Moreno 0001, Brian Roark |
INTERSPEECH | 4 |
| 2014 | Encoding linear models as weighted finite-state transducersabstractWe present algorithms, implemented as an extension to the OpenFst library, that yield a class of transducers that encode linear models for structured inference tasks like segmentation and tagging.This allows the use of general finite-state operations with such models.For instance, finite-state composition can be used to apply the model to lattice input (or other more general automata) and then the result automaton can be passed to subsequent processing such as general shortest path algorithms.We demonstrate the use of the library extension on graphemeto-phoneme conversion, encoding multiple varieties of linear models for that task, and achieve solid PER/WER gains over previous best reported results on g2p conversion of a publicly available dataset (CMU). Cyril Allauzen, Keith B. Hall, Michael Riley 0001, Brian Roark |
INTERSPEECH | 5 |
| 2014 | Computational analysis of trajectories of linguistic development in autismabstractDeficits in semantic and pragmatic expression are among the hallmark linguistic features of autism. Recent work in deriving computational correlates of clinical spoken language measures has demonstrated the utility of automated linguistic analysis for characterizing the language of children with autism. Most of this research, however, has focused either on young children still acquiring language or on small populations covering a wide age range. In this paper, we extract numerous linguistic features from narratives produced by two groups of children with and without autism from two narrow age ranges. We find that although many differences between diagnostic groups remain constant with age, certain pragmatic measures, particularly the ability to remain on topic and avoid digressions, seem to improve. These results confirm findings reported in the psychology literature while underscoring the need for careful consideration of the age range of the population under investigation when performing clinically oriented computational analysis of spoken language. Emily Tucker Prud'hommeaux, Eric Morley, Masoud Rouhizadeh, Laura Silverman, Jan P. H. van Santen, Brian Roark, Richard Sproat, Sarah Kauper, Rachel DeLaHunta |
SLT | 6 |
| 2014 | Applications of Lexicographic Semirings to Problems in Speech and Language ProcessingabstractThis paper explores lexicographic semirings and their application to problems in speech and language processing. Specifically, we present two instantiations of binary lexicographic semirings, one involving a pair of tropical weights, and the other a tropical weight paired with a novel string semiring we term the categorial semiring. The first of these is used to yield an exact encoding of backoff models with epsilon transitions. This lexicographic language model semiring allows for off-line optimization of exact models represented as large weighted finite-state transducers in contrast to implicit (on-line) failure transition representations. We present empirical results demonstrating that, even in simple intersection scenarios amenable to the use of failure transitions, the use of the more powerful lexicographic semiring is competitive in terms of time of intersection. The second of these lexicographic semirings is applied to the problem of extracting, from a lattice of word sequences tagged for part of speech, only the single best-scoring part of speech tagging for each word sequence. We do this by incorporating the tags as a categorial weight in the second component of a 〈Tropical, Categorial〉 lexicographic semiring, determinizing the resulting word lattice acceptor in that semiring, and then mapping the tags back as output labels of the word lattice transducer. We compare our approach to a competing method due to Povey et al. (2012). Richard Sproat, Mahsa Yarmohammadi, Izhak Shafran, Brian Roark |
Comput. Linguistics | 4 |
| 2013 | Smoothed marginal distribution constraints for language modeling
Brian Roark, Cyril Allauzen, Michael Riley 0001 |
ACL (1) | 1 |
| 2013 | Improved inference and autotyping in EEG-based BCI typing systemsabstractThe RSVP Keyboard™ is a brain-computer interface (BCI)-based typing system for people with severe physical disabilities, specifically those with locked-in syndrome (LIS). It uses signals from an electroencephalogram (EEG) combined with information from an n-gram language model to select letters to be typed. One characteristic of the system as currently configured is that it does not keep track of past EEG observations, i.e., observations of user intent made while the user was in a different part of a typed message. We present a principled approach for taking all past observations into account, and show that this method results in a 20% increase in simulated typing speed under a variety of conditions on realistic stimuli. We also show that this method allows for a principled and improved estimate of the probability of the backspace symbol, by which mis-typed symbols are corrected. Finally, we demonstrate the utility of automatically typing likely letters in certain contexts, a technique that achieves increased typing speed under our new method, though not under the baseline approach. Andrew Fowler, Brian Roark, Umut Orhan, Deniz Erdogmus, Melanie Fried-Oken |
ASSETS | 2 |
| 2013 | Pair Language Models for Deriving Alternative Pronunciations and Spellings from Pronunciation DictionariesabstractPronunciation dictionaries provide a readily available parallel corpus for learning to transduce between character strings and phoneme strings or vice versa.Translation models can be used to derive character-level paraphrases on either side of this transduction, allowing for the automatic derivation of alternative pronunciations or spellings.We examine finitestate and SMT-based methods for these related tasks, and demonstrate that the tasks have different characteristics -finding alternative spellings is harder than alternative pronunciations and benefits from round-trip algorithms when the other does not.We also show that we can increase accuracy by modeling syllable stress. Russell Beckley, Brian Roark |
EMNLP | 2 |
| 2013 | Investigation of MT-based ASR confusion models for semi-supervised discriminative language modelingabstractSemi-supervised discriminative language modeling uses simulated N-best lists instead of real ASR outputs as its training examples. In this study we apply two techniques in which artificial examples are generated using a WFST and an MT system trained on pairs of reference text and ASR output. We compare the performance of these techniques with the structured prediction and ranking variants of the WER-sensitive perceptron algorithm, and contrast with the supervised case where real ASR outputs are given as input. Choosing Turkish statistical morphs as n-gram features, we analyze the similarities between the hypotheses of these three setups and the number of utilized features. We show that the MT-based system yields the lowest WER, not only because the examples generated by this technique are more effective, but also because the ranking perceptron generalizes better with this setup. When trained on a combination of artificial WFST and MT data, the structured perceptron performs as well on an unseen test set as it does when trained on real ASR output. Erinç Dikici, Emily Tucker Prud'hommeaux, Brian Roark, Murat Saraclar |
INTERSPEECH | 3 |
| 2013 | Discriminative Joint Modeling of Lexical Variation and Acoustic Confusion for Automated Narrative Retelling Assessment
Maider Lehr, Izhak Shafran, Emily Tucker Prud'hommeaux, Brian Roark |
HLT-NAACL | 4 |
| 2013 | Distributional semantic models for the evaluation of disordered language
Masoud Rouhizadeh, Emily Tucker Prud'hommeaux, Brian Roark, Jan P. H. van Santen |
HLT-NAACL | 3 |
| 2013 | Speech and Language processing as assistive technologies
Kathleen F. McCoy, John L. Arnott, Leo Ferres, Melanie Fried-Oken, Brian Roark |
Comput. Speech Lang. | 5 |
| 2013 | Huffman scanning: Using language models within fixed-grid keyboard emulation
Brian Roark, Russell Beckley, Chris Gibbons, Melanie Fried-Oken |
Comput. Speech Lang. | 1 |
| 2012 | Semi-supervised discriminative language modeling for Turkish ASRabstractWe present our work on semi-supervised learning of discriminative language models where the negative examples for sentences in a text corpus are generated using confusion models for Turkish at various granularities, specifically, word, sub-word, syllable and phone levels. We experiment with different language models and various sampling strategies to select competing hypotheses for training with a variant of the perceptron algorithm. We find that morph-based confusion models with a sample selection strategy aiming to match the error distribution of the baseline ASR system gives the best performance. We also observe that substituting half of the supervised training examples with those obtained in a semi-supervised manner gives similar results. Arda Çelebi, Hasim Sak, Erinç Dikici, Murat Saraclar, Maider Lehr, Emily Tucker Prud'hommeaux, Puyang Xu, Nathan Glenn, Damianos Karakos, Sanjeev Khudanpur, Brian Roark, Kenji Sagae, Izhak Shafran, Dan Bikel, Chris Callison-Burch, Yuan Cao 0007, Keith B. Hall, Eva Hasler, Philipp Koehn, Adam Lopez, Matt Post, Darcey Riley |
ICASSP | 11 |
| 2012 | RSVP keyboard: An EEG based typing interfaceabstractHumans need communication. The desire to communicate remains one of the primary issues for people with locked-in syndrome (LIS). While many assistive and augmentative communication systems that use various physiological signals are available commercially, the need is not satisfactorily met. Brain interfaces, in particular, those that utilize event related potentials (ERP) in electroencephalography (EEG) to detect the intent of a person noninvasively, are emerging as a promising communication interface to meet this need where existing options are insufficient. Existing brain interfaces for typing use many repetitions of the visual stimuli in order to increase accuracy at the cost of speed. However, speed is also crucial and is an integral portion of peer-to-peer communication; a message that is not delivered timely often looses its importance. Consequently, we utilize rapid serial visual presentation (RSVP) in conjunction with language models in order to assist letter selection during the brain-typing process with the final goal of developing a system that achieves high accuracy and speed simultaneously. This paper presents initial results from the RSVP Keyboard system that is under development. These initial results on healthy and locked-in subjects show that single-trial or few-trial accurate letter selection may be possible with the RSVP Keyboard paradigm. Umut Orhan, Kenneth E. Hild II, Deniz Erdogmus, Brian Roark, Barry Oken, Melanie Fried-Oken |
ICASSP | 4 |
| 2012 | Hallucinated n-best lists for discriminative language modelingabstractThis paper investigates semi-supervised methods for discriminative language modeling, whereby n-best lists are “hallucinated” for given reference text and are then used for training n-gram language models using the perceptron algorithm. We perform controlled experiments on a very strong baseline English CTS system, comparing three methods for simulating ASR output, and compare the results with training with “real” n-best list output from the baseline recognizer. We find that methods based on extracting phrasal cohorts - similar to methods from machine translation for extracting phrase tables - yielded the largest gains of our three methods, achieving over half of the WER reduction of the fully supervised methods. Kenji Sagae, Maider Lehr, Emily Tucker Prud'hommeaux, Puyang Xu, Nathan Glenn, Damianos Karakos, Sanjeev Khudanpur, Brian Roark, Murat Saraclar, Izhak Shafran, Dan Bikel, Chris Callison-Burch, Yuan Cao 0007, Keith B. Hall, Eva Hasler, Philipp Koehn, Adam Lopez, Matt Post, Darcey Riley |
ICASSP | 8 |
| 2012 | Continuous space discriminative language modelingabstractDiscriminative language modeling is a structured classification problem. Log-linear models have been previously used to address this problem. In this paper, the standard dot-product feature representation used in log-linear models is replaced by a non-linear function parameterized by a neural network. Embeddings are learned for each word and features are extracted automatically through the use of convolutional layers. Experimental results show that as a stand-alone model the continuous space model yields significantly lower word error rate (1% absolute), while having a much more compact parameterization (60%-90% smaller). If the baseline scores are combined, our approach performs equally well. Puyang Xu, Sanjeev Khudanpur, Maider Lehr, Emily Tucker Prud'hommeaux, Nathan Glenn, Damianos Karakos, Brian Roark, Kenji Sagae, Murat Saraclar, Izhak Shafran, Dan Bikel, Chris Callison-Burch, Yuan Cao 0007, Keith B. Hall, Eva Hasler, Philipp Koehn, Adam Lopez, Matt Post, Darcey Riley |
ICASSP | 7 |
| 2012 | Deriving conversation-based features from unlabeled speech for discriminative language modeling
Damianos Karakos, Brian Roark, Izhak Shafran, Kenji Sagae, Maider Lehr, Emily Tucker Prud'hommeaux, Puyang Xu, Nathan Glenn, Sanjeev Khudanpur, Murat Saraclar, Dan Bikel, Mark Dredze, Chris Callison-Burch, Yuan Cao 0007, Keith B. Hall, Eva Hasler, Philipp Koehn, Adam Lopez, Matt Post, Darcey Riley |
INTERSPEECH | 2 |
| 2012 | Fully Automated Neuropsychological Assessment for Detecting Mild Cognitive ImpairmentabstractWe present an end-to-end system for automatically scoring spoken responses to a narrative recall test administered to seniors when screening for cognitive impairment. In Wechsler Logical Memory (WLM) test, a patient listens to a brief narrative, then retells the story once immediately and again after a brief delay. We transcribe the retellings automatically using an ASR system, align the transcripts to the source narrative, extract features that replicate the standard clinical scoring method, and then use the features for automatic assessment using a classifier. On a test corpus of 72 subjects, we empirically evaluate different ASR adaptation strategies and analyze the errors with respect to clinical assessment. Despite imperfect recognition, the system presented here yields classification accuracy comparable to that of manually assigned scores. Our results show that automatic assessment of neuropsychological tests such as the WLM is practical for screening large cohorts. Index Terms: clinical diagnostics, classifying mild cognitive impairment Maider Lehr, Emily Tucker Prud'hommeaux, Izhak Shafran, Brian Roark |
INTERSPEECH | 4 |
| 2012 | Phrasal Cohort Based Unsupervised Discriminative Language Modeling
Puyang Xu, Brian Roark, Sanjeev Khudanpur |
INTERSPEECH | 2 |
| 2012 | Finite-State Chart Constraints for Reduced Complexity Context-Free Parsing PipelinesabstractWe present methods for reducing the worst-case and typical-case complexity of a context-free parsing pipeline via hard constraints derived from finite-state pre-processing. We perform O(n) predictions to determine if each word in the input sentence may begin or end a multi-word constituent in chart cells spanning two or more words, or allow single-word constituents in chart cells spanning the word itself. These pre-processing constraints prune the search space for any chart-based parsing algorithm and significantly decrease decoding time. In many cases cell population is reduced to zero, which we term chart cell “closing.” We present methods for closing a sufficient number of chart cells to ensure provably quadratic or even linear worst-case complexity of context-free inference. In addition, we apply high precision constraints to achieve large typical-case speedups and combine both high precision and worst-case bound constraints to achieve superior performance on both short and long strings. These bounds on processing are achieved without reducing the parsing accuracy, and in some cases accuracy improves. We demonstrate that our method generalizes across multiple grammars and is complementary to other pruning techniques by presenting empirical results for both exact and approximate inference using the exhaustive CKY algorithm, the Charniak parser, and the Berkeley parser. We also report results parsing Chinese, where we achieve the best reported results for an individual model on the commonly reported data set. Brian Roark, Kristy Hollingshead, Nathan Bodenstab |
Comput. Linguistics | 1 |
| 2012 | Discriminative Language Modeling With Linguistic and Statistically Derived FeaturesabstractThis paper focuses on integrating linguistically motivated and statistically derived information into language modeling. We use discriminative language models (DLMs) as a complementary approach to the conventional$n$-gram language models to benefit from discriminatively trained parameter estimates for overlapping features. In our DLM approach, relevant information is encoded as features. Feature weights are discriminatively trained using training examples and used to re-rank the$N$-best hypotheses of the baseline automatic speech recognition (ASR) system. In addition to presenting a more complete picture of previously proposed feature sets that extract implicit information available at lexical and sub-lexical levels using both linguistic and statistical approaches, this paper attempts to incorporate semantic information in the form of topic sensitive features. We explore linguistic features to incorporate complex morphological and syntactic language characteristics of Turkish, an agglutinative language with rich morphology, into language modeling. We also apply DLMs to our sub-lexical-based ASR system where the vocabulary is composed of sub-lexical units. Obtaining implicit linguistic information from sub-lexical hypotheses is not as straightforward as word hypotheses, so we use statistical methods to derive useful information from sub-lexical units. DLMs with linguistic and statistical features yield significant, 0.8%–1.1% absolute, improvements over our baseline word-based and sub-word-based ASR systems. The explored features can be easily extended to DLM for other languages . Ebru Arisoy, Murat Saraclar, Brian Roark, Izhak Shafran |
IEEE Trans. Speech Audio Process. | 3 |
| 2011 | Beam-Width Prediction for Efficient Context-Free Parsing
Nathan Bodenstab, Aaron Dunlop, Keith B. Hall, Brian Roark |
ACL | 4 |
| 2011 | Alignment of spoken narratives for automated neuropsychological assessmentabstractNarrative recall tasks are commonly included in neurological examinations, as deficits in narrative memory are associated with disorders such as Alzheimer's dementia. We explore methods for automatically scoring narrative retellings via alignment to a source narrative. Standard alignment methods, designed for large bilingual corpora for machine translation, yield high alignment error rates (AER) on our small monolingual corpora. We present modifications to these methods that obtain a decrease in AER, an increase in scoring accuracy, and diagnostic classification performance comparable to that of manual methods, thus demonstrating the utility of these techniques for this task and other tasks relying on monolingual alignments. Emily Tucker Prud'hommeaux, Brian Roark |
ASRU | 2 |
| 2011 | Efficient determinization of tagged word lattices using categorial and lexicographic semiringsabstractSpeech and language processing systems routinely face the need to apply finite state operations (e.g., POS tagging) on results from intermediate stages (e.g., ASR output) that are naturally represented in a compact lattice form. Currently, such needs are met by converting the lattices into linear sequences (n-best scoring sequences) before and after applying the finite state operations. In this paper, we eliminate the need for this unnecessary conversion by addressing the problem of picking only the single-best scoring output labels for every input sequence. For this purpose, we define a categorial semiring that allows determinzation over strings and incorporate it into a 〈Tropical, Categorial〉 lexicographic semiring. Through examples and empirical evaluations we show how determinization in this lexicographic semiring produces the desired output. The proposed solution is general in nature and can be applied to multi-tape weighted transducers that arise in many applications. Izhak Shafran, Richard Sproat, Mahsa Yarmohammadi, Brian Roark |
ASRU | 4 |
| 2011 | Minimum Imputed-Risk: Unsupervised Discriminative Training for Machine Translation
Zhifei Li 0001, Jason Eisner, Sanjeev Khudanpur, Brian Roark |
EMNLP | 5 |
| 2011 | Extraction of Narrative Recall Patterns for Neuropsychological AssessmentabstractPoor narrative memory is associated with a variety of neurodegenerative and developmental disorders, such as autism and Alzheimer’s related dementia. Hence, narrative recall tasks are included in most standard neurological examinations. In this paper, we explore methods for automatically assessing the quality of retellings via alignment to the original narrative. Word alignments serve both to automate manual scoring and to derive other features related to narrative coherence that can be used for diagnostic classification. Despite relatively high word alignment error rates, the automatic alignments provide sufficient information to achieve nearly as accurate diagnostic classification as manual scores. Furthermore, additional features that become available with alignment provide utility in classifying subject groups. While the additional features we explore here did not provide additive gains in accuracy, they point the way to the development of many potentially useful features in this domain. Emily Tucker Prud'hommeaux, Brian Roark |
INTERSPEECH | 2 |
| 2011 | Spoken Language Derived Measures for Detecting Mild Cognitive ImpairmentabstractSpoken responses produced by subjects during neuropsychological exams can provide diagnostic markers beyond exam performance. In particular, characteristics of the spoken language itself can discriminate between subject groups. We present results on the utility of such markers in discriminating between healthy elderly subjects and subjects with mild cognitive impairment (MCI). Given the audio and transcript of a spoken narrative recall task, a range of markers are automatically derived. These markers include speech features such as pause frequency and duration, and many linguistic complexity measures. We examine measures calculated from manually annotated time alignments (of the transcript with the audio) and syntactic parse trees, as well as the same measures calculated from automatic (forced) time alignments and automatic parses. We show statistically significant differences between clinical subject groups for a number of measures. These differences are largely preserved with automation. We then present classification results, and demonstrate a statistically significant improvement in the area under the ROC curve (AUC) when using automatic spoken language derived features in addition to the neuropsychological test scores. Our results indicate that using multiple, complementary measures can aid in automatic detection of MCI. Brian Roark, Margaret Mitchell, John-Paul Hosom, Kristy Hollingshead, Jeffrey A. Kaye |
IEEE Trans. Speech Audio Process. | 1 |
| 2010 | Syntactic and sub-lexical features for Turkish discriminative language modelsabstractThis paper investigates syntactic and sub-lexical features in Turkish discriminative language models (DLMs). DLM is a feature-based language modeling approach. It reranks the ASR output with discriminatively trained feature parameters. Syntactic information is incorporated into DLM as part-of-speech (PoS) tag n-gram features and head-to-head dependency relations. Sub-lexical units are first utilized as language modeling units in the baseline recognizer. Then, sub-lexical features are used to rerank the sub-lexical hypotheses. We explore features, similar to syntactic features, on sub-lexical units to reveal the implicit morpho-syntactic information conveyed by these units. We find out that DLM yields more improvement for sub-lexical units than for words. Basic sub-lexical n-gram features result in 0.6% reduction over the baseline and morpho-syntactic features yield an additional 0.4% reduction on the test set. Ebru Arisoy, Murat Saraclar, Brian Roark, Izhak Shafran |
ICASSP | 3 |
| 2010 | Prenominal Modifier Ordering via Multiple Sequence Alignment
Aaron Dunlop, Margaret Mitchell, Brian Roark |
HLT-NAACL | 3 |
| 2009 | Deriving lexical and syntactic expectation-based measures for psycholinguistic modeling via incremental top-down parsing
Brian Roark, Asaf Bachrach, Carlos Cardenas, Christophe Pallier |
EMNLP | 1 |
| 2009 | Linear Complexity Context-Free Parsing Pipelines via Chart Constraints
Brian Roark, Kristy Hollingshead |
HLT-NAACL | 1 |
| 2008 | Classifying Chart Cells for Quadratic Complexity Context-Free Inference
Brian Roark, Kristy Hollingshead |
COLING | 1 |
| 2008 | Discriminative n-gram language modeling for TurkishabstractIn this paper Discriminative Language Models (DLMs) are applied to the Turkish Broadcast News transcription task. Turkish presents a challenge to Automatic Speech Recognition (ASR) systems due to its rich morphology. Therefore, in addition to word n-gram features, morphology based features like root n-grams and inflectional group n-grams are incorporated into DLMs in order to improve the language models. Various feature sets provide reductions in the word error rate (WER). Our best result is obtained with the inflectional group n-gram features. 1.0 % absolute improvement is achieved over the baseline model and this improvement is statistically significant at p<0.001 as measured by the NIST MAPSSWE significance test. Index Terms: discriminative language modeling, speech recognition, agglutinative languages Ebru Arisoy, Brian Roark, Izhak Shafran, Murat Saraclar |
INTERSPEECH | 2 |
| 2007 | The utility of parse-derived features for automatic discourse segmentation
Seeger Fisher, Brian Roark |
ACL | 2 |
| 2007 | Pipeline Iteration
Kristy Hollingshead, Brian Roark |
ACL | 2 |
| 2007 | The SRI/OGI 2006 spoken term detection systemabstractThis paper describes the system developed jointly at SRI and OGI for participation in the 2006 NIST Spoken Term Detection (STD) evaluation. We participated in the three genres of the English track: Broadcast News (BN), Conversational Telephone Speech (CTS), and Conference Meetings (MTG). The system consists of two phases. First, audio indexing, an offline phase, converts the input speech waveform into a searchable index. Second, term retrieval, possibly an online phase, returns a ranked list of occurrences for each search term. We used a word-based indexing approach, obtained with SRI’s large vocabulary Speech-to-Text (STT) system. Apart from describing the submitted system and its performance on the NIST evaluation metric, we study the tradeoffs between performance and system design. We examine performance versus indexing speed, effectiveness of different index ranking schemes on the NIST score, and the utility of approaches to deal with out-of-vocabulary (OOV) terms. Dimitra Vergyri, Izhak Shafran, Andreas Stolcke, Venkata Ramana Rao Gadde, Murat Akbacak, Brian Roark, Wen Wang 0001 |
INTERSPEECH | 6 |
| 2007 | Putting Linguistics into Speech Recognition: The Regulus Grammar Compiler Manny Rayner, Beth Ann Hockey, and Pierette Bouillon (NASA Ames Research Center and University of Geneva) Stanford, CA: CSLI Publications (CSLI studies in computational linguistics, edited by Ann Copestake), 2006, xiv+305 pp; hardbound, ISBN 1-57586-525-4abstractThis book provides a detailed overview of Regulus, an open-source toolkit for building and compiling grammars for spoken dialog systems, which has been used for a number of applications including Clarissa, a spoken language system in use on the International Space Station.The focus is on controlled-language user-initiative systems, in which the user drives the dialog, but is required to use constrained language.Such a spoken language interface allows for a constrained range of commands-for example, "open the pod-bay doors"-to be issued in circumstances where, ergonomically, other interface options are infeasible.The emphasis of the approach, given the kind of application that is the focus of the work, is thus less on robustness of speech recognition and more on depth of semantic processing and quality of the dialog management.It is an interesting book, one which succeeds in motivating key problems and presenting general approaches to solving them, with enough in the way of explicit details to allow even a complete novice in spoken language processing to implement simple dialog systems.The book is split into two parts.The first half is a very detailed tutorial on using Regulus to build and compile grammars.Regulus, although itself open-source, makes use of SICStus Prolog for the dialog processing and the Nuance Toolkit for speech recognition.Grammars in Regulus are specified with features requiring unification, but are compiled into context-free grammars for use with the Nuance speech recognition system.An abundance of implementation details and guidance for the reader (or grammar writer) is provided in this part of the book for building grammars as well as a dialog manager.The presentation includes example implementations handling such phenomena as ellipsis and corrections.In addition, details for building spoken language translation systems are presented within the same range of constrained language applications.The tutorial format is terrifically explicit, which will make this volume appropriate for undergraduate courses looking to provide students with hands-on exercises in building spoken dialog systems.One issue with the premise of an open-source toolkit that relies upon other software (SICStus Prolog, Nuance) for key parts of the application is that one is required to obtain and use that software.In the book, the authors note that an individual license of SICStus is available for a relatively small fee, and that Nuance has a program to license their Toolkit for research purposes.Unfortunately, since the writing of the book, corporate changes at Nuance have made obtaining such a research license more challenging, and this reviewer was only able to do so after several weeks of e-mail persistence.It would be very beneficial to future readers if the authors would scout out the current state of Brian Roark |
Comput. Linguistics | 1 |
| 2007 | Discriminative n-gram language modeling
Brian Roark, Murat Saraclar, Michael Collins 0001 |
Comput. Speech Lang. | 1 |
| 2006 | PCFGs with Syntactic and Prosodic Indicators of Speech RepairsabstractA grammatical method of combining two kinds of speech repair cues is presented. One cue, prosodic disjuncture, is detected by a decision tree-based ensemble classifier that uses acoustic cues to identify where normal prosody seems to be interrupted (Lickley, 1996). The other cue, syntactic parallelism, codifies the expectation that repairs continue a syntactic category that was left unfinished in the reparandum (Levelt, 1983). The two cues are combined in a Treebank PCFG whose states are split using a few simple tree transformations. Parsing performance on the Switchboard and Fisher corpora suggests that these two cues help to locate speech repairs in a synergistic way. John Hale, Izhak Shafran, Lisa Yung, Bonnie J. Dorr, Mary P. Harper, Anna Krasnyanskaya, Matthew Lease, Yang Liu 0004, Brian Roark, Matthew G. Snover, Robin Stewart |
ACL | 9 |
| 2006 | Reranking for Sentence Boundary Detection in Conversational SpeechabstractWe present a reranking approach to sentence-like unit (SU) boundary detection, one of the EARS metadata extraction tasks. Techniques for generating relatively small n-best lists with high oracle accuracy are presented. For each candidate, features are derived from a range of information sources, including the output of a number of parsers. Our approach yields significant improvements over the best performing system from the NIST RT-04F community evaluation Brian Roark, Yang Liu 0004, Mary P. Harper, Robin Stewart, Matthew Lease, Matthew G. Snover, Izhak Shafran, Bonnie J. Dorr, John Hale, Anna Krasnyanskaya, Lisa Yung |
ICASSP (1) | 1 |
| 2006 | SParseval: Evaluation Metrics for Parsing Speech
Brian Roark, Mary P. Harper, Eugene Charniak, Bonnie J. Dorr, Mark Johnson 0001, Jeremy G. Kahn, Yang Liu 0004, Mari Ostendorf, John Hale, Anna Krasnyanskaya, Matthew Lease, Izhak Shafran, Matthew G. Snover, Robin Stewart, Lisa Yung |
LREC | 1 |
| 2006 | Probabilistic Context-Free Grammar Induction Based on Structural Zeros
Mehryar Mohri, Brian Roark |
HLT-NAACL | 2 |
| 2006 | MAP adaptation of stochastic grammars
Michiel Bacchiani, Michael Riley 0001, Brian Roark, Richard Sproat |
Comput. Speech Lang. | 3 |
| 2006 | Utterance classification with discriminative language modeling
Murat Saraclar, Brian Roark |
Speech Commun. | 2 |
| 2005 | Discriminative Syntactic Language Modeling for Speech RecognitionabstractWe describe a method for discriminative training of a language model that makes use of syntactic features. We follow a reranking approach, where a baseline recogniser is used to produce 1000-best output for each acoustic input, and a second "reranking" model is then used to choose an utterance from these 1000-best lists. The reranking model makes use of syntactic features together with a parameter estimation method that is based on the perception algorithm. We describe experiments on the Switchboard speech recognition task. The syntactic features provide an additional 0.3% reduction in test-set error rate beyond the model of (Roark et al., 2004a; Roark et al., 2004b) (significant at p < 0.001), which makes use of a discriminatively trained n-gram model, giving a total reduction of 1.2% over the baseline Switchboard system. Michael Collins 0001, Brian Roark, Murat Saraclar |
ACL | 2 |
| 2005 | Joint Discriminative Language Modeling and Utterance ClassificationabstractThis paper investigates discriminative language modeling in a scenario with two kinds of observed errors: errors in ASR transcription and errors in utterance classification. Using the perceptron algorithm, we train joint language and class models either independently or simultaneously, under various parameter update conditions. On a large vocabulary customer service call-classification application, we show that simultaneous optimization of class, n-gram, and class/n-gram feature weights results in a significant WER reduction over a model using just n-gram features, while additionally significantly outperforming a deployed baseline in classification error rate. A range of parameter update approaches for the various feature sets are presented and evaluated. The resulting models are encoded as weighted finite-state automata, and are used by intersecting the model with word lattices. Murat Saraclar, Brian Roark |
ICASSP (1) | 2 |
| 2004 | Incremental Parsing with the Perceptron AlgorithmabstractThis paper describes an incremental parsing approach where parameters are estimated using a variant of the perceptron algorithm. A beam-search algorithm is used during both training and decoding phases of the method. The perceptron approach was implemented with the same feature set as that of an existing generative model (Roark, 2001a), and experimental results show that it gives competitive performance to the generative model on parsing the Penn treebank. We demonstrate that training a perceptron model to combine with the generative model during search provides a 2.1 percent F-measure improvement over the generative model alone, to 88.8 percent. Michael Collins 0001, Brian Roark |
ACL | 2 |
| 2004 | Discriminative Language Modeling with Conditional Random Fields and the Perceptron AlgorithmabstractThis paper describes discriminative language modeling for a large vocabulary speech recognition task. We contrast two parameter estimation methods: the perceptron algorithm, and a method based on conditional random fields (CRFs). The models are encoded as deterministic weighted finite state automata, and are applied by intersecting the automata with word-lattices that are the output from a baseline recognizer. The perceptron algorithm has the benefit of automatically selecting a relatively small feature set in just a couple of passes over the training data. However, using the feature set output from the perceptron algorithm (initialized with their weights), CRF training provides an additional 0.5% reduction in word error rate, for a total 1.8% absolute reduction from the baseline of 39.2%. Brian Roark, Murat Saraclar, Michael Collins 0001, Mark Johnson 0001 |
ACL | 1 |
| 2004 | A generalized construction of integrated speech recognition transducersabstractWe showed in previous work that weighted finite-state transducers provide a common representation for many components of a speech recognition system and described general algorithms for combining these representations to build a single optimized and compact transducer integrating all these components, directly mapping from HMM states to words. This approach works well for certain well-controlled input transducers, but presents some problems related to the efficiency of composition and the applicability of determinization and weight-pushing with more general transducers. We generalize our prior construction of the integrated speech recognition transducer to work with an arbitrary number of component transducers and, to a large extent, release the constraints imposed on the type of input transducers by providing more general solutions to these problems. This generalization allowed us to deal with cases where our prior optimization did not apply. Our experiments in the AT&T HMIHY 0300 task and an AT&T VoiceTone task show the efficiency of our generalized optimization technique. We report a 1.6 recognition speed-up in the HMIHY 0300 task, 1.8 speed-up in a VoiceTone task using a word-based language model, and 1.7 using a class-based model. Cyril Allauzen, Mehryar Mohri, Michael Riley 0001, Brian Roark |
ICASSP (1) | 4 |
| 2004 | Meta-data conditional language modelingabstractAutomatic speech recognition (ASR) often occurs in circumstances in which knowledge external to the speech signal, or meta-data, is given. For example, a company receiving a call from a customer might have access to a database record of that customer. Conditioning the ASR models directly on this information to improve the transcription accuracy is hampered because, generally, the meta-data takes on many values and a training corpus has little data for each meta-data condition. The paper presents an algorithm to construct language models conditioned on such metadata. It uses tree-based clustering of the the training data to derive automatically meta-data projections, useful as language model conditioning contexts. The algorithm was tested on a multiple domain voice mail transcription task. We compare the performance of an adapted system aware of the domain shift to a system that only has meta-data to infer that fact. The meta-data used were the caller ID strings associated with the voice mail messages. The meta-data adapted system matched the performance of the system adapted using the domain knowledge explicitly. Michiel Bacchiani, Brian Roark |
ICASSP (1) | 2 |
| 2004 | Improved name recognition with meta-data dependent name networksabstractA transcription system that requires accurate general name transcription is faced with the problem of covering the large number of names it may encounter, Without any prior knowledge, this requires a large increase in the size and complexity of the system due to the expansion of the lexicon. Furthermore, this increase will adversely affect the system performance due to the increased confusability. Here we propose a method that uses meta-data, available at runtime to ensure better name coverage without significantly increasing the system complexity. We tested this approach on a voicemail transcription task and assumed meta-data to be available in the form of a caller ID string (as it would show up on a caller ID enabled telephone) and the name of the mailbox owner. Networks representing possible spoken realization of those names are generated at runtime and included in the network of the decoder. The decoder network is built at training time using a class-dependent language model, with caller and mailbox name instances modeled as class tokens. The class tokens are replaced at test time with the name networks built from the meta-data. The proposed algorithm showed a reduction in the error rate of name tokens of 22.1%. Sameer Maskey, Michiel Bacchiani, Brian Roark, Richard Sproat |
ICASSP (1) | 3 |
| 2004 | Corrective language modeling for large vocabulary ASR with the perceptron algorithmabstractThis paper investigates error-corrective language modeling using the perceptron algorithm on word lattices. The resulting model is encoded as a weighted finite-state automaton, and is used by intersecting the model with word lattices, making it simple and inexpensive to apply during decoding. We present results for various training scenarios for the Switchboard task, including using n-gram features of different orders, and performing n-best extraction versus using full word lattices. We demonstrate the importance of making the training conditions as close as possible to testing conditions. The best approach yields a 1.3 percent improvement in first pass accuracy, which translates to 0.5 percent improvement after other rescoring passes. Brian Roark, Murat Saraclar, Michael Collins 0001 |
ICASSP (1) | 1 |
| 2004 | A General Weighted Grammar Library
Cyril Allauzen, Mehryar Mohri, Brian Roark |
CIAA | 3 |
| 2004 | Robust garden path parsingabstractThis paper presents modifications to a standard probabilistic context-free grammar that enable a predictive parser to avoid garden pathing without resorting to any ad-hoc heuristic repair. The resulting parser is shown to apply efficiently to both newspaper text and telephone conversations with complete coverage and excellent accuracy. The distribution over trees is peaked enough to allow the parser to find parses efficiently, even with the much larger search space resulting from overgeneration. Empirical results are provided for both Wall St. Journal and Switchboard test corpora. Brian Roark |
Nat. Lang. Eng. | 1 |
| 2003 | Generalized Algorithms for Constructing Statistical Language ModelsabstractRecent text and speech processing applications such as speech mining raise new and more general problems related to the construction of language models. We present and describe in detail several new and efficient algorithms to address these more general problems and report experimental results demonstrating their usefulness. We give an algorithm for computing efficiently the expected counts of any sequence in a word lattice output by a speech recognizer or any arbitrary weighted automaton; describe a new technique for creating exact representations of n-gram language models by weighted automata whose size is practical for offline use even for a vocabulary size of about 500,000 words and an n-gram order n = 6; and present a simple and more general technique for constructing class-based language models that allows each class to represent an arbitrary weighted automaton. An efficient implementation of our algorithms and techniques has been incorporated in a general software library for language modeling, the GRM Library, that includes many other text and grammar processing functionalities. Cyril Allauzen, Mehryar Mohri, Brian Roark |
ACL | 3 |
| 2003 | Unsupervised language model adaptationabstractThis paper investigates unsupervised language model adaptation, from ASR transcripts. N-gram counts from these transcripts can be used either to adapt an existing n-gram model or to build an n-gram model from scratch. Various experimental results are reported on a particular domain adaptation task, namely building a customer care application starting from a general voicemail transcription system. The experiments investigate the effectiveness of various adaptation strategies, including iterative adaptation and self-adaptation on the test data. They show an error rate reduction of 3.9% over the unadapted baseline performance, from 28% to 24.1%, using 17 hours of unsupervised adaptation material. This is 51% of the 7.7% adaptation gain obtained by supervised adaptation. Self-adaptation on the test data resulted in a 1.3% improvement over the baseline. Michiel Bacchiani, Brian Roark |
ICASSP (1) | 2 |
| 2003 | Supervised and unsupervised PCFG adaptation to novel domains
Brian Roark, Michiel Bacchiani |
HLT-NAACL | 1 |
| 2002 | Markov Parsing: Lattice Rescoring with a Statistical ParserabstractWe present a generalization of an incremental statistical parsing algorithm that allows for the re-scoring of lattices of word hypotheses, for use by a speech recognizer. This approach contrasts with other lattice parsing algorithms, which either do not provide scores for strings in the lattice (i.e. they just produce parse trees) or use search techniques (e.g. A-star) to find the best paths through the lattice, without re-scoring every arc. We show that a very large efficiency gain can be had in processing 1000-best lists without reducing word accuracy when the lists are encoded in lattices instead of trees. Further, this allows for processing arbitrary lattices without n-best extraction. This can lead to more interesting methods of combination with other models, both acoustic and language, through, for example, adaptation or confusion matrices. Brian Roark |
ACL | 1 |
| 2001 | Probabilistic Top-Down Parsing and Language ModelingabstractThis paper describes the functioning of a broad-coverage probabilistic top-down parser, and its application to the problem of language modeling for speech recognition. The paper first introduces key notions in language modeling and probabilistic parsing, and briefly reviews some previous approaches to using syntactic structure for language modeling. A lexicalized probabilistic top-down parser is then presented, which performs very well, in terms of both the accuracy of returned parses and the efficiency with which they are found, relative to the best broad-coverage statistical parsers. A new language model that utilizes probabilistic top-down parsing is then outlined, and empirical results show that it improves upon previous work in test corpus perplexity. Interpolation with a trigram model yields an exceptional improvement relative to the improvement observed by other models, demonstrating the degree to which the information captured by our parsing model is orthogonal to that captured by a trigram model. A small recognition experiment also demonstrates the utility of the model. Brian Roark |
Comput. Linguistics | 1 |
| 2000 | Compact non-left-recursive grammars using the selective left-corner transform and factoring
Mark Johnson 0001, Brian Roark |
COLING | 2 |
| 1999 | Efficient probabilistic top-down and left-corner parsingabstractThis paper examines efficient predictive broad, coverage parsing without dynamic programming. In contrast to bottom-up methods, depth-first top-down parsing produces partial parses that are fully connected trees spanning the entire left context, from which any kind of non-local dependency or partial semantic interpretation can in principle be read. We contrast two predictive parsing approaches, top-down and left-corner parsing, and find both to be viable. In addition, we find that enhancement with non-local information not only improves parser accuracy, but also substantially improves the search efficiency. Brian Roark, Mark Johnson 0001 |
ACL | 1 |