VLDB 2026 Research / reviewers in the wild / expert
Mari Ostendorf
dblp:85/2189
· DBLP profile ↗
215ranked-venue papers
18as first author
23since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 140 · 10 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 128 · 9 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 since 2021Security and privacy · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RadTimeline: Timeline Summarization for Longitudinal Radiological Lung Findings
Sitong Zhou, Meliha Yetisgen, Mari Ostendorf |
LREC | 3 |
| 2024 | Investigating the Influence of Stance-Taking on Conversational Timing of Task-Oriented Speech
Sara Ng, Gina-Anne Levow, Mari Ostendorf, Richard A. Wright |
INTERSPEECH | 3 |
| 2024 | OrchestraLLM: Efficient Orchestration of Language Models for Dialogue State TrackingabstractChia-Hsuan Lee, Hao Cheng, Mari Ostendorf. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Chia-Hsuan Lee 0001, Hao Cheng 0002, Mari Ostendorf |
NAACL-HLT | 3 |
| 2024 | Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language ModelsabstractHumans draw to facilitate reasoning: we draw auxiliary lines when solving geometry problems; we mark and circle when reasoning on maps; we use sketches to amplify our ideas and relieve our limited-capacity working memory. However, such actions are missing in current multimodal language models (LMs). Current chain-of-thought and tool-use paradigms only use text as intermediate reasoning steps. In this work, we introduce Sketchpad, a framework that gives multimodal LMs a visual sketchpad and tools to draw on the sketchpad. The LM conducts planning and reasoning according to the visual artifacts it has drawn. Different from prior work, which uses text-to-image models to enable LMs to draw, Sketchpad enables LMs to draw with lines, boxes, marks, etc., which is closer to human sketching and better facilitates reasoning. \name can also use specialist vision models during the sketching process (e.g., draw bounding boxes with object detection models, draw masks with segmentation models), to further enhance visual perception and reasoning. We experiment on a wide range of math tasks (including geometry, functions, graph, chess) and complex visual reasoning tasks. Sketchpad substantially improves performance on all tasks over strong base models with no sketching, yielding an average gain of 12.7% on math tasks, and 8.6% on vision tasks. GPT-4o with Sketchpad sets a new state of the art on all tasks, including V*Bench (80.3%), BLINK spatial reasoning (83.9%), and visual correspondence (80.8%). We will release all code and data. Yushi Hu, Dan Roth 0001, Mari Ostendorf, Luke Zettlemoyer, Noah A. Smith, Ranjay Krishna |
NeurIPS | 5 |
| 2024 | Just ASR + LLM? A Study on Speech Large Language Models' Ability to Identify And Understand Speaker in Spoken DialogueabstractIn recent years, we have observed a rapid advancement in speech language models (SpeechLLMs), catching up with humans’ listening and reasoning abilities. SpeechLLMs have demonstrated impressive spoken dialog question-answering (SQA) performance in benchmarks like Gaokao, the English listening test of the college entrance exam in China, which seemingly requires understanding both the spoken content and voice characteristics of speakers in a conversation. However, after carefully examining Gaokao’s questions, we find the correct answers to many questions can be inferred from the conversation transcript alone, i.e. without speaker segmentation and identification. Our evaluation of state-of-the-art models Qwen-Audio and WavLLM on both Gaokao and our proposed “What Do You Like?” dataset shows a significantly higher accuracy in these context-based questions than in identity-critical questions, which can only be answered reliably with correct speaker identification. The results and analysis suggest that when solving SQA, the current SpeechLLMs exhibit limited speaker awareness from the audio and behave similarly to an LLM reasoning from the conversation transcription without sound. We propose that tasks focused on identity-critical questions could offer a more accurate evaluation framework of SpeechLLMs in SQA. Junkai Wu, Xulin Fan, Bo-Ru Lu, Xilin Jiang, Nima Mesgarani, Mark Hasegawa-Johnson, Mari Ostendorf |
SLT | 7 |
| 2023 | Leveraging Multiple Sources in Automatic African American English Dialect Detection for Adults and ChildrenabstractThis paper1presents a novel system which utilizes acoustic, phonological, morphosyntactic, and prosodic information for binary automatic dialect detection of African American English. We train this system utilizing adult speech data and then evaluate on both children’s and adults’ speech with unmatched training and testing scenarios. The proposed system combines novel and state-of-the-art architectures, including a multi-source transformer language model pre-trained on Twitter text data and fine-tuned on ASR transcripts as well as an LSTM acoustic model trained on self-supervised learning representations, in order to learn a comprehensive view of dialect. We show robust, explainable performance across recording conditions for different features for adult speech, but fusing multiple features is important for good results on children’s speech. Alexander Johnson, Vishwas M. Shetty, Mari Ostendorf, Abeer Alwan |
ICASSP | 3 |
| 2023 | TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question AnsweringabstractDespite thousands of researchers, engineers, and artists actively working on improving text-to-image generation models, systems often fail to produce images that accurately align with the text inputs. We introduce TIFA (Text-to-Image Faithfulness evaluation with question Answering), an automatic evaluation metric that measures the faithfulness of a generated image to its text input via visual question answering (VQA). Specifically, given a text input, we automatically generate several question-answer pairs using a language model. We calculate image faithfulness by checking whether existing VQA models can answer these questions using the generated image. TIFA is a reference-free metric that allows for fine-grained and interpretable evaluations of generated images. TIFA also has better correlations with human judgments than existing metrics. Based on this approach, we introduce TIFA v1.0, a benchmark consisting of 4K diverse text inputs and 25K questions across 12 categories (object, counting, etc.). We present a comprehensive evaluation of existing text-to-image models using TIFA v1.0 and highlight the limitations and challenges of current models. For instance, we find that current text-to-image models, despite doing well on color and material, still struggle in counting, spatial relations, and composing multiple objects. We hope our benchmark will help carefully measure the research progress in text-to-image synthesis and provide valuable insights for further research.1 Yushi Hu, Benlin Liu, Jungo Kasai, Yizhong Wang, Mari Ostendorf, Ranjay Krishna, Noah A. Smith |
ICCV | 5 |
| 2023 | Binding Language Models in Symbolic Languages
Zhoujun Cheng, Tianbao Xie, Peng Shi 0010, Chengzu Li, Rahul Nadkarni, Yushi Hu, Caiming Xiong, Dragomir R. Radev, Mari Ostendorf, Luke Zettlemoyer, Noah A. Smith, Tao Yu 0009 |
ICLR | 9 |
| 2023 | Selective Annotation Makes Language Models Better Few-Shot Learners
Hongjin Su, Jungo Kasai, Chen Henry Wu, Jiayi Xin, Rui Zhang 0037, Mari Ostendorf, Luke Zettlemoyer, Noah A. Smith, Tao Yu 0009 |
ICLR | 8 |
| 2023 | Fine-Grained Human Feedback Gives Better Rewards for Language Model TrainingabstractLanguage models (LMs) often exhibit undesirable text generation behaviors, including generating false, toxic, or irrelevant outputs.
Reinforcement learning from human feedback (RLHF)---where human preference judgments on LM outputs are transformed into a learning signal---has recently shown promise in addressing these issues. However, such holistic feedback conveys limited information on long text outputs; it does not indicate which aspects of the outputs influenced user preference; e.g., which parts contain what type(s) of errors. In this paper, we use fine-grained human feedback (e.g., which sentence is false, which sub-sentence is irrelevant) as an explicit training signal. We introduce Fine-Grained RLHF, a framework that enables training and learning from reward functions that are fine-grained in two respects: (1) density, providing a reward after every segment (e.g., a sentence) is generated; and (2) incorporating multiple reward models associated with different feedback types (e.g., factual incorrectness, irrelevance, and information incompleteness). We conduct experiments on detoxification and long-form question answering to illustrate how learning with this reward function leads to improved performance, supported by both automatic and human evaluation. Additionally, we show that LM behaviors can be customized using different combinations of fine-grained reward models. We release all data, collected human feedback, and codes at https://FineGrainedRLHF.github.io. Zeqiu Wu, Yushi Hu, Nouha Dziri, Alane Suhr, Prithviraj Ammanabrolu, Noah A. Smith, Mari Ostendorf, Hannaneh Hajishirzi |
NeurIPS | 8 |
| 2023 | InSCIt: Information-Seeking Conversations with Mixed-Initiative InteractionsabstractAbstract In an information-seeking conversation, a user may ask questions that are under-specified or unanswerable. An ideal agent would interact by initiating different response types according to the available knowledge sources. However, most current studies either fail to or artificially incorporate such agent-side initiative. This work presents InSCIt, a dataset for Information-Seeking Conversations with mixed-initiative Interactions. It contains 4.7K user-agent turns from 805 human-human conversations where the agent searches over Wikipedia and either directly answers, asks for clarification, or provides relevant information to address user queries. The data supports two subtasks, evidence passage identification and response generation, as well as a human evaluation protocol to assess model performance. We report results of two systems based on state-of-the-art models of conversational knowledge identification and open-domain question answering. Both systems significantly underperform humans, suggesting ample room for improvement in future studies.1 Zeqiu Wu, Ryu Parish, Hao Cheng 0002, Sewon Min, Prithviraj Ammanabrolu, Mari Ostendorf, Hannaneh Hajishirzi |
Trans. Assoc. Comput. Linguistics | 6 |
| 2022 | CONQRR: Conversational Query Rewriting for Retrieval with Reinforcement LearningabstractCompared to standard retrieval tasks, passage retrieval for conversational question answering (CQA) poses new challenges in understanding the current user question, as each question needs to be interpreted within the dialogue context.Moreover, it can be expensive to retrain well-established retrievers such as search engines that are originally developed for nonconversational queries.To facilitate their use, we develop a query rewriting model CONQRR that rewrites a conversational question in the context into a standalone question.It is trained with a novel reward function to directly optimize towards retrieval using reinforcement learning and can be adapted to any off-theshelf retriever.CONQRR achieves state-ofthe-art results on a recent open-domain CQA dataset containing conversations from three different sources, and is effective for two different off-the-shelf retrievers.Our extensive analysis also shows the robustness of CON-QRR to out-of-domain dialogues as well as to zero query rewriting supervision. Zeqiu Wu, Yi Luan, Hannah Rashkin, David Reitter, Hannaneh Hajishirzi, Mari Ostendorf, Gaurav Tomar |
EMNLP | 6 |
| 2022 | Leveraging Prosody for Punctuation Prediction of Spontaneous SpeechabstractThesis (Master's)--University of Washington, 2022 Yeonjin Cho, Sara Ng, Trang Tran 0001, Mari Ostendorf |
INTERSPEECH | 4 |
| 2022 | Automatic Dialect Density Estimation for African American EnglishabstractIn this paper, we explore automatic prediction of dialect density of the African American English (AAE) dialect, where dialect density is defined as the percentage of words in an utterance that contain characteristics of the non-standard dialect.We investigate several acoustic and language modeling features, including the commonly used X-vector representation and Com-ParE feature set, in addition to information extracted from ASR transcripts of the audio files and prosodic information.To address issues of limited labeled data, we use a weakly supervised model to project prosodic and X-vector features into lowdimensional task-relevant representations.An XGBoost model is then used to predict the speaker's dialect density from these features and show which are most significant during inference.We evaluate the utility of these features both alone and in combination for the given task.This work, which does not rely on hand-labeled transcripts, is performed on audio segments from the CORAAL database.We show a significant correlation between our predicted and ground truth dialect density measures for AAE speech in this database and propose this work as a tool for explaining and mitigating bias in speech technology. Alexander Johnson, Kevin Everson, Vijay Ravi, Anissa Gladney, Mari Ostendorf, Abeer Alwan |
INTERSPEECH | 5 |
| 2022 | Spoken language interaction with robots: Recommendations for future researchabstractWith robotics rapidly advancing, more effective human–robot interaction is increasingly needed to realize the full potential of robots for society. While spoken language must be part of the solution, our ability to provide spoken language interaction capabilities is still very limited. In this article, based on the report of an interdisciplinary workshop convened by the National Science Foundation, we identify key scientific and engineering advances needed to enable effective spoken language interaction with robotics. We make 25 recommendations, involving eight general themes: putting human needs first, better modeling the social and interactive aspects of language, improving robustness, creating new methods for rapid adaptation, better integrating speech and language with other communication modalities, giving speech and language components access to rich representations of the robot’s current knowledge and state, making all components operate in real time, and improving research infrastructure and resources. Research and development that prioritizes these topics will, we believe, provide a solid foundation for the creation of speech-capable robots that are easy and effective for humans to work with. Matthew Marge, Carol Y. Espy-Wilson, Nigel G. Ward, Abeer Alwan, Yoav Artzi, Mohit Bansal, Gilmer L. Blankenship, Joyce Y. Chai, Hal Daumé III, Debadeepta Dey, Mary P. Harper, Thomas Howard, Casey Kennington, Ivana Kruijff-Korbayová, Dinesh Manocha, Cynthia Matuszek, Ross Mead, Raymond J. Mooney, Roger K. Moore, Mari Ostendorf, Heather Pon-Barry, Alexander I. Rudnicky, Matthias Scheutz, Robert St. Amant, Stefanie Tellex, David R. Traum, Zhou Yu 0005 |
Comput. Speech Lang. | 20 |
| 2021 | A Controllable Model of Grounded Response GenerationabstractCurrent end-to-end neural conversation models inherently lack the flexibility to impose semantic control in the response generation process, often resulting in uninteresting responses. Attempts to boost informativeness alone come at the expense of factual accuracy, as attested by pretrained language models' propensity to "hallucinate" facts. While this may be mitigated by access to background knowledge, there is scant guarantee of relevance and informativeness in generated responses. We propose a framework that we call controllable grounded response generation (CGRG), in which lexical control phrases are either provided by a user or automatically extracted by a control phrase predictor from dialogue context and grounding knowledge. Quantitative and qualitative results show that, using this framework, a transformer based model with a novel inductive attention mechanism, trained on a conversation-like Reddit dataset, outperforms strong generation baselines. Zeqiu Wu, Michel Galley, Chris Brockett, Yizhe Zhang 0002, Xiang Gao 0011, Chris Quirk, Rik Koncel-Kedziorski, Jianfeng Gao 0001, Hannaneh Hajishirzi, Mari Ostendorf, William B. Dolan |
AAAI | 10 |
| 2021 | Representations for Question Answering from Documents with Tables and TextabstractTables in Web documents are pervasive and can be directly used to answer many of the queries searched on the Web, motivating their integration in question answering. Very often information presented in tables is succinct and hard to interpret with standard language representations. On the other hand, tables often appear within textual context, such as an article describing the table. Using the information from an article as additional context can potentially enrich table representations. In this work we aim to improve question answering from tables by refining table representations based on information from surrounding text. We also present an effective method to combine text and table-based predictions for question answering from full documents, obtaining significant improvements on the Natural Questions dataset. Victoria Zayats, Kristina Toutanova, Mari Ostendorf |
EACL | 3 |
| 2021 | Dialogue State Tracking with a Language Model using Schema-Driven PromptingabstractTask-oriented conversational systems often use dialogue state tracking to represent the user's intentions, which involves filling in values of pre-defined slots.Many approaches have been proposed, often using task-specific architectures with special-purpose classifiers.Recently, good results have been obtained using more general architectures based on pretrained language models.Here, we introduce a new variation of the language modeling approach that uses schema-driven prompting to provide task-aware history encoding that is used for both categorical and non-categorical slots.We further improve performance by augmenting the prompting with schema descriptions, a naturally occurring source of indomain knowledge.Our purely generative system achieves state-of-the-art performance on MultiWOZ 2.2 and achieves competitive performance on two other benchmarks: Multi-WOZ 2.1 and M2M.The data and code will be available at https://github.com/ chiahsuan156/DST-as-Prompting. Chia-Hsuan Lee 0001, Hao Cheng 0002, Mari Ostendorf |
EMNLP (1) | 3 |
| 2021 | DIALKI: Knowledge Identification in Conversational Systems through Dialogue-Document ContextualizationabstractIdentifying relevant knowledge to be used in conversational systems that are grounded in long documents is critical to effective response generation.We introduce a knowledge identification model that leverages the document structure to provide dialogue-contextualized passage encodings and better locate knowledge relevant to the conversation.An auxiliary loss captures the history of dialogue-document connections.We demonstrate the effectiveness of our model on two document-grounded conversational datasets and provide analyses showing generalization to unseen documents and long dialogue contexts. Zeqiu Wu, Bo-Ru Lu, Hannaneh Hajishirzi, Mari Ostendorf |
EMNLP (1) | 4 |
| 2021 | Revisiting Parity of Human vs. Machine Conversational Speech Transcription
Courtney Mansfield, Sara Ng, Gina-Anne Levow, Richard A. Wright, Mari Ostendorf |
Interspeech | 5 |
| 2021 | Assessing the Use of Prosody in Constituency Parsing of Imperfect TranscriptsabstractThis work explores constituency parsing on automatically recognized transcripts of conversational speech. The neural parser is based on a sentence encoder that leverages word vectors contextualized with prosodic features, jointly learning prosodic feature extraction with parsing. We assess the utility of the prosody in parsing on imperfect transcripts, i.e. transcripts with automatic speech recognition (ASR) errors, by applying the parser in an N-best reranking framework. In experiments on Switchboard, we obtain 13-15% of the oracle N-best gain relative to parsing the 1-best ASR output, with insignificant impact on word recognition error rate. Prosody provides a significant part of the gain, and analyses suggest that it leads to more grammatical utterances via recovering function words. Trang Tran 0001, Mari Ostendorf |
Interspeech | 2 |
| 2021 | Extracting COVID-19 diagnoses and symptoms from clinical text: A new annotated corpus and neural event extraction framework
Kevin Lybarger, Mari Ostendorf, Meliha Yetisgen |
J. Biomed. Informatics | 2 |
| 2021 | Annotating social determinants of health using active learning, and characterizing determinants using neural event extraction
Kevin Lybarger, Mari Ostendorf, Meliha Yetisgen |
J. Biomed. Informatics | 2 |
| 2020 | A Novel Corpus With Detailed Annotations of Social Determinants of Health
Kevin Lybarger, Kylie Kerker, Jolie Shen, Erica Qiao, Özlem Uzuner, Mari Ostendorf, Meliha Yetisgen |
AMIA | 7 |
| 2020 | Mining Effective Negative Training Samples for Keyword SpottingabstractMax-pooling neural network architectures have been proven to be useful for keyword spotting (KWS), but standard training methods suffer from a class-imbalance problem when using all frames from negative utterances. To address the problem, we propose an innovative algorithm, Regional Hard-Example (RHE) mining, to find effective negative training samples, in order to control the ratio of negative vs. positive data. To maintain the diversity of the negative samples, multiple non-contiguous difficult frames per negative training utterance are dynamically selected during training, based on the model statistics at each training epoch. Further, to improve model learning, we introduce a weakly constrained max-pooling method for positive training utterances, which constrains max-pooling over the keyword ending frames only at early stages of training. Finally, data augmentation is combined to bring further improvement. We assess the algorithms by conducting experiments on wake-up word detection tasks with two different neural network architectures. The experiments consistently show that the proposed methods provide significant improvements compared to a strong baseline. At a false alarm rate of once per hour, our methods achieve 45-58% relative reduction in false rejection rates over a strong baseline. Jingyong Hou, Yangyang Shi, Mari Ostendorf, Mei-Yuh Hwang, Lei Xie 0001 |
ICASSP | 3 |
| 2020 | Analysis of Disfluency in Children's SpeechabstractDisfluencies are prevalent in spontaneous speech, as shown in many studies of adult speech. Less is understood about children's speech, especially in pre-school children who are still developing their language skills. We present a novel dataset with annotated disfluencies of spontaneous explanations from 26 children (ages 5--8), interviewed twice over a year-long period. Our preliminary analysis reveals significant differences between children's speech in our corpus and adult spontaneous speech from two corpora (Switchboard and CallHome). Children have higher disfluency and filler rates, tend to use nasal filled pauses more frequently, and on average exhibit longer reparandums than repairs, in contrast to adult speakers. Despite the differences, an automatic disfluency detection system trained on adult (Switchboard) speech transcripts performs reasonably well on children's speech, achieving an F1 score that is 10\% higher than the score on an adult out-of-domain dataset (CallHome). Trang Tran 0001, Morgan Tinkler, Gary Yeung, Abeer Alwan, Mari Ostendorf |
INTERSPEECH | 5 |
| 2019 | Automatic Identification of Social Determinants of Health from Clinical Records
Kevin Lybarger, Mari Ostendorf, Özlem Uzuner, Meliha Yetisgen |
AMIA | 2 |
| 2019 | On the Role of Style in Parsing Speech with Neural ModelsabstractThe differences in written text and conversational speech are substantial; previous parsers trained on treebanked text have given very poor results on spontaneous speech. For spoken language, the mismatch in style also extends to prosodic cues, though it is less well understood. This paper re-examines the use of written text in parsing speech in the context of recent advances in neural language processing. We show that neural approaches facilitate using written text to improve parsing of spontaneous speech, and that prosody further improves over this state-of-the-art result. Further, we find an asymmetric degradation from read vs. spontaneous mismatch, with spontaneous speech more generally useful for training parsers. Trang Tran 0001, Jiahong Yuan, Yang Liu 0004, Mari Ostendorf |
INTERSPEECH | 4 |
| 2019 | Disfluencies and Human Speech Transcription ErrorsabstractThis paper explores contexts associated with errors in transcrip-tion of spontaneous speech, shedding light on human perceptionof disfluencies and other conversational speech phenomena. Anew version of the Switchboard corpus is provided with disfluency annotations for careful speech transcripts, together with results showing the impact of transcription errors on evaluation of automatic disfluency detection. Victoria Zayats, Trang Tran 0001, Richard A. Wright, Courtney Mansfield, Mari Ostendorf |
INTERSPEECH | 5 |
| 2019 | Region Proposal Network Based Small-Footprint Keyword SpottingabstractWe apply an anchor-based region proposal network (RPN) for end-to-end keyword spotting (KWS). RPNs have been widely used for object detection in image and video processing; here, it is used to jointly model keyword classification and localization. The method proposes several anchors as rough locations of the keyword in an utterance and jointly learns classification and transformation to the ground truth region for each positive anchor. Additionally, we extend the keyword/non-keyword binary classification to detect multiple keywords. We verify our proposed method on a hotword detection data set with two hotwords. At a false alarm rate of one per hour, our method achieved more than 15% relative reduction in false rejection of the two keywords over multiple recent baselines. In addition, our method predicts the location of the keyword with over 90% overlap, which can be important for many applications. Jingyong Hou, Yangyang Shi, Mari Ostendorf, Mei-Yuh Hwang, Lei Xie 0001 |
IEEE Signal Process. Lett. | 3 |
| 2018 | Using Neural Multi-task Learning to Extract Substance Abuse Information from Clinical Notes
Kevin Lybarger, Meliha Yetisgen, Mari Ostendorf |
AMIA | 3 |
| 2018 | Multi-Task Identification of Entities, Relations, and Coreference for Scientific Knowledge Graph ConstructionabstractWe introduce a multi-task setup of identifying and classifying entities, relations, and coreference clusters in scientific articles.We create SCIERC, a dataset that includes annotations for all three tasks and develop a unified framework called Scientific Information Extractor (SCIIE) for with shared span representations.The multi-task setup reduces cascading errors between tasks and leverages cross-sentence relations through coreference links.Experiments show that our multi-task model outperforms previous models in scientific information extraction without using any domain-specific features.We further show that the framework supports construction of a scientific knowledge graph, which we use to analyze information in scientific literature. 1 Extracting nodes (entities) The SCIIE model extracts entities, their relations, and coreference Yi Luan, Luheng He, Mari Ostendorf, Hannaneh Hajishirzi |
EMNLP | 3 |
| 2018 | Domain Adversarial Training for Accented Speech RecognitionabstractIn this paper, we propose a domain adversarial training (DAT) algorithm to alleviate the accented speech recognition problem. In order to reduce the mismatch between labeled source domain data (“standard” accent) and unlabeled target domain data (with heavy accents), we augment the learning objective for a Kaldi TDNN network with a domain adversarial training (DAT) objective to encourage the model to learn accent-invariant features. In experiments with three Mandarin accents, we show that DAT yields up to 7.45% relative character error rate reduction when we do not have transcriptions of the accented speech, compared with the baseline trained on standard accent data only. We also find a benefit from DAT when used in combination with training from automatic transcriptions on the accented data. Furthermore, we find that DAT is superior to multi-task learning for accented speech recognition. Sining Sun, Ching-Feng Yeh, Mei-Yuh Hwang, Mari Ostendorf, Lei Xie 0001 |
ICASSP | 4 |
| 2018 | Training Augmentation with Adversarial Examples for Robust Speech RecognitionabstractThis paper explores the use of adversarial examples in training speech recognition systems to increase robustness of deep neural network acoustic models.During training, the fast gradient sign method is used to generate adversarial examples augmenting the original training data.Different from conventional data augmentation based on data transformations, the examples are dynamically generated based on current acoustic model parameters.We assess the impact of adversarial data augmentation in experiments on the Aurora-4 and CHiME-4 single-channel tasks, showing improved robustness against noise and channel variation.Further improvement is obtained when combining adversarial examples with teacher/student training, leading to a 23% relative word error rate reduction on Aurora-4. Sining Sun, Ching-Feng Yeh, Mari Ostendorf, Mei-Yuh Hwang, Lei Xie 0001 |
INTERSPEECH | 3 |
| 2018 | Parsing Speech: a Neural Approach to Integrating Lexical and Acoustic-Prosodic InformationabstractTrang Tran, Shubham Toshniwal, Mohit Bansal, Kevin Gimpel, Karen Livescu, Mari Ostendorf. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Trang Tran 0001, Shubham Toshniwal, Mohit Bansal, Kevin Gimpel, Karen Livescu, Mari Ostendorf |
NAACL-HLT | 6 |
| 2018 | Low-Rank RNN Adaptation for Context-Aware Language ModelingabstractA context-aware language model uses location, user and/or domain metadata (context) to adapt its predictions. In neural language models, context information is typically represented as an embedding and it is given to the RNN as an additional input, which has been shown to be useful in many applications. We introduce a more powerful mechanism for using context to adapt an RNN by letting the context vector control a low-rank transformation of the recurrent layer weight matrix. Experiments show that allowing a greater fraction of the model parameters to be adjusted has benefits in terms of perplexity and classification for several different types of context. Aaron Jaech, Mari Ostendorf |
Trans. Assoc. Comput. Linguistics | 2 |
| 2018 | Conversation Modeling on Reddit Using a Graph-Structured LSTMabstractThis paper presents a novel approach for modeling threaded discussions on social media using a graph-structured bidirectional LSTM (long-short term memory) which represents both hierarchical and temporal conversation structure. In experiments with a task of predicting popularity of comments in Reddit discussions, the proposed model outperforms a node-independent architecture for different sets of input features. Analyses show a benefit to the model over the full course of the discussion, improving detection in both early and late stages. Further, the use of language cues with the bidirectional tree state updates helps with identifying controversial comments. Victoria Zayats, Mari Ostendorf |
Trans. Assoc. Comput. Linguistics | 2 |
| 2017 | Automatically Detecting Likely Edits in Clinical Notes Created Using Automatic Speech Recognition
Kevin Lybarger, Mari Ostendorf, Meliha Yetisgen |
AMIA | 2 |
| 2017 | A Factored Neural Network Model for Characterizing Online Discussions in Vector SpaceabstractWe develop a novel factored neural model that learns comment embeddings in an unsupervised way leveraging the structure of distributional context in online discussion forums. The model links different context with related language factors in the embedding space, providing a way to interpret the factored embeddings. Evaluated on a community endorsement prediction task using a large collection of topic-varying Reddit discussions, the factored embeddings consistently achieve improvement over other text representations. Qualitative analysis shows that the model captures community style and topic, as well as response trigger patterns. Hao Cheng 0002, Hao Fang 0002, Mari Ostendorf |
EMNLP | 3 |
| 2017 | Scientific Information Extraction with Semi-supervised Neural TaggingabstractThis paper addresses the problem of extracting keyphrases from scientific articles and categorizing them as corresponding to a task, process, or material.We cast the problem as sequence tagging and introduce semi-supervised methods to a neural tagging model, which builds on recent advances in named entity recognition.Since annotated training data is scarce in this domain, we introduce a graph-based semi-supervised algorithm together with a data selection scheme to leverage unannotated articles.Both inductive and transductive semi-supervised learning strategies outperform state-of-the-art information extraction performance on the 2017 SemEval Task 10 ScienceIE task. Yi Luan, Mari Ostendorf, Hannaneh Hajishirzi |
EMNLP | 2 |
| 2017 | An Open Letter to the Members of the IEEE Industrial Electronics Technical Community
Peter Willett 0001, Mari Ostendorf, Michael P. Polis, Rob Reilly |
IEEE Trans. Ind. Informatics | 2 |
| 2016 | Deep Reinforcement Learning with a Natural Language Action SpaceabstractJi He, Jianshu Chen, Xiaodong He, Jianfeng Gao, Lihong Li, Li Deng, Mari Ostendorf. Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2016. Jianshu Chen, Xiaodong He 0001, Jianfeng Gao 0001, Lihong Li 0001, Li Deng 0001, Mari Ostendorf |
ACL (1) | 7 |
| 2016 | Deep Reinforcement Learning with a Combinatorial Action Space for Predicting Popular Reddit ThreadsabstractWe introduce an online popularity prediction and tracking task as a benchmark task for reinforcement learning with a combinatorial, natural language action space.A specified number of discussion threads predicted to be popular are recommended, chosen from a fixed window of recent comments to track.Novel deep reinforcement learning architectures are studied for effective modeling of the value function associated with actions comprised of interdependent sub-actions.The proposed model, which represents dependence between sub-actions through a bi-directional LSTM, gives the best performance across different experimental configurations and domains, and it also generalizes well with varying numbers of recommendation requests. Mari Ostendorf, Xiaodong He 0001, Jianshu Chen, Jianfeng Gao 0001, Lihong Li 0001, Li Deng 0001 |
EMNLP | 2 |
| 2016 | Characterizing the Language of Online Communities and its Relation to Community ReceptionabstractThis work investigates style and topic aspects of language in online communities: looking at both utility as an identifier of the community and correlation with community reception of content.Style is characterized using a hybrid word and part-of-speech tag n-gram language model, while topic is represented using Latent Dirichlet Allocation.Experiments with several Reddit forums show that style is a better indicator of community identity than topic, even for communities organized around specific topics.Further, there is a positive correlation between the community reception to a contribution and the style similarity to that community, but not so for topic similarity. Trang Tran 0001, Mari Ostendorf |
EMNLP | 2 |
| 2016 | Domain Adaptation of Recurrent Neural Networks for Natural Language UnderstandingabstractThe goal of this paper is to use multi-task learning to efficiently scale slot filling models for natural language understanding to handle multiple target tasks or domains. The key to scalability is reducing the amount of training data needed to learn a model for a new task. The proposed multi-task model delivers better performance with less data by leveraging patterns that it learns from the other tasks. The approach supports an open vocabulary, which allows the models to generalize to unseen words, which is particularly important when very little training data is used. A newly collected crowd-sourced data set, covering four different domains, is used to demonstrate the effectiveness of the domain adaptation and open vocabulary techniques. Aaron Jaech, Larry Heck, Mari Ostendorf |
INTERSPEECH | 3 |
| 2016 | Disfluency Detection Using a Bidirectional LSTMabstractWe introduce a new approach for disfluency detection using a Bidirectional Long-Short Term Memory neural network (BLSTM). In addition to the word sequence, the model takes as input pattern match features that were developed to reduce sensitivity to vocabulary size in training, which lead to improved performance over the word sequence alone. The BLSTM takes advantage of explicit repair states in addition to the standard reparandum states. The final output leverages integer linear programming to incorporate constraints of disfluency structure. In experiments on the Switchboard corpus, the model achieves state-of-the-art performance for both the standard disfluency detection task and the correction detection task. Analysis shows that the model has better detection of non-repetition disfluencies, which tend to be much harder to detect. Victoria Zayats, Mari Ostendorf, Hannaneh Hajishirzi |
INTERSPEECH | 2 |
| 2016 | Phonological Pun-derstandingabstractAaron Jaech, Rik Koncel-Kedziorski, Mari Ostendorf. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016. Aaron Jaech, Rik Koncel-Kedziorski, Mari Ostendorf |
HLT-NAACL | 3 |
| 2016 | Using Pronunciation-Based Morphological Subword Units to Improve OOV Handling in Keyword SearchabstractOut-of-vocabulary (OOV) keywords present a challenge for keyword search (KWS) systems especially in the low-resource setting. Previous research has centered around approaches that use a variety of subword units to recover OOV words. This paper systematically investigates morphology-based subword modeling approaches on seven low-resource languages. We show that using morphological subword units (morphs) in speech recognition decoding is substantially better than expanding word-decoded lattices into subword units including phones, syllables and morphs. As alternatives to grapheme-based morphs, we apply unsupervised morphology learning to sequences of phonemes, graphones, and syllables. Using one of these phone-based morphs is almost always better than using the grapheme-based morphs, but the particular choice varies with the language. By combining the different methods, a substantial gain is obtained over the best single case for all languages, especially for OOV performance. Yanzhang He, Peter Baumann 0003, Hao Fang 0002, Brian Hutchinson, Aaron Jaech, Mari Ostendorf, Eric Fosler-Lussier, Janet B. Pierrehumbert |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2015 | Open-Domain Name Error Detection using a Multi-Task RNNabstractOut-of-vocabulary name errors in speech recognition create significant problems for downstream language processing, but the fact that they are rare poses challenges for automatic detection, particularly in an open-domain scenario.To address this problem, a multi-task recurrent neural network language model for sentence-level name detection is proposed for use in combination with out-of-vocabulary word detection.The sentence-level model is also effective for leveraging external text data.Experiments show a 26% improvement in name-error detection F-score over a system using n-gram lexical features. Hao Cheng 0002, Hao Fang 0002, Mari Ostendorf |
EMNLP | 3 |
| 2015 | What Your Username Says About YouabstractUsernames are ubiquitous on the Internet, and they are often suggestive of user demographics.This work looks at the degree to which gender and language can be inferred from a username alone by making use of unsupervised morphology induction to decompose usernames into sub-units.Experimental results on the two tasks demonstrate the effectiveness of the proposed morphological features compared to a character n-gram baseline. Aaron Jaech, Mari Ostendorf |
EMNLP | 2 |
| 2015 | Talking to the crowd: What do people react to in online discussions?abstractThis paper addresses the question of how language use affects community reaction to comments in online discussion forums, and the relative importance of the message vs. the messenger.A new comment ranking task is proposed based on community annotated karma in Reddit discussions, which controls for topic and timing of comments.Experimental work with discussion threads from six subreddits shows that the importance of different types of language features varies with the community of interest. Aaron Jaech, Victoria Zayats, Hao Fang 0002, Mari Ostendorf, Hannaneh Hajishirzi |
EMNLP | 4 |
| 2015 | Investigating the role of 'yeah' in stance-dense conversationabstractThis study investigates characteristics of stance-related discourse function, stance strength, and polarity in uses of the word ‘yeah.’ In an annotated corpus of 20 talker dyads engaged in collaborative tasks, over 2300 ‘yeahs’ fall into six common stance-act categories. While agreement, usually with weak, positive stance, accounts for about three-quarters of the instances, opinion-offering, convincing, reluctance to accept an idea, backchannels, and no-stance represent other common stance-related uses. We assess combinations of acousticprosodic characteristics (duration, intensity, pitch) to identify those which differentiate these stance categories for ‘yeah’ and to determine how they relate to levels of stance strength and polarity. Differences in vowel duration and intensity help to differentiate these fine-grained functions of ‘yeah.’ Within the larger agreement category, we can further assess the effects of stance strength and polarity, finding that positive polarity is signaled by higher pitch, lower intensity, and longer vowel duration, while greater stance strength shows higher pitch and intensity. Finally, a small set of negative ‘yeahs’ is examined for more specific stance functions which may be distinguishable by differing pitch and intensity contours. Valerie Freeman, Gina-Anne Levow, Richard A. Wright, Mari Ostendorf |
INTERSPEECH | 4 |
| 2015 | Learning phrase patterns for ASR name error detection using semantic similarity
Alex Marin, Mari Ostendorf |
INTERSPEECH | 2 |
| 2015 | Aligning Sentences from Standard Wikipedia to Simple WikipediaabstractWilliam Hwang, Hannaneh Hajishirzi, Mari Ostendorf, Wei Wu. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015. William Hwang, Hannaneh Hajishirzi, Mari Ostendorf |
HLT-NAACL | 3 |
| 2015 | Unediting: Detecting Disfluencies Without Careful TranscriptsabstractVictoria Zayats, Mari Ostendorf, Hannaneh Hajishirzi. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015. Victoria Zayats, Mari Ostendorf, Hannaneh Hajishirzi |
HLT-NAACL | 2 |
| 2015 | Exponential Language Modeling Using Morphological Features and Multi-Task LearningabstractFor languages with fast vocabulary growth and limited resources, data sparsity leads to challenges in training a language model. One strategy for addressing this problem is to leverage morphological structure as features in the model. This paper explores different uses of unsupervised morphological features in both the history and prediction space for three word-based exponential models (maximum entropy, logbilinear, and recurrent neural net (RNN)). Multi-task training is introduced as a regularizing mechanism to improve performance in the continuous-space approaches. The models are compared to non-parametric baselines. From using the RNN with morphological features and multi-task learning, experiments with conversational speech from four languages show we can obtain consistent gains of 7-11% in perplexity reduction in a limited-resource scenario (10 hrs speech), and 12-18% when the training size is increased ( 80 hrs ). Results are mixed for all other approaches, compared to a modified Kneser-Ney baseline, but morphology is useful in continuous-space models compared to their word-only baseline. Multi-task learning improves both continuous-space models. Hao Fang 0002, Mari Ostendorf, Peter Baumann 0003, Janet B. Pierrehumbert |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2015 | A Sparse Plus Low-Rank Exponential Language Model for Limited Resource ScenariosabstractThis paper describes a new exponential language model that decomposes the model parameters into one or more low-rank matrices that learn regularities in the training data and one or more sparse matrices that learn exceptions (e.g., keywords). The low-rank matrices induce continuous-space representations of words and histories. The sparse matrices learn multi-word lexical items and topic/domain idiosyncrasies. This model generalizes the standard ℓ1-regularized exponential language model, and has an efficient accelerated first-order training algorithm. Language modeling experiments show that the approach is useful in scenarios with limited training data, including low resource languages and domain adaptation. Brian Hutchinson, Mari Ostendorf, Maryam Fazel |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2014 | Subword-based modeling for handling OOV words inkeyword spottingabstractThis work compares ASR decoding at different subword levels crossed with alternative keyword search strategies to handle the OOV issue for keyword spotting in the low-resource setting. We show that a morpheme-based subword modeling approach is effective in recovering OOV keywords within a Turkish low-resource keyword spotting task, where mixed word and morpheme decoding approach outperforms the traditional subword-based search from word-decoded lattices that are broken down to subword lattices. Furthermore, unsupervised learning of morphology works almost as well as a rule-based system designed for the language despite the low-resource condition. A staged keyword search strategy benefits from both methods of morphological analysis. Yanzhang He, Brian Hutchinson, Peter Baumann 0003, Mari Ostendorf, Eric Fosler-Lussier, Janet B. Pierrehumbert |
ICASSP | 4 |
| 2014 | Domain adaptation for parsing in automatic speech recognitionabstractThis paper addresses the problem of adapting a parser trained on out-of-domain data for use in automatic speech recognition (ASR) rescoring and error detection tasks. Using a self-training approach and adaptation with weakly-supervised data, we obtain improvements in ASR rescoring of confusion networks. Features extracted from the parser output are also used to improve detection of general ASR errors and out-of-vocabulary word regions in conjunction with a maximum entropy classifier. Alex Marin, Mari Ostendorf |
ICASSP | 2 |
| 2014 | Manipulating stance and involvement using collaborative tasks: an exploratory comparisonabstractThe ATAROS project aims to identify acoustic signals of stance-taking in order to inform the development of automatic stance recognition in natural speech. Due to the typically low frequency of stance-taking in existing corpora that have been used to investigate related phenomena such as subjectivity, we are creating an audio corpus of unscripted conversations between dyads as they complete collaborative tasks designed to elicit a high density of stance-taking at increasing levels of involvement. To validate our experimental design and provide a preliminary assessment of the corpus, we examine a fully transcribed and time-aligned portion to compare the speaking styles in two tasks, one expected to elicit low involvement and weak stances, the other high involvement and strong stances. We find that although overall measures such as task duration and total word count do not indicate consistent differences across tasks, speakers do display significant differences in speaking style. Factors such as increases in speaking rate, turn length, and disfluencies from weakto strong-stance tasks are consistent with increased involvement by the participants and provide evidence in support of the experimental design. Valerie Freeman, Julian Chan, Gina-Anne Levow, Richard A. Wright, Mari Ostendorf, Victoria Zayats |
INTERSPEECH | 5 |
| 2014 | Relating automatic vowel space estimates to talker intelligibilityabstractDifferences in pronunciation have been shown to underlie significant talker-dependent intelligibility differences. There are several dimensions of variability that are correlated with talker intelligibility including pitch range, vowel-space expansion, and rhythmic patterns. Prior work has shown that some of the better predictors of individual intelligibility are based on the talker’s F1 by F2 vowel space, but findings are based on handcorrected measurements on carefully balanced sets of vowels, making large scale analysis impractical. This paper proposes a novel method for automatic estimation of a talker’s vowel space using sparse expanded vowel space representations, including an approximate convex hull sampling, which are projected to a low dimensional space for intelligibility scoring. Both supervised and unsupervised mappings are used to generate an intelligibility score. Automatic intelligibility rankings are assessed in terms of correlation with an intelligibility score based on human transcription accuracy. We find that including a larger sample of vowels (beyond point vowels) leads to improved performance, obtaining correlations of roughly 0.6 for this feature alone, which is a strong result given that there are other factors that may also contribute to a talker’s intelligibility in addition to a talker’s vowel space area. Yi Luan, Richard A. Wright, Mari Ostendorf, Gina-Anne Levow |
INTERSPEECH | 3 |
| 2014 | Learning phrase patterns for text classification using a knowledge graph and unlabeled dataabstractThis paper explores a novel method for learning phrase pattern features for text classification, employing a mapping of selected words into a knowledge graph and self-training over unlabeled data. Using Support Vector Machine classification, we obtain improvements over lexical and fully-supervised phrase pattern features in domain and intent detection for language understanding, particularly in conjunction with the use of unlabeled data. Our best results are obtained using unlabeled data filtered for both model training and feature learning based on the confidence of the baseline classifiers. Alex Marin, Roman Holenstein, Ruhi Sarikaya, Mari Ostendorf |
INTERSPEECH | 4 |
| 2014 | Multi-domain disfluency and repair detectionabstractThis paper investigates automatic detection of different types of self-repairs in spontaneous speech under different social contexts, from casual conversations to government hearings. The work shows that a simple CRF-based model is effective for cross-domain training, which is important for contexts where annotated data is not available. The approach explicitly represents common types of disfluencies observed in multi-domain data both in the model state space and the features extracted. In addition, the model incorporates an expanded state space for recognizing the repair structure, unlike prior work that annotates only the reparandum. Victoria Zayats, Mari Ostendorf, Hannaneh Hajishirzi |
INTERSPEECH | 2 |
| 2014 | Effective data-driven feature learning for detecting name errors in automatic speech recognitionabstractThis paper addresses the problem of detecting name errors in automatic speech recognition (ASR) output. The highly skewed label distributions (i.e. name errors are infrequent), sparse training data, and large number of potential lexical features pose significant challenges for training name error classification systems. Data-driven feature learning is needed for handling multiple languages but is sensitive to over fitting. We address the problem by designing aggregate features using a related (sentence-level name detection) task, and reduce dimensionality of the lexical features using word classes. Experiments on conversational domain data in both English and Iraqi Arabic show that best results are obtained using all feature mapping methods plus feature selection using L1 regularization. Alex Marin, Mari Ostendorf |
SLT | 3 |
| 2014 | Recognition of stance strength and polarity in spontaneous speechabstractFrom activities as simple as scheduling a meeting to those as complex as balancing a national budget, people take stances in negotiations and decision making. While the related areas of subjectivity and sentiment analysis have received significant attention, work has focused almost exclusively on text, and much stance-taking activity is carried out verbally. This paper investigates automatic recognition of stance-taking in spontaneous speech. It first presents a new annotated corpus of spontaneous, conversational speech designed to elicit high densities of stance-taking at different strengths. Speaker spurts are annotated both for strength of stance-taking behavior and polarity of stance. Based on this annotated corpus, we develop classifiers for automatic recognition of stance-taking behavior in speech. We employ a range of lexical, speaking style, and prosodic features in a boosting framework. The classifiers achieve strong accuracies on both binary detection of stance and four-way recognition of stance strength, well above most common class assignment. Finally, we classify the polarity of stance-taking spurts, obtaining accuracies around 80%. The best classifiers rely primarily on word unigram features, with speaking style and prosodic features yielding lower accuracies but still well above common class assignment. Gina-Anne Levow, Valerie Freeman, Alena Hrynkevich, Mari Ostendorf, Richard A. Wright, Julian Chan, Yi Luan, Trang Tran 0001 |
SLT | 4 |
| 2014 | Editorial: Expanding the Technical Reach of our TransactionsabstractWe thank the strong support of both IEEE Signal Processing Society and the ACM Publication Boards for this successful merger.IEEE's TASLP is a very well-established publication, strong in both quality and quantity.It is closely linked to ICASSP and to a number of workshops, such as ASRU and SLT.The language area is a relatively recent addition to TASLP, incorporated in 2006.ACM's TSLP is a more recent publication.The quality has been very high, but the quantity has only sustained a quarterly publication.There is no ACM Special Interest Group (SIG) or conference connection in this area, so it has been more difficult to maintain a direct link to the research community.One of the main motivations for the merger is that IEEE's TASLP does not yet have a strong profile in the language processing community, and this is reflected in the submissions received and the composition of the Editorial Board.ACM's TSLP has a stronger profile and Editorial Board membership in this area.Thus, it is clear that a joint transactions will be stronger than either publication on its own.For several years, the IEEE Signal Processing Society has recognized the importance of information processing in the work of a wide range of researchers within the Society beyond the traditional scope of signal processing.This led to the technical scope of the IEEE TRANSACTIONS ON SPEECH AND AUDIO PROCESSING being expanded to include Language Processing in 2006.While serving as editor of the IEEE SIGNAL PROCESSING MAGAZINE, one of the authors wrote in 2008 and 2010 two editorials [1], [2] that elaborated on the need for expanding the technical reach of signal processing by including the new "understanding" or "interpretation" component of signals consisting of language/text and bio-sequence data.They are both symbolic in nature, which were outside of the traditional definition of "signal" with numerical values in nature.This editorial on the formation of the joint IEEE/ACM TASLP, which highlights the importance of text or written language processing, is a concrete embodiment of the goal of expanding the technical reach of signal processing.As part of the merger of the two transactions, we have taken the opportunity to restructure the editorial board.In addition to the role of Editor-in-Chief, there will be six Senior Area Editors to advise the Editor-in-Chief across the full range of topics covered by the merged journal. Li Deng 0001, Steve Renals, Marcello Federico, Mari Ostendorf |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2013 | "Can you give me another word for hyperbaric?": Improving speech translation using targeted clarification questionsabstractWe present a novel approach for improving communication success between users of speech-to-speech translation systems by automatically detecting errors in the output of automatic speech recognition (ASR) and statistical machine translation (SMT) systems. Our approach initiates system-driven targeted clarification about errorful regions in user input and repairs them given user responses. Our system has been evaluated by unbiased subjects in live mode, and results show improved success of communication between users of the system. Necip Fazil Ayan, Arindam Mandal, Michael W. Frandsen, Jing Zheng 0001, Peter Blasco, Andreas Kathol, Frédéric Béchet, Benoît Favre, Alex Marin, Tom Kwiatkowski, Mari Ostendorf, Luke Zettlemoyer, Philipp Salletmayr, Julia Hirschberg, Svetlana Stoyanchev |
ICASSP | 11 |
| 2013 | Exceptions in language as learned by the multi-factor sparse plus low-rank language modelabstractWord usage is influenced by diverse factors, including topic, genre and various speaker/author characteristics. To characterize these aspects of language, we introduce the “Multi-Factor Sparse Plus Low Rank” exponential language model, which allows supervised joint training of arbitrary overlapping factor-specific model components. This flexible architecture has the advantage of being highly interpretable. The elements of sparse parameter matrices can be viewed as factor-dependent corrections (e.g. topic- or speaker-dependent phenomena). In topic modeling experiments on conversational telephone speech, we obtain modest perplexity reductions over an n-gram baseline and demonstrate topic-dependent keyword extraction that leads to a 13% (absolute) improvement in precision over TFIDF. We also show how keywords can be jointly learned for speakers, roles and topics in a study of Supreme Court oral arguments. Brian Hutchinson, Mari Ostendorf, Maryam Fazel |
ICASSP | 2 |
| 2013 | A sequential repetition model for improved disfluency detection
Mari Ostendorf, Sangyun Hahn |
INTERSPEECH | 1 |
| 2013 | Atypical Prosodic Structure as an Indicator of Reading Level and Text Difficulty
Julie Medero, Mari Ostendorf |
HLT-NAACL | 2 |
| 2013 | Learning Phrase Patterns for Text ClassificationabstractThis paper introduces methods to discriminatively learn phrase patterns for use as features in text classification. An efficient solution is described using a recursive algorithm with a mutual information selection criterion. The algorithm automatically determines when word classes are useful in specific locations of a phrase pattern, allowing for variable specificity depending on the amount of labeled data available. Experiments are carried out on three text classification tasks in both English and Chinese, resulting in improved performance when adding the phrase patterns to the existingn-gram features. Bin Zhang 0009, Alex Marin, Brian Hutchinson, Mari Ostendorf |
IEEE Trans. Speech Audio Process. | 4 |
| 2013 | Graph-Based Query Strategies for Active LearningabstractThis paper proposes two new graph-based query strategies for active learning in a framework that is convenient to combine with semi-supervised learning based on label propagation. The first strategy selects instances independently to maximize the change to a maximum entropy model using label propagation results in a gradient length measure of model change. The second strategy involves a batch criterion that integrates label uncertainty with diversity and density objectives. Experiments on sentiment classification demonstrate that both methods consistently improve over a standard active learning baseline, and that the batch criterion also gives consistent improvement over semi-supervised learning alone. Mari Ostendorf |
IEEE Trans. Speech Audio Process. | 2 |
| 2012 | Detecting targets of alignment moves in multiparty discussionsabstractIn analyzing goal-oriented multiparty discussions, one challenge is to determine who is responding to whom when they are supporting or opposing a remark put forward by another participant. This paper looks at algorithms for detecting the target discussant, comparing findings of important features for three genres of text and spoken discussions. Comparing to the common baseline of “previous speaker,” we find gains from considering content and semantic similarity in target detection, but with substantial differences in accuracy and important features across genres. Alex Marin, Bin Zhang 0009, Mari Ostendorf |
ICASSP | 4 |
| 2012 | A Sparse Plus Low Rank Maximum Entropy Language ModelabstractThis work introduces a new maximum entropy language model that decomposes the model parameters into a low rank component that learns regularities in the training data and a sparse component that learns exceptions (e.g. multiword expressions). The low rank component corresponds to a continuous-space language model. This model generalizes the standard ℓ1regularized maximum entropy model, and has an efficient accelerated first-order training algorithm. In conversational speech language modeling experiments, we see perplexity reductions Brian Hutchinson, Mari Ostendorf, Maryam Fazel |
INTERSPEECH | 2 |
| 2012 | Using syntactic and confusion network structure for out-of-vocabulary word detectionabstractThis paper addresses the problem of detecting words that are out-of-vocabulary (OOV) for a speech recognition system to improve automatic speech translation. The detection system leverages confidence prediction techniques given a confusion network representation and parsing with OOV word tokens to identify spans associated with true OOV words. Working in a resource-constrained domain, we achieve OOV detection F-scores of 60-66 and reduce word error rate by 12% relative to the case where OOV words are not detected. Alex Marin, Tom Kwiatkowski, Mari Ostendorf, Luke Zettlemoyer |
SLT | 3 |
| 2012 | Joint reranking of parsing and word recognition with automatic segmentation
Jeremy G. Kahn, Mari Ostendorf |
Comput. Speech Lang. | 2 |
| 2012 | A Message from the Vice President of Publications on New Developments in Signal Processing Society PublicationsabstractElectronic publishing through IEEE Xplore offers the possibility of augmenting papers with multimedia examples - including audio, image, and video files - which provide concrete illustrations of the utility of a new algorithm. It is encouraged for all authors to take advantage of these opportunities, which are especially relevant to many problems in signal processing and can increase the impact of their work. To make the papers more easily reproducible, it is encouraged to submit code that can be accessed through IEEE Xplore with the paper. IEEE Xplore is also now augmenting the paper presentation with links to related content and citation information, and more developments are in the works to make the presentation more useful to readers. Mari Ostendorf |
IEEE Signal Process. Lett. | 1 |
| 2012 | A Message from the Vice President of Publications on New Developments in Signal Processing Society PublicationsabstractThe IEEE Signal Processing Society staff and volunteers are looking for ways to better serve their authors and readers. In addition, the IEEE is also more broadly pursuing innovations in content delivery. The author wishes to alert readers to the advances that are already in place and developments to look for in the future. Mari Ostendorf |
IEEE Trans. Speech Audio Process. | 1 |
| 2012 | A Message from the Vice President of Publications on New Developments in Signal Processing Society PublicationsabstractElectronic publishing through IEEE Xplore offers the possibility of augmenting papers with multimedia examples - including audio, image, and video files - which provide concrete illustrations of the utility of a new algorithm. It is encouraged for all authors to take advantage of these opportunities, which are especially relevant to many problems in signal processing and can increase the impact of their work. To make the papers more easily reproducible, it is encouraged to submit code that can be accessed through IEEE Xplore with the paper. IEEE Xplore is also now augmenting the paper presentation with links to related content and citation information, and more developments are in the works to make the presentation more useful to readers. Mari Ostendorf |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2012 | Editorial Message From the Vice President of Publications on New Developments in Signal Processing Society Publications
Mari Ostendorf |
IEEE Trans. Image Process. | 1 |
| 2011 | Analyzing conversations using rich phrase patternsabstractIndividual words are not powerful enough for many complex language classification problems. N-gram features include word context information, but are limited to contiguous word sequences. In this paper, we propose to use phrase patterns to extend n-grams for analyzing conversations, using a discriminative approach to learning patterns with a combination of words and word classes to address data sparsity issues. Improvements in performance are reported for two conversation analysis tasks: speaker role recognition and alignment classification. Bin Zhang 0009, Alex Marin, Brian Hutchinson, Mari Ostendorf |
ASRU | 4 |
| 2011 | Low Rank Language Models for Small Training SetsabstractSeveral language model smoothing techniques are available that are effective for a variety of tasks; however, training with small data sets is still difficult. This letter introduces the low rank language model, which uses a low rank tensor representation of joint probability distributions for parameter-tying and optimizes likelihood under a rank constraint. It obtains lower perplexity than standard smoothing techniques when the training set is small and also leads to perplexity reduction when used in domain adaptation via interpolation with a general, out-of-domain model. Brian Hutchinson, Mari Ostendorf, Maryam Fazel |
IEEE Signal Process. Lett. | 2 |
| 2010 | Unsupervised broadcast conversation speaker role labelingabstractWe present an approach to unsupervised speaker role labeling in talk show data that makes use of two complementary sets of features: structural features that encode the participation patterns of speakers, and lexical features, which capture characteristic phrases. Techniques for using multiple clusterings are explored, leading to more robust results. Experiments on English and Mandarin talk shows yield performance similar to that reported for broadcast news using supervised learning. Brian Hutchinson, Bin Zhang 0009, Mari Ostendorf |
ICASSP | 3 |
| 2010 | Automatic Generation of Personalized Annotation Tags for Twitter Users
Bin Zhang 0009, Mari Ostendorf |
HLT-NAACL | 3 |
| 2010 | Extracting Phrase Patterns with Minimum Redundancy for Unsupervised Speaker Role Classification
Bin Zhang 0009, Brian Hutchinson, Mari Ostendorf |
HLT-NAACL | 4 |
| 2010 | Detecting authority bids in online discussionsabstractThis paper looks at the problem of detecting a particular type of social behavior in discussions: attempts to establish credibility as an authority on a particular topic. Using maximum entropy modeling, we explore questions related to feature extraction and turn vs. discussion-level modeling in experiments with online discussion text given only a small amount of labeled training data. We also introduce a method for learning interaction words from unlabeled data. Preliminary experiments show that a word-based approach (as used in topic classification) can be used successfully for turn-level modeling, but is less effective at the discussion level. We also find that sentence complexity features are almost as useful as lexical features, and that interaction words are more robust than the full vocabulary when combined with other features. Alex Marin, Mari Ostendorf, Bin Zhang 0009, Jonathan T. Morgan, Meghan Oxley, Mark Zachry, Emily M. Bender |
SLT | 2 |
| 2009 | Part-of-speech histograms for genre classification of textabstractThis work addresses the problem of classifying the genre of text, which is useful for a variety of language processing problems. We propose statistics of POS histograms as classification features, coupled with a quadratic discriminant classifier. In experiments on six different text and speech genres, we demonstrate enhanced performance compared to standard techniques using word frequency count features and POS trigram features. Experiments on genres that were not seen in training show intuitive overlaps with the training classes. Sergey Feldman, Marius A. Marin, Mari Ostendorf, Maya R. Gupta |
ICASSP | 3 |
| 2009 | Acoustic-based pitch-accent detection in speech: Dependence on word identity and insensitivity to variations inword usageabstractPast work has produced fairly accurate automatic pitch-accent detectors, but it has often been noted that the accent class of a word is highly dependent on word identity, with some words and word types usually being accented and others not. We argue that a good accent detector should not only have high overall accuracy, but also be able to distinguish between accented and unaccented variants of the same word. We report on experiments with several classifiers trained on a hand-labeled corpus, using a large set of acoustic features. Results show that while the classifiers have a high overall accuracy, they perform disappointingly on words with atypical accent status or whose prior accent status is more uncertain. We further report on attempts to improve the performance on these sub-tasks via feature selection and engineering of the training set. Anna Margolis, Mari Ostendorf |
ICASSP | 2 |
| 2009 | Filtering web text to match target genresabstractIn language modeling for speech recognition, both the amount of training data and the match to the target task impact the goodness of the model, with the trade-off usually favoring more data. For conversational speech, having some genre-matched text is particularly important, but also hard to obtain. This paper proposes a new approach for genre detection and compares different alternatives for filtering Web text for genre to improve language models for use in automatic transcription of broadcast conversations (talk shows). Marius A. Marin, Sergey Feldman, Mari Ostendorf, Maya R. Gupta |
ICASSP | 3 |
| 2009 | Freshman design: A signal-processing approachabstractThis paper describes development and evaluation of a freshman design course focusing on signal processing and communications, entitled ldquoThe Digital World of Multimedia.rdquo The course covers basic concepts of sampling, frequency analysis, and filtering without calculus but with coverage of these topics in both audio and image processing. Six labs provide collaborative, experimental learning as well as design opportunities. Connections to modern technology are provided in guest lectures and reading assignments that leverage IEEE magazines, and the topics provide opportunities for discussing the societal impact of signal processing technology. Mari Ostendorf, Jayson Bowen, Anna Margolis |
ICASSP | 1 |
| 2009 | Transcribing human-directed speech for spoken language processing
Mari Ostendorf |
INTERSPEECH | 1 |
| 2009 | Improving the recognition of names by document-level clustering
Bin Zhang 0009, Jeremy G. Kahn, Mari Ostendorf |
INTERSPEECH | 4 |
| 2009 | Improving robustness of MLLR adaptation with speaker-clustered regression class trees
Arindam Mandal, Mari Ostendorf, Andreas Stolcke |
Comput. Speech Lang. | 2 |
| 2009 | A machine learning approach to reading level assessment
Sarah E. Petersen, Mari Ostendorf |
Comput. Speech Lang. | 2 |
| 2009 | Expected dependency pair match: predicting translation quality with expected syntactic structure
Jeremy G. Kahn, Matthew G. Snover, Mari Ostendorf |
Mach. Transl. | 3 |
| 2009 | Building A Highly Accurate Mandarin Speech Recognizer With Language-Independent Technologies and Language-Dependent ModulesabstractWe describe a system for highly accurate large-vocabulary Mandarin speech recognition. The prevailing hidden Markov model based technologies are essentially language independent and constitute the backbone of our system. These include minimum-phone-error discriminative training and maximum-likelihood linear regression adaptation, among others. Additionally, careful considerations are taken into account for Mandarin-specific issues including lexical word segmentation, tone modeling, phone set design, and automatic acoustic segmentation. Our system comprises two sets of acoustic models for the purposes of cross adaptation. The systems are designed to be complementary in terms of errors but with similar overall accuracy by using different phone sets and different combinations of discriminative learning. The outputs of the two subsystems are then rescored by an adapted n-gram language model. Final confusion network combination yielded 9.1% character error rate on the DARPA GALE 2007 official evaluation, the best Mandarin recognition system in that year. Mei-Yuh Hwang, Gang Peng 0001, Mari Ostendorf, Wen Wang 0001, Arlo Faria, Aaron Heidel |
IEEE Trans. Speech Audio Process. | 3 |
| 2008 | Punctuating speech for information extractionabstractThis paper studies the effect of automatic sentence boundary detection and comma prediction on entity and relation extraction in speech. We show that punctuating the machine generated transcript according to maximum F-measure of period and comma annotation results in suboptimal information extraction. Precisely, period and comma decision thresholds can be chosen in order to improve the entity value score and the relation value score by 4% relative. Error analysis shows that preventing noun-phrase splitting by generating longer sentences and fewer commas can be harmful for IE performance. Indeed, it seems that missed punctuation allows syntactic parsers to merge noun-phrases and prevent the extraction of correct information. Benoît Favre, Ralph Grishman, Dustin Hillard, Heng Ji 0001, Dilek Hakkani-Tür, Mari Ostendorf |
ICASSP | 6 |
| 2008 | Parsing-based objective functions for speech recognition in translation applicationsabstractThis paper looks at a parsing-based alternative to word error rate (WER) for optimizing recognition, SParseval, hypothesizing that it may be a better objective for applications such as translation. We find that SParseval is more correlated than WER with human measures of subsequent translation performance, but that optimizing explicitly for SParseval does not give a significant reduction in translation error as measured by automatic methods based on a single translation reference. However, anecdotal examples indicate that SParseval does improve automatic speech recognition (ASR) results, leaving open the possibility that it may be more useful in the future or for other language processing tasks. Dustin Hillard, Mei-Yuh Hwang, Mary P. Harper, Mari Ostendorf |
ICASSP | 4 |
| 2008 | Non-segmental duration feature extraction for prosodic classificationabstractThis paper presents a set of novel duration features for detecting pitch accent and phrase boundaries, which depend on articulatory timing rather than segmental duration information. The features are computed from the detected syllable nuclei and boundaries, using peaks and valleys in an energy contour but also leveraging information from a simple HMM phone manner class recognizer to increase recall. In experiments on the hand-segmented TIMIT corpus, we obtain greater than 90 % F-measure for vowel detection. In prosody detection experiments on the BU Radio News corpus, comparing to a segmental feature baseline, we obtain similar performance for pitch accent detection and slightly worse boundary detection from the new features without the need for phonetic alignments. Index Terms: prosody, prominence, pitch accent, boundary detection, duration features Amy Dashiell, Brian Hutchinson, Anna Margolis, Mari Ostendorf |
INTERSPEECH | 4 |
| 2008 | Cross-validation and aggregated EM training for robust parameter estimation
Takahiro Shinozaki, Mari Ostendorf |
Comput. Speech Lang. | 2 |
| 2007 | Building a highly accurate Mandarin speech recognizerabstractWe describe a highly accurate large-vocabulary continuous Mandarin speech recognizer, a collaborative effort among four research organizations. Particularly, we build two acoustic models (AMs) with significant differences but similar accuracy for the purposes of cross adaptation and system combination. This paper elaborates on the main differences between the two systems, where one recognizer incorporates a discriminatively trained feature while the other utilizes a discriminative feature transformation. Additionally we present an improved acoustic segmentation algorithm and topic-based language model (LM) adaptation. Coupled with increased acoustic training data, we reduced the character error rate (CER) of the DARPA GALE 2006 evaluation set to 15.3% from 18.4%. Mei-Yuh Hwang, Gang Peng 0001, Wen Wang 0001, Arlo Faria, Aaron Heidel, Mari Ostendorf |
ASRU | 6 |
| 2007 | Efficient use of overlap information in speaker diarizationabstractSpeaker overlap in meetings is thought to be a significant contributor to error in speaker diarization, but it is not clear if overlaps are problematic for speaker clustering and/or if errors could be addressed by assigning multiple labels in overlap regions. In this paper, we look at these issues experimentally, assuming perfect detection of overlaps, to assess the relative importance of these problems and the potential impact of overlap detection. With our best features, we find that detecting overlaps could potentially improve diarization accuracy by 15% relative, using a simple strategy of assigning speaker labels in overlap regions according to the labels of the neighboring segments. In addition, the use of cross-correlation features with MFCC's reduces the performance gap due to overlaps, so that there is little gain from removing overlapped regions before clustering. Scott Otterson, Mari Ostendorf |
ASRU | 2 |
| 2007 | Cross-Site and Intra-Site ASR System Combination: Comparisons on Lattice and 1-Best MethodsabstractWe evaluate system combination techniques for automatic speech recognition using systems from multiple sites who participated in the TC-STAR 2006 evaluation. Both lattice and 1-best combination techniques are tested for cross-site and intra-site tasks. For pairwise combinations the lattice based approaches can outperform 1-best ROVER with confidence scores, but 1-best ROVER results are equal (or even better) when combining three or four systems. Björn Hoffmeister, Dustin Hillard, Stefan Hahn, Ralf Schlüter, Mari Ostendorf, Hermann Ney |
ICASSP (4) | 5 |
| 2007 | Word-Level Tone Modeling for Mandarin Speech RecognitionabstractStandard HMM-based Mandarin speech recognition systems do not exploit the suprasegmental nature of tones, but explicit tone models can be incorporated with lattice rescoring. This work extends previous approaches to explicit tone modeling from the syllable level to the word level, incorporating a hierarchical backoff. Word-dependent tone models are trained to explicitly model the tone coarticulation within the word. For less frequent words, tonal-syllable-dependent tone models or plain tone models are used as backoff. More generally, context-dependent tone models can be used as backoff. The word-dependent tone modeling framework can be viewed as a generalization of the traditional context-independent and context-dependent tone modeling. Under this framework, different types of tone modeling strategies are compared experimentally on a Mandarin broadcast news speech recognition task, showing significant gains from the word-level tone modeling approach. Mari Ostendorf |
ICASSP (4) | 2 |
| 2007 | Cross-Validation EM Training for Robust Parameter EstimationabstractA new maximum likelihood training algorithm is proposed that compensates for weaknesses of the EM algorithm by using cross-validation likelihood in the expectation step to avoid overtraining. By using a set of sufficient statistics associated with a partitioning of the training data, as in parallel EM, the algorithm has the same order of computational requirements as the original EM algorithm. Analyses using a GMM with artificial data show the proposed algorithm is more robust for overtraining than the conventional EM algorithm. Large vocabulary recognition experiments on Mandarin broadcast news data show that the method makes better use of more parameters and gives lower recognition error rates than EM training. Takahiro Shinozaki, Mari Ostendorf |
ICASSP (4) | 2 |
| 2007 | Improving speech translation with automatic boundary predictionabstractThis paper investigates the influence of automatic sentence boundary and sub-sentence punctuation prediction on machine translation (MT) of automatically recognized speech.We use prosodic and lexical cues to determine sentence boundaries, and successfully combine two complementary approaches to sentence boundary prediction.We also introduce a new feature for segmentation prediction that directly considers the assumptions of the phrase translation model.In addition, we show how automatically predicted commas can be used to constrain reordering in MT search.We evaluate the presented methods using a state-of-the-art phrase-based statistical MT system on two large vocabulary tasks.We find that careful optimization of the segmentation parameters directly for translation quality improves the translation results in comparison to independent optimization for segmentation quality of the predicted source language sentence boundaries. Evgeny Matusov, Dustin Hillard, Mathew Magimai-Doss, Dilek Hakkani-Tür, Mari Ostendorf, Hermann Ney |
INTERSPEECH | 5 |
| 2007 | Automatic acoustic segmentation for speech recognition on broadcast recordingsabstractThis paper investigates the issue of automatic segmentation of speech recordings for broadcast news (BN) and broadcast conversation (BC) speech recognition. Our previous segmentation algorithm often exhibited high deletion errors, where some speech segments were misclassified as non-speech and thus were never passed on to the recognizer. In contrast with our previous segmentation models, which only differentiated between speech and non-speech segments, phonetic knowledge is applied to represent speech by using multiple models for different types of speech segments. Moreover, the “pronunciation ” of the speech segment has been modified to loosen the minimum duration constraint. This method makes use of language specific knowledge, while keeping the number of models low to achieve fast segmentation. Experimental results show that the new segmenter outperforms our previous segmenter significantly, particularly in reducing deletion errors. Index Terms: speech recognition, automatic acoustic segmentation, machine translation, broadcast news, broadcast conversation 1. Gang Peng 0001, Mei-Yuh Hwang, Mari Ostendorf |
INTERSPEECH | 3 |
| 2007 | Symbolic phonetic features for modeling of pronunciation variation
Rebecca Bates 0001, Mari Ostendorf, Richard A. Wright |
Speech Commun. | 2 |
| 2006 | Compensating for Word Posterior Estimation Bias in Confusion NetworksabstractThis paper looks at the problem of confidence estimation at the word network level, where multiple hypotheses from a recognizer are represented in a confusion network. Given features of the network, an SVM is used to estimate the probability that the correct word is missing from a candidate slot and then other word probabilities are normalized accordingly. The result is a reduction in overall bias of the estimated word posteriors and an improvement in the confidence estimate for the top word hypothesis in particular Dustin Hillard, Mari Ostendorf |
ICASSP (1) | 2 |
| 2006 | Advances in lecture recognition: the ISL RT-06s evaluation systemabstractThis paper describes the 2006 lecture recognition system developed at the Interactive Systems Laboratories (ISL), for individual head-microphone (IHM), single distant microphone (SDM), and multiple distant microphones (MDM) conditions.It was evaluated in RT-06S rich transcription meeting evaluation sponsored by the US National Institute of Standards and Technologies (NIST).We describe the principal differences between our current system and those submitted in previous years, namely, improved acoustic and language models, cross adaptation between systems with different front-ends and phoneme sets, and the use of various automatic speech segmentation algorithms.Our system achieved word error rates of 38.5% (53.4%) and 22.9% (32.2%), respectively, on the MDM and IHM conditions of the RT-05S (RT-06S) lecture evaluation set. Christian Fügen, Matthias Wölfel, John W. McDonough, Shajith Ikbal, Florian Kraft, Kornel Laskowski, Mari Ostendorf, Sebastian Stüker, Ken'ichi Kumatani |
INTERSPEECH | 7 |
| 2006 | Improved tone modeling for Mandarin broadcast news speech recognitionabstractTone has a crucial role in Mandarin speech in distinguishing ambiguous words. Most state-of-the-art Mandarin automatic speech recognition systems adopt embedded tone modeling, where tonal acoustic units are used and F0 features are appended to the spectral feature vector. In this paper, we combine the embedded aproach (using improved F0 smoothing) with explicit tone modeling in rescoring the output lattices. Oracle experiments indicate 32% relative improvement can be achieved by rescoring with perfect tone information. Recognition experiments on Mandarin broad-cast news show that, even with an accuracy of only 70%, the explicit tone classifier offers complementary knowledge and improves performance significantly. Through the combination of tone modeling techniques, the character error rate on the CTV test set can be improved from 13.0% to 11.5%. Man-Hung Siu, Mei-Yuh Hwang, Mari Ostendorf, Tan Lee |
INTERSPEECH | 4 |
| 2006 | Speaker clustered regression-class trees for MLLR adaptationabstractA speaker clustering algorithm is presented that is based on an eigenspace representation of Maximum Likelihood Linear Regression (MLLR) transformations and is used for training cluster-dependent regression-class trees for MLLR adaptation. It is shown that significant automatic speech recognition (ASR) system performance gains are possible by choosing the best regression-class tree structure for individual speakers. To take advantage of the potential gains, an algorithm for combining the MLLR mean transformations from cluster-specific trees is described that effectively results in a soft regression-class tree. In conversational speech recognition, only small overall improvements are obtained, but the number of speakers that have performance degradation due to adaptation is reduced by over 70%. Arindam Mandal, Mari Ostendorf, Andreas Stolcke |
INTERSPEECH | 2 |
| 2006 | Assessing the reading level of web pages
Sarah E. Petersen, Mari Ostendorf |
INTERSPEECH | 2 |
| 2006 | SParseval: Evaluation Metrics for Parsing Speech
Brian Roark, Mary P. Harper, Eugene Charniak, Bonnie J. Dorr, Mark Johnson 0001, Jeremy G. Kahn, Yang Liu 0004, Mari Ostendorf, John Hale, Anna Krasnyanskaya, Matthew Lease, Izhak Shafran, Matthew G. Snover, Robin Stewart, Lisa Yung |
LREC | 8 |
| 2006 | Agreement/Disagreement Classification: Exploiting Unlabeled Data using Contrast Classifiers
Sangyun Hahn, Richard E. Ladner, Mari Ostendorf |
HLT-NAACL | 3 |
| 2006 | Impact of Automatic Comma Prediction on POS/Name Tagging of speechabstractThis work looks at the impact of automatically predicted commas on part-of-speech (POS) and name tagging of speech recognition transcripts of Mandarin broadcast news. There is a significant gain in both POS and name tagging accuracy due to using automatically predicted commas over sentence boundary prediction alone. One difference between Mandarin and English is that there are two types of commas, and experiments here show that, while they can be reliably distinguished in automatic prediction, the distinction does not give a clear benefit for POS or name tagging. Dustin Hillard, Zhongqiang Huang, Heng Ji 0001, Ralph Grishman, Dilek Hakkani-Tür, Mary P. Harper, Mari Ostendorf, Wen Wang 0001 |
SLT | 7 |
| 2006 | Parse Structure and Segmentation for Improving speech RecognitionabstractSeparate avenues of prior work have shown that parsing language models lead to improved recognition performance, and that segmentation of speech into sentence-like units has an impact on parser performance. This paper brings these two findings together, showing that segmentation also impacts the quality of a syntax-based language model, such that larger reductions in word error rate are possible when using sentence-like segmentations rather than simple paused-based strategies. Further, we show that the same types of syntactic features used in parse reranking can also be used to reduce word error rate in an N-best rescoring framework. William P. McNeill, Jeremy G. Kahn, Dustin Hillard, Mari Ostendorf |
SLT | 4 |
| 2006 | Enriching speech recognition with automatic detection of sentence boundaries and disfluenciesabstractEffective human and automatic processing of speech requires recovery of more than just the words. It also involves recovering phenomena such as sentence boundaries, filler words, and disfluencies, referred to as structural metadata. We describe a metadata detection system that combines information from different types of textual knowledge sources with information from a prosodic classifier. We investigate maximum entropy and conditional random field models, as well as the predominant hidden Markov model (HMM) approach, and find that discriminative models generally outperform generative models. We report system performance on both broadcast news and conversational telephone speech tasks, illustrating significant performance differences across tasks and as a function of recognizer performance. The results represent the state of the art, as assessed in the NIST RT-04F evaluation Yang Liu 0004, Elizabeth Shriberg, Andreas Stolcke, Dustin Hillard, Mari Ostendorf, Mary P. Harper |
IEEE Trans. Speech Audio Process. | 5 |
| 2006 | Recent innovations in speech-to-text transcription at SRI-ICSI-UWabstractWe summarize recent progress in automatic speech-to-text transcription at SRI, ICSI, and the University of Washington. The work encompasses all components of speech modeling found in a state-of-the-art recognition system, from acoustic features, to acoustic modeling and adaptation, to language modeling. In the front end, we experimented with nonstandard features, including various measures of voicing, discriminative phone posterior features estimated by multilayer perceptrons, and a novel phone-level macro-averaging for cepstral normalization. Acoustic modeling was improved with combinations of front ends operating at multiple frame rates, as well as by modifications to the standard methods for discriminative Gaussian estimation. We show that acoustic adaptation can be improved by predicting the optimal regression class complexity for a given speaker. Language modeling innovations include the use of a syntax-motivated almost-parsing language model, as well as principled vocabulary-selection techniques. Finally, we address portability issues, such as the use of imperfect training transcripts, and language-specific adjustments required for recognition of Arabic and Mandarin Andreas Stolcke, Barry Y. Chen, Horacio Franco, Venkata Ramana Rao Gadde, Martin Graciarena, Mei-Yuh Hwang, Katrin Kirchhoff, Arindam Mandal, Nelson Morgan, Tim Ng, Mari Ostendorf, M. Kemal Sönmez, Anand Venkataraman, Dimitra Vergyri, Wen Wang 0001, Jing Zheng 0001, Qifeng Zhu 0001 |
IEEE Trans. Speech Audio Process. | 12 |
| 2005 | A Quantitative Analysis of Lexical Differences Between Genders in Telephone ConversationsabstractIn this work, we provide an empirical analysis of differences in word use between genders in telephone conversations, which complements the considerable body of work in sociolinguistics concerned with gender linguistic differences. Experiments are performed on a large speech corpus of roughly 12000 conversations. We employ machine learning techniques to automatically categorize the gender of each speaker given only the transcript of his/her speech, achieving 92% accuracy. An analysis of the most characteristic words for each gender is also presented. Experiments reveal that the gender of one conversation side influences lexical use of the other side. A surprising result is that we were able to classify male-only vs. female-only conversations with almost perfect accuracy. Constantinos Boulis, Mari Ostendorf |
ACL | 2 |
| 2005 | Reading Level Assessment Using Support Vector Machines and Statistical Language ModelsabstractReading proficiency is a fundamental component of language competency. However, finding topical texts at an appropriate reading level for foreign and second language learners is a challenge for teachers. This task can be addressed with natural language processing technology to assess reading level. Existing measures of reading level are not well suited to this task, but previous work and our own pilot experiments have shown the benefit of using statistical language models. In this paper, we also use support vector machines to combine features from traditional reading level measures, statistical language models, and other language processing tools to produce a better method of assessing reading level. Sarah E. Schwarm, Mari Ostendorf |
ACL | 2 |
| 2005 | Multi-Rate and Variable-Rate Modeling of Speech At Phone and Syllable Time ScalesabstractThis paper introduces a multi-rate extension of hidden Markov models (HMMs), for joint acoustic modeling of speech at multiple time scales. The approach complements the usual short-term, phone-based representation of speech with wide modeling units and long-term temporal features. We consider two alternatives for coarse scale, representing either phones, or syllable structure and lexical stress, and both fixed- and variable-rate dependencies between time scales. Experiments on conversational telephone speech (CTS) show that the proposed multi-rate approach significantly improves recognition accuracy over HMM- and other coupled HMM-based approaches (e.g. feature concatenation) for combining short- and long-term acoustic and linguistic information. Özgür Çetin, Mari Ostendorf |
ICASSP (1) | 2 |
| 2005 | DBN-Based Multi-stream Models for Mandarin Toneme RecognitionabstractA toneme in Mandarin Chinese is a tonal phone which consists of a base phone (main vowel) and a tone. To capture both, most recognition systems use two feature streams: the standard MFCC for the base phones, and pitch features for the tones. In this paper we propose the use of dynamic Bayesian networks for modeling the two streams in toneme recognition. We used the Graphical Model Toolkit to build and compare three different models: a standard HMM with concatenated features, and synchronous and asynchronous multi-stream systems. Stream-level model parameter tying is also exploited. The toneme recognition results show significant improvements by using the multi-stream models. Gang Ji, Tim Ng, Jeff A. Bilmes, Mari Ostendorf |
ICASSP (1) | 5 |
| 2005 | Structural metadata research in the EARS programabstractBoth human and automatic processing of speech require recognition of more than just words. In this paper we provide a brief overview of research on structural metadata extraction in the DARPA EARS rich transcription program. Tasks include detection of sentence boundaries, filler words, and disfluencies. Modeling approaches combine lexical, prosodic, and syntactic information, using various modeling techniques for knowledge source integration. The performance of these methods is evaluated by task, by data source (broadcast news versus spontaneous telephone conversations) and by whether transcriptions come from humans or from an (errorful) automatic speech recognizer. A representative sample of results shows that combining multiple knowledge sources (words, prosody, syntactic information) is helpful, that prosody is more helpful for news speech than for conversational speech, that word errors significantly impact performance, and that discriminative models generally provide benefit over maximum likelihood models. Important remaining issues, both technical and programmatic, are also discussed. Yang Liu 0004, Elizabeth Shriberg, Andreas Stolcke, Barbara Peskin, Jeremy Ang, Dustin Hillard, Mari Ostendorf, Marcus Tomalin, Philip C. Woodland, Mary P. Harper |
ICASSP (5) | 7 |
| 2005 | Web-Data Augmented Language Models for Mandarin Conversational Speech RecognitionabstractLack of data is a problem in training language models for conversational speech recognition, particularly for languages other than English. Experiments in English have successfully used Web-based text collection, targeted for a conversational style, to augment small sets of transcribed speech; we look at extending these techniques to Mandarin. In addition, we investigate different techniques for topic adaptation. Experiments in recognizing Mandarin telephone conversations show that the use of filtered Web data leads to a 28% reduction in perplexity and 7% reduction in character error rate, with most of the gain due to the general filtered Web data. Tim Ng, Mari Ostendorf, Mei-Yuh Hwang, Man-Hung Siu, Ivan Bulyko |
ICASSP (1) | 2 |
| 2005 | Human language technology: opportunities and challengesabstractIn recent years, there has been dramatic progress in both speech and language processing, in many cases leveraging some of the same underlying methods. This progress and the growing technical ties motivate efforts to combine speech and language technologies in spoken document processing applications. This paper outlines some of the issues involved, as well as the opportunities, presenting an overview of the special double session on this topic. Mari Ostendorf, Elizabeth Shriberg, Andreas Stolcke |
ICASSP (5) | 1 |
| 2005 | Using symbolic prominence to help design feature subsets for topic classification and clustering of natural human-human conversations
Constantinos Boulis, Mari Ostendorf |
INTERSPEECH | 2 |
| 2005 | Incorporating tone-related MLP posteriors in the feature representation for Mandarin ASRabstractTone has a crucial role in Mandarin speech in distinguishing ambiguous words. In most state-of-the-art Mandarin automatic speech recognition systems, tonal acoustic units are used and F0 features are appended to the spectral features (MFCC/PLP). However, a tone depends on the F0 contour of a time span much longer than a frame. Ideally, systems would compute the framelevel likelihood of a tone using more than the F0 and derivative values at the current frame. Inspired by the tandem approach, we propose to extract tone-related features for each frame by using longer acoustic context information in a multi-layer perceptron (MLP). The extracted tone-related posteriors are then appended to the spectral feature vector to form a new feature vector for back-end HMM systems. Results show that significant improvement can be achieved by adding these tone-related MLP posterior features in a Mandarin conversational telephone speech recognition task. In one configuration, the character error rate was reduced from 35.7 % to 33.2%. 1. Mei-Yuh Hwang, Mari Ostendorf |
INTERSPEECH | 3 |
| 2005 | Leveraging speaker-dependent variation of adaptationabstractThis work introduces an automatic procedure for determining the size of regression class trees for individual speakers using an ensemble of speaker-level features to control the number of transformations, if any, that should be estimated by maximum likelihood linear regression. Experiments with a state-of-the-art speech recognition system that uses this procedure show improvements in word error rate for conversational telephone speech. Arindam Mandal, Mari Ostendorf, Andreas Stolcke |
INTERSPEECH | 2 |
| 2005 | Data sampling for improved speech recognizer training
Takahiro Shinozaki, Mari Ostendorf, Les E. Atlas |
INTERSPEECH | 2 |
| 2005 | Improving out-of-vocabulary name resolution
David D. Palmer, Mari Ostendorf |
Comput. Speech Lang. | 2 |
| 2005 | Error-correction detection and response generation in a spoken dialogue system
Ivan Bulyko, Katrin Kirchhoff, Mari Ostendorf, J. Goldberg |
Speech Commun. | 3 |
| 2005 | A quantitative assessment of the importance of tone in mandarin speech recognitionabstractHow important is tone to Mandarin speech recognition? Because contextual information can resolve some lexical ambiguities, the question has been raised as to whether a strong language model obviates the need for explicit tone modeling. This letter addresses that question by evaluating the effect of tone information on word perplexity under different conditions, including knowledge of the correct tone sequence, perfect acoustic information, and imperfect syllable and tone information. Results show that conditioning on tone information can reduce word uncertainty in conversational speech by 11%-20%, depending on the accuracy of syllable and tone recognition. Tim Ng, Man-Hung Siu, Mari Ostendorf |
IEEE Signal Process. Lett. | 3 |
| 2004 | Multi-rate hidden Markov models and their application to machining tool-wear classificationabstractThe paper introduces a multi-rate hidden Markov model (multi-rate HMM) for multi-scale stochastic modeling of non-stationary processes. The multi-rate HMM decomposes the process variability into scale-based components, and characterizes both the intra-scale temporal evolution of the process and the inter-scale interactions. Applying these models to the machine tool-wear classification problem in a titanium milling task shows that multi-rate HMMs outperform HMMs in terms of both accuracy and confidence of predictions. Özgür Çetin, Mari Ostendorf |
ICASSP (5) | 2 |
| 2004 | The ICSI-SRI-UW metadata extraction systemabstractBoth human and automatic processing of speech require recognizing more than just the words. We describe a state-of-the-art system for automatic detection of “metadata” (information beyond the words) in both broadcast news and spontaneous telephone conversations, developed as part of the DARPA EARS Rich Transcription program. System tasks include sentence boundary detection, filler word detection, and detection/correction of disfluencies. To achieve best performance, we combine information from different types of language models (based on words, part-of-speech classes, and automatically induced classes) with information from a prosodic classifier. The prosodic classifier employs bagging and ensemble approaches to better estimate posterior probabilities. We use confusion networks to improve robustness to speech recognition errors. Most recently, we have investigated a maximum entropy approach for the sentence boundary detection task, yielding a gain over our standard HMM approach. We report results for these techniques on the official NIST Rich Transcription metadata tasks. Elizabeth Shriberg, Andreas Stolcke, Dustin Hillard, Mari Ostendorf, Barbara Peskin, Mary P. Harper, Yang Liu 0004 |
INTERSPEECH | 4 |
| 2004 | From switchboard to meetings: development of the 2004 ICSI-SRI-UW meeting recognition systemabstractWe describe the ICSI-SRI-UW team’s entry in the Spring 2004 NIST Meeting Recognition Evaluation. The system was derived from SRI’s 5xRT Conversational Telephone Speech (CTS) recognizer by adapting CTS acoustic and language models to the Meeting domain, adding noise reduction and delay-sum array processing for far-field recognition, and postprocessing for cross-talk suppression. A modified MAP adaptation procedure was developed to make best use of discriminatively trained (MMIE) prior models. These meeting-specific changes yielded an overall 9% and 22% relative improvement as compared to the original CTS system, and 16% and 29% relative improvement as compared to our 2002 Meeting Evaluation system, for the individual-headset and multiple-distant microphones conditions, respectively. Andreas Stolcke, Chuck Wooters, Ivan Bulyko, Martin Graciarena, Scott Otterson, Barbara Peskin, Mari Ostendorf, David Gelbart, Nikki Mirghafori, Tuomo W. Pirinen |
INTERSPEECH | 7 |
| 2004 | Detecting Structural Metadata with Decision Trees and Transformation-Based Learning
Joungbum Kim, Sarah E. Schwarm, Mari Ostendorf |
HLT-NAACL | 3 |
| 2004 | Combining Multiple Clustering Systems
Constantinos Boulis, Mari Ostendorf |
PKDD | 2 |
| 2004 | Adaptive language modeling with varied sources to cover new vocabulary itemsabstractN-gram language modeling typically requires large quantities of in-domain training data, i.e., data that matches the task in both topic and style. For conversational speech applications, particularly meeting transcription, obtaining large volumes of speech transcripts is often unrealistic; topics change frequently and collecting conversational-style training data is time-consuming and expensive. In particular, new topics introduce new vocabulary items which are not included in existing models. In this work, we use a variety of data sources (reflecting different sizes and styles), combined using mixture n-gram models. We study the impact of the different sources on vocabulary expansion and recognition accuracy, and investigate possible indicators of the usefulness of a data source. For the task of recognizing meeting speech, we obtain a 9% relative reduction in the overall word error rate and a 61% relative reduction in the word error rate for "new" words added to the vocabulary over a baseline language model trained from general conversational speech data. Sarah E. Schwarm, Ivan Bulyko, Mari Ostendorf |
IEEE Trans. Speech Audio Process. | 3 |
| 2003 | Cross-stream observation dependencies for multi-stream speech recognitionabstractThis paper extends prior work in multi-stream modeling by introducing cross-stream observation dependencies and a new discriminative criterion for selecting such dependencies. Experimental results combining short-term PLP features with longterm TRAP features show gains associated with a multi-stream model with partial state asynchrony over a baseline HMM. Frame-based analyses show significant discriminant information in the added cross-stream dependencies, but so far there are only small gains in recognition accuracy. 1. Özgür Çetin, Mari Ostendorf |
INTERSPEECH | 2 |
| 2003 | Getting More Mileage from Web Text Sources for Conversational Speech Language Modeling using Class-Dependent Mixtures
Ivan Bulyko, Mari Ostendorf, Andreas Stolcke |
HLT-NAACL | 2 |
| 2003 | Detection Of Agreement vs. Disagreement In Meetings: Training With Unlabeled Data
Dustin Hillard, Mari Ostendorf, Elizabeth Shriberg |
HLT-NAACL | 2 |
| 2003 | Parameter reduction schemes for loosely coupled HMMs
Harriet J. Nock, Mari Ostendorf |
Comput. Speech Lang. | 2 |
| 2003 | Acoustic model clustering based on syllable structure
Izhak Shafran, Mari Ostendorf |
Comput. Speech Lang. | 2 |
| 2003 | Multilevel Classification of Milling Tool Wear with Confidence EstimationabstractAn important problem during industrial machining operations is the detection and classification of tool wear. Past research in this area has demonstrated the effectiveness of various feature sets and binary classifiers. Here, the goal is to develop a classifier which makes use of the dynamic characteristics of tool wear in a metal milling application and which replaces the standard binary classification result with two outputs: a prediction of the wear level (quantized) and a gradient measure that is the posterior probability (or confidence) that the tool is worn given the observed feature sequence. The classifier tracks the dynamics of sensor data within a single cutting pass as well as the evolution of wear from sharp to dull. Different alternatives to parameter estimation with sparsely-labeled training data are proposed and evaluated. We achieve high accuracy across changing cutting conditions, even with a limited feature set drawn from a single sensor. Randall K. Fish, Mari Ostendorf, Gary D. Bernard, David A. Castañón |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2002 | Robust splicing costs and efficient search with BMM Models for concatenative speech synthesisabstractWith the growing popularity of corpus-based methods for concatenative speech synthesis, a large amount of interest has been placed on borrowing techniques from the ASR community. This paper explores the applications of Buried Markov Models (BMM) to speech synthesis. We show that BMMs are more efficient than HMMs as a synthesis model, and focus on using BMM dependencies for computing splicing costs. We also show how the computational complexity of the dynamic search can be significantly reduced by constraining the splicing points with a negligible loss in synthesis quality. Ivan Bulyko, Mari Ostendorf, Jeff A. Bilmes |
ICASSP | 2 |
| 2002 | Text normalization with varied data sources for conversational speech language modelingabstractCollecting sufficient language model training data for good speech recognition performance in a new domain is often difficult. However, there may be other sources of data that are matched in terms of topic or style, if not both. This paper looks at the use of text normalization tools to make these data more suitable for language model training, in conjunction with mixture models to combine data from different sources. We specifically address the task of recognizing meeting speech, showing a small reduction in word error rate over a baseline language model trained from conversational speech data. Sarah E. Schwarm, Mari Ostendorf |
ICASSP | 2 |
| 2002 | The 2001 GMTK-based SPINE ASR systemabstractThis paper provides a detailed description of the University of Washington automatic speech recognition (ASR) system for the 2001 DARPA SPeech In Noisy Environments (SPINE) task. Our system makes heavy use of the graphical modeling toolkit (GMTK), a general purpose graphical modeling-based ASR system that allows arbitrary parameter tying, flexible deterministic and stochastic dependencies between variables, and a generalized maximum likelihood parameter estimation algorithm. In our SPINE system, GMTK was used for acoustic model training whereas feature extraction, speaker adaptation, and first-pass decoding were performed by HTK. Our integrated GMTK/HTK system demonstrates the relative merits provided by each tool. Novel aspects of our SPINE system include the capturing of correlations among feature vectors via a globally-shared factored sparse inverse covariance matrix and generalized EM training. 1. Özgür Çetin, Harriet J. Nock, Katrin Kirchhoff, Jeff A. Bilmes, Mari Ostendorf |
INTERSPEECH | 5 |
| 2002 | Efficient integrated response generation from multiple targets using weighted finite state transducers
Ivan Bulyko, Mari Ostendorf |
Comput. Speech Lang. | 2 |
| 2002 | Graceful degradation of speech recognition performance over packet-erasure networksabstractThis paper explores packet loss recovery for automatic speech recognition (ASR) in spoken dialog systems, assuming an architecture in which a lightweight client communicates with a remote ASR server. Speech is transmitted with source and channel codes optimized for the ASR application, i.e., to minimize word error rate. Unequal amounts of forward error correction, depending on the data's effect on ASR performance, are assigned to protect against packet loss. Experiments with simulated packet loss in a range of loss conditions are conducted on the DARPA Communicator (air travel information) task. Results show that the approach provides robust ASR performance which degrades gracefully as packet loss rates increase. Transmitting at 5.2 Kbps with up to 200 ms added delay, leads to only a 7% relative degradation in word error rate even under extremely adverse network conditions. Constantinos Boulis, Mari Ostendorf, Eve A. Riskin, Scott Otterson |
IEEE Trans. Speech Audio Process. | 2 |
| 2001 | Joint prosody prediction and unit selection for concatenative speech synthesisabstractWe describe how prosody prediction can be efficiently integrated with the unit selection process in a concatenative speech synthesizer under a weighted finite-state transducer (WFST) architecture. WFSTs representing prosody prediction and unit selection can be composed during synthesis, thus effectively expanding the space of possible prosodic targets. We implemented a symbolic prosody prediction module and a unit selection database as the synthesis components of a travel planning system. Results of perceptual experiments show that by combining the steps of prosody prediction and unit selection we are able to achieve improved naturalness of synthetic speech compared to the sequential implementation. Ivan Bulyko, Mari Ostendorf |
ICASSP | 2 |
| 2001 | Joint use of dynamical classifiers and ambiguity plane featuresabstractThis paper argues for using ambiguity plane features within dynamic statistical models for classification problems. The relative contribution of the two model components are investigated in the context of acoustically monitoring cutter wear during milling of titanium, an application where it is known that standard static classification techniques work poorly. Experiments show that explicit modeling of long-term context via a hidden Markov model state improves performance, but mainly by using this to augment sparsely labelled training data. An additional performance gain is achieved by using the shorter-term context of ambiguity plane features. Mari Ostendorf, Les E. Atlas, Robert Fish, Özgür Çetin, Somsak Sukittanon, Gary D. Bernard |
ICASSP | 1 |
| 2001 | Unit selection for speech synthesis using splicing costs with weighted finite state transducersabstractIn this paper we describe how unit selection for concatenative speech synthesis can be implemented efficiently for sub-phonetic units using weighted finite state transducers (WFST). We also introduce splicing costs as a measure to indicate which unit boundaries are particularly good or poor joint points. Splicing costs extend the flexibility offered by the unit selection paradigm. Through a perceptual experiment we demonstrate an improvement in speech quality achieved by using splicing costs during unit selection. Ivan Bulyko, Mari Ostendorf |
INTERSPEECH | 2 |
| 2001 | Improved word confidence estimation using long range featuresabstractThis paper describes experiments in improving word confidence estimation using document- and task-level features of the hypothesized word sequence from a recognizer. The improved confidence estimates are shown to improve information extraction performance, specifically named entity (NE) recognition. The detected names can then be used to further improve confidence estimation in a multi-pass NE recognition framework. David D. Palmer, Mari Ostendorf |
INTERSPEECH | 2 |
| 2001 | Graceful degradation of speech recognition performance over lossy packet networksabstractThis paper explores packet loss recovery in client-server Automatic Speech Recognition (ASR) systems. A forward error correction (FEC) system is designed and tested over several channel loss models, at variable amounts of data acquisition delay. In experiments with simulated packet loss, the FEC system provides robust ASR performance which degrades gracefully as packet loss rates increase. Comparing this scheme to several alternatives under low and medium loss channel conditions, we found one approach (multiple transmission plus interpolation) that yielded similar performance, but the FEC system should scale better to lower bit rate conditions. Eve A. Riskin, Constantinos Boulis, Scott Otterson, Mari Ostendorf |
INTERSPEECH | 4 |
| 2001 | Obituary: M. W. Macon 1969-2001
Mari Ostendorf, Jan P. H. van Santen, Mark A. Clements |
Comput. Speech Lang. | 1 |
| 2001 | Normalization of non-standard words
Richard Sproat, Alan W. Black, Stanley F. Chen, Shankar Kumar, Mari Ostendorf, Christopher Richards |
Comput. Speech Lang. | 5 |
| 2000 | Hidden Markov models for monitoring machining tool-wearabstractAs summarized by Atlas, Bernard, and Narayanan (1996), the sensing of acoustic vibrations can remotely estimate the state of wear at the tool edge. This form of monitoring offers the potential to characterize, in real time, the efficiency of metal removal processes such as drilling and milling. For example, information about sudden increases in tool wear, if manifest as a change in acoustic vibration, could be valuable to a machine operator. The nature of this monitoring problem has some similarities to automatic speech recognition. For example, there is significant tool-to-tool variation in details of vibration and lifetime. Also, the easy adaptability of monitoring systems across manufacturing processes is important. In this work we model the evolution of vibration signals with the same technique which has shown to be successful in speech recognition: hidden Markov models (HMMs). We focus on the monitoring of milling processes at three different time scales and show the how HMMs can give accurate wear prediction. Les E. Atlas, Mari Ostendorf, Gary D. Bernard |
ICASSP | 2 |
| 2000 | Use of higher level linguistic structure in acoustic modeling for speech recognitionabstractCurrent speech recognition systems perform poorly on conversational speech as compared to read speech, largely because of the additional acoustic variability observed in conversational speech. Our hypothesis is that there are systematic effects, related to higher level structures, that are not being captured in the current acoustic models. In this paper we describe a method to extend standard clustering to incorporate such features in estimating acoustic models. We report recognition improvements obtained on the Switchboard task over triphones and pentaphones by the use of word- and syllable-level features. In addition, we report preliminary studies on clustering with prosodic information. Izhak Shafran, Mari Ostendorf |
ICASSP | 2 |
| 2000 | Integrating a context-dependent phrase grammar in the variable n-gram frameworkabstractFocuses on the learning of multi-word lexical units, or phrases, and how to model them within the variable n-gram framework. We introduce the notion of context-dependent phrases and suggest an algorithm for the unsupervised learning of phrases. Also, we propose an approach to integrate a phrase grammar and a variable n-gram without the need to explicitly handle multi-word lexical items. The combined variable n-gram phrase grammar improves recognition accuracy on the Switchboard corpus over both the baseline trigram and using a variable n-gram alone. Man-Hung Siu, Mari Ostendorf |
ICASSP | 2 |
| 2000 | Obituary: J. Allen 1934-2000
Mari Ostendorf |
Comput. Speech Lang. | 1 |
| 2000 | Editorial: New developments at CSL
Mari Ostendorf, Stephen Young |
Comput. Speech Lang. | 1 |
| 2000 | Robust information extraction from automatically generated speech transcriptions
David D. Palmer, Mari Ostendorf, John D. Burger |
Speech Commun. | 2 |
| 2000 | Variable n-grams and extensions for conversational speech language modelingabstractRecent progress in variable n-gram language modeling provides an efficient representation of n-gram models and makes training of higher order n grams possible. We apply the variable n-gram design algorithm to conversational speech, extending the algorithm to learn skips and context-dependent classes to handle conversational speech characteristics such as filler words, repetitions, and other disfluencies. Experiments show that using the extended variable n-gram results in a language model that captures 4-gram context with less than half the parameters of a standard trigram while also improving the test perplexity and recognition accuracy. Man-Hung Siu, Mari Ostendorf |
IEEE Trans. Speech Audio Process. | 2 |
| 1999 | Predicting gradient F0 variation: pitch range and accent prominenceabstractThis paper uses an information-based approach to conduct feature types selection for language modeling in a systematic manner.We describe a quantitative analysis of the information gain and the information redundancy for various combinations of feature types inspired by both dependency structure and bigram structure through analyzing an English treebank corpus and taking word prediction as the object.The experiments yield several conclusions on the predictive value of several feature types and feature types combinations for word prediction, which are expected to provide reliable reference for feature type selection in language modeling. Ivan Bulyko, Mari Ostendorf |
EUROSPEECH | 2 |
| 1999 | A new metric for stochastic language model evaluationabstractThough Perplexity shows good correlation with word error rate within simple n-gram framework like Wall Street Journal task, it has been reported that perplexity have poor correlation with WER when more complicated LM is used. In this paper, a global measure for language model evaluation is proposed which achieveshigher correlation between word accuracy. The metric is based on difference of LM score between a word in the evaluation text and the word that gives the maximum score at that context. Two experiments were carried out to investigate the correlation between word accuracy and the proposed measure. In the first experiment, LMs in this paper were created using n-gram adaptation by n-gram count mixture. 47 LMs were created for the experiments by changing mixture weight and vocabulary cut-off threshold. Correlation betwen perplexity and word accuracy was very poor (correlation coefficient -0.36). On the other hand, the proposed metric gave much higher correlation (correlation coefficient 0.82). In the second experiment, a simple mixture trigram model was employed to recognize Switchboard task data. The highest correlation between word accuracy and the proposed method was 0.81, which was much higher than the correlation between PP and accucary 0.59. Akinori Ito, Masaki Kohda, Mari Ostendorf |
EUROSPEECH | 3 |
| 1999 | Robust information extraction from spoken language dataabstractThis paper presents our new approach to model tone coarticulation of Chinese continuous speech for tone recognition. We suggest that coarticulation effects between two neighboring tones are rather unstable, since they may be uni-directional, bi-directional, or none despite of the same phonetic contexts. Instability is suggested due to non-local prosodic events like prosodic phrase boundaries or stress effects. Hence, we propose that context dependent tone models should be estimated according to the exact underlying coarticulation effects. To simplify label work for coarticulation effects, label sets as few as 3 labels are adopted. Also F0 contours of tone nuclei are used to facilitate human to discriminate tone coarticulation effects. A new search algorithm for the output candidates was also proposed to adopt the new modeling method of tone coarticulation effects. Preliminary experiments on a female’s utterances of data corpus HKU96 showed the effectiveness of the new approach. David D. Palmer, Mari Ostendorf, John D. Burger |
EUROSPEECH | 2 |
| 1999 | Relevance weighting for combining multi-domain data for n-gram language modelingabstractStandard statistical language modeling techniques suffer from sparse-data problems in tasks where large amounts of domain-specific text are not available. In this paper, we focus on improving the estimation of domain-dependent n -gram models by the selective use of out-of-domain text data. Previous approaches for estimating language models from multi-domain data have not accounted for the characteristic variations of style and content across domains. In contrast, this work aims at differentially weighting subsets of the out-of-domain data according to style and/or content similarity to the given task, where “style" is represented by part-of-speech statistics and “content" by the particular choice of vocabulary items. In addition to n -gram estimation, the differential weights can be used for lexicon design. Recognition experiments are based on the Switchboard corpus of spontaneous conversations, with out-of-domain text drawn from the Wall Street Journal and Broadcast News corpora. The similarity weighting approach gives a 3–5% reduction in word error rate over a domain-specific n -gram language model, providing some of the largest language modeling gains reported for the Switchboard task in recent years. Rukmini Iyer, Mari Ostendorf |
Comput. Speech Lang. | 2 |
| 1999 | Joint lexicon, acoustic unit inventory and model designabstractAlthough most parameters in a speech recognition system are estimated from data by the use of an objective function, the unit inventory and lexicon are generally hand crafted and therefore unlikely to be optimal. This paper proposes a joint solution to the related problems of learning a unit inventory and corresponding lexicon from data. On a speaker-independent read speech task with a 1k vocabulary, the proposed algorithm outperforms phone-based systems at both high and low complexities. Obwohl die meisten Parameter eines Spracherkennungssystems aus Daten geschätzt werden, ist die Wahl der akustischen Grundeinheiten und des Lexikons normalerweise nicht automatisch und deshalb wahrscheinlich nicht optimal. Dieser Artikel stellt einen kombinierten Ansatz für die Lösung dieser verwandten Probleme dar – das Lernen von akustischen Grundeinheiten und des zugehörigen Lexikons aus Daten. Experimente mit sprecher-unabhängigen gelesenen Sprachdaten mit einem Vokabular von 1000 Wörtern zeigen, daß der vorgestellte Ansatz besser ist als ein System niedriger oder höherer Komplexität, das auf Phonemen basiert ist. Bien que la plupart des paramètres dans un système de reconnaissance de la parole soient estimés à partie des données en utilisant une fonction objective, l'inventaire des unités acoustiques et le lexique sont généralement créés à la main, et donc susceptibles de ne pas être optimeux. Cette étude propose une solution conjointe aux problèmes interdépendants que sont l'apprentissage à partir des données d'un inventaire des unités acoustiques et du lexique correspondant. Nous avons testé l'algorithme proposé sur des échantillons lus, en reconnaissance indépendantes du locuteur avec un vocabulaire de 1k: il surpassé les systèmes phonétiques en faible ou forte complexité. Michiel Bacchiani, Mari Ostendorf |
Speech Commun. | 2 |
| 1999 | Reducing the effects of linear channel distortion on continuous speech recognitionabstractLinear channel compensation in speech recognition typically involves estimating an additive shift in the cepstral domain. This paper explores both Bayesian and maximum likelihood techniques to transform either the features or the model parameters. Experiments on the Macrophone corpus show error rate reductions of up to 16% over cepstral mean subtraction for short utterances. Rebecca Bates 0001, Mari Ostendorf |
IEEE Trans. Speech Audio Process. | 2 |
| 1999 | Modeling long distance dependence in language: topic mixtures versus dynamic cache modelsabstractStandard statistical language models use n-grams to capture local dependencies, or use dynamic modeling techniques to track dependencies within an article. In this paper, we investigate a new statistical language model that captures topic-related dependencies of words within and across sentences. First, we develop a topic-dependent, sentence-level mixture language model which takes advantage of the topic constraints in a sentence or article. Second, we introduce topic-dependent dynamic adaptation techniques in the framework of the mixture model, using n-gram caches and content word unigram caches. Experiments with the static (or unadapted) mixture model on the North American Business (NAB) task show a 21% reduction in perplexity and a 3-4% improvement in recognition accuracy over a general n-gram model, giving a larger gain than that obtained with supervised dynamic cache modeling. Further experiments on the Switchboard corpus also showed a small improvement in performance with the sentence-level mixture model. Cache modeling techniques introduced in the mixture framework contributed a further 14% reduction in perplexity and a small improvement in recognition accuracy on the NAB task for both supervised and unsupervised adaptation. Rukmini Iyer, Mari Ostendorf |
IEEE Trans. Speech Audio Process. | 2 |
| 1999 | A dynamical system model for generating fundamental frequency for speech synthesisabstractHigher quality speech synthesis is required for widespread use of text to-speech (TTS) technology, and prosody is one component of synthesis technology with the greatest need for improvement. This paper describes a new approach to generation of two important cues to prosodic patterns-fundamental frequency (F/sub 0/) and energy contours-given symbolic prosodic labels and text. Specifically, the approach represents vectors of F/sub 0/ and energy with a dynamical system model, which allows automatic estimation of parameters from labeled speech. Parameters at different time scales in the model are structured to capture segment, syllable, phrase and discourse level effects based on linguistic research. F/sub 0/ generation experiments with the dynamical system model show improved synthetic speech quality over the hybrid target/filter approach. Kenneth N. Ross, Mari Ostendorf |
IEEE Trans. Speech Audio Process. | 2 |
| 1998 | Using automatically-derived acoustic sub-word units in large vocabulary speech recognition
Michiel Bacchiani, Mari Ostendorf |
ICSLP | 2 |
| 1998 | Prosody prediction for speech synthesis using transformational rule-based learningabstractCurrent speech synthesis systems produce intelligible output under conditions with low background noise and low cognitive load. However, the quality is far from natural and intelligibility degrades significantly under less than ideal conditions. One widely agreed upon area of improvement in speech synthesis output is prosody. Prosody includes the acoustic characteristics of speech that communicate important syntactic, semantic, and discourse information about the utterance. The acoustic correlate of prosody are the pauses, fundamental frequency contours, energy, and duration changes of utterances. Typically, prosody synthesis is a two step process, where symbolic prosodic labels such as phrase boundaries and relative emphasis are predicted from annotated text and then the acoustic correlates are predicted from these labels combined with phonetic information. The goal of this research is to improve the prediction of symbolic prosodic labels for text-to-speech systems, specifically, location of phrase boundaries and phrase-level emphasis (i.e. pitch accents). To date, the most successful algorithms for predicting symbolic prosodic labels are based on either handwritten rules or statistical methods. This research will adopt and modify an alternative algorithm: transformational rule-based learning (TRBL), which has had success in many natural language processing tasks. This learning algorithm is automatically trainable like statistical methods, but is less sensitive to sparse training data conditions than these methods. A second contribution of the thesis is an analysis of the interaction of phrase and accent symbols in prediction. Previous approaches have predicted these prosodic events in a serial fashion, but the order is not agreed upon. In this study, we compare seria... Cameron S. Fordyce, Mari Ostendorf |
ICSLP | 2 |
| 1998 | SABLE: a standard for TTS markupabstractCurrently, speech synthesizers are controlled by a multitude of proprietary tag sets. These tag sets vary substantially across synthesizers and are an inhibitor to the adoption of speech synthesis technology by developers. SABLE is an XML/SGML-based markup scheme for text-to-speech synthesis, developed to address the need for a common TTS control paradigm. This paper presents an overview of the SABLE specification, and provides links to sites where further information on SABLE can be accessed. Richard Sproat, Andrew J. Hunt, Mari Ostendorf, Paul Taylor 0001, Alan W. Black, Kevin A. Lenzo, Mike Edgington |
ICSLP | 3 |
| 1998 | Automatic detection of sentence boundaries and disfluencies based on recognized wordsabstractWe study the problem of detecting linguistic events at interword boundaries, such as sentence boundaries and disfluency locations, in speech transcribed by an automatic recognizer. Recovering such events is crucial to facilitate speech understanding and other natural language processing tasks. Our approach is based on a combination of prosodic cues modeled by decision trees, and word-based event N-gram language models. Several model combination approaches are investigated. The techniques are evaluated on conversational speech from the Switchboard corpus. Model combination is shown to give a significant win over individual knowledge sources. 1. INTRODUCTION Current automatic speech recognition systems output a string of words. Most natural language understanding systems, however, require structural information such as punctuation, which is present in text but not overtly indicated in spoken language. Similarly, for speech understanding and information extraction, it is important to fi... Andreas Stolcke, Elizabeth Shriberg, Rebecca Bates 0001, Mari Ostendorf, Dilek Hakkani-Tür, Madelaine Plauché, Gökhan Tür |
ICSLP | 4 |
| 1998 | A comparison of constrained trajectory segment models for large vocabulary speech recognitionabstractThis paper compares parametric and nonparametric constrained-mean trajectory segment models for large vocabulary speech recognition, extending distribution clustering techniques to handle polynomial mean trajectory models for robust parameter estimation. The parametric model has fewer free parameters and gives similar recognition performance to the nonparametric model, but has higher recognition costs. Ashvin Kannan, Mari Ostendorf |
IEEE Trans. Speech Audio Process. | 2 |
| 1997 | Adaptation of polynomial trajectory segment models for large vocabulary speech recognitionabstractSegment models are a generalization of HMMs that can represent feature dynamics and/or correlation in time. We develop the theory of Bayesian and maximum-likelihood adaptation for a segment model characterized by a polynomial mean trajectory. We show how adaptation parameters can be shared and adaptation detail can be controlled at run-time based on the amount of adaptation data available. Results on the Switchboard corpus show error reductions for unsupervised transcription mode adaptation and supervised batch mode adaptation. Ashvin Kannan, Mari Ostendorf |
ICASSP | 2 |
| 1997 | Transforming out-of-domain estimates to improve in-domain language models
Rukmini Iyer, Mari Ostendorf |
EUROSPEECH | 2 |
| 1997 | Modeling dependency in adaptation of acoustic models using multiscale tree processes
Ashvin Kannan, Mari Ostendorf |
EUROSPEECH | 2 |
| 1997 | Variable n-gram language modeling and extensions for conversational speech
Man-Hung Siu, Mari Ostendorf |
EUROSPEECH | 2 |
| 1997 | HMM topology design using maximum likelihood successive state splittingabstractModelling contextual variations of phones is widely accepted as an important aspect of a continuous speech recognition system, and HMM distribution clustering has been sucessfully used to obtain robust models of context through distribution tying. However, as systems move to the challenge of spontaneous speech, temporal variation also becomes important. This paper describes a method fordesigning HMM topologies that learn both temporal and contextual variation, extending previous work on successive state splitting (SSS). The new approach uses a maximum likelihood criterion consistently at each step, overcoming the previous SSS limitation to speaker-dependent training. Initial experiments show both performance gains and training cost reduction over SSS with the reformulated algorithm. Mari Ostendorf, Harald Singer |
Comput. Speech Lang. | 1 |
| 1997 | Prosodic and lexical indications of discourse structure in human-machine interactionsabstractFrom a discourse perspective, utterances may vary in at least two important respects: (i) they can occupy a different hierarchical position in a larger-scale information unit and (ii) they can represent different types of speech acts. Spoken language systems will improve if they adequately take into account both discourse segmentation and utterance purpose. An important question then is how such discourse-structural features can be detected. Analyses of monologues and human-human dialogues have shown that a good indicator of these factors is prosody, defined as the set of suprasegmental speech features. This paper explores whether speakers also use prosody to highlight discourse structure in a particular type of human-machine interaction, viz., information query in a travel-planning domain. More specifically, it investigates if speakers signal (i) the start of a new topic by marking the initial utterance of a discourse segment, and (ii) whether an utterance is a normal request for information or part of a correction sub-dialogue. The study reveals that in human-machine interactions, both discourse segmentation and utterance purpose can have particular prosodic correlates, although speakers also mark this information through choice of wording. Therefore, it is useful to explore in the future the possibilities of incorporating prosody in spoken language systems as a cue to discourse structure. Auf Grund der Struktur einer Rede können sich sprachliche Äuβerungen in mindestens zweierlei Hinsicht voneinander unterscheiden: (i) in ihrer hierarchischen Position in einer gröβeren Informationseinheit, und (ii) in Hinblick auf den Sprechakt, den sie vertreten. Eine adäquate Inkorporierung der Diskurssegmentierung und der Sprechakte sollte dazu beitragen, sprachliche Dialogsysteme zu verbessern. Eine wichtige Frage ist dabei, wie die strukturellen Eigenschaften einer Rede gefunden werden können. Untersuchungen von Monologen und Dialogen zwischen Menschen haben gezeigt daβ Prosodie, definiert als die Gesamtheit suprasegmenteller Spracheigenschaften, ein guter Indikator dieser Faktoren ist. Dieser Beitrag erforscht, am Beispiel eines Reiseinformationssystems, ob Sprecher Prosodie in gleicher Weise auch bei Mensch-Maschine Kommunikation benutzen. Das spezifische Ziel ist herauszufinden, ob Sprecher mit prosodischen Mitteln andeuten (i) ob eine Äuβerung die erste in einer gröβeren Informationseinheit ist, und (ii) ob eine Äuβerung eine einfache Bitte um Information ist, oder eine Korrektur. Diese Arbeit weist nach, daβ die Diskurseinheiten und Sprechakte prosodische Korrelate haben in Mensch-Maschine Dialogen, obwohl Sprecher auch lexikalische Mittel anwenden. Es ist deshalb sinnvoll, in der Zukunft die Inkorporierung prosodischer Eigenschaften in sprachlichen Dialogsystemen weiter zu untersuchen. Les énoncés d'un discours peuvent varier selon deux dimensions essentielles: (i) leur position vis-à-vis de la hiérarchie des unités du discours, (ii) leur contenu, en termes d'actes du discours. L'amélioration des systèmes de dialogue oral peut donc être substantielle s'il est convenablement tenu compte des caractéristiques structurelles du discours et de l'intention d'un énoncé. Une question importante concerne donc la détection de ces éléments structurels. L'analyse de monologues et de dialogues homme-homme montre que la prosodie est un indicateur fidèle de ces facteurs. Cet article étudie dans une application de renseignements pour voyages touristiques ou d'affaires si les locuteurs utilisent la prosodie, pour souligner la structure du discours. Plus particulièrement, il explore si les locuteurs indiquent (i) un changement de thème en marquant prosodiquement le premier énoncé d'une partie du discours, et (ii) si un énoncé donné exprime une demande d'information ou une correction dans un sous-dialogue. Cette étude montre qu'en situation de dialogue homme-machine, la segmentation du discours et l'intention d'un énoncé peuvent être spécifiquement corrélées à des indices prosodiques, bien que les locuteurs puissent également avoir recours au choix d'une phraséologie particulière. L'intégration d'indices prosodiques dans les systèmes de dialogue oral semble donc être une issue prometteuse. Marc Swerts, Mari Ostendorf |
Speech Commun. | 2 |
| 1997 | Using out-of-domain data to improve in-domain language modelsabstractStandard statistical language modeling techniques suffer from sparse data problems when applied to real tasks in speech recognition, where large amounts of domain-dependent text are not available. We investigate new approaches to improve sparse application-specific language models by combining domain dependent and out-of-domain data, including a back-off scheme that effectively leads to context-dependent multiple interpolation weights, and a likelihood-based similarity weighting scheme to discriminatively use data to train a task-specific language model. Experiments with both approaches on a spontaneous speech recognition task (switchboard), lead to reduced word error rate over a domain-specific n-gram language model, giving a larger gain than that obtained with previous brute-force data combination approaches. Rukmini Iyer, Mari Ostendorf, Herbert Gish |
IEEE Signal Process. Lett. | 2 |
| 1996 | Design of a speech recognition system based on acoustically derived segmental unitsabstractThe design of a speech recognition system based on acoustically-derived, segmental units can be divided in three steps: unit design, lexicon building and pronunciation modeling. We formulate an iterative unit design procedure which consistently uses a maximum likelihood (ML) objective in successive application of resegmentation and model re-estimation. The lexicon building allows multi-word entries in the lexicon but restricts the number of these entries in order to avoid a too costly search. Selected multi-word lexical entries are those with high frequency (such as function words) and those which consistently exhibit cross-word phone assimilation. The stochastic pronunciation model represents the likelihood of a particular acoustic segment sequence given the phonetic baseform of a lexical item, where the sequence of baseform phones are treated as a Markov state sequence and each state can emit multiple segments. Michiel Bacchiani, Mari Ostendorf, Yoshinori Sagisaka, Kuldip K. Paliwal |
ICASSP | 2 |
| 1996 | A dependence tree model of phone correlationabstractThis paper introduces a new statistical model for acoustic observations in speech recognition. The model represents intra-utterance phone correlation and can be viewed as a probabilistic model of an utterance or a speaker. The phone correlation is modeled by a dependence tree, using the mutual information between pairs of phones as a measure of correlation. The experiments presented focus on robust tree topology design and on robust estimation for a given topology. With appropriate design algorithms, the dependence trees are shown to provide a better model for independent test data than an independent phone model. Orith Ronen, Mari Ostendorf |
ICASSP | 2 |
| 1996 | Maximum likelihood successive state splittingabstractModeling contextual variations of phones is widely accepted as an important aspect of a continuous speech recognition system, and much research has been devoted to finding robust models of context for HMM systems. In particular, decision tree clustering has been used to tie output distributions across pre-defined states, and successive state splitting (SSS) has been used to define parsimonious HMM topologies. We describe a new HMM design algorithm, called maximum likelihood successive state splitting (ML-SSS), that combines advantages of both these approaches. Specifically, an HMM topology is designed using a greedy search for the best temporal and contextual splits using a constrained EM algorithm. In Japanese phone recognition experiments, ML-SSS shows recognition performance gains and training cost reduction over SSS under several training conditions. Harald Singer, Mari Ostendorf |
ICASSP | 2 |
| 1996 | Modeling long distance dependence in language: topic mixtures vs. dynamic cache models
Rukmini Iyer, Mari Ostendorf |
ICSLP | 2 |
| 1996 | Modeling disfluencies in conversational speech
Man-Hung Siu, Mari Ostendorf |
ICSLP | 2 |
| 1996 | Prediction of abstract prosodic labels for speech synthesisabstractHigher quality speech synthesis is required to make text-to-speech technologies useful in more applications, and prosody is one component of synthesis technology with the greatest need for improvement. This paper describes computational models for the prediction of abstract prosodic labels for synthesis—accent location, symbolic tones and relative prominence level—from text that is tagged with part-of-speech labels and marked for prosodic constituent structure. Specifically, the model uses multiple levels of a prosodic hierarchy and at each level combines decision tree probability functions with Markov sequence assumptions. An advantage of decision trees is the ability to incorporate linguistic knowledge in an automatic training framework, which is needed for building systems that reflect particular speaking styles. Studies of accent and tone variability across speakers are reported and used to motivate new evaluation metrics. Prediction experiments show an improvement in accuracy of prominence location prediction over simple decision trees, with accuracy similar to the level of variability observed across speakers. Kenneth N. Ross, Mari Ostendorf |
Comput. Speech Lang. | 2 |
| 1996 | From HMM's to segment models: a unified view of stochastic modeling for speech recognitionabstractMany alternative models have been proposed to address some of the shortcomings of the hidden Markov model (HMM), which is currently the most popular approach to speech recognition. In particular, a variety of models that could be broadly classified as segment models have been described for representing a variable-length sequence of observation vectors in speech recognition applications. Since there are many aspects in common between these approaches, including the general recognition and training problems, it is useful to consider them in a unified framework. The paper describes a general stochastic model that encompasses most of the models proposed in the literature, pointing out similarities of the models in terms of correlation and parameter tying assumptions, and drawing analogies between segment models and HMMs. In addition, we summarize experimental results assessing different modeling assumptions and point out remaining open questions. Mari Ostendorf, Vassilios Digalakis, Owen Kimball |
IEEE Trans. Speech Audio Process. | 1 |
| 1995 | Lattice-based search strategies for large vocabulary speech recognitionabstractThe design of search algorithms is an important issue in large vocabulary speech recognition, especially as more complex models are developed for improving recognition accuracy. Multi-pass search strategies have been used as a means of applying simple models early on to prune the search space for subsequent passes using more expensive knowledge sources. The pruned search space is typically represented by an N-best sentence list or a word lattice. Here, we investigate three alternatives for lattice search: N-best rescoring, a lattice dynamic programming search algorithm and a lattice local search algorithm. Both the lattice dynamic programming and lattice local search algorithms are shown to achieve comparable performance to the N-best search algorithm while running as much as 10 times faster on a 20 k word lexicon; the local search algorithm has the additional advantage of accommodating sentence-level knowledge sources. F. Richardson, Mari Ostendorf, Jan Robin Rohlicek |
ICASSP | 2 |
| 1995 | A dynamical system model for recognizing intonation patterns
Kenneth N. Ross, Mari Ostendorf |
EUROSPEECH | 2 |
| 1995 | Parameter estimation of dependence tree models using the EM algorithmabstractA dependence tree is a model for the joint probability distribution of an n-dimensional random vector, which requires a relatively small number of free parameters by making Markov-like assumptions on the tree. The authors address the problem of maximum likelihood estimation of dependence tree models with missing observations, using the expectation-maximization algorithm. The solution involves computing observation probabilities with an iterative "upward-downward" algorithm, which is similar to an algorithm proposed for belief propagation in causal trees, a special case of Bayesian networks.> Orith Ronen, Jan Robin Rohlicek, Mari Ostendorf |
IEEE Signal Process. Lett. | 3 |
| 1995 | The challenge of spoken language systems: research directions for the ninetiesabstractA spoken language system combines speech recognition, natural language processing and human interface technology. It functions by recognizing the person's words, interpreting the sequence of words to obtain a meaning in terms of the application, and providing an appropriate response back to the user. Potential applications of spoken language systems range from simple tasks, such as retrieving information from an existing database (traffic reports, airline schedules), to interactive problem solving tasks involving complex planning and reasoning (travel planning, traffic routing), to support for multilingual interactions. We examine eight key areas in which basic research is needed to produce spoken language systems: (1) robust speech recognition; (2) automatic training and adaptation; (3) spontaneous speech; (4) dialogue models; (5) natural language response generation; (6) speech synthesis and speech generation; (7) multilingual systems; and (8) interactive multimodal systems. In each area, we identify key research challenges, the infrastructure needed to support research, and the expected benefits. We conclude by reviewing the need for multidisciplinary research, for development of shared corpora and related resources, for computational support and far rapid communication among researchers. The successful development of this technology will increase accessibility of computers to a wide range of users, will facilitate multinational communication and trade, and will create new research specialties and jobs in this rapidly expanding area.> Ronald A. Cole, Lynette Hirschman, Les E. Atlas, Mary E. Beckman, Alan Biermann, Marcia A. Bush, Mark A. Clements, Jordan Cohen, Oscar Garcia, Brian A. Hanson, Hynek Hermansky, Steve Levinson, Kathy McKeown, Nelson Morgan, David G. Novick, Mari Ostendorf, Sharon L. Oviatt, Patti Price, Harvey F. Silverman, Judy Spitz, Alex Waibel, Clifford J. Weinstein, Stephen A. Zahorian, Victor Zue |
IEEE Trans. Speech Audio Process. | 16 |
| 1994 | A Hierarchical Stochastic Model for Automatic Prediction of Prosodic Boundary Location
Mari Ostendorf, Nanette Veilleux |
Comput. Linguistics | 1 |
| 1994 | Maximum likelihood clustering of Gaussians for speech recognitionabstractDescribes a method for clustering multivariate Gaussian distributions using a maximum likelihood criterion. The authors point out possible applications of model clustering, and then use the approach to determine classes of shared covariances for contest modeling in speech recognition, achieving an order of magnitude reduction in the number of covariance parameters, with no loss in recognition performance.> Ashvin Kannan, Mari Ostendorf, Jan Robin Rohlicek |
IEEE Trans. Speech Audio Process. | 2 |
| 1994 | Automatic labeling of prosodic patternsabstractThis paper describes a general algorithm for labeling prosodic patterns in speech, which provides a mechanism for mapping sequences of observations (vectors of acoustic correlates) to prosodic labels using decision trees and a Markov sequence model. Important and novel features of the approach are that it allows many dissimilar correlates to be treated in a unified manner to provide more robust labeling, and that it is designed to be a post-word-recognition processing step. Application of the algorithm is illustrated with experimental results for labeling prosodic phrasing and phrasal prominence in two corpora of professionally read speech. The labels produced by the automatic algorithm exhibit agreement with hand-labeled prominence and phrasing that is close to the agreement between different human labelers.> Colin W. Wightman, Mari Ostendorf |
IEEE Trans. Speech Audio Process. | 2 |
| 1993 | A comparison of trajectory and mixture modeling in segment-based word recognition
Ashvin Kannan, Mari Ostendorf |
ICASSP (2) | 2 |
| 1993 | Probabilistic parse scoring with prosodic information
Nanette Veilleux, Mari Ostendorf |
ICASSP (2) | 2 |
| 1993 | Parse scoring with prosodic information: an analysis/synthesis approach
Mari Ostendorf, Colin W. Wightman, Nanette Veilleux |
Comput. Speech Lang. | 1 |
| 1993 | ML estimation of a stochastic linear system with the EM algorithm and its application to speech recognitionabstractA nontraditional approach to the problem of estimating the parameters of a stochastic linear system is presented. The method is based on the expectation-maximization algorithm and can be considered as the continuous analog of the Baum-Welch estimation algorithm for hidden Markov models. The algorithm is used for training the parameters of a dynamical system model that is proposed for better representing the spectral dynamics of speech for recognition. It is assumed that the observed feature vectors of a phone segment are the output of a stochastic linear dynamical system, and it is shown how the evolution of the dynamics as a function of the segment length can be modeled using alternative assumptions. A phoneme classification task using the TIMIT database demonstrates that the approach is the first effective use of an explicit model for statistical dependence between frames of speech.> Vassilios Digalakis, Jan Robin Rohlicek, Mari Ostendorf |
IEEE Trans. Speech Audio Process. | 3 |
| 1992 | A Bayesian approach to speaker adaptation for the stochastic segment modelabstractSpeaker adaptation is frequently used to achieve good speech recognition performance without the high costs associated with training a speaker-dependent model. The main goal of this study is to investigate speaker adaptation for recognizers using multivariate Gaussian densities, specifically, the stochastic segment model. A Bayesian approach is followed, with estimation of the parameters of a speaker-adapted model based on prior densities obtained from speaker-independent data. Experimental results achieve 16% error reduction using mean adaptation with roughly 3 min of speech, nearly half the difference between speaker-independent and speaker-dependent recognition rates.> Burhan Necioglu, Mari Ostendorf, Jan Robin Rohlicek |
ICASSP | 2 |
| 1992 | Context modeling with the stochastic segment modelabstractThe authors describe an approach, the stochastic segment model, for context modeling in continuous speech recognition for models based on multivariate Gaussian distributions. Typically, robust context models in hidden Markov models (HMMs) are obtained by using mixture distributions; here the authors tie covariance parameters across classes of similar context. The specific classes over which parameters are tied can be based on models with less context or determined by clustering, where they have investigated both hand-specified linguistically motivated clusters and automatic k-means clustering. Experimental results on phoneme classification show that clustering improves performance, and word recognition results show that error reduction over context-independent models using this approach is comparable to that achieved with discrete hidden-Markov models using mixture distributions.> Mari Ostendorf, Ibrahim Bechwati, Owen Kimball |
ICASSP | 1 |
| 1992 | Automatic recognition of intonational featuresabstractThe authors report the initial development of an algorithm to automatically detect boundary tones and prominences in continuous speech. Utilizing phoneme durations given by a speech recognizer, the authors use a tree quantizer and hidden Markov model to label these intonational features. In speaker-independent tests on a corpus of professionally read speech, 77% of the boundary tones were correctly detected while 3% of the detections were false alarms. For prominences, the corresponding numbers were 86% and 14%.> Colin W. Wightman, Mari Ostendorf |
ICASSP | 2 |
| 1992 | Factors affecting pitch accent placement
Kenneth N. Ross, Mari Ostendorf, Stefanie Shattuck-Hufnagel |
ICSLP | 2 |
| 1992 | TOBI: a standard for labeling English prosody
Kim E. A. Silverman, Mary E. Beckman, John F. Pitrelli, Mari Ostendorf, Colin W. Wightman, Patti Price, Janet B. Pierrehumbert, Julia Hirschberg |
ICSLP | 4 |
| 1992 | Parse scoring with prosodic informationabstractThe relative size and location of prosodic phrase boundaries provides an important cue for resolving syntactic ambiguity, and can be used to improve the accuracy of automatic speech understanding. This paper describes an approach to scoring candidate sentence hypotheses and associated parses using prosodic phrase cues. Specifically, for each hypothesized parse, prosodic breaks are automatically detected and the probability of these breaks given the parse is computed based on a stochastic model of the prosody/syntax relationship. The parse probability can be used to rank sentence hypotheses and associated parses, optionally in combination with other scores. Both the prosodic break recognition algorithm and the prosody/syntax model can be automatically trained and can therefore be designed specifically for different speaking styles or task domains, given appropriate labeled data. We have demonstrated the potential of this approach in experiments with a corpus of ambiguous sentences spoke... Nanette Veilleux, Mari Ostendorf, Colin W. Wightman |
ICSLP | 2 |
| 1991 | A dynamical system approach to continuous speech recognitionabstractAn dynamical system model is proposed for better representing the spectral dynamics of speech for recognition. It is assumed that the observed feature vectors of a phone segment are the output of a stochastic linear dynamical system, and two alternative assumptions regarding the relationship of the segment length and the evolution of the dynamics are considered. Training is equivalent to the identification of a stochastic linear system, and a nontraditional approach based on the estimate-maximize algorithm is followed. This model is evaluated on a phoneme classification task using the TIMIT database. It is shown that the classification performance obtained using the proposed model is significantly better than that obtained using either an independent-frame or a Gauss-Markov assumption on the observed frames.> Vassilios Digalakis, Jan Robin Rohlicek, Mari Ostendorf |
ICASSP | 3 |
| 1991 | Automatic recognition of prosodic phrasesabstractThe authors report on the development of two algorithms to automatically detect prosodic phrases. The first algorithm uses simple frame-based likelihood classifiers to detect breaths and silences, which yields a large percentage of major phrase breaks. To label other levels of prosodic structure, a second algorithm is introduced that uses phoneme durations given by a speech recognizer in conjunction with a tree quantizer and hidden Markov model to label a hierarchy of prosodic phrase breaks. This second algorithm yields phrase break predictions that have good correlation with hand labels, and correctly detects more than 90% of the major phrase boundaries.> Colin W. Wightman, Mari Ostendorf |
ICASSP | 2 |
| 1990 | Isolated word intonation recognition using hidden Markov modelsabstractA method is described for recognition of intonation patterns based on discrete distribution hidden Markov models (HMMs) and vector quantization techniques. Fundamental frequency and energy features, were used to determine the best combination of feature processing and quantization techniques for recognition of statement, question, command, calling, and continuation intonation patterns in isolated words. A recognition accuracy of 89% was achieved for the best-case speaker- and word-independent performance. Recognition performance of human listeners on a 100-word subset yielded 77% accuracy, compared to 83% using HMMs on the same subset.> John Butzberger, Mari Ostendorf, Patti Price, Stefanie Shattuck-Hufnagel |
ICASSP | 2 |
| 1990 | Joint quantizer design and parameter estimation for discrete hidden Markov modelsabstractAn approach that involves designing a vector quantizer to maximize the mutual information between the hidden Markov model (HMM) states and the quantized observations is presented. The iterative design of the quantizer and the HMM parameters is shown to be jointly a maximum-likelihood estimate. Methodologies for using the maximum mutual information (MMI) criterion for quantizer design are described, and some initial results are presented to demonstrate that the MMI criterion yields improved speech recognition performance.> Mari Ostendorf, Jan Robin Rohlicek |
ICASSP | 1 |
| 1990 | Markov modeling of prosodic phrase structureabstractA simple computational model for predicting phrase boundaries from text is described. The hierarchical structure of the model is based on linguistic theory, and the model itself is probabilistic. The probabilistic component is important for several reasons: it can capture the fact that the same sequence of words can be produced with a variety of prosodic contours (some more likely than others), it allows for automatic training of the model, and it provides a mechanism for combining a prosodic component with other knowledge sources. Results are presented for the model, trained and evaluated on a database of FM radio news.> Nanette Veilleux, Mari Ostendorf, Patti Price, Stefanie Shattuck-Hufnagel |
ICASSP | 2 |
| 1990 | The use of relative duration in syntactic disambiguation
Patti Price, Colin W. Wightman, Mari Ostendorf, John Bear |
ICSLP | 3 |
| 1988 | Stochastic segment modelling using the estimate-maximize algorithm [speech recognition]abstractA probabilistic model called the stochastic segment model is introduced that describes the statistical dependence of all the frames of a speech segment. The model uses a time-warping transformation to map the sequence of observed frames to the appropriate frames of the segment model. The joint density of the observed frames is then given by the joint density of the selected model frames. The automatic training and recognition algorithms are discussed and a few preliminary recognition results are presented.> Salim E. Roucos, Mari Ostendorf, Herbert Gish, Alan Derr |
ICASSP | 2 |