VLDB 2026 Research / reviewers in the wild / expert
Eric Fosler-Lussier
dblp:80/6326 · also Eric Fosler
· DBLP profile ↗
133ranked-venue papers
11as first author
20since 2021 · last 2026
0000-0001-8004-5169ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 95 · 9 first-author · 12 since 2021Artificial intelligence and machine learning · 76 · 4 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | VISTA: Verification In Sequential Turn-based AssessmentabstractHallucination-defined here as generated statements unsupported or contradicted by available evidence or conversational context-remains a major obstacle to using conversational AI systems in settings that demand factual reliability.Existing metrics evaluate isolated responses or treat unverifiable content as errors, limiting their use for multi-turn dialogue.We introduce VISTA (Verification In Sequential Turn-based Assessment), a framework for evaluating conversational factuality via claim-level verification and sequential consistency tracking.VISTA decomposes each turn into atomic claims, verifies them against trusted sources and dialogue history, and categorizes unverifiable statements (subjective, contradicted, lacking evidence, or abstaining).Across eight large language models and four dialogue factuality benchmarks (AIS, BEGIN, FAITHDIAL, and FADE), VISTA substantially improves hallucination detection over FActScore and LLM-as-Judge baselines.Human evaluation confirms that VISTA's decomposition improves annotator agreement and reveals inconsistencies in existing benchmarks.Further analyses show that incorporating dialogue context into verification substantially improves contradiction detection, and that VISTA reliably identifies abstentions.By modeling factuality as a dynamic property of conversation, VISTA offers a more transparent, human-aligned measure of truthfulness in dialogue systems. Ashley Lewis, Andrew Perrault, Eric Fosler-Lussier, Michael White 0001 |
ACL (1) | 3 |
| 2025 | A Non-autoregressive Model for Joint STT and TTSabstractIn this paper, we take a step towards jointly modeling automatic speech recognition (STT) and speech synthesis (TTS) in a fully non-autoregressive way. We develop a novel multimodal framework capable of handling the speech and text modalities as input either individually or together. The proposed model can also be trained with unpaired speech or text data owing to its multimodal nature. We further propose an iterative refinement strategy to improve the STT and TTS performance of our model such that the partial hypothesis at the output can be fed back to the input of our model, thus iteratively improving both STT and TTS predictions. We show that our joint model can effectively perform both STT and TTS tasks, outperforming the STT-specific baseline in all tasks and performing competitively with the TTS-specific baseline across a wide range of evaluation metrics. Vishal Sunder, Brian Kingsbury, George Saon, Samuel Thomas 0001, Slava Shechtman, Hagai Aronowitz, Eric Fosler-Lussier, Luis A. Lastras |
ICASSP | 7 |
| 2025 | End-to-End Diarization utilizing Attractor Deep Clustering
David Palzer, Matthew Maciejewski, Eric Fosler-Lussier |
INTERSPEECH | 3 |
| 2025 | The Ohio Child Speech CorpusabstractThis paper reports on the creation and composition of a new corpus of children's speech, the Ohio Child Speech Corpus, which is publicly available on the Talkbank-CHILDES website. The audio corpus contains speech samples from 303 children ranging in age from 4 – 9 years old, all of whom participated in a seven-task elicitation protocol conducted in a science museum lab. In addition, an interactive social robot controlled by the researchers joined the sessions for approximately 60% of the children, and the corpus itself was collected in the peri‑pandemic period. Two analyses are reported that highlighted these last two features. One set of analyses found that the children spoke significantly more in the presence of the robot relative to its absence, but no effects of speech complexity (as measured by MLU) were found for the robot's presence. Another set of analyses compared children tested immediately post-pandemic to children tested a year later on two school-readiness tasks, an Alphabet task and a Reading Passages task. This analysis showed no negative impact on these tasks for our highly-educated sample of children just coming off of the pandemic relative to those tested later. These analyses demonstrate just two possible types of questions that this corpus could be used to investigate. Sharifa Alghowinem, Abeer Alwan, Kristina Bowdrie, Cynthia Breazeal, Cynthia G. Clopper, Eric Fosler-Lussier, Izabela A. Jamsek, Devan Lander, Rajiv Ramnath, Jory Ross |
Speech Commun. | 7 |
| 2024 | Improving Neural Diarization through Speaker Attribute Attractors and Local Dependency ModelingabstractIn recent years, end-to-end approaches have made notable progress in addressing the challenge of speaker diarization, which involves segmenting and identifying speakers in multi-talker recordings. One such approach, Encoder-Decoder Attractors (EDA), has been proposed to handle variable speaker counts as well as better guide the network during training. In this study, we extend the attractor paradigm by moving beyond direct speaker modeling and instead focus on representing more detailed ‘speaker attributes’ through a multi-stage process of intermediate representations. Additionally, we enhance the architecture by replacing transformers with conformers, a convolution-augmented transformer, to model local dependencies. Experiments demonstrate improved diarization performance on the CALLHOME dataset. David Palzer, Matthew Maciejewski, Eric Fosler-Lussier |
ICASSP | 3 |
| 2024 | End-To-End Real Time Tracking of Children's Reading with Pointer NetworkabstractIn this work, we explore how a real time reading tracker can be built efficiently for children’s voices. While previously proposed reading trackers focused on ASR-based cascaded approaches, we propose a fully end-to-end model making it less prone to lags in voice tracking. We employ a pointer network that directly learns to predict positions in the ground truth text conditioned on the streaming speech. To train this pointer network, we generate ground truth training signals by using forced alignment between the read speech and the text being read on the training set. Exploring different forced alignment models, we find a neural attention based model is at least as close in alignment accuracy to the Montreal Forced Aligner, but surprisingly is a better training signal for the pointer network. Our results are reported on one adult speech data (TIMIT) and two children’s speech datasets (CMU Kids and Reading Races). Our best model can accurately track adult speech with 87.8% accuracy and the much harder and disfluent children’s speech with 77.1% accuracy on CMU Kids data and a 65.3% accuracy on the Reading Races dataset. Vishal Sunder, Beulah Karrolla, Eric Fosler-Lussier |
ICASSP | 3 |
| 2024 | Improving Transducer-Based Spoken Language Understanding With Self-Conditioned CTC and Knowledge TransferabstractIn this paper, we propose to improve end-to-end (E2E) spoken language understand (SLU) in an RNN transducer model (RNN-T) by incorporating a joint self-conditioned CTC automatic speech recognition (ASR) objective. Our proposed model is akin to an E2E differentiable cascaded model which performs ASR and SLU sequentially and we ensure that the SLU task is conditioned on the ASR task by having CTC self conditioning. This novel joint modeling of ASR and SLU improves SLU performance significantly over just using SLU optimization. We further improve the performance by aligning the acoustic embeddings of this model with the semantically richer BERT model. Our proposed knowledge transfer strategy makes use of a bag-of-entity prediction layer on the aligned embeddings and the output of this is used to condition the RNN-T based SLU decoding. These techniques show significant improvement over several strong baselines and can perform at par with large models like Whisper with significantly fewer parameters. Vishal Sunder, Eric Fosler-Lussier |
SLT | 2 |
| 2024 | A randomized prospective study of a hybrid rule- and data-driven virtual patientabstractAbstract Randomized prospective studies represent the gold standard for experimental design. In this paper, we present a randomized prospective study to validate the benefits of combining rule-based and data-driven natural language understanding methods in a virtual patient dialogue system. The system uses a rule-based pattern matching approach together with a machine learning (ML) approach in the form of a text-based convolutional neural network, combining the two methods with a simple logistic regression model to choose between their predictions for each dialogue turn. In an earlier, retrospective study, the hybrid system yielded a nearly 50% error reduction on our initial data, in part due to the differential performance between the two methods as a function of label frequency. Given these gains, and considering that our hybrid approach is unique among virtual patient systems, we compare the hybrid system to the rule-based system by itself in a randomized prospective study. We evaluate 110 unique medical student subjects interacting with the system over 5,296 conversation turns, to verify whether similar gains are observed in a deployed system. This prospective study broadly confirms the findings from the earlier one but also highlights important deficits in our training data. The hybrid approach still improves over either rule-based or ML approaches individually, even handling unseen classes with some success. However, we observe that live subjects ask more out-of-scope questions than expected. To better handle such questions, we investigate several modifications to the system combination component. These show significant overall accuracy improvements and modest F1 improvements on out-of-scope queries in an offline evaluation. We provide further analysis to characterize the difficulty of the out-of-scope problem that we have identified, as well as to suggest future improvements over the baseline we establish here. Adam Stiff, Michael White 0001, Eric Fosler-Lussier, Lifeng Jin, Evan Jaffe, Douglas Danforth |
Nat. Lang. Eng. | 3 |
| 2023 | Fine-Grained Textual Knowledge Transfer to Improve RNN Transducers for Speech Recognition and UnderstandingabstractRNN Tranducer (RNN-T) technology is very popular for building deployable models for end-to-end (E2E) automatic speech recognition (ASR) and spoken language understanding (SLU). Since these are E2E models operating on speech directly, there remains a potential to improve their performance using purely text based models like BERT, which have strong language understanding capabilities. In this paper, we propose a new training criteria for RNN-T based E2E ASR and SLU to transfer BERT’s knowledge into these systems. In the first stage of our proposed mechanism, we improve ASR performance by using a fine-grained, tokenwise knowledge transfer from BERT. In the second stage, we fine-tune the ASR model for SLU such that the above knowledge is explicitly utilized by the RNN-T model for improved performance. Our techniques improve ASR performance on the Switchboard and CallHome test sets of the NIST Hub5 2000 evaluation and on the recently released SLURP dataset on which we achieve a new state-of-the-art performance. For SLU, we show significant improvements on the SLURP slot filling task, outperforming HuBERT-base and reaching a performance close to HuBERTlarge. Compared to large transformer based speech models like HuBERT, our model is significantly more compact and uses only 300 hours of speech pretraining data. Vishal Sunder, Samuel Thomas 0001, Hong-Kwang Jeff Kuo, Brian Kingsbury, Eric Fosler-Lussier |
ICASSP | 5 |
| 2023 | End-to-End Word-Level Disfluency Detection and Classification in Children's Reading AssessmentabstractDisfluency detection and classification on children’s speech has a great potential for teaching reading skills. Word-level assessment of children’s speech can help teachers to effectively gauge their students’ progress. Hence, we propose a novel attention-based model to perform word-level disfluency detection and classification in a fully end-to-end (E2E) manner making it fast and easy to use. We develop a word-level disfluency annotation scheme using which we annotate a dataset of children read speech, the reading races dataset (READR). We also annotate disfluencies in the existing CMU Kids corpus. The proposed model significantly outperforms traditional cascaded baselines, which use forced alignments, on both datasets. To deal with the inevitable class-imbalance in the datasets, we propose a novel technique called HiDeC (Hierarchical Detection and Classification) which yields a detection improvement of 23% and 16% and a classification improvement of 3.8% and 19.3% relative F1-score on the READR and CMU Kids datasets respectively. Lavanya Venkatasubramaniam, Vishal Sunder, Eric Fosler-Lussier |
ICASSP | 3 |
| 2023 | ConvKT: Conversation-Level Knowledge Transfer for Context Aware End-to-End Spoken Language Understanding
Vishal Sunder, Eric Fosler-Lussier, Samuel Thomas 0001, Hong-Kwang Jeff Kuo, Brian Kingsbury |
INTERSPEECH | 2 |
| 2023 | Bootstrapping a Conversational Guide for Colonoscopy PrepabstractPulkit Arya, Madeleine Bloomquist, Subhankar Chakraborty, Andrew Perrault, William Schuler, Eric Fosler-Lussier, Michael White. Proceedings of the 24th Meeting of the Special Interest Group on Discourse and Dialogue. 2023. Pulkit Arya, Madeleine Bloomquist, Subhankar Chakraborty, Andrew Perrault, William Schuler, Eric Fosler-Lussier, Michael White 0001 |
SIGDIAL | 6 |
| 2023 | Guest editorial: Special issue on advances in deep learning based speech processing
Eric Fosler-Lussier, Emmanuel Vincent 0001 |
Neural Networks | 3 |
| 2022 | Towards End-to-End Integration of Dialog History for Improved Spoken Language UnderstandingabstractDialog history plays an important role in spoken language understanding (SLU) performance in a dialog system. For end-to-end (E2E) SLU, previous work has used dialog history in text form, which makes the model dependent on a cascaded automatic speech recognizer (ASR). This rescinds the benefits of an E2E system which is intended to be compact and robust to ASR errors. In this paper, we propose a hierarchical conversation model that is capable of directly using dialog history in speech form, making it fully E2E. We also distill semantic knowledge from the available gold conversation transcripts by jointly training a similar text-based conversation model with an explicit tying of acoustic and semantic embeddings. We also propose a novel technique that we call DropFrame to deal with the long training time incurred by adding dialog history in an E2E manner. On the HarperValleyBank dialog dataset, our E2E history integration outperforms a history independent baseline by 7.7% absolute F1 score on the task of dialog action recognition. Our model performs competitively with the state-of-the-art history based cascaded baseline, but uses 48% fewer parameters. In the absence of gold transcripts to fine-tune an ASR model, our model outperforms this baseline by a significant margin of 10% absolute F1 score. Vishal Sunder, Samuel Thomas 0001, Hong-Kwang Jeff Kuo, Jatin Ganhotra, Brian Kingsbury, Eric Fosler-Lussier |
ICASSP | 6 |
| 2022 | Tokenwise Contrastive Pretraining for Finer Speech-to-BERT Alignment in End-to-End Speech-to-Intent SystemsabstractRecent advances in End-to-End (E2E) Spoken Language Understanding (SLU) have been primarily due to effective pretraining of speech representations.One such pretraining paradigm is the distillation of semantic knowledge from state-of-the-art text-based models like BERT to speech encoder neural networks.This work is a step towards doing the same in a much more efficient and fine-grained manner where we align speech embeddings and BERT embeddings on a token-by-token basis.We introduce a simple yet novel technique that uses a cross-modal attention mechanism to extract token-level contextual embeddings from a speech encoder such that these can be directly compared and aligned with BERT based contextual embeddings.This alignment is performed using a novel tokenwise contrastive loss.Fine-tuning such a pretrained model to perform intent recognition using speech directly yields state-of-the-art performance on two widely used SLU datasets.Our model improves further when fine-tuned with additional regularization using SpecAugment especially when speech is noisy, giving an absolute improvement as high as 8% over previous results. Vishal Sunder, Eric Fosler-Lussier, Samuel Thomas 0001, Hong-Kwang Jeff Kuo, Brian Kingsbury |
INTERSPEECH | 2 |
| 2022 | Hallucination of Speech Recognition Errors With Sequence to Sequence LearningabstractPrior work in this domain has focused on modeling errors at the phonetic level, while using a lexicon to convert the phones to words, usually accompanied by an FST Language model. We present novel end-to-end models to directly predict hallucinated ASR word sequence outputs, conditioning on an input word sequence as well as a corresponding phoneme sequence. This improves prior published results for recall of errors from an in-domain ASR system’s transcription of unseen data, as well as an out-of-domain ASR system’s transcriptions of audio from an unrelated task, while additionally exploring an in-between scenario when limited characterization data from the test ASR system is obtainable. To verify the extrinsic validity of the method, we also use our hallucinated ASR errors to augment training for a spoken question classifier, finding that they enable robustness to real ASR errors in a downstream task, when scarce or even zero task-specific audio was available at train-time. Prashant Serai, Vishal Sunder, Eric Fosler-Lussier |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2021 | Learning Latent Structures for Cross Action Phrase Relations in Wet Lab ProtocolsabstractChaitanya Kulkarni, Jany Chan, Eric Fosler-Lussier, Raghu Machiraju. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Chaitanya Kulkarni, Jany Chan, Eric Fosler-Lussier, Raghu Machiraju |
ACL/IJCNLP (1) | 3 |
| 2021 | Diachronic Analysis of the Evolution of COVID-19 Scientific Literature
Denis Newman-Griffis, Venkatesh Sivaraman, Adam Perer, Eric Fosler-Lussier, Harry Hochheiser |
AMIA | 4 |
| 2021 | Handling Class Imbalance in Low-Resource Dialogue Systems by Combining Few-Shot Classification and InterpolationabstractUtterance classification performance in low-resource dialogue systems is constrained by an inevitably high degree of data imbalance in class labels. We present a new end-to-end pairwise learning framework that is designed specifically to tackle this phenomenon by inducing a few-shot classification capability in the utterance representations and augmenting data through an interpolation of utterance representations. Our approach is a general purpose training methodology, agnostic to the neural architecture used for encoding utterances. We show significant improvements in macro-F1 score over standard cross-entropy training for three different neural architectures, demonstrating improvements on a Virtual Patient dialogue dataset as well as a low-resourced emulation of the Switchboard dialogue act classification dataset. Vishal Sunder, Eric Fosler-Lussier |
ICASSP | 2 |
| 2021 | Ambiguity in medical concept normalization: An analysis of types and coverage in electronic health record datasetsabstractOBJECTIVES: Normalizing mentions of medical concepts to standardized vocabularies is a fundamental component of clinical text analysis. Ambiguity-words or phrases that may refer to different concepts-has been extensively researched as part of information extraction from biomedical literature, but less is known about the types and frequency of ambiguity in clinical text. This study characterizes the distribution and distinct types of ambiguity exhibited by benchmark clinical concept normalization datasets, in order to identify directions for advancing medical concept normalization research. MATERIALS AND METHODS: We identified ambiguous strings in datasets derived from the 2 available clinical corpora for concept normalization and categorized the distinct types of ambiguity they exhibited. We then compared observed string ambiguity in the datasets with potential ambiguity in the Unified Medical Language System (UMLS) to assess how representative available datasets are of ambiguity in clinical language. RESULTS: We found that <15% of strings were ambiguous within the datasets, while over 50% were ambiguous in the UMLS, indicating only partial coverage of clinical ambiguity. The percentage of strings in common between any pair of datasets ranged from 2% to only 36%; of these, 40% were annotated with different sets of concepts, severely limiting generalization. Finally, we observed 12 distinct types of ambiguity, distributed unequally across the available datasets, reflecting diverse linguistic and medical phenomena. DISCUSSION: Existing datasets are not sufficient to cover the diversity of clinical concept ambiguity, limiting both training and evaluation of normalization methods for clinical text. Additionally, the UMLS offers important semantic information for building and evaluating normalization methods. CONCLUSIONS: Our findings identify 3 opportunities for concept normalization research, including a need for ambiguity-specific clinical datasets and leveraging the rich semantics of the UMLS in new methods and evaluation measures for normalization. Denis Newman-Griffis, Guy Divita, Bart Desmet, Ayah Zirikly, Carolyn P. Rosé, Eric Fosler-Lussier |
J. Am. Medical Informatics Assoc. | 6 |
| 2020 | Ambiguity in Medical Concept Normalization: An Analysis of Types and Coverage in Electronic Health Record Datasets
Denis Newman-Griffis, Guy Divita, Bart Desmet, Ayah Zirikly, Carolyn P. Rosé, Eric Fosler-Lussier |
AMIA | 6 |
| 2020 | Contextualized Embeddings for Enriching Linguistic Analyses on PolitenessabstractLinguistic analyses in natural language processing (NLP) have often been performed around the static notion of words where the context (surrounding words) is not considered.For example, previous analyses on politeness have focused on comparing the use of static words such as personal pronouns across (im)polite requests without taking the context of those words into account.Current word embeddings in NLP do capture context and thus can be leveraged to enrich linguistic analyses.In this work, we introduce a model which leverages the pre-trained BERT model to cluster contextualized representations of a word based on (1) the context in which the word appears and (2) the labels of items the word occurs in.Using politeness as case study, this model is able to automatically discover interpretable, fine-grained context patterns of words, some of which align with existing theories on politeness.Our model further discovers novel finer-grained patterns associated with (im)polite language.For example, the word please can occur in impolite contexts that are predictable from BERT clustering.The approach proposed here is validated by showing that features based on fine-grained patterns inferred from the clustering improve over politeness-word baselines. Ahmad Aljanaideh, Eric Fosler-Lussier, Marie-Catherine de Marneffe |
COLING | 2 |
| 2020 | Phonetic Feedback for Speech Enhancement with and Without Parallel Speech DataabstractWhile deep learning systems have gained significant ground in speech enhancement research, these systems have yet to make use of the full potential of deep learning systems to provide high-level feedback. In particular, phonetic feedback is rare in speech enhancement research even though it includes valuable top-down information. We use the technique of mimic loss to provide phonetic feedback to an off-the-shelf enhancement system, and find gains in objective intelligibility scores on CHiME-4 data. This technique takes a frozen acoustic model trained on clean speech to provide valuable feedback to the enhancement model, even in the case where no parallel speech data is available. Our work is one of the first to show intelligibility improvement for neural enhancement systems without parallel speech data, and we show phonetic feedback can improve a state-of-the-art neural enhancement system trained with parallel speech data. Peter Plantinga, Deblin Bagchi, Eric Fosler-Lussier |
ICASSP | 3 |
| 2020 | End to End Speech Recognition Error Prediction with Sequence to Sequence LearningabstractSimulating the errors made by a speech recognizer on plain text has proven useful to help train downstream NLP tasks to be robust to real ASR errors at test time. Prior work in this domain has focused on modeling confusions at the phonetic level, and using a lexicon to convert from words to phones and back, usually accompanied by an FST Language model. We present a novel end to end model to simulate ASR errors. Our approach trains a convolutional sequence to sequence model to take as direct input a word sequence and predict a word sequence as an output. The end to end modeling improves prior published results for recall of recognition errors made by a Switchboard ASR system on unseen Fisher data; we also demonstrate cross-domain robustness by predicting errors made by an unrelated cloud-based ASR system on a Virtual Patient task. Prashant Serai, Adam Stiff, Eric Fosler-Lussier |
ICASSP | 3 |
| 2020 | How Self-Attention Improves Rare Class Performance in a Question-Answering Dialogue AgentabstractContextualized language modeling using deep Transformer networks has been applied to a variety of natural language processing tasks with remarkable success.However, we find that these models are not a panacea for a questionanswering dialogue agent corpus task, which has hundreds of classes in a long-tailed frequency distribution, with only thousands of data points.Instead, we find substantial improvements in recall and accuracy on rare classes from a simple one-layer RNN with multi-headed self-attention and static word embeddings as inputs.While much research has used attention weights to illustrate what input is important for a task, the complexities of our dialogue corpus offer a unique opportunity to examine how the model represents what it attends to, and we offer a detailed analysis of how that contributes to improved performance on rare classes.A particularly interesting phenomenon we observe is that the model picks up implicit meanings by splitting different aspects of the semantics of a single word across multiple attention heads. Adam Stiff, Eric Fosler-Lussier |
SIGdial | 3 |
| 2019 | Automated classification of mobility activities in free text clinical narratives
Denis Newman-Griffis, Ayah Zirikly, Pei-Shu Ho, Jonathan Camacho, Maryanne Sacco, Alex Marr, Albert M. Lai, Eric Fosler-Lussier |
AMIA | 8 |
| 2019 | Towards Real-Time Mispronunciation Detection in Kids' SpeechabstractModern mispronunciation detection and diagnosis systems have seen significant gains in accuracy due to the introduction of deep learning. However, these systems have not been evaluated for the ability to be run in real-time, an important factor in applications that provide rapid feedback. In particular, the state-of-the-art uses bi-directional recurrent networks, where a uni-directional network may be more appropriate. Teacher-student learning is a natural approach to use to improve a uni-directional model, but when using a CTC objective, this is limited by poor alignment of outputs to evidence. We address this limitation by trying two loss terms for improving the alignments of our models. One loss is an “alignment loss” term that encourages outputs only when features do not resemble silence. The other loss term uses a uni-directional model as teacher model to align the bi-directional model. Our proposed model uses these aligned bi-directional models as teacher models. Experiments on the CSLU kids' corpus show that these changes decrease the latency of the outputs, and improve the detection rates, with a trade-off between these goals. Peter Plantinga, Eric Fosler-Lussier |
ASRU | 2 |
| 2019 | Improving Speech Recognition Error Prediction for Modern and Off-the-shelf Speech RecognizersabstractModeling the errors of a speech recognizer can help simulate errorful recognized speech data from plain text, which has proven useful for tasks like discriminative language modeling, improving robustness of NLP systems, where limited or even no audio data is available at train time. Previous work typically considered replicating behavior of GMM-HMM based systems, but the behavior of more modern posterior-based neural network acoustic models is not the same and requires adjustments to the error prediction model. In this work, we extend a prior phonetic confusion based model for predicting speech recognition errors in two ways: first, we introduce a sampling-based paradigm that better simulates the behavior of a posterior-based acoustic model. Second, we investigate replacing the confusion matrix with a sequence-to-sequence model in order to introduce context dependency into the prediction. We evaluate the error predictors in two ways: first by predicting the errors made by a Switchboard ASR system on unseen data (Fisher), and then using that same predictor to estimate the behavior of an unrelated cloud-based ASR system on a novel task. Sampling greatly improves predictive accuracy within a 100-guess paradigm, while the sequence model performs similarly to the confusion matrix. Prashant Serai, Eric Fosler-Lussier |
ICASSP | 3 |
| 2019 | Improving Human-computer Interaction in Low-resource Settings with Text-to-phonetic Data AugmentationabstractOff-the-shelf speech recognition systems can yield useful results and accelerate application development, but general-purpose systems applied to specialized domains can introduce acoustically small-but semantically catastrophic-errors. Furthermore, sufficient audio data may not be available to develop custom acoustic models for niche tasks. To address these problems, we propose a concept to improve performance in text classification tasks that use speech transcripts as input, without any in-domain audio data. Our method augments available typewritten text training data with inferred phonetic information so that the classifier will learn semantically important acoustic regularities, making it more robust to transcription errors from the general purpose ASR. We successfully pilot our method in a speech-based virtual patient used for medical training, recovering up to 62% of errors incurred by feeding a small test set of speech transcripts to a classification model trained on typescript. Adam Stiff, Prashant Serai, Eric Fosler-Lussier |
ICASSP | 3 |
| 2019 | Spatial and Channel Attention Based Convolutional Neural Networks for Modeling Noisy SpeechabstractIn recent years, Residual Networks (ResNets) have significantly increased the modeling power of convolutional neural networks (CNNs) by introducing residual connections. In this paper, we explore the incorporation of spatial and channel attention into the structure of ResNets for noisy speech recognition tasks. In our experiments, we implemented spatial attention as a bottom-up top-down structure where the input features are first down sampled and then up sampled to generate attention maps. At each block of the ResNet, the generated CNN features are composed with spatial attention maps over the temporal-frequency space, learning to attend to salient acoustic features and suppress noise. Our model also includes channel attention that attends to different channels of feature maps. ResNet blocks with spatial and channel attention modules can be easily stacked to construct deeper networks. We show that the proposed network structure has the ability to suppress noisy signals in speech audio without requiring parallel clean speech for training, and achieve promising WER reductions on CHiME2 and CHiME3. Sirui Xu 0001, Eric Fosler-Lussier |
ICASSP | 2 |
| 2018 | Spectral Feature Mapping with MIMIC Loss for Robust Speech RecognitionabstractFor the task of speech enhancement, local learning objectives are agnostic to phonetic structures helpful for speech recognition. We propose to add a global criterion to ensure de-noised speech is useful for downstream tasks like ASR. We first train a spectral classifier on clean speech to predict senone labels. Then, the spectral classifier is joined with our speech enhancer as a noisy speech recognizer. This model is taught to imitate the output of the spectral classifier alone on clean speech. This mimic loss is combined with the traditional local criterion to train the speech enhancer to produce de-noised speech. Feeding the de-noised speech to an off-the-shelf Kaldi training recipe for the CHiME- 2 corpus shows significant improvements in WER. Deblin Bagchi, Peter Plantinga, Adam Stiff, Eric Fosler-Lussier |
ICASSP | 4 |
| 2018 | Application of Progressive Neural Networks for Multi-Stream Wfst Combination in One-Pass DecodingabstractMany state-of-the-art automatic speech recognition (ASR) systems adopt system combination techniques to improve recognition performance. In this paper, we investigate the possibility of transferring knowledge between models for different noisy speech domains and integrating these models via system combination. The first contribution of our work is the use of progressive neural networks for modeling the acoustic features of noisy speech. We train progressive neural networks on subdivided noisy data to achieve knowledge transfer between different noise conditions. Our second contribution is an improved multi-stream WFST framework that combines the output of the progressive networks at longer timescales (e.g., word hypotheses). The score fusion is performed by a trained LSTM at the word boundary on the decoding lattice. By adopting both knowledge transfer and system combination techniques, we achieve improved performance compared with independently trained deep neural networks. Sirui Xu 0001, Eric Fosler-Lussier |
ICASSP | 2 |
| 2018 | An Exploration of Mimic Architectures for Residual Network Based Spectral MappingabstractSpectral mapping uses a deep neural network (DNN) to map directly from noisy speech to clean speech. Our previous study [1] found that the performance of spectral mapping improves greatly when using helpful cues from an acoustic model trained on clean speech. The mapper network learns to mimic the input favored by the spectral classifier and cleans the features accordingly. In this study, we explore two new innovations: we replace a DNN-based spectral mapper with a residual network that is more attuned to the goal of predicting clean speech. We also examine how integrating long term context in the mimic criterion (via wide-residual biL-STM networks) affects the performance of spectral mapping compared to DNNs. Our goal is to derive a model that can be used as a preprocessor for any recognition system; the features derived from our model are passed through the standard Kaldi ASR pipeline and achieve a WER of 9.3%, which is the lowest recorded word error rate for CHiME-2 dataset using only feature adaptation. Peter Plantinga, Deblin Bagchi, Eric Fosler-Lussier |
SLT | 3 |
| 2017 | Cross-Lingual Transfer Learning for POS Tagging without Cross-Lingual ResourcesabstractTraining a POS tagging model with crosslingual transfer learning usually requires linguistic knowledge and resources about the relation between the source language and the target language.In this paper, we introduce a cross-lingual transfer learning model for POS tagging without ancillary resources such as parallel corpora.The proposed cross-lingual model utilizes a common BLSTM that enables knowledge transfer from other languages, and private BLSTMs for language-specific representations.The cross-lingual model is trained with language-adversarial training and bidirectional language modeling as auxiliary objectives to better represent language-general information while not losing the information about a specific target language.Evaluating on POS datasets from 14 languages in the Universal Dependencies corpus, we show that the proposed transfer learning model improves the POS tagging performance of the target languages without exploiting any linguistic knowledge between the source language and the target language. Joo-Kyung Kim, Young-Bum Kim, Ruhi Sarikaya, Eric Fosler-Lussier |
EMNLP | 4 |
| 2016 | Automatic data source identification for clinical trial eligibility criteria resolution
Chaitanya P. Shivade, Courtney Hebert, Kelly Regan-Fendt, Eric Fosler-Lussier, Albert M. Lai |
AMIA | 4 |
| 2016 | Experiences with Shared Resources for Research and Education in Speech and Language Processingabstract\n Contains fulltext :\n 161889.pdf (Publisher’s version ) (Open Access)\n Rebecca Bates 0001, Eric Fosler-Lussier, Florian Metze, Martha A. Larson, Gina-Anne Levow, Emily Mower Provost |
INTERSPEECH | 2 |
| 2016 | A WFST Framework for Single-Pass Multi-Stream Decoding
Sirui Xu 0001, Eric Fosler-Lussier |
INTERSPEECH | 2 |
| 2016 | Articulatory feature-based pronunciation modeling
Karen Livescu, Preethi Jyothi, Eric Fosler-Lussier |
Comput. Speech Lang. | 3 |
| 2016 | Speech Production in Speech Technologies: Introduction to the CSL Special Issue
Karen Livescu, Frank Rudzicz, Eric Fosler-Lussier, Mark Hasegawa-Johnson, Jeff A. Bilmes |
Comput. Speech Lang. | 3 |
| 2016 | Using Pronunciation-Based Morphological Subword Units to Improve OOV Handling in Keyword SearchabstractOut-of-vocabulary (OOV) keywords present a challenge for keyword search (KWS) systems especially in the low-resource setting. Previous research has centered around approaches that use a variety of subword units to recover OOV words. This paper systematically investigates morphology-based subword modeling approaches on seven low-resource languages. We show that using morphological subword units (morphs) in speech recognition decoding is substantially better than expanding word-decoded lattices into subword units including phones, syllables and morphs. As alternatives to grapheme-based morphs, we apply unsupervised morphology learning to sequences of phonemes, graphones, and syllables. Using one of these phone-based morphs is almost always better than using the grapheme-based morphs, but the particular choice varies with the language. By combining the different methods, a substantial gain is obtained over the best single case for all languages, especially for OOV performance. Yanzhang He, Peter Baumann 0003, Hao Fang 0002, Brian Hutchinson, Aaron Jaech, Mari Ostendorf, Eric Fosler-Lussier, Janet B. Pierrehumbert |
IEEE ACM Trans. Audio Speech Lang. Process. | 7 |
| 2015 | Speech Adaptation in Extended Ambient Intelligence EnvironmentsabstractThis Blue Sky presentation focuses on a major shift toward a notion of “ambient intelligence” that transcends general applications targeted at the general population. The focus is on highly personalized agents that accommodate individual differences and changes over time. This notion of Extended Ambient Intelligence (EAI) concerns adaptation to a person’s preferences and experiences, as well as changing capabilities, most notably in an environment where conversational engagement is central. An important step in moving this research forward is the accommodation of different degrees of cognitive capability (including speech processing) that may vary over time for a given user—whether through improvement or through deterioration. We suggest that the application of divergence detection to speech patterns may enable adaptation to a speaker’s increasing or decreasing level of speech impairment over time. Taking an adaptive approach toward technology development in this arena may be a first step toward empowering those with special needs so that they may live with a high quality of life. It also represents an important step toward a notion of ambient intelligence that is personalized beyond what can be achieved by mass-produced, one-size-fits-all software currently in use on mobile devices. Bonnie J. Dorr, Lucian Galescu, Ian Perera, Kristy Hollingshead, David J. Atkinson 0001, Micah Clark, William J. Clancey, Yorick Wilks, Eric Fosler-Lussier |
AAAI | 9 |
| 2015 | Combining spectral feature mapping and multi-channel model-based source separation for noise-robust automatic speech recognitionabstractAutomatic Speech Recognition systems suffer from severe performance degradation in the presence of myriad complicating factors such as noise, reverberation, multiple speech sources, multiple recording devices, etc. Previous challenges have sparked much innovation when it comes to designing systems capable of handling these complications. In this spirit, the CHiME-3 challenge presents system builders with the task of recognizing speech in a real-world noisy setting wherein speakers talk to an array of 6 microphones in a tablet. In order to address these issues, we explore the effectiveness of first applying a model-based source separation mask to the output of a beamformer that combines the source signals recorded by each microphone, followed by a DNN-based front end spectral mapper that predicts clean filterbank features. The source separation algorithm MESSL (Model-based EM Source Separation and Localization) has been extended from two channels to multiple channels in order to meet the demands of the challenge. We report on interactions between the two systems, cross-cut by the use of a robust beamforming algorithm called BeamformIt. Evaluations of different system settings reveal that combining MESSL and the spectral mapper together on the baseline beamformer algorithm boosts the performance substantially. Deblin Bagchi, Michael I. Mandel, Zhongqiu Wang 0001, Yanzhang He, Andrew R. Plummer, Eric Fosler-Lussier |
ASRU | 6 |
| 2015 | Knowledge Graph Inference for spoken dialog systemsabstractWe propose Inference Knowledge Graph, a novel approach of remapping existing, large scale, semantic knowledge graphs into Markov Random Fields in order to create user goal tracking models that could form part of a spoken dialog system. Since semantic knowledge graphs include both entities and their attributes, the proposed method merges the semantic dialog-state-tracking of attributes and the database lookup of entities that fulfill users' requests into one single unified step. Using a large semantic graph that contains all businesses in Bellevue, WA, extracted from Microsoft Satori, we demonstrate that the proposed approach can return significantly more relevant entities to the user than a baseline system using database lookup. Paul A. Crook, Ruhi Sarikaya, Eric Fosler-Lussier |
ICASSP | 4 |
| 2015 | Deep neural network based spectral feature mapping for robust speech recognitionabstractAutomatic speech recognition (ASR) systems suffer from performance degradation under noisy and reverberant conditions. In this work, we explore a deep neural network (DNN) based approach for spectral feature mapping from corrupted speech to clean speech. The DNN based mapping substantially reduces interference and produces estimated clean spectral features for ASR training and decoding. We experiment with several different feature mapping approaches and demonstrate that a DNN trained to predict clean log filterbank coefficients from noisy spectrogram directly can be extremely effective. The experiments show that the ASR systems with these cleaned features perform well under joint noisy and reverberant conditions, and achieve the state-of-the-art results on the CHiME-2 corpus with stereo (corrupted and clean) data. Yanzhang He, Deblin Bagchi, Eric Fosler-Lussier, DeLiang Wang |
INTERSPEECH | 4 |
| 2015 | Segmental conditional random fields with deep neural networks as acoustic models for first-pass word recognitionabstractDiscriminative segmental models, such as segmental conditional random fields (SCRFs), have been successfully applied to speech recognition recently in lattice rescoring to integrate detectors across different levels of units, such as phones and words. However, the lattice generation has been constrained by a baseline decoder, typically a frame-based hybrid HMMDNN system, which still suffers from the well-known frame independent assumption. In this paper, we propose to use SCRFs with DNNs directly as the acoustic model, a one-pass unified framework that can utilize local phone classifiers, phone transitions and long-span features, in direct word decoding to model phones or sub-phonetic segments with variable length. We describe a WFST-based approach to utilize the proposed acoustic model efficiently with the language model in first-pass word recognition. Our evaluation on the WSJ corpus shows our SCRF-DNN system outperforms a hybrid HMM-DNN system and a frame-level CRF-DNN system using the same label space. Yanzhang He, Eric Fosler-Lussier |
INTERSPEECH | 2 |
| 2015 | The speech recognition virtual kitchen turns one
Florian Metze, Eric Riebling, Eric Fosler-Lussier, Andrew R. Plummer, Rebecca Bates 0001 |
INTERSPEECH | 3 |
| 2015 | Corpus-based discovery of semantic intensity scalesabstractChaitanya Shivade, Marie-Catherine de Marneffe, Eric Fosler-Lussier, Albert M. Lai. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015. Chaitanya P. Shivade, Marie-Catherine de Marneffe, Eric Fosler-Lussier, Albert M. Lai |
HLT-NAACL | 3 |
| 2015 | Textual inference for eligibility criteria resolution in clinical trialsabstractClinical trials are essential for determining whether new interventions are effective. In order to determine the eligibility of patients to enroll into these trials, clinical trial coordinators often perform a manual review of clinical notes in the electronic health record of patients. This is a very time-consuming and exhausting task. Efforts in this process can be expedited if these coordinators are directed toward specific parts of the text that are relevant for eligibility determination. In this study, we describe the creation of a dataset that can be used to evaluate automated methods capable of identifying sentences in a note that are relevant for screening a patient's eligibility in clinical trials. Using this dataset, we also present results for four simple methods in natural language processing that can be used to automate this task. We found that this is a challenging task (maximum F-score=26.25), but it is a promising direction for further research. Chaitanya P. Shivade, Courtney Hebert, Marcelo A. Lopetegui, Marie-Catherine de Marneffe, Eric Fosler-Lussier, Albert M. Lai |
J. Biomed. Informatics | 5 |
| 2015 | Comparison of UMLS terminologies to identify risk of heart disease using clinical notesabstractThe second track of the 2014 i2b2 challenge asked participants to automatically identify risk factors for heart disease among diabetic patients using natural language processing techniques for clinical notes. This paper describes a rule-based system developed using a combination of regular expressions, concepts from the Unified Medical Language System (UMLS), and freely-available resources from the community. With a performance (F1=90.7) that is significantly higher than the median (F1=87.20) and close to the top performing system (F1=92.8), it was the best rule-based system of all the submissions in the challenge. We also used this system to evaluate the utility of different terminologies in the UMLS towards the challenge task. Of the 155 terminologies in the UMLS, 129 (76.78%) have no representation in the corpus. The Consumer Health Vocabulary had very good coverage of relevant concepts and was the most useful terminology for the challenge task. While segmenting notes into sections and lists has a significant impact on the performance, identifying negations and experiencer of the medical event results in negligible gain. Chaitanya P. Shivade, Pranav Malewadkar, Eric Fosler-Lussier, Albert M. Lai |
J. Biomed. Informatics | 3 |
| 2014 | Cross-narrative Temporal Ordering of Medical EventsabstractCross-narrative temporal ordering of medical events is essential to the task of generating a comprehensive timeline over a patient's history.We address the problem of aligning multiple medical event sequences, corresponding to different clinical narratives, comparing the following approaches: (1) A novel weighted finite state transducer representation of medical event sequences that enables composition and search for decoding, and (2) Dynamic programming with iterative pairwise alignment of multiple sequences using global and local alignment algorithms.The cross-narrative coreference and temporal relation weights used in both these approaches are learned from a corpus of clinical narratives.We present results using both approaches and observe that the finite state transducer approach performs performs significantly better than the dynamic programming one by 6.8% for the problem of multiple-sequence alignment. Preethi Raghavan, Eric Fosler-Lussier, Noémie Elhadad, Albert M. Lai |
ACL (1) | 2 |
| 2014 | Subword-based modeling for handling OOV words inkeyword spottingabstractThis work compares ASR decoding at different subword levels crossed with alternative keyword search strategies to handle the OOV issue for keyword spotting in the low-resource setting. We show that a morpheme-based subword modeling approach is effective in recovering OOV keywords within a Turkish low-resource keyword spotting task, where mixed word and morpheme decoding approach outperforms the traditional subword-based search from word-decoded lattices that are broken down to subword lattices. Furthermore, unsupervised learning of morphology works almost as well as a rule-based system designed for the language despite the low-resource condition. A staged keyword search strategy benefits from both methods of morphological analysis. Yanzhang He, Brian Hutchinson, Peter Baumann 0003, Mari Ostendorf, Eric Fosler-Lussier, Janet B. Pierrehumbert |
ICASSP | 5 |
| 2014 | The speech recognition virtual kitchen: launch party
Andrew R. Plummer, Eric Riebling, Florian Metze, Eric Fosler-Lussier, Rebecca Bates 0001 |
INTERSPEECH | 5 |
| 2014 | A discriminative sequence model for dialog state tracking using user goal change detectionabstractDue to the dominating influence of Partially Observable Markov Decision Process (POMDP) framework used in spoken dialog systems, most previously proposed dialog state tracking methods favor generative models. However, in this work we adopt a discriminative approach to model the evolution of the belief state within a spoken dialog system - more specifically, we use Conditional Random Fields (CRFs). Although we are not the first to apply CRFs to dialog state tracking, the proposed approach considers the dialog state tracking task as a sequence tagging problem, in the hope of capturing the evolving user goals during a dialog. Equipped with an incremental decoding strategy as well as user goal change detection, our results show that both sequence modeling and goal change information could bring advantage to the task. Eric Fosler-Lussier |
SLT | 2 |
| 2014 | Syllable based keyword search: Transducing syllable lattices to word latticesabstractThis paper presents a weighted finite state transducer (WFST) based syllable decoding and transduction framework for keyword search (KWS). Acoustic context dependent phone models are trained from word forced alignments. Then syllable decoding is done with lattices generated using a syllable lexicon and language model (LM). To process out-of-vocabulary (OOV) keywords, pronunciations are produced using a grapheme-to-syllable (G2S) system. A syllable to word lexical transducer containing both in-vocabulary (IV) and OOV keywords is then constructed and composed with a keyword-boosted LM transducer. The composed transducer is then used to transduce syllable lattices to word lattices for final KWS. We show that our method can effectively perform KWS on both IV and OOV keywords, and yields up to 0.03 Actual Term-Weighted Value (ATWV) improvement over searching keywords directly in subword lattices. Word Error Rates (WER) and KWS results are reported for three different languages. James Hieronymus, Yanzhang He, Eric Fosler-Lussier, Steven Wegmann |
SLT | 4 |
| 2014 | A review of approaches to identifying patient phenotype cohorts using electronic health recordsabstractOBJECTIVE: To summarize literature describing approaches aimed at automatically identifying patients with a common phenotype. MATERIALS AND METHODS: We performed a review of studies describing systems or reporting techniques developed for identifying cohorts of patients with specific phenotypes. Every full text article published in (1) Journal of American Medical Informatics Association, (2) Journal of Biomedical Informatics, (3) Proceedings of the Annual American Medical Informatics Association Symposium, and (4) Proceedings of Clinical Research Informatics Conference within the past 3 years was assessed for inclusion in the review. Only articles using automated techniques were included. RESULTS: Ninety-seven articles met our inclusion criteria. Forty-six used natural language processing (NLP)-based techniques, 24 described rule-based systems, 41 used statistical analyses, data mining, or machine learning techniques, while 22 described hybrid systems. Nine articles described the architecture of large-scale systems developed for determining cohort eligibility of patients. DISCUSSION: We observe that there is a rise in the number of studies associated with cohort identification using electronic medical records. Statistical analyses or machine learning, followed by NLP techniques, are gaining popularity over the years in comparison with rule-based systems. CONCLUSIONS: There are a variety of approaches for classifying patients into a particular phenotype. Different techniques and data sources are used, and good performance is reported on datasets at respective institutions. However, no system makes comprehensive use of electronic medical records addressing all of their known weaknesses. Chaitanya P. Shivade, Preethi Raghavan, Eric Fosler-Lussier, Peter J. Embí, Noémie Elhadad, Stephen B. Johnson, Albert M. Lai |
J. Am. Medical Informatics Assoc. | 3 |
| 2013 | Discriminative articulatory models for spoken term detection in low-resource conversational settingsabstractWe study spoken term detection (STD) - the task of determining whether and where a given word or phrase appears in a given segment of speech - using articulatory feature-based pronunciation models. The models are motivated by the requirements of STD in low-resource settings, in which it may not be feasible to train a large-vocabulary continuous speech recognition system, as well as by the need to address pronunciation variation in conversational speech. Our STD system is trained to maximize the expected area under the receiver operating characteristic curve, often used to evaluate STD performance. In experimental evaluations on the Switchboard corpus, we find that our approach outperforms a baseline HMM-based system across a number of training set sizes, as well as a discriminative phone-based model in some settings. Rohit Prabhavalkar, Karen Livescu, Eric Fosler-Lussier, Joseph Keshet |
ICASSP | 3 |
| 2013 | Discriminative training of WFST factors with application to pronunciation modelingabstractOne of the most popular speech recognition architectures consists of multiple components (like the acoustic, pronunciation and language models) that are modeled as weighted finite state transducer (WFST) factors in a cascade. These factor WFSTs are typically trained in isolation and combined efficiently for decoding. Recent work has explored jointly estimating parameters for these models using considerable amounts of training data. We propose an alternative approach to selectively train factor WFSTs in such an architecture, while still leveraging information from the entire cascade. This technique allows us to effectively estimate parameters of a factor WFST using relatively small amounts of data, if the factor is small. Our approach involves an online training paradigm for linear models adapted for discriminatively training one or more WFSTs in a cascade. We apply this method to train a pronunciation model for recognition on conversational speech, resulting in significant improvements in recognition performance over the baseline model. Index Terms: Pronunciation models, weighted finite state transducers, large-margin training. Preethi Jyothi, Eric Fosler-Lussier, Karen Livescu |
INTERSPEECH | 2 |
| 2013 | The speech recognition virtual kitchen
Florian Metze, Eric Fosler-Lussier, Rebecca Bates 0001 |
INTERSPEECH | 2 |
| 2013 | Conditional Random Fields in Speech, Audio, and Language ProcessingabstractConditional random fields (CRFs) are probabilistic sequence models that have been applied in the last decade to a number of applications in audio, speech, and language processing. In this paper, we provide a tutorial overview of CRF technologies, pointing to other resources for more in-depth discussion; in particular, we describe the common linear-chain model as well as a number of common extensions within the CRF family of models. An overview of the mathematical techniques used in training and evaluating these models is also provided, as well as a discussion of the relationships with other probabilistic models. Finally, we survey recent work in speech, audio, and language processing to show how the same CRF technology can be deployed in different scenarios. Eric Fosler-Lussier, Yanzhang He, Preethi Jyothi, Rohit Prabhavalkar |
Proc. IEEE | 1 |
| 2013 | A Direct Masking Approach to Robust ASRabstractRecently, much work has been devoted to the computation of binary masks for speech segregation. Conventional wisdom in the field of ASR holds that these binary masks cannot be used directly; the missing energy significantly affects the calculation of the cepstral features commonly used in ASR. We show that this commonly held belief may be a misconception; we demonstrate the effectiveness of directly using the masked data on both a small and large vocabulary dataset. In fact, this approach, which we term the direct masking approach, performs comparably to two previously proposed missing feature techniques. We also investigate the reasons why other researchers may have not come to this conclusion; variance normalization of the features is a significant factor in performance. This work suggests a much better baseline than unenhanced speech for future work in missing feature ASR. William Hartmann, Arun Narayanan, Eric Fosler-Lussier, DeLiang Wang |
IEEE Trans. Speech Audio Process. | 3 |
| 2012 | Inter-Annotator Reliability of Medical Events, Coreferences and Temporal Relations in Clinical Narratives by Annotators with Varying Levels of Clinical Expertise
Preethi Raghavan, Eric Fosler-Lussier, Albert M. Lai |
AMIA | 2 |
| 2012 | ASR-driven top-down binary mask estimation using spectral priorsabstractTypical mask estimation algorithms use low-level features to estimate the interfering noise or instantaneous SNR. We propose a simple top-down approach to mask estimation. The estimated mask is based on a specific hypothesis of the underlying speech without using information about the interference or the instantaneous SNR. In this pilot study, we observe a 9% reduction in word error over a baseline recognition system on the Aurora4 corpus, though much greater gains could theoretically be achieved through improvements to the model selection process. We also present SNR improvement results showing our method performs as well as a standard MMSE-based method, demonstrating that speech recognition can aid speech enhancement. Thus, the relationship between recognition and enhancement need not be one way: linguistic information can play a significant role in speech enhancement. William Hartmann, Eric Fosler-Lussier |
ICASSP | 2 |
| 2012 | Improved Model Selection for the ASR-Driven Binary MaskabstractIn a previous study, we proposed an alternative masking criterion for binary mask estimation based on the under-lying linguistic information. We estimated this mask by selecting from a set of candidate masks at each frame based on the hypotheses from an ASR system. Our pre-vious system provided an 8 % reduction in WER. In this work, we present an improved method for selecting the correct candidate mask at each frame, increasing the re-duction in WER to 14%. Our new method uses a discrim-inative sequence model and provides a framework that can incorporate other mask estimations as features. Index Terms: speech recognition, binary mask estima-tion William Hartmann, Eric Fosler-Lussier |
INTERSPEECH | 2 |
| 2012 | Efficient Segmental Conditional Random Fields for One-Pass Phone Recognition
Yanzhang He, Eric Fosler-Lussier |
INTERSPEECH | 2 |
| 2012 | Discriminatively learning factorized finite state pronunciation models from dynamic Bayesian networksabstractThis paper describes an approach to efficiently construct, and discriminatively train, a weighted finite state transducer (WFST) representation for an articulatory feature-based model of pronunciation. This model is originally implemented as a dynamic Bayesian network (DBN). The work is motivated by a desire to (1) incorporate such a pronunciation model in WFSTbased recognizers, and to (2) learn discriminative models that are more general than the DBNs. The approach is quite general, though here we show how it applies to a specific model. We use the conditional independence assumptions imposed by the DBN to efficiently convert it into a sequence of WFSTs (factor FSTs) which, when composed, yield the same model as the DBN. We then introduce a linear model of the arc weights of the factor FSTs and discriminatively learn its weights using the averaged perceptron algorithm. We demonstrate the approach using a lexical access task in which we recognize a word given its surface realization. Our experimental results using a phonetically transcribed subset of the Switchboard corpus show that the discriminatively learned model performs significantly better than the original DBN. Preethi Jyothi, Eric Fosler-Lussier, Karen Livescu |
INTERSPEECH | 2 |
| 2012 | The Speech Recognition Virtual Kitchen: An Initial Prototype
Florian Metze, Eric Fosler-Lussier |
INTERSPEECH | 2 |
| 2012 | Associative and Semantic Features Extracted From Web-Harvested Corpora
Elias Iosif, Maria Giannoudaki, Eric Fosler-Lussier, Alexandros Potamianos |
LREC | 3 |
| 2012 | Ranking-based readability assessment for early primary children's literature
Eric Fosler-Lussier, Robert Lofthus |
HLT-NAACL | 2 |
| 2012 | Exploring Semi-Supervised Coreference Resolution of Medical Concepts using Semantic and Temporal Features
Preethi Raghavan, Eric Fosler-Lussier, Albert M. Lai |
HLT-NAACL | 2 |
| 2011 | A factored conditional random field model for articulatory feature forced transcriptionabstractWe investigate joint models of articulatory features and apply these models to the problem of automatically generating articulatory transcriptions of spoken utterances given their word transcriptions. The task is motivated by the need for larger amounts of labeled articulatory data for both speech recognition and linguistics research, which is costly and difficult to obtain through manual transcription or physical measurement. Unlike phonetic transcription, in our task it is important to account for the fact that the articulatory features can desynchronize. We consider factored models of the articulatory state space with an explicit model of articulator asynchrony. We compare two types of graphical models: a dynamic Bayesian network (DBN), based on previously proposed models; and a conditional random field (CRF), which we develop here. We demonstrate how task-specific constraints can be leveraged to allow for efficient exact inference in the CRF. On the transcription task, the CRF outperforms the DBN, with relative improvements of 2.2% to 10.0%. Rohit Prabhavalkar, Eric Fosler-Lussier, Karen Livescu |
ASRU | 2 |
| 2011 | Investigations into the incorporation of the Ideal Binary Mask in ASRabstractWhile much work has been dedicated to exploring how best to incorporate the Ideal Binary Mask (IBM) in automatic speech recognition (ASR) for noisy signals, we demonstrate that the simple use of masked speech can outperform standard spectral reconstruction methods. We explore the effects of both the accuracy of the mask estimation and the strength of the language model on our results. The relative performance of these techniques is directly tied to the accuracy of the estimated mask. Although the use of masked speech fails when significant numbers of errors are present, the maximum performance for spectral reconstruction techniques also drops significantly. This implies improvements in mask estimation can provide greater gains in ASR performance than improvements in the incorporation of the IBM in ASR. Previous work may have ignored the direct use of masked speech due to its poor performance on tasks without a strong language model. William Hartmann, Eric Fosler-Lussier |
ICASSP | 2 |
| 2011 | Lexical access experiments with context-dependent articulatory feature-based modelsabstractWe address the problem of pronunciation variation in conversational speech with a context-dependent articulatory feature-based model. The model is an extension of previous work using dynamic Bayesian networks, which allow for easy factorization of a state into multiple variables representing the articulatory features. We build context-dependent decision trees for the articulatory feature distributions, which are incorporated into the dynamic Bayesian networks, and experiment with different sets of context variables. We evaluate our models on a lexical access task using a phonetically transcribed subset of the Switchboard corpus. We find that our models outperform a context-dependent phonetic baseline. Preethi Jyothi, Karen Livescu, Eric Fosler-Lussier |
ICASSP | 3 |
| 2011 | Robust speech recognition using multiple prior models for speech reconstructionabstractPrior models of speech have been used in robust automatic speech recognition to enhance noisy speech. Typically, a single prior model is trained by pooling the entire training data. In this paper we propose to train multiple prior models of speech instead of a single prior model. The prior models can be trained based on distinct characteristics of speech. In this study, they are trained based on voicing characteristics. The trained prior models are then used to reconstruct noisy speech. Significant improvements are obtained on the Aurora-4 robust speech recognition task when multiple priors are used; in conjunction with an uncertainty transform technique, multiple priors yield a 13.7% absolute improvement in the average word error rate over directly recognizing noisy speech. Arun Narayanan, Xiaojia Zhao, DeLiang Wang, Eric Fosler-Lussier |
ICASSP | 4 |
| 2010 | Backpropagation training for multilayer conditional random field based phone recognitionabstractConditional random fields (CRFs) have recently found increased popularity in automatic speech recognition (ASR) applications. CRFs have previously been shown to be effective combiners of posterior estimates from multilayer perceptrons (MLPs) in phone and word recognition tasks. In this paper, we describe a novel hybrid Multilayer-CRF structure (ML-CRF), where a MLP-like hidden layer serves as input to the CRF; moreover, we propose a technique for directly training the ML-CRF to optimize a conditional log-likelihood based criterion, based on error backpropagation. The proposed technique thus allows for the implicit learning of suitable feature functions for the CRF. We present results for initial phone recognition experiments on the TIMIT database that indicate that our proposed method is a promising approach for training CRFs. Rohit Prabhavalkar, Eric Fosler-Lussier |
ICASSP | 2 |
| 2010 | Machine learning for text selection with expressive unit-selection voicesabstractWe show that a ranking model produced by machine learning outperforms two baselines when applied to the task of selecting texts for use in creating a unit-selection synthesis voice with good domain coverage. The model learns to predict the estimated utility of an utterance based on features relating it to the utterances selected so far and a corpus of target utterances. Our analyses indicate that our discriminative approach continues to work well even though the presence of rich prosodic and nonprosodic features significantly expands the search space beyond what has previously been handled by greedy methods. Index Terms: speech synthesis, unit selection, machine learning Dominic Espinosa, Michael White 0001, Eric Fosler-Lussier, Chris Brew |
INTERSPEECH | 3 |
| 2010 | Discriminative language modeling using simulated ASR errorsabstractIn this paper, we approach the problem of discriminatively training language models using a weighted finite state transducer (WFST) framework that does not require acoustic training data. The phonetic confusions prevalent in the recognizer are modeled using a confusion matrix that takes into account information from the pronunciation model (word-based phone confusion log likelihoods) and information from the acoustic model (distances between the phonetic acoustic models). This confusion matrix, within the WFST framework, is used to generate confusable word graphs that serve as inputs to the averaged perceptron algorithm to train the parameters of the discriminative language model. Experiments on a large vocabulary speech recognition task show significant word error rate reductions when compared to a baseline using a trigram model trained with the maximum likelihood criterion. Preethi Jyothi, Eric Fosler-Lussier |
INTERSPEECH | 2 |
| 2010 | Learning speaker normalization using semisupervised manifold alignmentabstractAs a child acquires language, he or she: perceives acous-tic information in his or her surrounding environment; identi-fies portions of the ambient acoustic information as language-related; and associates that language-related information with his or her perception of his or her own language-related acous-tic productions. The present work models the third task. We use a semisupervised alignment algorithm based on manifold learning. We discuss the concepts behind this approach, and the application of the algorithm to this task. We present experimen-tal evidence indicating the usefulness of manifold alignment in learning speaker normalization. Index Terms: speaker normalization, manifold alignment, lan-guage acquisition Andrew R. Plummer, Mary E. Beckman, Mikhail Belkin, Eric Fosler-Lussier, Benjamin Munson |
INTERSPEECH | 4 |
| 2010 | Combining monaural and binaural evidence for reverberant speech segregationabstractMost existing binaural approaches to speech segregation rely on spatial filtering. In environments with minimal reverberation and when sources are well separated in space, spatial filtering can achieve excellent results. However, in everyday environments performance degrades substantially. To address these limitations, we incorporate monaural analysis within a binaural segregation system. We use monaural cues to perform both local and across frequency grouping of mixture components, allowing for a more robust application of spatial filtering. We propose a novel framework in which we combine monaural grouping evidence and binaural localization evidence in a linear model for the estimation of the ideal binary mask. Results indicate that with appropriately designed features that capture both monaural and binaural evidence, an extremely simple model achieves a signal-to-noise ratio improvement of up to 3.6 dB relative to using spatial filtering alone. Index Terms: Speech segregation, binaural localization, monaural grouping, linear model John Woodruff, Rohit Prabhavalkar, Eric Fosler-Lussier, DeLiang Wang |
INTERSPEECH | 3 |
| 2010 | Investigations into the Crandem Approach to Word Recognition
Rohit Prabhavalkar, Preethi Jyothi, William Hartmann, Jeremy Morris, Eric Fosler-Lussier |
HLT-NAACL | 5 |
| 2009 | Transition features for CRF-based speech recognition and boundary detectionabstractIn this paper, we investigate a variety of spectral and time domain features for explicitly modeling phonetic transitions in speech recognition. Specifically, spectral and energy distance metrics, as well as, time derivatives of phonological descriptors and MFCCs are employed. The features are integrated in an extended Conditional Random Fields statistical modeling framework that supports general-purpose transition models. For evaluation purposes, we measure both phonetic recognition task accuracy and precision/recall of boundary detection. Results show that when transition features are used in a CRF-based recognition framework, recognition performance improves significantly due to the reduction of phone deletions. The boundary detection performance also improves mainly for transitions among silence, stop, and fricative phonetic classes. Spiros Dimopoulos, Eric Fosler-Lussier, Alexandros Potamianos |
ASRU | 2 |
| 2009 | Multiple time resolution analysis of speech signal using MCE training with application to speech recognitionabstractIn this paper, we propose two methods of multiple time-resolution analysis of speech and their application to automatic speech recognition (ASR). Constant frame-rate multi-scale analysis is proposed based on a box of multi-scale features. Then a variable rate analysis is proposed based on the selection of the optimal temporal resolution on the fly by a properly trained non-linear classifier unit. The classifier's parameters are trained using the discriminative method of minimum classification error (MCE) training. We use the recently proposed conditional random fields (CRF) phonetic recognition system that effectively combines highly correlated features. Results are reported on a frame-wise classification task and also on TIMIT phone recognition task. Results show that (i) CRFs can effectively combine multi-scale features and (ii) MCE trained variable rate CRFs are competitive with the ldquoboxrdquo combination method. Spiros Dimopoulos, Alexandros Potamianos, Eric Fosler-Lussier |
ICASSP | 3 |
| 2009 | Investigating phonetic information reduction and lexical confusabilityabstractIn the presence of pronunciation variation and the masking ef-fects of additive noise, we investigate the role of phonetic in-formation reduction and lexical confusability on ASR perfor-mance. Contrary to previous work [1], we show that place of ar-ticulation as a representation for unstressed segments performs at least as well as manner of articulation in the presence of ad-ditive noise. Methods of phonetic reduction introduce lexical confusibility which negatively impact performance. By limiting this confusability, recognizers that employ high levels of pho-netic reduction (40.1%) can perform as well a baseline system in the presence of nonstationary noise. Index Terms: spoken language analysis, speech recognition, articulatory features William Hartmann, Eric Fosler-Lussier |
INTERSPEECH | 2 |
| 2009 | Evaluating parameters for mapping adult vowels to imitative babblingabstractWe design a neural network model of first language acquisition to explore the relationship between child and adult speech sounds. The model learns simple vowel categories using a produce-and-perceive babbling algorithm in addition to listening to ambient speech. The model is similar to that of Westermann & Miranda (2004), but adds a dynamic aspect in that it adapts in both the articulatory and acoustic domains to changes in the child’s speech patterns. The training data is designed to replicate infant speech sounds and articulatory configurations. By exploring a range of articulatory and acoustic dimensions, we see how the child might learn to draw correspondences between his or her own speech and that of a caretaker, whose productions are quite different from the child’s. We also design an imitation evaluation paradigm that gives insight into the strengths and weaknesses of the model. Index Terms: language acquisition, neural networks, selforganizing maps, language development Ilana Heintz, Mary E. Beckman, Eric Fosler-Lussier, Lucie Ménard |
INTERSPEECH | 3 |
| 2009 | A comparison of audio-free speech recognition error prediction methodsabstractPredicting possible speech recognition errors can be invaluable for a number of Automatic Speech Recognition (ASR) applications. In this study, we extend a Weighted Finite State Transducer (WFST) framework for error prediction to facilitate a comparison between two approaches of predicting confusable words: examining recognition errors on the training set to learn phone confusions and utilizing distances between the phonetic acoustic models for the prediction task. We also expand the framework to deal with continuous word recognition and we can accurately predict 60% of the misrecognized sentences (with an average words-per-sentence count of 15) and a little over 70% of the total number of errors from the unseen test data where no acoustic information related to the test data is utilized. Index Terms: Finite State Transducer, Automatic Speech Recognition, Error prediction Preethi Jyothi, Eric Fosler-Lussier |
INTERSPEECH | 2 |
| 2009 | CRANDEM: conditional random fields for word recognition
Jeremy Morris, Eric Fosler-Lussier |
INTERSPEECH | 2 |
| 2009 | Monaural segregation of voiced speech using discriminative random fields
Rohit Prabhavalkar, Zhaozhang Jin, Eric Fosler-Lussier |
INTERSPEECH | 3 |
| 2009 | Discriminative Input Stream Combination for Conditional Random Field Phone RecognitionabstractIn recent studies, we and others have found that conditional random fields (CRFs) can be effectively used to perform phone classification and recognition tasks by combining non-Gaussian distributed representations of acoustic input. In previous work by I. Heintz (latent phonetic analysis: Use of singular value decomposition to determine features for CRF phone recognition,Proc.ICASSP, pp. 4541-4544, 2008), we experimented with combining phonological feature posterior estimators and phone posterior estimators within a CRF framework; we found that treating posterior estimates as terms in a ldquophoneme information retrievalrdquo task allowed for a more effective use of multiple posterior streams than directly feeding these acoustic representations to the CRF recognizer. In this paper, we examine some of the design choices in our previous work, and extend our results to up to six acoustic feature streams. We concentrate on feature design, rather than feature selection, to find the best way of combining features for introduction into a log-linear model. We improve upon our previous work to find that several different dimensionality reduction techniques (SVD, PARAFAC2, KLT), followed by a nonlinear transform provided by a multilayer perceptron, provides a significant gain in phone recognition accuracy on the TIMIT task. Ilana Heintz, Eric Fosler-Lussier, Chris Brew |
IEEE Trans. Speech Audio Process. | 2 |
| 2008 | Crandem systems: Conditional random field acoustic models for hidden Markov modelsabstractIn recent years, Conditional Random Fields (CRFs) have been examined as a statistical model for speech recognition. In this paper, we explore the use of features derived via CRFs as inputs to a Tandem-style HMM ASR system (that is, a Crandem system). We present a model for deriving frame-level posterior features via CRFs to use in Crandem modeling and additionally provide experimental results that show the Crandem system can slightly significantly outperform both a comparable Tandem system and a comparable CRF system on the task of phone recognition. Eric Fosler-Lussier, Jeremy Morris |
ICASSP | 1 |
| 2008 | Latent phonetic analysis: Use of singular value decomposition to determine features for CRF phone recognitionabstractWe exploit an analogy between document retrieval and phone recognition, and adapt the method of latent semantic analysis for the latter task. By mapping into a space of reduced dimensionality, we hope to uncover previously unexploited relationships between posterior estimates of phonetic events and the parts of phones represented by HMM states. We find that features defined over the reduced space complement those previously known, such as, for example, phonological features. We are able to effectively combine all of these features in a phone recognition task by using the constraint-based framework of conditional random fields (CRFs), which allows the use of large and highly redundant feature spaces. Ilana Heintz, Eric Fosler-Lussier, Chris Brew |
ICASSP | 2 |
| 2008 | Investigations into phonological attribute classifier representations for CRF phone recognition
Prateeti Mohapatra, Eric Fosler-Lussier |
INTERSPEECH | 2 |
| 2008 | SCARE: a Situated Corpus with Annotated Referring Expressions
Laura Stoia, Darla Magdalena Shockley, Donna K. Byron, Eric Fosler-Lussier |
LREC | 4 |
| 2008 | Conditional Random Fields for Integrating Local Discriminative ClassifiersabstractConditional random fields (CRFs) are a statistical framework that has recently gained in popularity in both the automatic speech recognition (ASR) and natural language processing communities because of the different nature of assumptions that are made in predicting sequences of labels compared to the more traditional hidden Markov model (HMM). In the ASR community, CRFs have been employed in a method similar to that of HMMs, using the sufficient statistics of input data to compute the probability of label sequences given acoustic input. In this paper, we explore the application of CRFs to combine local posterior estimates provided by multilayer perceptrons (MLPs) corresponding to the frame-level prediction of phone classes and phonological attribute classes. We compare phonetic recognition using CRFs to an HMM system trained on the same input features and show that the monophone label CRF is able to achieve superior performance to a monophone-based HMM and performance comparable to a 16 Gaussian mixture triphone-based HMM; in both of these cases, the CRF obtains these results with far fewer free parameters. The CRF is also able to better combine these posterior estimators, achieving a substantial increase in performance over an HMM-based triphone system by mixing the two highly correlated sets of phone class and phonetic attribute class posteriors. Jeremy Morris, Eric Fosler-Lussier |
IEEE Trans. Speech Audio Process. | 2 |
| 2007 | Further Experiments with Detector-Based Conditional Random Fields in Phonetic RecognitionabstractIn our prior work with conditional random fields (CRFs), we have shown that it is possible to achieve results in the phonetic recognition task with a CRF that approach the results of a similarly trained HMM system (but with many fewer parameters), and we have shown that using two different feature sets that are supposedly redundant gives an improvement in the performance of the CRF. In this paper, we explore two new areas with our CRF model. First, we show that by using two feature sets that are just transforms of each other, we achieve an improvement of results in the CRF model. Second, we show that by adding a single pass of realignment to our CRF model training, we achieve an accuracy result in the phone recognition task that is superior to that of an HMM system trained with triphone labels, despite only training the CRF on monophone labels with no explicit triphonic context. Jeremy Morris, Eric Fosler-Lussier |
ICASSP (4) | 2 |
| 2007 | The buckeye corpus of speech: updates and enhancementsabstractThis paper describes recent progress in the development of the Buckeye Corpus of Speech, a phonetically labeled corpus of conversational American English speech, first described in [1]. With the publication of the second phase of transcription, the corpus has nearly doubled in size from the first release. We briefly give an overview of the corpus, report on additional stud-ies of inter-labeler agreement, and describe a new GUI designed to facilitate searching the annotated speech files. Index Terms: corpora, transcription, phonetics, search tool 1. Eric Fosler-Lussier, Laura Dilley, Na'im R. Tyson, Mark A. Pitt |
INTERSPEECH | 1 |
| 2007 | An overview on automatic speech attribute transcription (ASAT)abstractAutomatic Speech Attribute Transcription (ASAT), an ITR project sponsored under the NSF grant (IIS-04-27113), is a cross-institute effort involving Georgia Institute of Technology, The Ohio State University, University of California at Berkeley, and Rutgers University. This project approaches speech recognition from a more linguistic perspective: unlike traditional ASR systems, humans detect acoustic and auditory cues, weigh and combine them to form theories, and then process these cognitive hypotheses until linguistically and pragmatically consistent speech understanding is achieved. A major goal of the ASAT paradigm is to develop a detection-based approach to automatic speech recognition (ASR) based on attribute detection and knowledge integration. We report on progress of the ASAT project, present a sharable platform for community collaboration, and highlight areas of potential interdisciplinary ASR research. Index Terms: attributes, events, features, detection, speech recognition, speech attribute transcription, utterance verification Mark A. Clements, Sorin Dusan, Eric Fosler-Lussier, Keith Johnson, Biing-Hwang Juang, Lawrence R. Rabiner |
INTERSPEECH | 4 |
| 2007 | Information Seeking Spoken Dialogue Systems- Part I: Semantics and PragmaticsabstractIn this paper, the semantic and pragmatic modules of a spoken dialogue system development platform are presented and evaluated. The main goal of this research is to create spoken dialogue system modules that are portable across applications domains and interaction modalities. We propose a hierarchical semantic representation that encodes all information supplied by the user over multiple dialogue turns and can efficiently represent and be used to argue with ambiguous or conflicting information. Implicit in this semantic representation is a pragmatic module, consisting of context tracking, pragmatic analysis and pragmatic scoring submodules, which computes pragmatic confidence scores for all system beliefs. These pragmatic scores are obtained by combining semantic and pragmatic evidence from the various sub-modules (taking into account the modality of input) and are used to rank-order attribute-value pairs in the semantic representation, as well as identifying and resolving ambiguities. These modules were implemented and evaluated within a travel reservation dialogue system under the auspices of the DARPA Communicator project, as well as for a movie information application. Formal evaluation of the semantic and pragmatic modules has shown that by incorporating pragmatic analysis and scoring, the quality of the system improves for over 20% of the dialogue fragments examined Egbert Ammicht, Eric Fosler-Lussier, Alexandros Potamianos |
IEEE Trans. Multim. | 2 |
| 2007 | Information Seeking Spoken Dialogue Systems- Part II: Multimodal DialogueabstractFor pt.1see ibid., vol. 9, p. 3 (2007). In this paper, the task and user interface modules of a multimodal dialogue system development platform are presented. The main goal of this work is to provide a simple, application-independent solution to the problem of multimodal dialogue design for information seeking applications. The proposed system architecture clearly separates the task and interface components of the system. A task manager is designed and implemented that consists of two main submodules: the electronic form module that handles the list of attributes that have to be instantiated by the user, and the agenda module that contains the sequence of user and system tasks. Both the electronic forms and the agenda can be dynamically updated by the user. Next a spoken dialogue module is designed that implements the speech interface for the task manager. The dialogue manager can handle complex error correction and clarification user input, building on the semantics and pragmatic modules presented in Part I of this paper. The spoken dialogue system is evaluated for a travel reservation task of the DARPA Communicator research program and shown to yield over 90% task completion and good performance for both objective and subjective evaluation metrics. Finally, a multimodal dialogue system which combines graphical and speech interfaces, is designed, implemented and evaluated. Minor modifications to the unimodal semantic and pragmatic modules were required to build the multimodal system. It is shown that the multimodal system significantly outperforms the unimodal speech-only system both in terms of efficiency (task success and time to completion) and user satisfaction for a travel reservation task Alexandros Potamianos, Eric Fosler-Lussier, Egbert Ammicht, Manolis Perakakis |
IEEE Trans. Multim. | 2 |
| 2006 | Noun Phrase Generation for Situated Dialogs
Laura Stoia, Darla Magdalena Shockley, Donna K. Byron, Eric Fosler-Lussier |
INLG | 4 |
| 2006 | Combining phonetic attributes using conditional random fieldsabstractA Conditional Random Field is a mathematical model for sequences that is similar in many ways to a Hidden Markov Model, but is discriminative rather than generative in nature. In this paper, we explore the application of the CRF model to ASR processing of discriminative phonetic features by building a system that performs first-pass phonetic recognition using discriminatively trained phonetic features. With this system, we show that this CRF model achieves an accuracy level in a phone recognition task that is superior to a similarly trained HMM model. Index Terms: speech recognition, conditional random fields. Jeremy Morris, Eric Fosler-Lussier |
INTERSPEECH | 2 |
| 2006 | Integrating phonetic boundary discrimination explicitly into HMM systems
Yu Wang 0001, Eric Fosler-Lussier |
INTERSPEECH | 2 |
| 2006 | The OSU Quake 2004 corpus of two-party situated problem-solving dialogs
Donna K. Byron, Eric Fosler-Lussier |
LREC | 2 |
| 2006 | Sentence Planning for Realtime Navigational Instruction
Laura Stoia, Donna K. Byron, Darla Magdalena Shockley, Eric Fosler-Lussier |
HLT-NAACL | 4 |
| 2006 | Unsupervised Combination of Metrics for Semantic Class InductionabstractIn this paper, unsupervised algorithms for combining semantic similarity metrics are proposed for the problem of automatic class induction. The automatic class induction algorithm is based on the work of Pargellis et al,. The semantic similarity metrics that are evaluated and combined are based on narrow- and wide-context vector- product similarity. The metrics are combined using linear weights that are computed 'on the fly' and are updated at each iteration of the class induction algorithm, forming a corpus-independent metric. Specifically, the weight of each metric is selected to be inversely proportional to the inter-class similarity of the classes induced by that metric and for the current iteration of the algorithm. The proposed algorithms are evaluated on two corpora: a semantically heterogeneous news domain (HR-Net) and an application-specific travel reservation corpus (ATIS). It is shown, that the (unsupervised) adaptive weighting scheme outperforms the (supervised) fixed weighting scheme. Up to 50% relative error reduction is achieved by the adaptive weighting scheme. Elias Iosif, Athanasios Tegos, Apostolos Pangos, Eric Fosler-Lussier, Alexandros Potamianos |
SLT | 4 |
| 2005 | Phonetic ignorance is bliss: investigating the effects of phonetic information reduction on ASR performance
Eric Fosler-Lussier, C. Anton Rytting, Soundararajan Srinivasan |
INTERSPEECH | 1 |
| 2005 | A framework for predicting speech recognition errors
Eric Fosler-Lussier, Ingunn Amdal, Hong-Kwang Jeff Kuo |
Speech Commun. | 1 |
| 2005 | Editorial
Eric Fosler-Lussier, William J. Byrne, Daniel Jurafsky |
Speech Commun. | 1 |
| 2004 | Auto-induced semantic classes
Andrew N. Pargellis, Eric Fosler-Lussier, Alexandros Potamianos, Augustine Tsai |
Speech Commun. | 2 |
| 2003 | Discourse Segmentation of Multi-Party ConversationabstractWe present a domain-independent topic segmentation algorithm for multi-party speech. Our feature-based algorithm combines knowledge about content using a text-based algorithm as a feature and about form using linguistic and acoustic cues about topic shifts extracted from speech. This segmentation algorithm uses automatically induced decision rules to combine the different features. The embedded text-based algorithm builds on lexical cohesion and has performance comparable to state-of-the-art algorithms based on lexical information. A significant error reduction is obtained by combining the two knowledge sources. Michel Galley, Kathy McKeown, Eric Fosler-Lussier, Hongyan Jing |
ACL | 3 |
| 2003 | Minimum verification error training for topic verificationabstractWe propose a new formulation of minimum verification error training and apply it to the problem of topic verification as an example. In topic verification, a decision is made as to whether a document truly belongs to a particular topic of interest. Such a decision typically depends on a comparison between a model for the desired topic and a model for background topics, using a decision threshold. We propose modeling the background topics as a cohort model consisting of a weighted combination of the M closest topics discovered from the training data. The weights and the decision threshold are optimized using the generalized probabilistic descent algorithm to explicitly minimize the verification error rate, which is defined to be a weighted sum of the Type I (false rejection) and Type II (false acceptance) errors. Hong-Kwang Jeff Kuo, Imed Zitouni, Eric Fosler-Lussier |
ICASSP (1) | 4 |
| 2002 | Discriminative training of language models for speech recognitionabstractIn this paper we describe how discriminative training can be applied to language models for speech recognition. Language models are important to guide the speech recognition search, particularly in compensating for mistakes in acoustic decoding. A frequently used measure of the quality of language models is the perplexity; however, what is more important for accurate decoding is not necessarily having the maximum likelihood hypothesis, but rather the best separation of the correct string from the competing, acoustically confusible hypotheses. Discriminative training can help to improve language models for the purpose of speech recognition by improving the separation of the correct hypothesis from the competing hypotheses. We describe the algorithm and demonstrate modest improvements in word and sentence error rates on the DARPA Communicator task without any increase in language model complexity. Hong-Kwang Jeff Kuo, Eric Fosler-Lussier, Hui Jiang 0001 |
ICASSP | 2 |
| 2002 | Adaptive language models for spoken dialogue systemsabstractIn this paper, we investigate both generative and statistical approaches for language modeling in spoken dialogue systems. Semantic class-based finite state and n-gram grammars are used for improving coverage and modeling accuracy when little training data is available. We have implemented dialogue-state specific language model adaptation to reduce perplexity and improve the efficiency of grammars for spoken dialogue systems. A novel algorithm for combining state-independent n-gram and state-dependent finite state grammars using acoustic confidence scores is proposed. Using this combination strategy, a relative word error reduction of 12% is achieved for certain dialogue states within a travel reservation task. Finally, semantic class multigrams are proposed and briefly evaluated for language modeling in dialogue systems. Roger Argiles Solsona, Eric Fosler-Lussier, Hong-Kwang Jeff Kuo, Alexandros Potamianos, Imed Zitouni |
ICASSP | 2 |
| 2002 | Discriminative training for call classification and routing
Hong-Kwang Jeff Kuo, Imed Zitouni, Eric Fosler-Lussier, Egbert Ammicht |
INTERSPEECH | 4 |
| 2002 | Using part-of-speech tags, context thresholding, and trigram contexts to improve the auto-induction of semantic classes
Andrew N. Pargellis, Eric Fosler-Lussier, Augustine Tsai |
INTERSPEECH | 2 |
| 2002 | Connectionist speech recognition of Broadcast News
Anthony J. Robinson, Gary D. Cook, Daniel P. W. Ellis, Eric Fosler-Lussier, Steve Renals, D. A. G. Williams |
Speech Commun. | 4 |
| 2001 | Using semantic class information for rapid development of language models within ASR dialogue systemsabstractWhen dialogue system developers tackle a new domain, much effort is required; the development of different parts of the system usually proceeds independently. Yet it may be profitable to coordinate development efforts between different modules. We focus our efforts on extending small amounts of language model training data by integrating semantic classes that were created for a natural language understanding module. By converting finite state parses of a training corpus into a probabilistic context free grammar and subsequently generating artificial data from the context free grammar, we can significantly reduce perplexity and automatic speech recognition (ASR) word error for situations with little training data. Experiments are presented using data from the ATIS and DARPA Communicator travel corpora. Eric Fosler-Lussier, Hong-Kwang Jeff Kuo |
ICASSP | 1 |
| 2001 | Ambiguity representation and resolution in spoken dialogue systemsabstractSpoken natural language often contains ambiguities that must be addressed by a spoken dialogue system. In this work, we present the internal semantic representation and resolution strategy of a dialogue system designed to understand ambiguous input. These mechanisms are domain independent; almost all task-specific knowledge is represented in parameterizable data structures. The system derives candidate descriptions of what the user said from raw input data, context-tracking and scoring. These candidates are chosen on the basis of a pragmatic analysis of responses elicited by an extensive implicit and explicit confirmation dialogue strategy, combined with specific error correction capabilities available to the user. This new ambiguity resolution strategy greatly improves dialogue interaction, eliminating about half of the errors in dialogues from a travel reservation task. Egbert Ammicht, Alexandros Potamianos, Eric Fosler-Lussier |
INTERSPEECH | 3 |
| 2001 | OASIS natural language call steering trialabstractA recent trial of natural language call steering on live UK calls to the operator is described along with its results. The characteristics of the problem are described along with the acoustic, language, semantic and dialogue modelling approaches employed. Natural language call steering is found to be viable, with recognition and semantic accuracy the current limiting factors. Peter J. Durston, Mark Farrell, David Attwater, James Allen, Hong-Kwang Jeff Kuo, Mohamed Afify, Eric Fosler-Lussier |
INTERSPEECH | 7 |
| 2001 | Hybrid natural language generation for spoken dialogue systemsabstractThe natural language generation component of most dialogue systems is based on templates. Template-based generators are hard to maintain and reuse, and the sentences they produce lack the variability and robustness needed by conversational systems. In this paper, a flexible and domain-independent natural language generator for spoken dialogue systems is proposed which combines fixed surface expressions with freely generated text. The generation algorithm follows a hybrid approach, combining finite state machine (FSM) grammars and corpus-based language models. In this approach, the FSM grammar (a reversible parser grammar) is constrained by a word and concept Ò-gram that takes terminals and non-terminal co-occurrences into account. The Ò-gram grammar helps prevent inappropriate derivations, therefore improving the quality of the generated texts. The proposed algorithm achieves faster than real-time performance because of the limited number of derivations. Michel Galley, Eric Fosler-Lussier, Alexandros Potamianos |
INTERSPEECH | 2 |
| 2001 | Metrics for measuring domain independence of semantic classesabstractThe design of dialogue systems for a new domain requires se-mantic classes (concepts) to be identified and defined. This process could be made easier by importing relevant concepts from previously studied domains to the new one. We pro-pose two methodologies, based on comparison of semantic classes across domains, for determining which concepts are domain-independent, and which are specific to the new task. The concept-comparison technique uses a context-dependent Kullback-Leibler distance measurement to compare all pairwise combinations of semantic classes, one from each domain. The concept-projection method uses a similar metric to project a sin-gle semantic class from one domain into the lexical environment of another. Initial results show that both methods are good in-dicators of the degree of domain independence for a wide range of concepts, manually generated for three different tasks: Car-men (children’s game), Movie (information retrieval) and Travel (flight reservations). 1. Andrew N. Pargellis, Eric Fosler-Lussier, Alexandros Potamianos |
INTERSPEECH | 2 |
| 2000 | A comparison of data-derived and knowledge-based modeling of pronunciation variationabstractThis paper focuses on modeling pronunciation variation in two different ways: data-derived and knowledge-based. The knowledge-based approach consists of using phonological rules to generate variants. The data-derived approach consists of performing phone recognition, followed by various pruning and smoothing methods to alleviate some of the errors in the phone recognition. Using phonological rules led to a small improvement in WER; whereas, using a data-derived approach in which the phone recognition was smoothed using simple decision trees (d-trees) prior to lexicon generation led to a significant improvement compared to the baseline. Furthermore, we found that 10% of variants generated by the phonological rules were also found using phone recognition, and this increased to 23% when the phone recognition output was smoothed by using d-trees. In addition, we propose a metric to measure confusability in the lexicon and we found that employing this confusion metric to prune variants results in roughly the same improvement as using the d-tree method. Mirjam Wester, Eric Fosler-Lussier |
INTERSPEECH | 2 |
| 2000 | Erratum to: "Effects of speaking rate and word frequency on pronunciations in convertional speech": [Speech Communication 29 (1999) 137-158]
Eric Fosler-Lussier, Nelson Morgan |
Speech Commun. | 1 |
| 1999 | Multi-level decision trees for static and dynamic pronunciation modelsabstractWe have been focusing on improving pronunciation models for automatic transcription of television and radio news reports by modeling phone, syllable, and word pronunciation distributions with decision trees. These models were employed in two separate sets of experiments. First, decision trees facilitated selection of word pronunciations derived automatically from data for use in a standard speech recognizer dictionary. We have seen a small but significant improvement with these automatically constructed dictionaries in our one-pass decoding system. In a second set of experiments, we allowed decision tree models to determine the probability of word pronunciations dynamically, dependent on the linguistic context of the word during recognition. Dynamic models provided an additional insignificant decrease in error, but improvements were focused within the spontaneous speech portion of the test set. 1. INTRODUCTION One goal of recent research within the ASR community has been to provide s... Eric Fosler-Lussier |
EUROSPEECH | 1 |
| 1999 | Effects of speaking rate and word frequency on pronunciations in convertional speech
Eric Fosler-Lussier, Nelson Morgan |
Speech Commun. | 1 |
| 1998 | Combining multiple estimators of speaking rateabstractWe report progress in the development of a measure of speaking rate that is computed from the acoustic signal. The newest form of our analysis incorporates multiple estimates of rate; besides the spectral moment for a full-band energy envelope that we have previously reported, we also used pointwise correlation between pairs of compressed sub-band energy envelopes. The complete measure, called mrate, has been compared to a reference syllable rate derived from a manually transcribed subset of the Switchboard database. The correlation with transcribed syllable rate is significantly higher than our earlier measure; estimates are typically within 1-2 syllables/second of the reference syllable rate. We conclude by assessing the use of mrate as a detector for rapid speech. Nelson Morgan, Eric Fosler-Lussier |
ICASSP | 2 |
| 1998 | Reduction of English function words in switchboardabstractThe causes of pronunciation reduction in 8458 occurrences of ten frequent English function words in a four-hour sample from conversations from the Switchboard corpus were examined. Using ordinary linear and logistic regression models, we examined the length of the words, the form of their vowel (basic, full, or reduced) , and final obstruent deletion. For all of these we found strong, independent effects of speaking rate, predictability, the form of the following word, and planning problem disfluencies. The results bear on issues in speech recognition, models of speech production, and conversational analysis. 1. INTRODUCTION This study reports the results of an investigation of some factors affecting the reduction or lenition of ten of the most frequent English words, namely I, and, the, that, a, you, to, of, it, and in, in the Switchboard corpus of conversational speech. Frequent function words are of particular interest because they are not only subject to the contextual and stylis... Daniel Jurafsky, Alan Bell, Eric Fosler-Lussier, Cynthia Girand, William D. Raymond |
ICSLP | 3 |
| 1997 | Speech recognition using on-line estimation of speaking rateabstractIn this paper, we describe a rate of speech estimator that is derived directly from the acoustic signal. This measure has been developed as an alternative to lexical measures of speaking rate such as phones or syllables per second, which, in previous work, we estimated using a first recognition pass; the accuracy of our earlier lexical rate estimate depended on the quality of recognition. Here we show that our new measure is a good predictor of word error rate, and in addition, correlates moderately well with lexical speech rate. We also show that a simple modification of the model transition probabilities based on this measure can reduce the error rate almost as much as using lexical phones per second calculated from manually transcribed data. When we categorized test utterances based on speaking rate thresholds computed from the training set, we observed that a different transition probability value was required to minimize the error rate in each speaking rate bin. However, the reduc... Nelson Morgan, Eric Fosler-Lussier, Nikki Mirghafori |
EUROSPEECH | 2 |
| 1996 | On Reversing the Generation Process in Optimality TheoryabstractOptimality Theory, a constraint-based phonology and morphology paradigm, has allowed linguists to make elegant analyses of many phenomena, including infixation and reduplication. In this work-in-progress, we build on the work of Ellison (1994) to investigate the possibility of using OT as a parsing tool that derives underlying forms from surface forms. Eric Fosler-Lussier |
ACL | 1 |
| 1996 | Towards robustness to fast speech in ASRabstractPsychoacoustic studies show that human listeners are sensitive to speaking rate variations. Automatic speech recognition (ASR) systems are even more affected by the changes in rate, as double to quadruple word recognition error rates of average speakers have been observed for fast speakers on many ASR systems. In our earlier work (see Proceedings of EUROSPEECH95, p.491-4, 1995), we studied the causes of higher error and concluded that both the acoustic-phonetic and the phonological differences are sources of higher word error rates. In this work, we have studied various measures for quantifying rate of speech (ROS) and used simple methods for estimating the speaking rate of a novel utterance using ASR technology. We have also implemented mechanisms that make our ASR system more robust to fast speech. Using our ROS estimator to identify fast sentences in the test set, our rate-dependent system has 24.5% fewer errors on the fastest sentences and 6.2% fewer errors on all sentences of the WSJ93 evaluation set relative to the baseline HMM/MLP system. Nikki Mirghafori, Eric Fosler-Lussier, Nelson Morgan |
ICASSP | 2 |
| 1995 | Learning Phonological Rule Probabilities from Speech Corpora with Exploratory Computational PhonologyabstractThis paper presents an algorithm for learning the probabilities of optional phonological rules from corpora. The algorithm is based on using a speech recognition system to discover the surface pronunciations of words in speech corpora; using an automatic system obviates expensive phonetic labeling by hand. We describe the details of our algorithm and show the probabilities the system has learned for ten common phonological rules which model reductions and coarticulation effects. These probabilities were derived from a corpus of 7203 sentences of read speech from the Wall Street Journal, and are shown to be a reasonably close match to probabilities from phonetically hand-transcribed data (TIMIT). Finally, we analyze the probability differences between rule use in male versus female speech, and suggest that the differences are caused by differing average rates of speech. Gary N. Tajchman, Daniel Jurafsky, Eric Fosler-Lussier |
ACL | 3 |
| 1995 | Using a stochastic context-free grammar as a language model for speech recognitionabstractThis paper describes a number of experiments in adding new grammatical knowledge to the Berkeley Restaurant Project (BeRP), our medium-vocabulary (1300 word), speaker-independent, spontaneous continuous-speech understanding system. We describe an algorithm for using a probabilistic Earley parser and a stochastic context-free grammar (SCFG) to generate word transition probabilities at each frame for a Viterbi decoder. We show that using an SCFG as a language model improves the word error rate from 34.6% (bigram) to 29.6% (SCFG), and the semantic sentence recognition error from from 39.0% (bigram) to 34.1% (SCFG). In addition, we get a further reduction to 28.8% word error by mixing the bigram and SCFG LMs. We also report on our preliminary results from using discourse-context information in the LM. Daniel Jurafsky, Chuck Wooters, Jonathan Segal, Andreas Stolcke, Eric Fosler-Lussier, Gary N. Tajchman, Nelson Morgan |
ICASSP | 5 |
| 1995 | Fast speakers in large vocabulary continuous speech recognition: analysis & antidotes
Nikki Mirghafori, Eric Fosler-Lussier, Nelson Morgan |
EUROSPEECH | 2 |
| 1995 | Building multiple pronunciation models for novel words using exploratory computational phonologyabstractIn this paper we describe a completely automatic algorithm that builds multiple pronunciation word models by expanding baseform pronunciations with a set of candidate phonological rules. We show how to train the probabilities of these phonological rules, and how to use these probabilities to assign pronunciation probabilities to words not seen in the training corpus. The algorithm we propose is an instance of the class of techniques we call Exploratory Computational Phonology. 1. INTRODUCTION One well-known difficulty in understanding speakerindependent continuous speech is variability in the pronunciation of words. This variability occurs across speakers and also across different contexts for a single speaker. In order to model this variation, recognition systems often use a richer lexicon in which each word has multiple pronunciations. Using a multiple-pronunciation lexicon requires setting a probability for each pronunciation. The minimal algorithm, for example, would assign each ... Gary N. Tajchman, Eric Fosler-Lussier, Daniel Jurafsky |
EUROSPEECH | 2 |
| 1994 | The berkeley restaurant projectabstractThis paper describes the architecture and performance of the Berkeley Restaurant Project (BeRP), a medium-vocabulary, speaker-independent, spontaneous continuous speech understanding system currently under development at ICSI. BeRP serves as a testbed for a number of our speech-related research projects, including robust feature extraction, connectionist phonetic likelihood estimation, automatic induction of multiplepronunciation lexicons, foreign accent detection and modeling, advanced language models, and lip-reading. In addition, it has proved quite usable in its function as a database frontend, even though many of our subjects are non-native speakers of English. Daniel Jurafsky, Chuck Wooters, Gary N. Tajchman, Jonathan Segal, Andreas Stolcke, Eric Fosler-Lussier, Nelson Morgan |
ICSLP | 6 |