VLDB 2026 Research / reviewers in the wild / expert
Izhak Shafran
dblp:66/3591
· DBLP profile ↗
67ranked-venue papers
9as first author
12since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 47 · 5 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 46 · 6 first-author · 7 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Retrieval Augmented End-to-End Spoken Dialog ModelsabstractWe recently developed a joint speech and language model (SLM [1]) which fuses a pretrained foundational speech model and a large language model (LLM), while preserving the in-context learning capability intrinsic to the pretrained LLM. In this paper, we apply SLM to dialog applications where the dialog states are inferred directly from the audio signal.Task-oriented dialogs often contain domain-specific entities, i.e., restaurants, hotels, train stations, and city names, which are difficult to recognize, however, critical for the downstream applications. Inspired by the RAG (retrieval-augmented generation) models, we propose a retrieval augmented SLM (ReSLM) that overcomes this weakness. We first train a retriever to retrieve text entities given audio inputs. The retrieved entities are then added as text inputs to the underlying LLM to bias model predictions. We evaluated ReSLM on speech MultiWoz task (DSTC-11 Challenge), and found that the retrieval augmentation boosts model performance, achieving joint goal accuracy (38.6% vs 32.7%), slot error rate (20.6% vs 24.8%) and ASR word error rate (5.5% vs 6.7%). While demonstrated on dialog state tracking, our approach is broadly applicable to speech tasks requiring custom contextual information or domain-specific entities. Mingqiu Wang, Izhak Shafran, Hagen Soltau, Wei Han 0002, Yuan Cao 0007, Laurent El Shafey |
ICASSP | 2 |
| 2024 | RoboVQA: Multimodal Long-Horizon Reasoning for RoboticsabstractWe present a scalable, bottom-up and intrinsically diverse data collection scheme that can be used for high-level reasoning with long and medium horizons and that has 2.2x higher throughput compared to traditional narrow top-down step-by-step collection. We collect realistic data by performing any user requests within the entirety of 3 office buildings and using multiple embodiments (robot, human, human with grasping tool). With this data, we show that models trained on all embodiments perform better than ones trained on the robot data only, even when evaluated solely on robot episodes. We explore the economics of collection costs and find that for a fixed budget it is beneficial to take advantage of the cheaper human collection along with robot collection. We release a large and highly diverse (29,520 unique instructions) dataset dubbed RoboVQA containing 829,502 (video, text) pairs for robotics-focused visual question answering. We also demonstrate how evaluating real robot experiments with an intervention mechanism enables performing tasks to completion, making it deployable with human oversight even if imperfect while also providing a single performance metric. We demonstrate a single video-conditioned model named RoboVQA-VideoCoCa trained on our dataset that is capable of performing a variety of grounded high-level reasoning tasks in broad realistic settings with a cognitive intervention rate 46% lower than the zeroshot state of the art visual language model (VLM) baseline and is able to guide real robots through long-horizon tasks. The performance gap with zero-shot state-of-the-art models indicates that a lot of grounded data remains to be collected for real-world deployment, emphasizing the critical need for scalable data collection approaches. Finally, we show that video VLMs significantly outperform single-image VLMs with an average error rate reduction of 19% across all VQA tasks. Thanks to video conditioning and dataset diversity, the model can be used as general video value functions (e.g. success and affordance) in situations where actions needs to be recognized rather than states, expanding capabilities and environment understanding for robots. Data and videos are available at robovqa.github.io Pierre Sermanet, Tianli Ding, Jeffrey Zhao, Fei Xia 0002, Debidatta Dwibedi, Keerthana Gopalakrishnan, Christine Chan, Gabriel Dulac-Arnold, Sharath Maddineni, Nikhil J. Joshi, Peter R. Florence, Wei Han 0002, Robert Baruch, Yao Lu 0006, Suvir Mirchandani, Peng Xu 0010, Pannag R. Sanketi, Karol Hausman, Izhak Shafran, Brian Ichter, Yuan Cao 0007 |
ICRA | 19 |
| 2023 | Detecting Speech Abnormalities With a Perceiver-Based Sequence Classifier that Leverages a Universal Speech ModelabstractWe propose a Perceiver-based sequence classifier to detect abnormalities in speech reflective of several neurological disorders. We combine this classifier with a Universal Speech Model (USM) that is trained on 12 million hours of diverse audio recordings. Our model compresses long sequences into a small set of class-specific latent representations and a factorized projection is used to predict different attributes of the disordered input speech. The benefit of our approach is that it allows us to model different regions of the input for different classes and is at the same time data efficient. We evaluated the proposed model extensively on a curated corpus from the Mayo Clinic. Our model outperforms standard transformer (80.9%) and perceiver (81.8%) models and achieves an average accuracy of 83.1%. With limited task-specific data, we find that pretraining is important and surprisingly pretraining with the un-related automatic speech recognition (ASR) task is also beneficial. Encodings from the middle layers provide a mix of both acoustic and phonetic information and achieve best prediction results compared to just using the final layer encodings (83.1% vs 79.6%). The results are promising and with further refinements may help clinicians detect speech abnormalities without needing access to highly specialized speech-language pathologists. Hagen Soltau, Izhak Shafran, Alex Ottenwess, Joseph R. Duffy, Rene L. Utianski, Leland Barnard, John L. Stricker, Daniela A. Wiepert, David T. Jones, Hugo Botha |
ASRU | 2 |
| 2023 | SLM: Bridge the Thin Gap Between Speech and Text Foundation ModelsabstractWe present a joint Speech and Language Model (SLM), a multitask, multilingual, and dual-modal model that takes advantage of pretrained foundational speech and language models. SLM freezes the pretrained foundation models to maximally preserves their capabilities, and only trains a simple adapter with just 1% (156M) of the foundation models’ parameters. This adaptation not only leads SLM to achieve strong performance on conventional tasks such as automatic speech recognition (ASR) and automatic speech translation (AST), but also unlocks the novel capability of zero-shot instruction-following for more diverse tasks. Given a speech input and a text instruction, SLM is able to perform unseen generation tasks including contextual biasing ASR using real-time context, dialog generation, speech continuation, and question answering. Our approach demonstrates that the representational gap between pretrained speech and language models is narrower than one would expect, and can be bridged by a simple adaptation mechanism. As a result, SLM is not only efficient to train, but also inherits strong capabilities already present in foundation models of different modalities. Mingqiu Wang, Wei Han 0002, Izhak Shafran, Zelin Wu, Chung-Cheng Chiu, Yuan Cao 0007, Nanxin Chen, Yu Zhang 0033, Hagen Soltau, Paul K. Rubenstein, Lukas Zilka, Golan Pundak, Nikhil Siddhartha, Johan Schalkwyk |
ASRU | 3 |
| 2023 | AnyTOD: A Programmable Task-Oriented Dialog SystemabstractJeffrey Zhao, Yuan Cao, Raghav Gupta, Harrison Lee, Abhinav Rastogi, Mingqiu Wang, Hagen Soltau, Izhak Shafran, Yonghui Wu. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Jeffrey Zhao, Yuan Cao 0007, Harrison Lee 0001, Abhinav Rastogi, Mingqiu Wang, Hagen Soltau, Izhak Shafran |
EMNLP | 8 |
| 2023 | ReAct: Synergizing Reasoning and Acting in Language Models
Shunyu Yao 0006, Jeffrey Zhao, Nan Du 0002, Izhak Shafran, Karthik Narasimhan, Yuan Cao 0007 |
ICLR | 5 |
| 2023 | Speech Aware Dialog System Technology Challenge (DSTC11)
Hagen Soltau, Izhak Shafran, Mingqiu Wang, Abhinav Rastogi, Jeffrey Zhao, Ye Jia, Wei Han 0002, Yuan Cao 0007, Aramys Miranda |
INTERSPEECH | 2 |
| 2023 | Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsabstractLanguage models are increasingly being deployed for general problem solving across a wide range of tasks, but are still confined to token-level, left-to-right decision-making processes during inference. This means they can fall short in tasks that require exploration, strategic lookahead, or where initial decisions play a pivotal role. To surmount these challenges, we introduce a new framework for language model inference, Tree of Thoughts (ToT), which generalizes over the popular Chain of Thought approach to prompting language models, and enables exploration over coherent units of text (thoughts) that serve as intermediate steps toward problem solving. ToT allows LMs to perform deliberate decision making by considering multiple different reasoning paths and self-evaluating choices to decide the next course of action, as well as looking ahead or backtracking when necessary to make global choices.
Our experiments show that ToT significantly enhances language models’ problem-solving abilities on three novel tasks requiring non-trivial planning or search: Game of 24, Creative Writing, and Mini Crosswords. For instance, in Game of 24, while GPT-4 with chain-of-thought prompting only solved 4\% of tasks, our method achieved a success rate of 74\%. Code repo with all prompts: https://github.com/princeton-nlp/tree-of-thought-llm. Shunyu Yao 0006, Jeffrey Zhao, Izhak Shafran, Thomas L. Griffiths 0001, Yuan Cao 0007, Karthik Narasimhan |
NeurIPS | 4 |
| 2022 | RNN Transducers for Named Entity Recognition with constraints on alignment for understanding medical conversations
Hagen Soltau, Izhak Shafran, Mingqiu Wang, Laurent El Shafey |
INTERSPEECH | 2 |
| 2022 | Unsupervised Slot Schema Induction for Task-oriented DialogabstractDian Yu, Mingqiu Wang, Yuan Cao, Izhak Shafran, Laurent Shafey, Hagen Soltau. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Mingqiu Wang, Yuan Cao 0007, Izhak Shafran, Laurent El Shafey, Hagen Soltau |
NAACL-HLT | 4 |
| 2021 | Word-Level Confidence Estimation for RNN TransducersabstractConfidence estimate is an often requested feature in applications such as medical transcription where errors can impact patient care and the confidence estimate could be used to alert medical professionals to verify potential errors in recognition. In this paper, we present a lightweight neural confidence model tailored for Automatic Speech Recognition (ASR) system with Recurrent Neural Network Transducers (RNN-T). Compared to other existing approaches, our model utilizes: (a) the time information associated with recognized words, which reduces the computational complexity, and (b) a simple and elegant trick for mapping between sub-word and word sequences. The mapping addresses the non-unique tokenization and token deletion problems while amplifying differences between confusable words. Through extensive empirical evaluations on two different long-form test sets, we demonstrate that the model achieves a performance of 0.4 Normalized Cross Entropy (NCE) and 0.05 Expected Calibration Error (ECE). It is robust across different ASR configurations, including target types (graphemes vs. morphemes), traffic conditions (streaming vs. non-streaming), and encoder types. We further discuss the importance of evaluation metrics to reflect practical applications and highlight the need for further work in improving Area Under the Curve (AUC) for Negative Precision Rate (NPV) and True Negative Rate (TNR). Mingqiu Wang, Hagen Soltau, Laurent El Shafey, Izhak Shafran |
ASRU | 4 |
| 2021 | Understanding Medical Conversations: Rich Transcription, Confidence Scores & Information ExtractionabstractIn this paper, we describe novel components for extracting clinically relevant information from medical conversations which will be available as Google APIs. We describe a transformer-based Recurrent Neural Network Transducer (RNN-T) model tailored for long-form audio, which can produce rich transcriptions including speaker segmentation, speaker role labeling, punctuation and capitalization. On a representative test set, we compare performance of RNN-T models with different encoders, units and streaming constraints. Our transformer-based streaming model performs at about 20% WER on the ASR task, 6% WDER on the diarization task, 43% SER on periods, 52% SER on commas, 43% SER on question marks and 30% SER on capitalization. Our recognizer is paired with a confidence model that utilizes both acoustic and lexical features from the recognizer. The model performs at about 0.37 NCE. Finally, we describe a RNN-T based tagging model. The performance of the model depends on the ontologies, with F-scores of 0.90 for medications, 0.76 for symptoms, 0.75 for conditions, 0.76 for diagnosis, and 0.61 for treatments. While there is still room for improvement, our results suggest that these models are sufficiently accurate for practical applications. Hagen Soltau, Mingqiu Wang, Izhak Shafran, Laurent El Shafey |
Interspeech | 3 |
| 2020 | The Medical Scribe: Corpus Development and Model Performance AnalysesabstractThere is a growing interest in creating tools to assist in clinical note generation using the audio of provider-patient encounters. Motivated by this goal and with the help of providers and medical scribes, we developed an annotation scheme to extract relevant clinical concepts. We used this annotation scheme to label a corpus of about 6k clinical encounters. This was used to train a state-of-the-art tagging model. We report ontologies, labeling results, model performances, and detailed analyses of the results. Our results show that the entities related to medications can be extracted with a relatively high accuracy of 0.90 F-score, followed by symptoms at 0.72 F-score, and conditions at 0.57 F-score. In our task, we not only identify where the symptoms are mentioned but also map them to canonical forms as they appear in the clinical notes. Of the different types of errors, in about 19-38% of the cases, we find that the model output was correct, and about 17-32% of the errors do not impact the clinical note. Taken together, the models developed in this work are more useful than the F-scores reflect, making it a promising approach for practical applications. Izhak Shafran, Nan Du 0002, Amanda Perry, Lauren Keyes, Mark Knichel, Ashley Domin, Yuhui Chen, Mingqiu Wang, Laurent El Shafey, Hagen Soltau, Justin S. Paul |
LREC | 1 |
| 2019 | Extracting Symptoms and their Status from Clinical ConversationsabstractThis paper describes novel models tailored for a new application, that of extracting the symptoms mentioned in clinical conversations along with their status. Lack of any publicly available corpus in this privacy-sensitive domain led us to develop our own corpus, consisting of about 3K conversations annotated by professional medical scribes. We propose two novel deep learning approaches to infer the symptom names and their status: (1) a new hierarchical span-attribute tagging (SA-T) model, trained using curriculum learning, and (2) a variant of sequence-to-sequence model which decodes the symptoms and their status from a few speaker turns within a sliding window over the conversation. This task stems from a realistic application of assisting medical providers in capturing symptoms mentioned by patients from their clinical conversations. To reflect this application, we define multiple metrics. From inter-rater agreement, we find that the task is inherently difficult. We conduct comprehensive evaluations on several contrasting conditions and observe that the performance of the models range from an F-score of 0.5 to 0.8 depending on the condition. Our analysis not only reveals the inherent challenges of the task, but also provides useful directions to improve the models. Nan Du 0002, Kai Chen 0010, Anjuli Kannan, Yuhui Chen, Izhak Shafran |
ACL (1) | 6 |
| 2019 | Learning to Infer Entities, Properties and their Relations from Clinical ConversationsabstractNan Du, Mingqiu Wang, Linh Tran, Gang Lee, Izhak Shafran. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Nan Du 0002, Mingqiu Wang, Gang Lee, Izhak Shafran |
EMNLP/IJCNLP (1) | 5 |
| 2019 | Joint Speech Recognition and Speaker Diarization via Sequence TransductionabstractSpeech applications dealing with conversations require not only recognizing the spoken words, but also determining who spoke when. The task of assigning words to speakers is typically addressed by merging the outputs of two separate systems, namely, an automatic speech recognition (ASR) system and a speaker diarization (SD) system. The two systems are trained independently with different objective functions. Often the SD systems operate directly on the acoustics and are not constrained to respect word boundaries and this deficiency is overcome in an ad hoc manner. Motivated by recent advances in sequence to sequence learning, we propose a novel approach to tackle the two tasks by a joint ASR and SD system using a recurrent neural network transducer. Our approach utilizes both linguistic and acoustic cues to infer speaker roles, as opposed to typical SD systems, which only use acoustic cues. We evaluated the performance of our approach on a large corpus of medical conversations between physicians and patients. Compared to a competitive conventional baseline, our approach improves word-level diarization error rate from 15.8% to 2.2%. Laurent El Shafey, Hagen Soltau, Izhak Shafran |
INTERSPEECH | 3 |
| 2018 | Complex Evolution Recurrent Neural Networks (ceRNNs)abstractUnitary Evolution Recurrent Neural Networks (uRNNs) have three attractive properties: (a) the unitary property, (b) the complex-valued nature, and (c) their efficient linear operators [1]. The literature so far does not address - how critical is the unitary property of the model? Furthermore, uRNNs have not been evaluated on large tasks. To study these shortcomings, we propose the complex evolution Recurrent Neural Networks (ceRNNs), which is similar to uRNNs but drops the unitary property selectively. On a simple multivariate linear regression task, we illustrate that dropping the constraints improves the learning trajectory. In copy memory task, ceRNNs and uRNNs perform identically, demonstrating that their superior performance over LSTMs is due to complex-valued nature and their linear operators. In a large scale real-world speech recognition, we find that pre-pending a uRNN degrades the performance of our baseline LSTM acoustic models, while pre-pending a ceRNN improves the performance over the baseline by 0.8% absolute WER. Izhak Shafran, Tom Bagby, R. J. Skerry-Ryan |
ICASSP | 1 |
| 2018 | Improvements to harmonic model for extracting better speech features in clinical applications
Meysam Asgari, Izhak Shafran |
Comput. Speech Lang. | 2 |
| 2017 | Adaptive Multichannel Dereverberation for Automatic Speech RecognitionabstractUtilizing an adaptive multichannel technique to mitigate reverberation present in received audio signals, prior to providing corresponding audio data to one or more additional component(s), such as automatic speech recognition (ASR) components. Implementations disclosed herein are “adaptive”, in that they utilize a filter, in the reverberation mitigation, that is online, causal and varies depending on characteristics of the input. Implementations disclosed herein are “multichannel”, in that a corresponding audio signal is received from each of multiple audio transducers (also referred to herein as “microphones”) of a client device, and the multiple audio signals (e.g., frequency domain representations thereof) are utilized in updating of the filter—and dereverberation occurs for audio data corresponding to each of the audio signals (e.g., frequency domain representations thereof) prior to the audio data being provided to ASR component(s) and/or other component(s). Joe Caroselli, Izhak Shafran, Arun Narayanan, Richard Rose |
INTERSPEECH | 2 |
| 2017 | Acoustic Modeling for Google Home
Bo Li 0028, Tara N. Sainath, Arun Narayanan, Joe Caroselli, Michiel Bacchiani, Ananya Misra, Izhak Shafran, Hasim Sak, Golan Pundak, Kean K. Chin, Khe Chai Sim, Ron J. Weiss, Kevin W. Wilson, Ehsan Variani, Chanwoo Kim 0001, Olivier Siohan, Mitch Weintraub, Erik McDermott, Richard Rose, Matt Shannon |
INTERSPEECH | 7 |
| 2017 | Multichannel Signal Processing With Deep Neural Networks for Automatic Speech RecognitionabstractMultichannel automatic speech recognition (ASR) systems commonly separate speech enhancement, including localization, beamforming, and postfiltering, from acoustic modeling. In this paper, we perform multichannel enhancement jointly with acoustic modeling in a deep neural network framework. Inspired by beamforming, which leverages differences in the fine time structure of the signal at different microphones to filter energy arriving from different directions, we explore modeling the raw time-domain waveform directly. We introduce a neural network architecture, which performs multichannel filtering in the first layer of the network, and show that this network learns to be robust to varying target speaker direction of arrival, performing as well as a model that is given oracle knowledge of the true target speaker direction. Next, we show how performance can be improved by factoring the first layer to separate the multichannel spatial filtering operation from a single channel filterbank which computes a frequency decomposition. We also introduce an adaptive variant, which updates the spatial filter coefficients at each time frame based on the previous inputs. Finally, we demonstrate that these approaches can be implemented more efficiently in the frequency domain. Overall, we find that such multichannel neural networks give a relative word error rate improvement of more than 5% compared to a traditional beamforming-based multichannel ASR system and more than 10% compared to a single channel waveform model. Tara N. Sainath, Ron J. Weiss, Kevin W. Wilson, Bo Li 0028, Arun Narayanan, Ehsan Variani, Michiel Bacchiani, Izhak Shafran, Andrew W. Senior, Kean K. Chin, Ananya Misra, Chanwoo Kim 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 8 |
| 2016 | Robust speech recognition using multivariate copula modelsabstractIn this paper, we continue our investigation into copula models for real-valued multivariate features with the goal of compensating for the mismatch in the training and the testing conditions. Previously, we reported results on UCI classification tasks where our method consistently outperformed other competing classifiers [1]. Here, we extend this work from classification to recognition and elaborate further on the mathematical properties of our models in the form of lemmas. We report results on the Aurora 4 automatic speech recognition (ASR) task which contains utterances with wide range of background noise that are not well represented in the training data. Our results show that the proposed copula-based models improve the accuracy by about 7% (11.6 vs 12.4) over a comparable baseline. Alireza Bayestehtashk, Izhak Shafran, Amir Babaeian |
ICASSP | 2 |
| 2016 | Reducing the Computational Complexity of Multimicrophone Acoustic Models with Integrated Feature Extraction
Tara N. Sainath, Arun Narayanan, Ron J. Weiss, Ehsan Variani, Kevin W. Wilson, Michiel Bacchiani, Izhak Shafran |
INTERSPEECH | 7 |
| 2016 | Complex Linear Projection (CLP): A Discriminative Approach to Joint Feature Extraction and Acoustic Modeling
Ehsan Variani, Tara N. Sainath, Izhak Shafran, Michiel Bacchiani |
INTERSPEECH | 3 |
| 2015 | Efficient and accurate multivariate class conditional densities using copulaabstractUnivariate densities can be modeled accurately and efficiently using nonparametric kernel density estimators, which unfortunately cannot be easily extended to the multivariate case. As an alternative, Gaussian mixture model is used to approximate underlying multivariate distributions, especially because its estimation is relatively straight forward through EM algorithm. However, the multivariate Gaussian mixture model imposes a particular form on the marginal, a Gaussian mixture model. This is a strong assumption on the marginal and is violated in many practical applications. We propose a simple generative classification model based on the copula model that takes advantage of the accuracy of the nonparametric univariate density estimator and the multivariate dependencies captured in the Gaussian mixture model, thus alleviating the aforementioned limitations. We compare the performance of our models with previous classification benchmarks from UCI repository and show that for the same number of parameters the proposed models consistently outperforms Gaussian mixture models. We find that these generative models perform as well or better than Support Vector Machine (SVM). Alireza Bayestehtashk, Izhak Shafran |
ICASSP | 2 |
| 2015 | Context dependent phone models for LSTM RNN acoustic modellingabstractLong Short Term Memory Recurrent Neural Networks (LSTM RNNs), combined with hidden Markov models (HMMs), have recently been show to outperform other acoustic models such as Gaussian mixture models (GMMs) and deep neural networks (DNNs) for large scale speech recognition. We argue that using multi-state HMMs with LSTM RNN acoustic models is an unnecessary vestige of GMM-HMM and DNN-HMM modelling since LSTM RNNs are able to predict output distributions through continuous, instead of piece-wise stationary, modelling of the acoustic trajectory. We demonstrate equivalent results for context independent whole-phone or 3-state models and show that minimum-duration modelling can lead to improved results. We go on to show that context dependent whole-phone models can perform as well as context dependent states, given a minimum duration model. Andrew W. Senior, Hasim Sak, Izhak Shafran |
ICASSP | 3 |
| 2015 | Fully automated assessment of the severity of Parkinson's disease from speech
Alireza Bayestehtashk, Meysam Asgari, Izhak Shafran, James McNames |
Comput. Speech Lang. | 3 |
| 2014 | Automatic measurement of affective valence and arousal in speechabstractMethods are proposed for measuring affective valence and arousal in speech. The methods apply support vector regression to prosodic and text features to predict human valence and arousal ratings of three stimulus types: speech, delexicalized speech, and text transcripts. Text features are extracted from transcripts via a lookup table listing per-word valence and arousal values and computing per-utterance statistics from the per-word values. Prediction of arousal ratings of delexicalized speech and of speech from prosodic features was successful, with accuracy levels not far from limits set by the reliability of the human ratings. Prediction of valence for these stimulus types as well as prediction of both dimensions for text stimuli proved more difficult, even though the corresponding human ratings were as reliable. Text based features did add, however, to the accuracy of prediction of valence for speech stimuli. We conclude that arousal of speech can be measured reliably, but not valence, and that improving the latter requires better lexical features. Meysam Asgari, Géza Kiss, Jan P. H. van Santen, Izhak Shafran, Xubo Song |
ICASSP | 4 |
| 2014 | Discriminative pronunciation modeling for dialectal speech recognitionabstractSpeech recognizers are typically trained with data from a stan-dard dialect and do not generalize to non-standard dialects. Mis-match mainly occurs in the acoustic realization of words, which is represented by acoustic models and pronunciation lexicon. Standard techniques for addressing this mismatch are generative in nature and include acoustic model adaptation and expansion of lexicon with pronunciation variants, both of which have lim-ited effectiveness. We present a discriminative pronunciation model whose parameters are learned jointly with parameters from the language models. We tease apart the gains from mod-eling the transitions of canonical phones, the transduction from surface to canonical phones, and the language model. We report experiments on African American Vernacular English (AAVE) using NPR’s StoryCorps corpus. Our models improve the per-formance over the baseline by about 2.1 % on AAVE, of which 0.6 % can be attributed to the pronunciation model. The model learns the most relevant phonetic transformations for AAVE speech. Index Terms: large vocabulary speech recognition, dialec-tal speech recognition, pronunciation modeling, discriminative training 1. Maider Lehr, Kyle Gorman, Izhak Shafran |
INTERSPEECH | 3 |
| 2014 | Applications of Lexicographic Semirings to Problems in Speech and Language ProcessingabstractThis paper explores lexicographic semirings and their application to problems in speech and language processing. Specifically, we present two instantiations of binary lexicographic semirings, one involving a pair of tropical weights, and the other a tropical weight paired with a novel string semiring we term the categorial semiring. The first of these is used to yield an exact encoding of backoff models with epsilon transitions. This lexicographic language model semiring allows for off-line optimization of exact models represented as large weighted finite-state transducers in contrast to implicit (on-line) failure transition representations. We present empirical results demonstrating that, even in simple intersection scenarios amenable to the use of failure transitions, the use of the more powerful lexicographic semiring is competitive in terms of time of intersection. The second of these lexicographic semirings is applied to the problem of extracting, from a lattice of word sequences tagged for part of speech, only the single best-scoring part of speech tagging for each word sequence. We do this by incorporating the tags as a categorial weight in the second component of a 〈Tropical, Categorial〉 lexicographic semiring, determinizing the resulting word lattice acceptor in that semiring, and then mapping the tags back as output labels of the word lattice transducer. We compare our approach to a competing method due to Povey et al. (2012). Richard Sproat, Mahsa Yarmohammadi, Izhak Shafran, Brian Roark |
Comput. Linguistics | 3 |
| 2014 | Inferring social nature of conversations from words: Experiments on a corpus of everyday telephone conversations
Anthony P. Stark, Izhak Shafran, Jeffrey A. Kaye |
Comput. Speech Lang. | 2 |
| 2013 | Adaptive H-Extrema for Automatic Immunogold Particle Detection
Guillaume Thibault, Kristiina Iljin, Christopher Arthur, Izhak Shafran, Joe W. Gray |
CIARP (2) | 4 |
| 2013 | Parsimonious multivariate copula model for density estimationabstractThe most common approaches for estimating multivariate density assume a parametric form for the joint distribution. The choice of this parametric form imposes constraints on the marginal distributions. Copula models disentangle the choice of marginals from the joint distributions, making it a powerful model for multivariate density estimation. However, so far, they have been widely studied mostly for low dimensional multivariate. In this paper, we investigate a popular Copula model - the Gaussian Copula model - for high dimensional settings. They however require estimation of a full correlation matrix which can cause data scarcity in this setting. One approach to address this problem is to impose constraints on the parameter space. In this paper, we present Toeplitz correlation structure to reduce the number of Gaussian Copula parameter. To increase the flexibility of our model, we also introduce mixture of Gaussian Copula as a natural extension of the Gaussian Copula model. Through empirical evaluation of likelihood on held-out data, we study the trade-off between correlation constraints and mixture flexibility, and report results on wine data sets from the UCI Repository as well as our corpus of monkey vocalizations. We find that mixture of Gaussian Copula with Toeplitz correlation structure models the data consistently better than Gaussian mixture models with equivalent number of parameters. Alireza Bayestehtashk, Izhak Shafran |
ICASSP | 2 |
| 2013 | Robust and accurate features for detecting and diagnosing autism spectrum disordersabstract). From the fundamental frequencies and the reconstructed noise-free signal, we compute other derived features such as Harmonic-to-Noise Ratio (HNR), shimmer, and jitter. In previous work, we found that these features detect voiced segments and speech more accurately than other algorithms and that they are useful in rating the severity of a subject's Parkinson's disease [3]. Here, we employ these features, along with standard features such as energy, cepstral, and spectral features. With these features, we detect ASD using a regression and identify the sub-type using a classifier. We find that our features improve the performance, measured in terms of unweighted average recall (UAR), of detecting autism spectrum disorder by 2.3% and classifying the disorder into four categories by 2.8% over the baseline results. Meysam Asgari, Alireza Bayestehtashk, Izhak Shafran |
INTERSPEECH | 3 |
| 2013 | Improving the accuracy and the robustness of harmonic model for pitch estimationabstractAccurate and robust estimation of pitch plays a central role in speech processing. Various methods in time, frequency and cepstral domain have been proposed for generating pitch candidates. Most algorithms excel when the background noise is minimal or for specific types of background noise. In this work, our aim is to improve the robustness and accuracy of pitch estimation across a wide variety of background noise conditions. For this we have chosen to adopt, the harmonic model of speech, a model that has gained considerable attention recently. We address two major weakness of this model. The problem of pitch halving and doubling, and the need to specify the number of harmonics. We exploit the energy of frequency in the neighborhood to alleviate halving and doubling. Using a model complexity term with a BIC criterion, we chose the optimal number of harmonics. We evaluated our proposed pitch estimation method with other state of the art techniques on Keele data set in terms of gross pitch error and fine pitch error. Through extensive experiments on several noisy conditions, we demonstrate that the proposed improvements provide substantial gains over other popular methods under different noise levels and environments. Index Terms: fundamental frequency estimation, robust pitch estimation Meysam Asgari, Izhak Shafran |
INTERSPEECH | 2 |
| 2013 | Discriminative Joint Modeling of Lexical Variation and Acoustic Confusion for Automated Narrative Retelling Assessment
Maider Lehr, Izhak Shafran, Emily Tucker Prud'hommeaux, Brian Roark |
HLT-NAACL | 2 |
| 2012 | Semi-supervised discriminative language modeling for Turkish ASRabstractWe present our work on semi-supervised learning of discriminative language models where the negative examples for sentences in a text corpus are generated using confusion models for Turkish at various granularities, specifically, word, sub-word, syllable and phone levels. We experiment with different language models and various sampling strategies to select competing hypotheses for training with a variant of the perceptron algorithm. We find that morph-based confusion models with a sample selection strategy aiming to match the error distribution of the baseline ASR system gives the best performance. We also observe that substituting half of the supervised training examples with those obtained in a semi-supervised manner gives similar results. Arda Çelebi, Hasim Sak, Erinç Dikici, Murat Saraclar, Maider Lehr, Emily Tucker Prud'hommeaux, Puyang Xu, Nathan Glenn, Damianos Karakos, Sanjeev Khudanpur, Brian Roark, Kenji Sagae, Izhak Shafran, Dan Bikel, Chris Callison-Burch, Yuan Cao 0007, Keith B. Hall, Eva Hasler, Philipp Koehn, Adam Lopez, Matt Post, Darcey Riley |
ICASSP | 13 |
| 2012 | Hallucinated n-best lists for discriminative language modelingabstractThis paper investigates semi-supervised methods for discriminative language modeling, whereby n-best lists are “hallucinated” for given reference text and are then used for training n-gram language models using the perceptron algorithm. We perform controlled experiments on a very strong baseline English CTS system, comparing three methods for simulating ASR output, and compare the results with training with “real” n-best list output from the baseline recognizer. We find that methods based on extracting phrasal cohorts - similar to methods from machine translation for extracting phrase tables - yielded the largest gains of our three methods, achieving over half of the WER reduction of the fully supervised methods. Kenji Sagae, Maider Lehr, Emily Tucker Prud'hommeaux, Puyang Xu, Nathan Glenn, Damianos Karakos, Sanjeev Khudanpur, Brian Roark, Murat Saraclar, Izhak Shafran, Dan Bikel, Chris Callison-Burch, Yuan Cao 0007, Keith B. Hall, Eva Hasler, Philipp Koehn, Adam Lopez, Matt Post, Darcey Riley |
ICASSP | 10 |
| 2012 | Continuous space discriminative language modelingabstractDiscriminative language modeling is a structured classification problem. Log-linear models have been previously used to address this problem. In this paper, the standard dot-product feature representation used in log-linear models is replaced by a non-linear function parameterized by a neural network. Embeddings are learned for each word and features are extracted automatically through the use of convolutional layers. Experimental results show that as a stand-alone model the continuous space model yields significantly lower word error rate (1% absolute), while having a much more compact parameterization (60%-90% smaller). If the baseline scores are combined, our approach performs equally well. Puyang Xu, Sanjeev Khudanpur, Maider Lehr, Emily Tucker Prud'hommeaux, Nathan Glenn, Damianos Karakos, Brian Roark, Kenji Sagae, Murat Saraclar, Izhak Shafran, Dan Bikel, Chris Callison-Burch, Yuan Cao 0007, Keith B. Hall, Eva Hasler, Philipp Koehn, Adam Lopez, Matt Post, Darcey Riley |
ICASSP | 10 |
| 2012 | Deriving conversation-based features from unlabeled speech for discriminative language modeling
Damianos Karakos, Brian Roark, Izhak Shafran, Kenji Sagae, Maider Lehr, Emily Tucker Prud'hommeaux, Puyang Xu, Nathan Glenn, Sanjeev Khudanpur, Murat Saraclar, Dan Bikel, Mark Dredze, Chris Callison-Burch, Yuan Cao 0007, Keith B. Hall, Eva Hasler, Philipp Koehn, Adam Lopez, Matt Post, Darcey Riley |
INTERSPEECH | 3 |
| 2012 | Fully Automated Neuropsychological Assessment for Detecting Mild Cognitive ImpairmentabstractWe present an end-to-end system for automatically scoring spoken responses to a narrative recall test administered to seniors when screening for cognitive impairment. In Wechsler Logical Memory (WLM) test, a patient listens to a brief narrative, then retells the story once immediately and again after a brief delay. We transcribe the retellings automatically using an ASR system, align the transcripts to the source narrative, extract features that replicate the standard clinical scoring method, and then use the features for automatic assessment using a classifier. On a test corpus of 72 subjects, we empirically evaluate different ASR adaptation strategies and analyze the errors with respect to clinical assessment. Despite imperfect recognition, the system presented here yields classification accuracy comparable to that of manually assigned scores. Our results show that automatic assessment of neuropsychological tests such as the WLM is practical for screening large cohorts. Index Terms: clinical diagnostics, classifying mild cognitive impairment Maider Lehr, Emily Tucker Prud'hommeaux, Izhak Shafran, Brian Roark |
INTERSPEECH | 3 |
| 2012 | Interspeech Pathology Challenge: Investigations into Speaker and Sentence Specific EffectsabstractIn this paper, we report our experiments on Interspeech 2012 Speaker Trait Pathology challenge task [2]. Specifically, we investigate two factors that impact the acoustic properties of the utterances collected in this task. Although the task treats utterances as independent data points, multiple utterances are recorded from individual speakers. Furthermore, the utterances correspond to readings of 17 given written sentences. In one experiment, we attempt to reduce variation due to speaker through dimensionality reduction. While these experiments showed promising results on development set, the performance did not translate to the evaluation test. In another, we learn classifiers conditioned on the sentences to capture sentence-specific signatures. This approach showed improved performance over the baseline on development set and the improvement translated to marginal gains on evaluation set. These experiments demonstrates the need to pay attention to the independence assumptions while collecting and defining clinical tasks. Anthony P. Stark, Alireza Bayestehtashk, Meysam Asgari, Izhak Shafran |
INTERSPEECH | 4 |
| 2012 | Hello, Who is Calling?: Can Words Reveal the Social Nature of Conversations?
Anthony P. Stark, Izhak Shafran, Jeffrey A. Kaye |
HLT-NAACL | 2 |
| 2012 | Robust detection of voiced segments in samples of everyday conversations using unsupervised HMMSabstractWe investigate methods for detecting voiced segments in everyday conversations from ambient recordings. Such recordings contain high diversity of background noise, making it difficult or infeasible to collect representative labelled samples for estimating noise-specific HMM models. The popular utility get-f0 and its derivatives compute normalized cross-correlation for detecting voiced segments, which unfortunately is sensitive to different types of noise. Exploiting the fact that voiced speech is not just periodic but also rich in harmonic, we model voiced segments by adopting harmonic models, which have recently gained considerable attention. In previous work, the parameters of the model were estimated independently for each frame using maximum likelihood criterion. However, since the distribution of harmonic coefficients depend on articulators of speakers, we estimate the model parameters more robustly using a maximum a posteriori criterion. We use the likelihood of voicing, computed from the harmonic model, as an observation probability of an HMM and detect speech using this unsupervised HMM. The one caveat of the harmonic model is that they fail to distinguish speech from other stationary harmonic noise. We rectify this weakness by taking advantage of the non-stationary property of speech. We evaluate our models empirically on a task of detecting speech on a large corpora of everyday speech and demonstrate that these models perform significantly better than standard voice detection algorithm employed in popular tools. Meysam Asgari, Izhak Shafran, Alireza Bayestehtashk |
SLT | 2 |
| 2012 | Discriminative Language Modeling With Linguistic and Statistically Derived FeaturesabstractThis paper focuses on integrating linguistically motivated and statistically derived information into language modeling. We use discriminative language models (DLMs) as a complementary approach to the conventional$n$-gram language models to benefit from discriminatively trained parameter estimates for overlapping features. In our DLM approach, relevant information is encoded as features. Feature weights are discriminatively trained using training examples and used to re-rank the$N$-best hypotheses of the baseline automatic speech recognition (ASR) system. In addition to presenting a more complete picture of previously proposed feature sets that extract implicit information available at lexical and sub-lexical levels using both linguistic and statistical approaches, this paper attempts to incorporate semantic information in the form of topic sensitive features. We explore linguistic features to incorporate complex morphological and syntactic language characteristics of Turkish, an agglutinative language with rich morphology, into language modeling. We also apply DLMs to our sub-lexical-based ASR system where the vocabulary is composed of sub-lexical units. Obtaining implicit linguistic information from sub-lexical hypotheses is not as straightforward as word hypotheses, so we use statistical methods to derive useful information from sub-lexical units. DLMs with linguistic and statistical features yield significant, 0.8%–1.1% absolute, improvements over our baseline word-based and sub-word-based ASR systems. The explored features can be easily extended to DLM for other languages . Ebru Arisoy, Murat Saraclar, Brian Roark, Izhak Shafran |
IEEE Trans. Speech Audio Process. | 4 |
| 2011 | Efficient determinization of tagged word lattices using categorial and lexicographic semiringsabstractSpeech and language processing systems routinely face the need to apply finite state operations (e.g., POS tagging) on results from intermediate stages (e.g., ASR output) that are naturally represented in a compact lattice form. Currently, such needs are met by converting the lattices into linear sequences (n-best scoring sequences) before and after applying the finite state operations. In this paper, we eliminate the need for this unnecessary conversion by addressing the problem of picking only the single-best scoring output labels for every input sequence. For this purpose, we define a categorial semiring that allows determinzation over strings and incorporate it into a 〈Tropical, Categorial〉 lexicographic semiring. Through examples and empirical evaluations we show how determinization in this lexicographic semiring produces the desired output. The proposed solution is general in nature and can be applied to multi-tape weighted transducers that arise in many applications. Izhak Shafran, Richard Sproat, Mahsa Yarmohammadi, Brian Roark |
ASRU | 1 |
| 2011 | Supervised and unsupervised feature selection for inferring social nature of telephone conversations from their contentabstractThe ability to reliably infer the nature of telephone conversations opens up a variety of applications, ranging from designing context-sensitive user interfaces on smartphones, to providing new tools for social psychologists and social scientists to study and understand social life of different subpopulations within different contexts. Using a unique corpus of everyday telephone conversations collected from eight residences over the duration of a year, we investigate the utility of popular features, extracted solely from the content, in classifying business-oriented calls from others. Through feature selection experiments, we find that the discrimination can be performed robustly for a majority of the calls using a small set of features. Remarkably, features learned from unsupervised methods, specifically latent Dirichlet allocation, perform almost as well as with as those from supervised methods. The unsupervised clusters learned in this task shows promise of finer grain inference of social nature of telephone conversations. Anthony P. Stark, Izhak Shafran, Jeffrey A. Kaye |
ASRU | 2 |
| 2011 | Discriminatively estimated discrete, parametric and smoothed-discrete duration models for speech recognitionabstractDuration of phonemic segments provide important cues for distinguishing words in languages such as Arabic. Recently, we proposed a discriminatively estimated joint acoustic, duration and language model for large vocabulary speech recognition. In that work, we found simple discrete models to be effective for modeling duration, albeit they were neither smoothed nor parsimonious. These limitations are ad dressed here with two alternative models parametric and smoothed-discrete models. Unlike previous work on para metric duration model, we estimate their parameters discriminatively and derive an analytical expression for estimating the parameters of a log-normal distribution using a recent approach. On a large vocabulary Arabic task, we empirically evaluated different segmental units and durations models. Our results show bigrams of clustered states modeled with smoothed-discrete duration models are relatively more accurate and efficient than other models considered. Maider Lehr, Izhak Shafran |
ICASSP | 2 |
| 2011 | Learning a Discriminative Weighted Finite-State Transducer for Speech RecognitionabstractWeighted finite-state transducers (WFSTs) have been widely adopted as efficient representations of a general speech recognition model. The WFST for speech recognizer is typically assembled or composed from the several components-the language model, the pronunciation mapping and the acoustic model-which are estimated separately without any end-to-end optimization. This paper examines how the weights of such transducers can be learned in a manner that captures the interaction between the components. The paths in the transducer are represented asn-grams defined over the input and output sequences whose linear weights are learned using a discriminative criterion. The resulting linear model factors into two weighted finite-state acceptors (WFSAs) which can be applied as corrections to the input and the output side of the initial WFST. This formulation allows duration cues to be incorporated seamlessly. Empirical results on a large vocabulary Arabic GALE task demonstrate that the proposed model improves word error rate substantially, with a gain of 1.5%-1.7% absolute. Through a series of experiments, we analyze the contributions from and interactions between acoustic, duration, and language components to find that duration cues play an important role in a large-vocabulary Arabic speech recognition task. Although this paper focuses on speech recognition, the proposed framework for learning the weights of a finite transducer is more general in nature and can be applied to other tasks such as utterance classification. Maider Lehr, Izhak Shafran |
IEEE Trans. Speech Audio Process. | 2 |
| 2010 | Syntactic and sub-lexical features for Turkish discriminative language modelsabstractThis paper investigates syntactic and sub-lexical features in Turkish discriminative language models (DLMs). DLM is a feature-based language modeling approach. It reranks the ASR output with discriminatively trained feature parameters. Syntactic information is incorporated into DLM as part-of-speech (PoS) tag n-gram features and head-to-head dependency relations. Sub-lexical units are first utilized as language modeling units in the baseline recognizer. Then, sub-lexical features are used to rerank the sub-lexical hypotheses. We explore features, similar to syntactic features, on sub-lexical units to reveal the implicit morpho-syntactic information conveyed by these units. We find out that DLM yields more improvement for sub-lexical units than for words. Basic sub-lexical n-gram features result in 0.6% reduction over the baseline and morpho-syntactic features yield an additional 0.4% reduction on the test set. Ebru Arisoy, Murat Saraclar, Brian Roark, Izhak Shafran |
ICASSP | 4 |
| 2010 | Discriminatively estimated joint acoustic, duration, and language model for speech recognitionabstractWe introduce a discriminative model for speech recognition that integrates acoustic, duration and language components. In the framework of finite state machines, a general model for speech recognition G is a finite state transduction from acoustic state sequences to word sequences (e.g., search graph in many speech recognizers). The lattices from a baseline recognizer can be viewed as an a posteriori version of G after having observed an utterance. So far, discriminative language models have been proposed to correct the output side of G and is applied on the lattices. The acoustic state sequences on the input side of these lattice can also be exploited to improve the choice of the best hypotheses through the lattice. Taking this view, the model proposed in this paper jointly estimates the parameters for acoustic and language components in a discriminative setting. The resulting model can be factored as corrections for the input and the output sides of the general model G. This formulation allows us to incorporate duration cues seamlessly. Empirical results on a large vocabulary Arabic GALE task demonstrate that the proposed model improves word error rate substantially, with a gain of 1.6% absolute. Through a series of experiments we analyze the contributions from and interactions between acoustic, duration and language components to find that duration cues play an important role in Arabic task. Maider Lehr, Izhak Shafran |
ICASSP | 2 |
| 2009 | Classifying clear and conversational speech based on acoustic featuresabstractThis paper reports an investigation of features relevant for classifying two speaking styles, namely, conversational speaking style and clear (e.g. hyper-articulated) speaking style. Spectral and prosodic features were automatically extracted from speech and classified using decision tree classifiers and multilayer perceptrons to achieve accuracies of about 71 % and 77% respectively. More interestingly, we found that out of the 56 features only about 9 features are needed to capture the most predictive power. While perceptual studies have shown that spectral cues are more useful than prosodic features for intelligibility [1], here we find prosodic features are more important for classification. Index Terms: binary classification, acoustic features, decision tree classifier, multilayer perceptron, Akiko Amano-Kusumoto, John-Paul Hosom, Izhak Shafran |
INTERSPEECH | 3 |
| 2008 | Discriminative n-gram language modeling for TurkishabstractIn this paper Discriminative Language Models (DLMs) are applied to the Turkish Broadcast News transcription task. Turkish presents a challenge to Automatic Speech Recognition (ASR) systems due to its rich morphology. Therefore, in addition to word n-gram features, morphology based features like root n-grams and inflectional group n-grams are incorporated into DLMs in order to improve the language models. Various feature sets provide reductions in the word error rate (WER). Our best result is obtained with the inflectional group n-gram features. 1.0 % absolute improvement is achieved over the baseline model and this improvement is statistically significant at p<0.001 as measured by the NIST MAPSSWE significance test. Index Terms: discriminative language modeling, speech recognition, agglutinative languages Ebru Arisoy, Brian Roark, Izhak Shafran, Murat Saraclar |
INTERSPEECH | 3 |
| 2007 | Exploiting prosody for PCFGs with latent annotationsabstractWe propose novel methods for integrating prosody in syntax using generative models. By adopting a grammar whose constituents have latent annotations, the influence of prosody on syntax can be learned from data. In one method, prosody is utilized to seed the latent annotations of a grammar which is then refined using EM iterations. In an orthogonal approach, we integrate prosody into grammar more explicitly using a model that jointly observes words and associated prosody. We evaluate the two methods by parsing speech data from the Switchboard corpus. The results are compared against baseline results from a model that does not use prosody. The experiments show that prosody improves a grammar in terms of accuracy as well as the parsimonious use of parameters. 1. Markus Dreyer, Izhak Shafran |
INTERSPEECH | 2 |
| 2007 | The SRI/OGI 2006 spoken term detection systemabstractThis paper describes the system developed jointly at SRI and OGI for participation in the 2006 NIST Spoken Term Detection (STD) evaluation. We participated in the three genres of the English track: Broadcast News (BN), Conversational Telephone Speech (CTS), and Conference Meetings (MTG). The system consists of two phases. First, audio indexing, an offline phase, converts the input speech waveform into a searchable index. Second, term retrieval, possibly an online phase, returns a ranked list of occurrences for each search term. We used a word-based indexing approach, obtained with SRI’s large vocabulary Speech-to-Text (STT) system. Apart from describing the submitted system and its performance on the NIST evaluation metric, we study the tradeoffs between performance and system design. We examine performance versus indexing speed, effectiveness of different index ranking schemes on the NIST score, and the utility of approaches to deal with out-of-vocabulary (OOV) terms. Dimitra Vergyri, Izhak Shafran, Andreas Stolcke, Venkata Ramana Rao Gadde, Murat Akbacak, Brian Roark, Wen Wang 0001 |
INTERSPEECH | 2 |
| 2006 | PCFGs with Syntactic and Prosodic Indicators of Speech RepairsabstractA grammatical method of combining two kinds of speech repair cues is presented. One cue, prosodic disjuncture, is detected by a decision tree-based ensemble classifier that uses acoustic cues to identify where normal prosody seems to be interrupted (Lickley, 1996). The other cue, syntactic parallelism, codifies the expectation that repairs continue a syntactic category that was left unfinished in the reparandum (Levelt, 1983). The two cues are combined in a Treebank PCFG whose states are split using a few simple tree transformations. Parsing performance on the Switchboard and Fisher corpora suggests that these two cues help to locate speech repairs in a synergistic way. John Hale, Izhak Shafran, Lisa Yung, Bonnie J. Dorr, Mary P. Harper, Anna Krasnyanskaya, Matthew Lease, Yang Liu 0004, Brian Roark, Matthew G. Snover, Robin Stewart |
ACL | 2 |
| 2006 | Corrective Models for Speech Recognition of Inflected Languages
Izhak Shafran, Keith B. Hall |
EMNLP | 1 |
| 2006 | Reranking for Sentence Boundary Detection in Conversational SpeechabstractWe present a reranking approach to sentence-like unit (SU) boundary detection, one of the EARS metadata extraction tasks. Techniques for generating relatively small n-best lists with high oracle accuracy are presented. For each candidate, features are derived from a range of information sources, including the output of a number of parsers. Our approach yields significant improvements over the best performing system from the NIST RT-04F community evaluation Brian Roark, Yang Liu 0004, Mary P. Harper, Robin Stewart, Matthew Lease, Matthew G. Snover, Izhak Shafran, Bonnie J. Dorr, John Hale, Anna Krasnyanskaya, Lisa Yung |
ICASSP (1) | 7 |
| 2006 | Discriminative Classifiers for Language RecognitionabstractMost language recognition systems consist of a cascade of three stages: (1) tokenizers that produce parallel phone streams, (2) phonotactic models that score the match between each phone stream and the phonotactic constraints in the target language, and (3) a final stage that combines the scores from the parallel streams appropriately [1]. This paper reports a series of contrastive experiments to assess the impact of replacing the second and third stages with large-margin discriminative classifiers. In addition, it investigates how sounds that are not represented in the tokenizers of the first stage can be approximated with composite units that utilize cross-stream dependencies obtained via multi-string alignments. This leads to a discriminative framework that can potentially incorporate a richer set of features such as prosodic and lexical cues. Experiments are reported on the NIST LRE 1996 and 2003 task and the results show that the new techniques give substantial gains over a competitive PPRLM baseline. Izhak Shafran, Jean-Luc Gauvain |
ICASSP (1) | 2 |
| 2006 | SParseval: Evaluation Metrics for Parsing Speech
Brian Roark, Mary P. Harper, Eugene Charniak, Bonnie J. Dorr, Mark Johnson 0001, Jeremy G. Kahn, Yang Liu 0004, Mari Ostendorf, John Hale, Anna Krasnyanskaya, Matthew Lease, Izhak Shafran, Matthew G. Snover, Robin Stewart, Lisa Yung |
LREC | 12 |
| 2005 | A Comparison of Classifiers for Detecting Emotion from SpeechabstractAccurate detection of emotion from speech has clear benefits for the design of more natural human-machine speech interfaces or for the extraction of useful information from large quantities of speech data. The task consists of assigning, out of a fixed set, an emotion category, e.g., anger, fear, or satisfaction, to a speech utterance. In recent work, several classifiers have been proposed for automatic detection of a speaker's emotion using spoken words as the input. These classifiers were designed independently and tested on separate corpora, making it difficult to compare their performance. This paper presents three classifiers, two popular classifiers from the literature modeling the word content via n-gram sequences, one based on an interpolated language model, another on a mutual information-based feature-selection approach, and compares them with a discriminant kernel-based technique that we recently adopted. We have implemented these three classification algorithms and evaluated their performance by applying them to a corpus collected from a spoken-dialog system that was widely deployed across the USA. The results show that our kernel-based classifier achieves an accuracy of 80.6%, and outperforms both the interpolated language model classifier, which achieved a classification accuracy of 70.1%, and the classifier using mutual information-based feature selection (78.8%). Izhak Shafran, Mehryar Mohri |
ICASSP (1) | 1 |
| 2005 | Accent detection and speech recognition for Shanghai-accented MandarinabstractAs speech recognition systems are used in ever more applications, it is crucial for the systems to be able to deal with accented speakers. Various techniques, such as acoustic model adaptation and pronunciation adaptation, have been reported to improve the recognition of non-native or accented speech. In this paper, we propose a new approach that combines accent detection, accent discriminative acoustic features, acoustic adaptation and model selection for accented Chinese speech recognition. Experimental results show that this approach can improve the recognition of accented speech. 1. Yanli Zheng, Richard Sproat, Liang Gu, Izhak Shafran, Haolang Zhou, Daniel Jurafsky, Rebecca Starr, Su-Youn Yoon |
INTERSPEECH | 4 |
| 2005 | Automatic Detection and Segmentation of Robot-Assisted Surgical Motions
Henry C. Lin 0001, Izhak Shafran, Todd E. Murphy, Allison M. Okamura, David D. Yuh, Gregory D. Hager |
MICCAI | 2 |
| 2004 | Task-specific minimum Bayes-risk decoding using learned edit distanceabstractThis paper extends the minimum Bayes-risk framework to incorporate a loss function specific to the task and the ASR system. The errors are modeled as a noisy channel and the parameters are learned from the data. The resulting loss function is used in the risk criterion for decoding. Experiments on a large vocabulary conversational speech recognition system demonstrate significant gains of about 1% absolute over MAP hypothesis and about 0.6% absolute over untrained loss function. The approach is general enough to be applicable to other sequence recognition problems such as in Optical Character Recognition (OCR) and in analysis of biological sequences. Izhak Shafran, William J. Byrne |
INTERSPEECH | 1 |
| 2003 | Robust speech detection and segmentation for real-time ASR applicationsabstractThis paper provides a solution for robust speech detection that can be applied across a variety of tasks. The solution is based on an algorithm that performs non-parametric estimation of the background noise spectrum using minimum statistics of the smoothed short-time Fourier transform (STFT). It is shown that the new algorithm can operate effectively under varying signal-to-noise ratios. Results are reported on two tasks - HMIHY and SPINE - which differ in their speaking style, background noise type and bandwidth. With a computational cost of less than 2% real-time on a 1GHz P-3 machine and a latency of 400 ms, it is suitable for real-time ASR applications. Izhak Shafran, Richard Rose |
ICASSP (1) | 1 |
| 2003 | Acoustic model clustering based on syllable structure
Izhak Shafran, Mari Ostendorf |
Comput. Speech Lang. | 1 |
| 2000 | Use of higher level linguistic structure in acoustic modeling for speech recognitionabstractCurrent speech recognition systems perform poorly on conversational speech as compared to read speech, largely because of the additional acoustic variability observed in conversational speech. Our hypothesis is that there are systematic effects, related to higher level structures, that are not being captured in the current acoustic models. In this paper we describe a method to extend standard clustering to incorporate such features in estimating acoustic models. We report recognition improvements obtained on the Switchboard task over triphones and pentaphones by the use of word- and syllable-level features. In addition, we report preliminary studies on clustering with prosodic information. Izhak Shafran, Mari Ostendorf |
ICASSP | 1 |