EDBT 2026 Demo / reviewers in the wild / expert
Ehsan Variani
dblp:68/10649
· DBLP profile ↗
31ranked-venue papers
15as first author
13since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 27 · 13 first-author · 11 since 2021Artificial intelligence and machine learning · 17 · 7 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Massive Sound Embedding Benchmark (MSEB)abstractAudio is a critical component of multimodal perception, and any truly intelligent system must demonstrate a wide range of auditory capabilities. These capabilities include transcription, classification, retrieval, reasoning, segmentation, clustering, reranking, and reconstruction. Fundamentally, each task involves transforming a raw audio signal into a meaningful 'embedding'—be it a single vector, a sequence of continuous or discrete representations, or another structured form—which then serves as the basis for generating the task's final response. To accelerate progress towards robust machine auditory intelligence, we present the Massive Sound Embedding Benchmark (MSEB): an extensible framework designed to evaluate the auditory components of any multimodal system. In its first release, MSEB offers a comprehensive suite of eight core tasks, with more planned for the future, supported by diverse datasets, including the new, large-scale Simple Voice Questions (SVQ) dataset. Our initial experiments establish clear performance headrooms, highlighting the significant opportunity to improve real-world multimodal experiences where audio is a core signal. We encourage the research community to use MSEB to assess their algorithms and contribute to its growth. The library is publicly hosted at https://github.com/google-research/mseb. Georg Heigold, Ehsan Variani, Tom Bagby, Cyril Allauzen, Ji Ma 0004, Shankar Kumar, Michael Riley 0001 |
NeurIPS | 2 |
| 2023 | JEIT: Joint End-to-End Model and Internal Language Model Training for Speech RecognitionabstractWe propose JEIT, a joint end-to-end (E2E) model and internal language model (ILM) training method to inject large-scale unpaired text into ILM during E2E training which improves rare-word speech recognition. With JEIT, the E2E model computes an E2E loss on audio-transcript pairs while its ILM estimates a cross-entropy loss on unpaired text. The E2E model is trained to minimize a weighted sum of E2E and ILM losses. During JEIT, ILM absorbs knowledge from unpaired text while the E2E training serves as regularization. Unlike ILM adaptation methods, JEIT does not require a separate adaptation step and avoids the need for Kullback-Leibler divergence regularization of ILM. We also show that modular hybrid autoregressive transducer (MHAT) performs better than HAT in the JEIT framework, and is much more robust than HAT during ILM adaptation. To push the limit of unpaired text injection, we further propose a combined JEIT and JOIST training (CJJT) that benefits from modality matching, encoder text injection and ILM training. Both JEIT and CJJT can foster a more effective LM fusion. With 100B unpaired sentences, JEIT/CJJT improves rare-word recognition accuracy by up to 16.4% over a model trained without unpaired text. Zhong Meng, Rohit Prabhavalkar, Tara N. Sainath, Tongzhou Chen, Ehsan Variani, Yu Zhang 0033, Bo Li 0028, Andrew Rosenberg, Bhuvana Ramabhadran |
ICASSP | 6 |
| 2023 | Alignment Entropy RegularizationabstractExisting training criteria in automatic speech recognition (ASR) permit the model to freely explore more than one time alignments between the feature and label sequences. In this paper, we use entropy to measure a model’s uncertainty, i.e. how it chooses to distribute the probability mass over the set of allowed alignments. Furthermore, we evaluate the effect of entropy regularization in encouraging the model to distribute the probability mass only on a smaller subset of allowed alignments. Experiments show that entropy regularization enables a much simpler decoding method without sacrificing word error rate, and provides better time alignment quality. Ehsan Variani, David Rybach, Cyril Allauzen, Michael Riley 0001 |
ICASSP | 1 |
| 2023 | Last: Scalable Lattice-Based Speech Modelling in JaxabstractWe introduce LAST, a LAttice-based Speech Transducer library in JAX. With an emphasis on flexibility, ease-of-use, and scalability, LAST implements differentiable weighted finite state automaton (WFSA) algorithms needed for training & inference that scale to a large WFSA such as a recognition lattice over the entire utterance. Despite these WFSA algorithms being well-known in the literature, new challenges arise from performance characteristics of modern architectures, and from nuances in automatic differentiation. We describe a suite of generally applicable techniques employed in LAST to address these challenges, and demonstrate their effectiveness with benchmarks on TPUv3 and V100 GPU. Ehsan Variani, Tom Bagby, Michael Riley 0001 |
ICASSP | 2 |
| 2022 | Multilingual Second-Pass Rescoring for Automatic Speech Recognition SystemsabstractSecond-pass rescoring is a well known technique to improve the performance of Automatic Speech Recognition (ASR) systems. Neural Oracle Search (NOS), which selects the most likely hypothesis from an N-best hypothesis list by integrating information from multiple sources, such as the input acoustic representations, N-best hypotheses, additional first-pass statistics, and unpaired textual information through an external language model, has shown success in rescoring for RNN-T first-pass models. Multilingual first-pass speech recognition models often outperform their monolingual counterparts when trained on related or low-resource languages. In this paper, we investigate the use of the NOS rescoring model on a first-pass multilingual model and show that similar to the first-pass model, the rescoring model can be made multilingual. Our first-pass multilingual model does not require a language-id and we make a realistic assumption that an estimate of the language-id would be available for second-pass rescoring. We conduct comprehensive experiments on two sets of languages, one consisting of related low-resource languages, and the other with a high-resource language added to the first set to analyze the performance of the multilingual NOS rescorer under different settings. Our experimental results show that, multilingual NOS can improve the first-pass multilingual model resulting in average word error rate reduction of 9.4% in the first case, and 8.4% in the second, and out-performing the monolingual counterparts in both cases. Neeraj Gaur, Tongzhou Chen, Ehsan Variani, Parisa Haghani, Bhuvana Ramabhadran, Pedro J. Moreno 0001 |
ICASSP | 3 |
| 2022 | UserLibri: A Dataset for ASR Personalization Using Only Text
Theresa Breiner, Swaroop Ramaswamy, Ehsan Variani, Shefali Garg, Rajiv Mathews, Khe Chai Sim, Kilol Gupta, Mingqing Chen, Lara McConnaughey |
INTERSPEECH | 3 |
| 2022 | On Adaptive Weight Interpolation of the Hybrid Autoregressive Transducer
Ehsan Variani, Michael Riley 0001, David Rybach, Cyril Allauzen, Tongzhou Chen, Bhuvana Ramabhadran |
INTERSPEECH | 1 |
| 2022 | Improving Rare Word Recognition with LM-aware MWER TrainingabstractLanguage models (LMs) significantly improve the recognition accuracy of end-to-end (E2E) models on words rarely seen during training, when used in either the shallow fusion or the rescoring setups. In this work, we introduce LMs in the learning of hybrid autoregressive transducer (HAT) models in the discriminative training framework, to mitigate the training versus inference gap regarding the use of LMs. For the shallow fusion setup, we use LMs during both hypotheses generation and loss computation, and the LM-aware MWER-trained model achieves 10\% relative improvement over the model trained with standard MWER on voice search test sets containing rare words. For the rescoring setup, we learn a small neural module to generate per-token fusion weights in a data-dependent manner. This model achieves the same rescoring WER as regular MWER-trained model, but without the need for sweeping fusion weights. Tongzhou Chen, Tara N. Sainath, Ehsan Variani, Rohit Prabhavalkar, W. Ronny Huang, Bhuvana Ramabhadran, Neeraj Gaur, Sepand Mavandadi, Cal Peyser, Trevor Strohman, Yanzhang He, David Rybach |
INTERSPEECH | 4 |
| 2022 | Global Normalization for Streaming Speech Recognition in a Modular FrameworkabstractWe introduce the Globally Normalized Autoregressive Transducer (GNAT) for addressing the label bias problem in streaming speech recognition. Our solution admits a tractable exact computation of the denominator for the sequence-level normalization. Through theoretical and empirical results, we demonstrate that by switching to a globally normalized model, the word error rate gap between streaming and non-streaming speech-recognition models can be greatly reduced (by more than 50% on the Librispeech dataset). This model is developed in a modular framework which encompasses all the common neural speech recognition models. The modularity of this framework enables controlled comparison of modelling choices and creation of new models. A JAX implementation of our models has been open sourced. Ehsan Variani, Michael Riley 0001, David Rybach, Matt Shannon, Cyril Allauzen |
NeurIPS | 1 |
| 2022 | Modular Hybrid Autoregressive TransducerabstractText-only adaptation of a transducer model remains challenging for end-to-end speech recognition since the transducer has no clearly separated acoustic model (AM), language model (LM) or blank model. In this work, we propose a modular hybrid autoregressive transducer (MHAT) that has structurally separated label and blank decoders to predict label and blank distributions, respectively, along with a shared acoustic encoder. The encoder and label decoder outputs are directly projected to AM and internal LM scores and then added to compute label posteriors. We train MHAT with an internal LM loss and a HAT loss to ensure that its internal LM becomes a standalone neural LM that can be effectively adapted to text. Moreover, text adaptation of MHAT fosters a much better LM fusion than internal LM subtraction-based methods. On Google's large-scale production data, a multi-domain MHAT adapted with 100B sentences achieves relative WER reductions of up to 12.4% without LM fusion and 21.5% with LM fusion from 400K-hour trained HAT. Zhong Meng, Tongzhou Chen, Rohit Prabhavalkar, Yu Zhang 0033, Gary Wang, Kartik Audhkhasi, Jesse Emond, Trevor Strohman, Bhuvana Ramabhadran, W. Ronny Huang, Ehsan Variani, Pedro J. Moreno 0001 |
SLT | 11 |
| 2021 | Cascaded Encoders for Unifying Streaming and Non-Streaming ASRabstractEnd-to-end (E2E) automatic speech recognition (ASR) models, by now, have shown competitive performance on several benchmarks. These models are structured to either operate in streaming or non-streaming mode. This work presents cascaded encoders for building a single E2E ASR model that can operate in both these modes simultaneously. The proposed model consists of streaming and non-streaming encoders. Input features are first processed by the streaming encoder; the non-streaming encoder operates exclusively on the output of the streaming encoder. A single decoder then learns to decode either using the output of the streaming or the non-streaming encoder. Results show that this model achieves similar word error rates (WER) as a standalone streaming model when operating in streaming mode, and obtains 10% – 27% relative improvement when operating in non-streaming mode. Our results also show that the proposed approach outperforms existing E2E two-pass models, especially on long-form speech. Arun Narayanan, Tara N. Sainath, Ruoming Pang, Chung-Cheng Chiu, Rohit Prabhavalkar, Ehsan Variani, Trevor Strohman |
ICASSP | 7 |
| 2021 | A Hybrid Seq-2-Seq ASR Design for On-Device and Server Applications
Cyril Allauzen, Ehsan Variani, Michael Riley 0001, David Rybach, Hao Zhang 0010 |
Interspeech | 2 |
| 2021 | An Efficient Streaming Non-Recurrent On-Device End-to-End Model with Improvements to Rare-Word Modeling
Tara N. Sainath, Yanzhang He, Arun Narayanan, Rami Botros, Ruoming Pang, David Rybach, Cyril Allauzen, Ehsan Variani, James Qin, Quoc-Nam Le-The, Shuo-Yiin Chang, Bo Li 0028, Anmol Gulati, Chung-Cheng Chiu, Diamantino Caseiro, Wei Li 0133, Qiao Liang 0001, Pat Rondon |
Interspeech | 8 |
| 2020 | Neural Oracle Search on N-BEST HypothesesabstractIn this paper, we propose a neural search algorithm to select the most likely hypothesis using a sequence of acoustic representations and multiple hypotheses as input. The algorithm provides a sequence level score for each audio-hypothesis pair that is obtained by integrating information from multiple sources, such as the input acoustic representations, N-best hypotheses, additional 1st-pass statistics, and unpaired textual information through an external language model. These scores are then used to map the search problem of identifying the most likely hypothesis to a sequence classification problem. The definition of the proposed algorithm is broad enough to allow its use as an alternative to beam search in the 1st-pass or as a 2nd-pass, rescoring step. This algorithm achieves up to 12% relative reductions in Word Error Rate (WER) across several languages over state-of-the-art baselines with relatively few additional parameters. We also propose the use of a binary classifier gating function that can learn to trigger the 2nd-pass neural search model when the 1-best hypothesis is not the oracle hypothesis, thereby avoiding extra computation. Ehsan Variani, Tongzhou Chen, James Apfel, Bhuvana Ramabhadran, Seungji Lee, Pedro J. Moreno 0001 |
ICASSP | 1 |
| 2020 | Hybrid Autoregressive Transducer (HAT)abstractThis paper proposes and evaluates the hybrid autoregressive transducer (HAT) model, a time-synchronous encoder-decoder model that preserves the modularity of conventional automatic speech recognition systems. The HAT model provides a way to measure the quality of the internal language model that can be used to decide whether inference with an external language model is beneficial or not. We evaluate our proposed model on a large-scale voice search task. Our experiments show significant improvements in WER compared to the state-of-the-art approaches1. Ehsan Variani, David Rybach, Cyril Allauzen, Michael Riley 0001 |
ICASSP | 1 |
| 2019 | A Density Ratio Approach to Language Model Fusion in End-to-End Automatic Speech RecognitionabstractThis article describes a density ratio approach to integrating external Language Models (LMs) into end-to-end models for Automatic Speech Recognition (ASR). Applied to a Recurrent Neural Network Transducer (RNN-T) ASR model trained on a given domain, a matched in-domain RNN-LM, and a target domain RNN-LM, the proposed method uses Bayes' Rule to define RNN-T posteriors for the target domain, in a manner directly analogous to the classic hybrid model for ASR based on Deep Neural Networks (DNNs) or LSTMs in the Hidden Markov Model (HMM) framework (Bourlard & Morgan, 1994). The proposed approach is evaluated in cross-domain and limited-data scenarios, for which a significant amount of target domain text data is used for LM training, but only limited (or no) {audio, transcript} training data pairs are used to train the RNN-T. Specifically, an RNN-T model trained on paired audio & transcript data from YouTube is evaluated for its ability to generalize to Voice Search data. The Density Ratio method was found to consistently outperform the dominant approach to LM and end-to-end ASR integration, Shallow Fusion. Erik McDermott, Hasim Sak, Ehsan Variani |
ASRU | 3 |
| 2019 | West: Word Encoded Sequence TransducersabstractMost of the parameters in large vocabulary models are used in embedding layer to map categorical features to vectors and in softmax layer for classification weights. This is a bottleneck in memory constraint on-device training applications like federated learning and on-device inference applications like automatic speech recognition (ASR). One way of compressing the embedding and softmax layers is to substitute larger units such as words with smaller sub-units such as characters. However, often the sub-unit models perform poorly compared to the larger unit models. We propose WEST, an algorithm for encoding categorical features and output classes with a sequence of random or domain dependent sub-units and demonstrate that this transduction can lead to significant compression without compromising performance. Ehsan Variani, Ananda Theertha Suresh, Mitch Weintraub |
ICASSP | 1 |
| 2018 | Sampled Connectionist Temporal ClassificationabstractThis article introduces and evaluates Sampled Connectionist Temporal Classification (CTC) which connects the CTC criterion to the Cross Entropy (CE) objective through sampling. Instead of computing the logarithm of the sum of the alignment path likelihoods, at each training step the sampled CTC only computes the CE loss between the sampled alignment path and model posteriors. It is shown that the sampled CTC objective is an unbiased estimator of an upper bound for the CTC loss, thus minimization of the sampled CTC is equivalent to the minimization of the upper bound of the CTC objective. The definition of the sampled CTC objective has the advantage that it is scalable computationally to the massive datasets using accelerated computation machines. The sampled CTC is compared with CTC in two large-scale speech recognition tasks and it is shown that sampled CTC can achieve similar WER performance of the best CTC baseline in about one fourth of the training time of the CTC baseline. Ehsan Variani, Tom Bagby, Kamel Lahouel, Erik McDermott, Michiel Bacchiani |
ICASSP | 1 |
| 2018 | Efficient Implementation of the Room Simulator for Training Deep Neural Network Acoustic ModelsabstractIn this paper, we describe how to efficiently implement an acoustic room simulator to generate large-scale simulated data for training deep neural networks.Even though Google Room Simulator in [1] was shown to be quite effective in reducing the Word Error Rates (WERs) for far-field applications by generating simulated far-field training sets, it requires a very large number of FFTs.Room Simulator used approximately 80 % of CPU usage in our CPU/GPU training architecture [2].In this work, we implement an efficient OverLap Addition (OLA) based filtering using the open-source FFTW3 library.Further, we investigate the effects of the Room Impulse Response (RIR) lengths.Experimentally, we conclude that we can cut the tail portions of RIRs whose power is less than 20 dB below the maximum power without sacrificing the speech recognition accuracy.However, we observe that cutting RIR tail more than this threshold harms the speech recognition accuracy for rerecorded test sets.Using these approaches, we were able to reduce CPU usage for the room simulator portion down to 9.69 % in CPU/GPU training architecture.Profiling result shows that we obtain 22.4 times speed-up on a single machine and 37.3 times speed up on Google's distributed training infrastructure. Chanwoo Kim 0001, Ehsan Variani, Arun Narayanan, Michiel Bacchiani |
INTERSPEECH | 2 |
| 2017 | Acoustic Modeling for Google Home
Bo Li 0028, Tara N. Sainath, Arun Narayanan, Joe Caroselli, Michiel Bacchiani, Ananya Misra, Izhak Shafran, Hasim Sak, Golan Pundak, Kean K. Chin, Khe Chai Sim, Ron J. Weiss, Kevin W. Wilson, Ehsan Variani, Chanwoo Kim 0001, Olivier Siohan, Mitch Weintraub, Erik McDermott, Richard Rose, Matt Shannon |
INTERSPEECH | 14 |
| 2017 | End-to-End Training of Acoustic Models for Large Vocabulary Continuous Speech Recognition with TensorFlow
Ehsan Variani, Tom Bagby, Erik McDermott, Michiel Bacchiani |
INTERSPEECH | 1 |
| 2017 | Multichannel Signal Processing With Deep Neural Networks for Automatic Speech RecognitionabstractMultichannel automatic speech recognition (ASR) systems commonly separate speech enhancement, including localization, beamforming, and postfiltering, from acoustic modeling. In this paper, we perform multichannel enhancement jointly with acoustic modeling in a deep neural network framework. Inspired by beamforming, which leverages differences in the fine time structure of the signal at different microphones to filter energy arriving from different directions, we explore modeling the raw time-domain waveform directly. We introduce a neural network architecture, which performs multichannel filtering in the first layer of the network, and show that this network learns to be robust to varying target speaker direction of arrival, performing as well as a model that is given oracle knowledge of the true target speaker direction. Next, we show how performance can be improved by factoring the first layer to separate the multichannel spatial filtering operation from a single channel filterbank which computes a frequency decomposition. We also introduce an adaptive variant, which updates the spatial filter coefficients at each time frame based on the previous inputs. Finally, we demonstrate that these approaches can be implemented more efficiently in the frequency domain. Overall, we find that such multichannel neural networks give a relative word error rate improvement of more than 5% compared to a traditional beamforming-based multichannel ASR system and more than 10% compared to a single channel waveform model. Tara N. Sainath, Ron J. Weiss, Kevin W. Wilson, Bo Li 0028, Arun Narayanan, Ehsan Variani, Michiel Bacchiani, Izhak Shafran, Andrew W. Senior, Kean K. Chin, Ananya Misra, Chanwoo Kim 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2016 | Reducing the Computational Complexity of Multimicrophone Acoustic Models with Integrated Feature Extraction
Tara N. Sainath, Arun Narayanan, Ron J. Weiss, Ehsan Variani, Kevin W. Wilson, Michiel Bacchiani, Izhak Shafran |
INTERSPEECH | 4 |
| 2016 | Complex Linear Projection (CLP): A Discriminative Approach to Joint Feature Extraction and Acoustic Modeling
Ehsan Variani, Tara N. Sainath, Izhak Shafran, Michiel Bacchiani |
INTERSPEECH | 1 |
| 2015 | A Gaussian Mixture Model layer jointly optimized with discriminative features within a Deep Neural Network architectureabstractThis article proposes and evaluates a Gaussian Mixture Model (GMM) represented as the last layer of a Deep Neural Network (DNN) architecture and jointly optimized with all previous layers using Asynchronous Stochastic Gradient Descent (ASGD). The resulting “Deep GMM” architecture was investigated with special attention to the following issues: (1) The extent to which joint optimization improves over separate optimization of the DNN-based feature extraction layers and the GMM layer; (2) The extent to which depth (measured in number of layers, for a matched total number of parameters) helps a deep generative model based on the GMM layer, compared to a vanilla DNN model; (3) Head-to-head performance of Deep GMM architectures vs. equivalent DNN architectures of comparable depth, using the same optimization criterion (frame-level Cross Entropy (CE)) and optimization method (ASGD); (4) Expanded possibilities for modeling offered by the Deep GMM generative model. The proposed Deep GMMs were found to yield Word Error Rates (WERs) competitive with state-of-the-art DNN systems, at the cost of pre-training using standard DNNs to initialize the Deep GMM feature extraction layers. An extension to Deep Subspace GMMs is described, resulting in additional gains. Ehsan Variani, Erik McDermott, Georg Heigold |
ICASSP | 1 |
| 2015 | NON-adaptive policies for 20 questions target localizationabstractThe problem of target localization with noise is addressed. The target is a sample from a continuous random variable with known distribution and the goal is to locate it with minimum mean squared error distortion. The localization scheme or policy proceeds by queries, or questions, weather or not the target belongs to some subset as it is addressed in the 20-question framework. These subsets are not constrained to be intervals and the answers to the queries are noisy. While this situation is well studied for adaptive querying, this paper is focused on the non adaptive querying policies based on dyadic questions. The asymptotic minimum achievable distortion under such policies is derived. Furthermore, a policy named the Aurelian1is exhibited which achieves asymptotically this distortion. Ehsan Variani, Kamel Lahouel, Avner Bar-Hen, Bruno Jedynak |
ISIT | 1 |
| 2014 | Deep neural networks for small footprint text-dependent speaker verificationabstractIn this paper we investigate the use of deep neural networks (DNNs) for a small footprint text-dependent speaker verification task. At development stage, a DNN is trained to classify speakers at the framelevel. During speaker enrollment, the trained DNN is used to extract speaker specific features from the last hidden layer. The average of these speaker features, or d-vector, is taken as the speaker model. At evaluation stage, a d-vector is extracted for each utterance and compared to the enrolled speaker model to make a verification decision. Experimental results show the DNN based speaker verification system achieves good performance compared to a popular i-vector system on a small footprint text-dependent speaker verification task. In addition, the DNN based system is more robust to additive noise and outperforms the i-vector system at low False Rejection operating points. Finally the combined system outperforms the i-vector system by 14% and 25% relative in equal error rate (EER) for clean and noisy conditions respectively. Ehsan Variani, Erik McDermott, Ignacio López-Moreno, Javier Gonzalez-Dominguez |
ICASSP | 1 |
| 2013 | Mean temporal distance: Predicting ASR error from temporal properties of speech signalabstractExtending previous work on prediction of phoneme recognition error from unlabeled data that were corrupted by unpredictable factors, the current work investigates a simple but effective method of estimating ASR performance by computing a function M(Δt), which represents the mean distance between speech feature vectors evaluated over certain finite time interval, determined as a function of temporal distance Δt between the vectors. It is shown that M(Δt) is a function of signal-to-noise ratio of speech signal. Comparing M(Δt) curves, derived on data used for training of the classifier, and on test utterances, allows for predicting error on the test data. Another interesting observation is that M(Δt) remains approximately constant, as temporal separation Δt exceeds certain critical interval (about 200 ms), indicating the extent of coarticulation in speech sounds. Hynek Hermansky, Ehsan Variani, Vijayaditya Peddinti |
ICASSP | 2 |
| 2013 | Multi-stream recognition of noisy speech with performance monitoringabstractA prototype multi-stream system with a performance monitor for stream selection is proposed to recognize speech in un-known noise. The speech signal is decomposed into seven band-limited streams. Posterior probabilities of phonemes are estimated by a multi-layer perceptron (MLP) in each of these band-limited streams. Estimated posterior vectors of all 127 combinations (processing streams) of the seven band-limited streams form inputs to a second-stage MLP that esti-mates posterior probabilities of phonemes in each processing stream. A performance monitor is designed to predict the re-liability of individual processing streams based on the outputs from these streams. The top N streams that are least affected by noise are selected and their outputs are averaged to yield the final posterior probability vector used in Viterbi search for the best phoneme sequence. Experimental results show that the proposed technique is effective in dealing with noise. Index Terms: Multi-stream speech recognition, Performance monitoring Ehsan Variani, Feipeng Li, Hynek Hermansky |
INTERSPEECH | 1 |
| 2012 | Estimating Classifier Performance in Unknown NoiseabstractWe propose and investigate a non-parametric method for identifying regions of speech that have unexpected distortions not seen in the training data. The method does not require knowledge of correct labels and relies only on divergence between statistics of the test and training data. Our experiments show that the proposed method re-quires a relatively small amount of test data of the order of several seconds to stabilize, and correlates well with recognition error observed on the test data. Index Terms: Unexpected distortions, confidence esti-mation, machine recognition of speech Ehsan Variani, Hynek Hermansky |
INTERSPEECH | 1 |
| 2011 | VTLN in the MFCC Domain: Band-Limited versus Local InterpolationabstractWe propose a new easy-to-implement method to compute a Lin-ear Transform (LT) to perform Vocal Tract Length Normalization (VTLN) on truncated Mel Frequency Cepstral Coefficients (MFCCs) normally used in distributed speech recognition. The method is based on a Local Interpolation which is independent of the Mel filter design. Local Interpolation (LILT) VTLN is theoretically and experimentally compared to a global scheme based on band-limited interpolation (BLI-VTLN) and the conventional frequency warp-ing scheme (FFT-VTLN). Investigating the interoperability of these methods shows that the performance of LILT-VTLN is on par with FFT-VTLN and BLI-VTLN. Models trained with LILT- and BLI-VTLN performance degrades if FFT-VTLN is used as a front-end. The degradation for LILT-VTLN is slightly less, indicating that it produces models that are a better match for FFT-VTLN. Index Terms — Automatic speech recognition, VTLN, fre-quency warping, linear transform Ehsan Variani, Thomas Schaaf |
INTERSPEECH | 1 |