VLDB 2026 Research / reviewers in the wild / expert
Upendra V. Chaudhari
dblp:25/4485
· DBLP profile ↗
38ranked-venue papers
15as first author
0since 2021 · last 2017
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 31 · 13 first-authorArtificial intelligence and machine learning · 22 · 10 first-authorSystems, architecture and hardware · 2Databases, data management, data science and information retrieval · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Speech recognition and synthesis · 44% Efficient and distributed learning · 31% Optimization for machine learning · 25% | |
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
High-performance computing · 48% Parallel and multicore computing · 36% Distributed systems · 8% | |
| Databases, data mining, and information retrieval
2 papers |
Information retrieval · 70% Data mining · 30% |
Topics — the 19 heaviest of 21, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Parallel and multicore computing › parallel computing › parallel machine learning
data-parallel training |
0.3 | 1 | 2017 | Parallel Deep Neural Network Training for Big Data on Blue Gene/Q · IEEE Trans. Parallel Distributed Syst. 2017 |
Machine learning › Efficient and distributed learning › distributed training
data parallel training |
0.2 | 1 | 2014 | Parallel Deep Neural Network Training for Big Data on Blue Gene/Q · SC 2014 |
Machine learning › Efficient and distributed learning
distributed training |
0.2 | 1 | 2014 | Parallel Deep Neural Network Training for Big Data on Blue Gene/Q · SC 2014 |
Machine learning › Optimization for machine learning › second-order optimization
hessian-free optimization |
0.2 | 1 | 2014 | Parallel Deep Neural Network Training for Big Data on Blue Gene/Q · SC 2014 |
Machine learning › Optimization for machine learning
second-order optimization |
0.2 | 1 | 2014 | Parallel Deep Neural Network Training for Big Data on Blue Gene/Q · SC 2014 |
High-performance computing › supercomputer architecture
blue gene/q |
0.2 | 1 | 2014 | Parallel Deep Neural Network Training for Big Data on Blue Gene/Q · SC 2014 |
High-performance computing
supercomputing |
0.2 | 1 | 2014 | Parallel Deep Neural Network Training for Big Data on Blue Gene/Q · SC 2014 |
Natural language and speech › Speech recognition and synthesis
acoustic modeling |
0.1 | 1 | 2012 | Hidden Markov Acoustic Modeling With Bootstrap and Restructuring for Low-Resourced Languages · IEEE Trans. Speech Audio Process. 2012 |
Natural language and speech › Speech recognition and synthesis › acoustic modeling
hidden markov model acoustic modeling |
0.1 | 1 | 2012 | Hidden Markov Acoustic Modeling With Bootstrap and Restructuring for Low-Resourced Languages · IEEE Trans. Speech Audio Process. 2012 |
Natural language and speech › Speech recognition and synthesis › automatic speech recognition
low-resource speech recognition |
0.1 | 1 | 2012 | Hidden Markov Acoustic Modeling With Bootstrap and Restructuring for Low-Resourced Languages · IEEE Trans. Speech Audio Process. 2012 |
Information retrieval › document retrieval
spoken document retrieval |
0.1 | 1 | 2012 | Matching Criteria for Vocabulary-Independent Search · IEEE Trans. Speech Audio Process. 2012 |
Machine learning › Efficient and distributed learning › distributed training
distributed DNN training |
0.1 | 1 | 2017 | Parallel Deep Neural Network Training for Big Data on Blue Gene/Q · IEEE Trans. Parallel Distributed Syst. 2017 |
Data mining › predictive modeling › classification
ensemble learning |
0.1 | 1 | 2006 | Resource Management for Networked Classifiers in Distributed Stream Mining Systems · ICDM 2006 |
Distributed systems › stream processing
distributed stream mining |
0.1 | 1 | 2006 | Resource Management for Networked Classifiers in Distributed Stream Mining Systems · ICDM 2006 |
Cloud and datacenter computing
resource management |
0.1 | 1 | 2006 | Resource Management for Networked Classifiers in Distributed Stream Mining Systems · ICDM 2006 |
Natural language and speech › Speech recognition and synthesis
automatic speech recognition |
0.0 | 1 | 2012 | Matching Criteria for Vocabulary-Independent Search · IEEE Trans. Speech Audio Process. 2012 |
Natural language and speech › Speech recognition and synthesis › automatic speech recognition › continuous speech recognition
large vocabulary continuous speech recognition |
0.0 | 1 | 2012 | Hidden Markov Acoustic Modeling With Bootstrap and Restructuring for Low-Resourced Languages · IEEE Trans. Speech Audio Process. 2012 |
Natural language and speech › Speech recognition and synthesis
speaker recognition |
0.0 | 1 | 2003 | Multigrained modeling with pattern specific maximum likelihood transformations for text-independent speaker recognition · IEEE Trans. Speech Audio Process. 2003 |
Natural language and speech › Speech recognition and synthesis › speaker recognition
text-independent speaker recognition |
0.0 | 1 | 2003 | Multigrained modeling with pattern specific maximum likelihood transformations for text-independent speaker recognition · IEEE Trans. Speech Audio Process. 2003 |
Methods — techniques the papers use, named apart from their topics
hessian-free optimization · 1.0data parallelism · 1.0second-order optimization · 0.6phonetic confusion matrix · 0.3edit distance · 0.3conditional random field · 0.3model restructuring · 0.1gaussian clustering · 0.1bootstrap aggregation · 0.1resource allocation · 0.1optimization · 0.1gaussian mixture model · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2017 | Parallel Deep Neural Network Training for Big Data on Blue Gene/QabstractDeep Neural Networks (DNNs) have recently been shown to significantly outperform existing machine learning techniques in several pattern recognition tasks. DNNs are the state-of-the-art models used in image recognition, object detection, classification and tracking, and speech and language processing applications. The biggest drawback to DNNs has been the enormous cost in computation and time taken to train the parameters of the networks-often a tenfold increase relative to conventional technologies. Such training time costs can be mitigated by the application of parallel computing algorithms and architectures. However, these algorithms often run into difficulties because of the cost of inter-processor communication bottlenecks. In this paper, we describe how to enable Parallel Deep Neural Network Training on the IBM Blue Gene/Q (BG/Q) computer system. Specifically, we explore DNN training using the data-parallel Hessian-free 2nd order optimization algorithm. Such an algorithm is particularly well-suited to parallelization across a large set of loosely coupled processors. BG/Q, with its excellent inter-processor communication characteristics, is an ideal match for this type of algorithm. The paper discusses how issues regarding programming model and data-dependent imbalances are addressed. Results on large-scale speech tasks show that the performance on BG/Q scales linearly up to 4,096 processes with no loss in accuracy. This allows us to train neural networks using billions of training examples in a few hours. I-Hsin Chung, Tara N. Sainath, Bhuvana Ramabhadran, Michael Picheny, John A. Gunnels, Vernon Austel, Upendra V. Chaudhari, Brian Kingsbury |
IEEE Trans. Parallel Distributed Syst. | 7 |
| 2014 | Parallel deep neural network training for LVCSR tasks using blue gene/QabstractWhile Deep Neural Networks (DNNs) have achieved tremendous success for LVCSR tasks, training these networks is slow. To date, the most common approach to train DNNs is via stochastic gradient descent (SGD), serially on a single GPU machine. Serial training, coupled with the large number of training parameters and speech data set sizes, makes DNN training very slow for LVCSR tasks. While 2nd order, data-parallel methods have also been explored, these methods are not always faster on CPU clusters due to the large communication cost between processors. In this work, we explore using a specialized hardware/software approach, utilizing a Blue Gene/Q (BG/Q) system, which has thousands of processors and excellent interprocessor communication. We explore using the 2nd order Hessian-free (HF) algorithm for DNN training with BG/Q, for both cross-entropy and sequence training of DNNs. Results on three LVCSR tasks indicate that using HF with BG/Q offers up to an 11x speedup, as well as an improved word error rate (WER), compared to SGD on a GPU. Tara N. Sainath, I-Hsin Chung, Bhuvana Ramabhadran, Michael Picheny, John A. Gunnels, Brian Kingsbury, George Saon, Vernon Austel, Upendra V. Chaudhari |
INTERSPEECH | 9 |
| 2014 | Parallel Deep Neural Network Training for Big Data on Blue Gene/QabstractDeep Neural Networks (DNNs) have recently been shown to significantly outperform existing machine learning techniques in several pattern recognition tasks. DNNs are the state-of-the-art models used in image recognition, object detection, classification and tracking, and speech and language processing applications. The biggest drawback to DNNs has been the enormous cost in computation and time taken to train the parameters of the networks - often a tenfold increase relative to conventional technologies. Such training time costs can be mitigated by the application of parallel computing algorithms and architectures. However, these algorithms often run into difficulties because of the cost of inter-processor communication bottlenecks. In this paper, we describe how to enable Parallel Deep Neural Network Training on the IBM Blue Gene/Q (BG/Q) computer system. Specifically, we explore DNN training using the data parallel Hessian-free 2nd order optimization algorithm. Such an algorithm is particularly well-suited to parallelization across a large set of loosely coupled processors. BG/Q, with its excellent inter-processor communication characteristics, is an ideal match for this type of algorithm. The paper discusses how issues regarding programming model and data-dependent imbalances are addressed. Results on large-scale speech tasks show that the performance on BG/Q scales linearly up to 4096 processes with no loss in accuracy. This allows us to train neural networks using billions of training examples in a few hours. I-Hsin Chung, Tara N. Sainath, Bhuvana Ramabhadran, Michael Picheny, John A. Gunnels, Vernon Austel, Upendra V. Chaudhari, Brian Kingsbury |
SC | 7 |
| 2013 | The IBM speech-to-speech translation system for smartphone: Improvements for resource-constrained tasks
Bowen Zhou 0002, Songfang Huang, Martin Cmejrek, Wei Zhang 0057, Jia Cui, Bing Xiang, Gregg Daggett, Upendra V. Chaudhari, Sameer Maskey, Etienne Marcheret |
Comput. Speech Lang. | 10 |
| 2012 | Constructing ensembles of dissimilar acoustic models using hidden attributes of training dataabstractOne of the objectives in acoustic modeling is to realize robust statistical models against the wide variety of acoustic conditions that are present in real world environments. As large amounts of training data become available, modeling subsets of the data with similar acoustic qualities can be done accurately and multiple acoustic models are jointly used as a form of system combination or model selection. In this paper, we propose a method to partition the training data for constructing ensembles of acoustic models using metadata attributes such as SNR, speaking rate, and duration via a binary tree. The metadata attribute used at each binary split in the decision tree is obtained using a metric proposed in this paper that is cosine-similarity based. The resulting multiple models are combined using voting techniques such as n-best ROVER. The proposed method improved the recognition accuracy by up to 4% relative over the state-of-the-art system on a large vocabulary continuous speech recognition voice search task. Takashi Fukuda, Ryuki Tachibana, Upendra V. Chaudhari, Bhuvana Ramabhadran, Puming Zhan |
ICASSP | 3 |
| 2012 | Matching Criteria for Vocabulary-Independent SearchabstractThis paper investigates a variety of progressively more complex similarity measures for vocabulary independent search in phone based audio transcripts. English audio data is segmented and decoded to produce a sequence of phones that represent the data. These sequences are then parsed intoN-grams which are used to index the data. The audio segments define the documents to be retrieved and are thus localized in time. Search is performed by expanding text based queries into phone sequences andN-grams, followed by matching these against the index. The baseline similarity measure combines elements found in the literature and uses edit distance with a phonetic confusion matrix to determine the similarity of query and indexN-grams. Comparable performance to other approaches in the literature is achieved. Extensions to the baseline are developed using a constrained form of the similarity measure together with the ability to account for higher order confusions, namely of phone bi-grams and tri-grams. Results show improved performance across a variety of system configurations. We then generalize further and use the framework of conditional random fields (CRFs) to model confusions. Whereas others in the literature have used CRFs to model parameters of an edit distance that incorporates deletions, substitutions, and insertions, our approach focuses on using CRFs to model context dependent phone level confusions directly. The CRF is trained on parallel phonetic transcripts, which provides a general framework for modeling the errors that a recognition system may make, taking contextual effects into consideration. Results obtained on both in and out of vocabulary (OOV) search tasks improve most notably for OOV, showing 5%-6% relative improvement. Finally, we investigate the degree to which the information captured in the three approaches is complementary and show that system combination can further improve performance. Upendra V. Chaudhari, Michael Picheny |
IEEE Trans. Speech Audio Process. | 1 |
| 2012 | Hidden Markov Acoustic Modeling With Bootstrap and Restructuring for Low-Resourced LanguagesabstractThis paper proposes an acoustic modeling approach based on bootstrap and restructuring to dealing with data sparsity for low-resourced languages. The goal of the approach is to improve the statistical reliability of acoustic modeling for automatic speech recognition (ASR) in the context of speed, memory and response latency requirements for real-world applications. In this approach, randomized hidden Markov models (HMMs) estimated from the bootstrapped training data are aggregated for reliable sequence prediction. The aggregation leads to an HMM with superior prediction capability at cost of a substantially larger size. For practical usage the aggregated HMM is restructured by Gaussian clustering followed by model refinement. The restructuring aims at reducing the aggregated HMM to a desirable model size while maintaining its performance close to the original aggregated HMM. To that end, various Gaussian clustering criteria and model refinement algorithms have been investigated in the full covariance model space before the conversion to the diagonal covariance model space in the last stage of the restructuring. Large vocabulary continuous speech recognition (LVCSR) experiments on Pashto and Dari have shown that acoustic models obtained by the proposed approach can yield superior performance over the conventional training procedure with almost the same run-time memory consumption and decoding speed. Peder A. Olsen, Pierre L. Dognin, Upendra V. Chaudhari, John R. Hershey, Bowen Zhou 0006 |
IEEE Trans. Speech Audio Process. | 6 |
| 2011 | An investigation of heuristic, manual and statistical pronunciation derivation for PashtoabstractIn this paper, we study the issue of generating pronunciations for training and decoding with an ASR system for Pashto in the context of a Speech to Speech Translation system developed for TRANSTAC. As with other low resourced languages, a limited amount of acoustic training data was available with a corresponding set of manually produced vowelized pronunciations. We augment this data with other sources, but lack pronunciations for unseen words in the new audio and associated text. Four methods are investigated for generating these pronunciations, or baseforms: an heuristic grapheme to phoneme map, manual annotation, and two methods based on statistical models. The first of these uses a joint Maximum Entropy N-gram model while the other is based on a log-linear Statistical Machine Translation model. We report results on a state of the art, discriminatively trained, ASR system and show that the manual and statistical methods provide an improvement over the grapheme to phoneme map. Moreover, we demonstrate that the automatic statistical methods can perform as well or better than manual generation by native speakers, even in the case where we have a significant number of high quality, manually generated pronunciations beyond those provided by the TRANSTAC program. Upendra V. Chaudhari, Bowen Zhou 0006 |
ASRU | 1 |
| 2011 | Frame-level AnyBoost for LVCSR with the MMI CriterionabstractThis paper propose a variant of AnyBoost for a large vocabulary continuous speech recognition (LVCSR) task. AnyBoost is an efficient algorithm to train an ensemble of weak learners by gradient descent for an objective function.We present a novel training procedure that trains acoustic models via the MMI criterion using data that is weighted proportional to the summation of the posterior functions of previous round of weak learners. Optimized for system combination by n-best ROVER at runtime, data weights for a new weak learner are computed as a weighted summation of posteriors of previous weak learners. We compare a frame-based version and a sentence-based version of our proposed algorithm with a frame-based AdaBoost algorithm. We will present results on a voice search task trained with different amounts of data with gains of 5.1% to 7.5% relative in WER can be obtained by three rounds of boosting. Ryuki Tachibana, Takashi Fukuda, Upendra V. Chaudhari, Bhuvana Ramabhadran, Puming Zhan |
ASRU | 3 |
| 2010 | A comparative study on system combination schemes for LVCSRabstractWe present a comparative study on combination schemes for large vocabulary continuous speech recognition by incorporating long-span class posterior probability features into conventional short-time cepstral features. System combination can improve the overall speech recognition performance when multiple systems exhibit different error patterns and multiple knowledge sources encode complementary information. A variety of combination approaches are investigated in this paper, e.g., feature concatenation single stream system, model combination multi-stream system, lattice rescoring and ROVER. These techniques work at different levels of a LVCSR system and have different computational cost. We compared their performance and analyzed their advantages and disadvantages on large vocabulary English broadcast news transcription tasks. Experimental results showed that model combination with independent tree consistently outperforms ROVER, feature concatenation and lattice rescoring. In addition, the phoneme posterior probability features do provide complementary information to short-time cepstral features. Chengyuan Ma, Hong-Kwang Jeff Kuo, Hagen Soltau, Upendra V. Chaudhari, Lidia Mangu |
ICASSP | 5 |
| 2010 | The IBM 2008 GALE Arabic speech transcription systemabstractThis paper describes the Arabic broadcast transcription system fielded by IBM in the GALE Phase 3.5 machine translation evaluation. Key advances compared to our Phase 2.5 system include improved discriminative training, the use of Subspace Gaussian Mixture Models (SGMM), neural network acoustic features, variable frame rate decoding, training data partitioning experiments, unpruned n-gram language models and neural network language models. These advances were instrumental in achieving a word error rate of 8.9% on the evaluation test set. George Saon, Hagen Soltau, Upendra V. Chaudhari, Stephen M. Chu, Brian Kingsbury, Hong-Kwang Jeff Kuo, Lidia Mangu, Daniel Povey |
ICASSP | 3 |
| 2010 | Acoustic modeling with bootstrap and restructuring for low-resourced languages
Pierre L. Dognin, Upendra V. Chaudhari, Bowen Zhou 0006 |
INTERSPEECH | 4 |
| 2009 | Articulatory feature detection with Support Vector Machines for integration into ASR and phone recognitionabstractWe study the use of support vector machines (SVM) for detecting the occurrence of articulatory features in speech audio data and using the information contained in the detector outputs to improve phone and speech recognition. Our expectation is that an SVM should be able to appropriately model the separation of the classes which may have complex distributions in feature space. We show that performance improves markedly when using discriminatively trained speaker dependent parameters for the SVM inputs, and compares quite well to results in the literature using other classifiers, namely artificial neural networks (ANN). Further, we show that the resulting detector outputs can be successfully integrated into a state of the art speech recognition system, with consequent performance gains. Notably, we test our system on English broadcast news data from dev04f. Upendra V. Chaudhari, Michael Picheny |
ASRU | 1 |
| 2009 | Improved vocabulary independent search with approximate match based on Conditional Random FieldsabstractWe investigate the use of Conditional Random Fields (CRF) to model confusions and account for errors in the phonetic decoding derived from Automatic Speech Recognition output. The goal is to improve the accuracy of approximate phonetic match, given query terms and an indexed database of documents, in a vocabulary independent audio search system. Audio data is ingested, segmented, decoded to produce a sequence of phones, and subsequently indexed using phone N-grams. Search is performed by expanding queries into phone sequences and matching against the index. The approximate match score is derived from a CRF, trained on parallel transcripts, which provides a general framework for modeling the errors that a recognition system may make taking contextual effects into consideration. Our approach differs from other work in the field in that we focus on using CRFs to model context dependent phone level confusions, rather than on explicitly modeling parameters of an edit distance. While, the results we obtain on both in and out of vocabulary (OOV) search tasks improve on previous work which incorporated high order phone confusions, the gains for OOV are more impressive. Upendra V. Chaudhari, Michael Picheny |
ASRU | 1 |
| 2008 | MAAI: Media analytics for actionable intelligenceabstractGovernment agencies, corporations, and police departments are plagued by information overload. The inability of fully analyze fragments of data scattered across the organizations reduces productivity, and more and more of these fragments are being gathered every day thanks to tools like the Internet and digital audio/video recorders. However, since much of this information is stored in computer systems and networks, it is possible to develop tools that will automatically analyze, assemble and associate the fragments so that precious human resources can focus on the analysis and interpretation of just those fragments that may contain valuable insight. Media analytics for actionable intelligence (MAAI) is a combination of automatic, semi-automatic and manual tools which provide three basic levels of analysis: automatic services, manual feedback, and higher level mining (e.g. timeline, social network, hypothesis generation) which allow investigators/analysts to act more efficiently and accurately. Upendra V. Chaudhari, Sarah Conrod, Alexander Faisman, Giridharan Iyengar, Dimitri Kanevsky, Mark Kogan, Ganesh N. Ramaswamy, Paola Virga |
ICASSP | 1 |
| 2008 | Discriminative graph training for ultra-fast low-footprint speech indexing
Upendra V. Chaudhari, Hong-Kwang Jeff Kuo, Brian Kingsbury |
INTERSPEECH | 1 |
| 2007 | Improvements in phone based audio search via constrained match with high order confusion estimatesabstractThis paper investigates an approximate similarity measure for searching in phone based audio transcripts. The baseline method combines elements found in the literature to form an approach based on a phonetic confusion matrix that is used to determine the similarity of an audio document and a query, both of which are parsed into phoneN-grams. Experimental results show comparable performance to other approaches in the literature. Extensions of the approach are developed based on a constrained form of the similarity measure that can take into consideration the system dependent errors that can occur. This is done by accounting for higher order confusions, namely of phone bi-grams and tri-grams. Results show improved performance across a variety of system configurations. Upendra V. Chaudhari, Michael Picheny |
ASRU | 1 |
| 2007 | Fast audio search using vector space modellingabstractMany techniques for retrieving arbitrary content from audio have been developed to leverage the important challenge of providing fast access to very large volumes of multimedia data. We present a two-stage method for fast audio search, where a vector-space modelling approach is first used to retrieve a short list of candidate audio segments for a query. The list of candidate segments is then searched using a word-based index for known words and a phone-based index for out-of-vocabulary words. We explore various system configurations and examine trade-offs between speed and accuracy. We evaluate our audio search system according to the NIST 2006 Spoken Term Detection evaluation initiative. We find that we can obtain a 30-times speedup for the search phase of our system with a 10% relative loss in accuracy. Brett Matthews, Upendra V. Chaudhari, Bhuvana Ramabhadran |
ASRU | 2 |
| 2007 | Optimized one-bit quantization for adapted GMM-based speaker verification
Ivy H. Tseng, Olivier Verscheure, Deepak S. Turaga, Upendra V. Chaudhari |
INTERSPEECH | 4 |
| 2006 | QUANTization for Adapted GMM-Based Speaker VerificationabstractState-of-the-art speaker verification systems are built around the likelihood ratio test, using Gaussian mixture models (GMM) for likelihood functions, a universal background model (UBM) for alternative speaker representation, and a form of Bayesian adaptation to derive speaker models from the UBM. This work tackles optimal quantizer design of the speech cepstral features (MFCCs) for such systems. The problem is posed as the minimization of loss of log-likelihood ratio between the quantized and unquantized speech features. First we show that the conventional mean squared error (MSE) quantizer for the top-scoring UBM Gaussian is optimal under practical assumptions. Then we derive the optimal bit allocation strategy across the dimensions of the feature vectors. Finally we demonstrate the validity of the approach against various quantization and bit allocation schemes by running experiments on the appropriately modified IBM speaker verification system. Experimental results on the HUB4 corpora show negligible impact on verification performance for bit rates as low as less than 1 bit per dimension on average in contrast to 32 bits per dimension in the original system Ivy H. Tseng, Olivier Verscheure, Deepak S. Turaga, Upendra V. Chaudhari |
ICASSP (1) | 4 |
| 2006 | Resource Management for Networked Classifiers in Distributed Stream Mining SystemsabstractNetworks of classifiers are capturing the attention of system and algorithmic researchers because they offer improved accuracy over single model classifiers, can be distributed over a network of servers for improved scalability, and can be adapted to available system resources. This work provides a principled approach for the optimized allocation of system resources across a networked chain of classifiers. We begin with an illustrative example of how complex classification tasks can be decomposed into a network of binary classifiers. We formally define a global performance metric by recursively collapsing the chain of classifiers into one combined classifier. The performance metric trades off the end-to-end probabilities of detection and false alarm, both of which depend on the resources allocated to each individual classifier. We formulate the optimization problem and present optimal resource allocation results for both simulated and state-of-the-art classifier chains operating on telephony data. Deepak S. Turaga, Olivier Verscheure, Upendra V. Chaudhari, Lisa Amini |
ICDM | 3 |
| 2006 | Efficient Speaker Detection via Target Dependent Data ReductionabstractSystems designed to extract time-critical information from large volumes of unstructured data must include the ability, both from an architectural and algorithmic point of view, to filter out unimportant data that might otherwise overwhelm the available resources. This paper presents an approach for data filtering to reduce computation in the context of a distributed speech processing architecture designed to detect or identify speakers. Here, filtering means either dropping and ignoring data or passing it on for further processing. The goal of the paper is to show that when the filter is designed to select and pass on a subset of the input data that best preserves the ability to recognize a specific desired speaker, or group of speakers, a large percentage of the data can be ignored while being able to preserve most of the accuracy Upendra V. Chaudhari, Olivier Verscheure, Juan M. Huerta, Xiang Li 0071, Ganesh N. Ramaswamy, Lisa Amini |
ICME | 1 |
| 2005 | Blind Change Detection for Audio SegmentationabstractAutomatic segmentation of audio streams according to speaker identities and environmental and channel conditions has become an important preprocessing step for speech recognition, speaker recognition, and audio data mining. In most previous approaches, the automatic segmentation was evaluated in terms of the performance of the final system, like the word error rate for speech recognition systems. In many applications, like online audio indexing, and information retrieval systems, the actual boundaries of the segments are required. We present an approach based on the cumulative sum (CuSum) algorithm for automatic segmentation which minimizes the missing probability for a given false alarm rate. We compare the CuSum algorithm to the Bayesian information criterion (BIC) algorithm, and a generalization of the Kolmogorov-Smirnov test for automatic segmentation of audio streams. We present a two-step variation of the three algorithms which improves the performance significantly. We present also a novel approach that combines hypothesized boundaries from the three algorithms to achieve the final segmentation of the audio stream. Our experiments, on the 1998 Hub4 broadcast news, show that a variation of the CuSum algorithm significantly outperforms the other two approaches and that combining the three approaches using a voting scheme improves the performance slightly compared to using the a two-step variation of the CuSum algorithm alone. Mohamed Kamal Omar, Upendra V. Chaudhari, Ganesh N. Ramaswamy |
ICASSP (1) | 2 |
| 2005 | Adaptive speech analytics: system, infrastructure, and behavior
Upendra V. Chaudhari, Ganesh N. Ramaswamy, Edward A. Epstein, Sasha Caskey, Mohamed Kamal Omar |
INTERSPEECH | 1 |
| 2004 | Policy analysis framework for conversational biometrics
Upendra V. Chaudhari, Ganesh N. Ramaswamy |
INTERSPEECH | 1 |
| 2003 | Audio-visual speaker recognition using time-varying stream reliability predictionabstractWe examine a time-varying, context dependent, information fusion methodology for multi-stream authentication based on audio and video data collected simultaneously during a user's interaction with a system. Scores obtained from the two data streams are combined based on the relative local richness, as compared to the training data or derived model, and on the stability of each stream. The results show that the proposed technique outperforms the use of video or audio data alone as well as the use of fused data streams (via concatenation). Of particular note is that the performance improvements are achieved for clean, high quality speech, whereas previous efforts focused on degraded speech conditions. Upendra V. Chaudhari, Ganesh N. Ramaswamy, Gerasimos Potamianos, Chalapathy Neti |
ICASSP (5) | 1 |
| 2003 | The IBM system for the NIST-2002 cellular speaker verification evaluationabstractThis paper presents an overview of the architecture and algorithms implemented in IBM's text-independent speaker verification system developed for the 2002 NIST speaker recognition evaluation, particularly for the 1-speaker detection task using cellular test data. We describe individual components including a Gaussianization front-end, celluar-codec post-processing, modeling, discriminative optimization and scoring steps. A combination of multiple, data-perturbed systems using a discriminative objective so as to achieve optimum performance for a low false alarm operating region obtained the top performance in the NIST 2002 1-speaker detection task. Ganesh N. Ramaswamy, Jirí Navrátil 0001, Upendra V. Chaudhari, Ran D. Zilca |
ICASSP (2) | 3 |
| 2003 | Information fusion and decision cascading for audio-visual speaker recognition based on time-varying stream reliability predictionabstractWe examine the techniques for multi-modal biometric information fusion for verification and identification of speakers, where the reliability of each data stream, either audio of video, is modeled with parameters that are time-varying and depend on the context created by its local behavior. The complementary nature and the time dependent relative reliability of audio and video data is studied in the context of verification and identification, on data collected during a user's interaction with an automated system. Of significance is that this data is not corrupted artificially. Particular focus is directed to verification and its ability to refine identification decisions, by indicating a level of confidence in the system decisions. Results show more striking effects for verification, when using time-dependent fusion, than for identification. Upendra V. Chaudhari, Ganesh N. Ramaswamy, Gerasimos Potamianos, Chalapathy Neti |
ICME | 1 |
| 2003 | Impact of audio segmentation and segment clustering on automated transcription accuracy of large spoken archivesabstractThis paper addresses the influence of audio segmentation and segment clustering on automatic transcription accuracy for large spoken archives. The work forms part of the ongoing MALACH project, which is developing advanced techniques for supporting access to the world’s largest digital archive of video oral histories collected in many languages from over 52000 survivors and witnesses of the Holocaust. We present several audio-only and audio-visual segmentation schemes, including two novel schemes: the first is iterative and audio-only, the second uses audio-visual synchrony. Unlike most previous work, we evaluate these schemes in terms of their impact upon recognition accuracy. Results on English interviews show the automatic segmentation schemes give performance comparable to (exhorbitantly expensive and impractically lengthy) manual segmentation when using a single pass decoding strategy based on speaker-independent models. However, when using a multiple pass decoding strategy with adaptation, results are sensitive to both initial audio segmentation and the scheme for clustering segments prior to adaptation: the combination of our best automatic segmentation and clustering scheme has an error rate 8% worse (relative) to manual audio segmentation and clustering due to the occurrence of “speaker-impure ” segments. 1. Bhuvana Ramabhadran, Jing Huang 0019, Upendra V. Chaudhari, Giridharan Iyengar, Harriet J. Nock |
INTERSPEECH | 3 |
| 2003 | An architecture for rapid decoding of large vocabulary conversational speechabstractThis paper addresses the question of how to design a large vocabulary recognition system so that it can simultaneously handle a sophisticated language model, perform state-ofthe-art speaker adaptation, and run in one times real time 1 (1 RT). The architecture we propose is based on classical HMM Viterbi decoding, but uses an extremely fast initial speaker-independent decoding to estimate VTL warp factors, feature-space and model-space MLLR transformations that are used in a final speaker-adapted decoding. We present results on past Switchboard evaluation data that indicate that this strategy compares favorably to published unlimited-time systems (running in several hundred times real-time). Coincidentally, this is the system that IBM fielded in the 2003 EARS Rich Transcription evaluation. 1. George Saon, Geoffrey Zweig, Brian Kingsbury, Lidia Mangu, Upendra V. Chaudhari |
INTERSPEECH | 5 |
| 2003 | Multigrained modeling with pattern specific maximum likelihood transformations for text-independent speaker recognitionabstractWe present a transformation-based, multigrained data modeling technique in the context of text independent speaker recognition, aimed at mitigating difficulties caused by sparse training and test data. Both identification and verification are addressed, where we view the entire population as divided into the target population and its complement, which we refer to as the background population. First, we present our development of maximum likelihood transformation based recognition with diagonally constrained Gaussian mixture models and show its robustness to data scarcity with results on identification. Then for each target and background speaker, a multigrained model is constructed using the transformation based extension as a building block. The training data is labeled with an HMM based phone labeler. We then make use of a graduated phone class structure to train the speaker model at various levels of detail. This structure is a tree with the root node containing all the phones. Subsequent levels partition the phones into increasingly finer grained linguistic classes. This method affords the use of fine detail where possible, i.e., as reflected in the amount of training data distributed to each tree node. We demonstrate the effectiveness of the modeling with verification experiments in matched and mismatched conditions. Upendra V. Chaudhari, Jirí Navrátil 0001, Stéphane H. Maes |
IEEE Trans. Speech Audio Process. | 1 |
| 2002 | Short-time Gaussianization for robust speaker verificationabstractIn this paper, a novel approach for robust speaker verification, namely short-time Gaussianization, is proposed. Short-time Gaussianization is initiated by a global linear transformation of the features, followed by a short-time windowed cumulative distribution function (CDF) matching. First, the linear transformation in the feature space leads to local independence or decorrelation. Then the CDF matching is applied to segments of speech localized in time and tries to warp a given feature so that its CDF matches normal distribution. It is shown that one of the recent techniques used for speaker recognition, feature warping [l] can be formulated within the framework of Gaussianization. Compared to the baseline system with cepstral mean subtraction (CMS), around 20% relative improvement in both equal error rate(EER) and minimum detection cost function (DCF) is obtained on NIST 2001 cellular phone data evaluation. Bing Xiang, Upendra V. Chaudhari, Jirí Navrátil 0001, Ganesh N. Ramaswamy, Ramesh A. Gopinath |
ICASSP | 2 |
| 2002 | The sphericity measure for cellular speaker verificationabstractThis paper provides a description of the properties of the symmetric sphericity measure between two Gaussian classes, and presents experimental results of its use in the context of text independent speaker verification, performed on cellular speech. A novel geometric interpretation of the sphericity measure is presented, emphasizing its robustness to variations in scaling in feature space compared to other distortion measures, along with an associated score normalization procedure. The experimental results clearly indicate the superiority of this method over the prevailing likelihood ratio approach, motivating further use of the sphericity measure in a multi-component modeling scheme. In addition, the direct use of this method with single Gaussian models is computationally extremely simple, and may provide acceptable performance in certain cases. Ran D. Zilca, Upendra V. Chaudhari, Ganesh N. Ramaswamy |
ICASSP | 2 |
| 2001 | Very large population text-independent speaker identification using transformation enhanced multi-grained modelsabstractPresents results on speaker identification with a population size of over 10000 speakers. Speaker modeling is accomplished via our transformation enhanced multigrained models. Pursuing two goals, the first is to study the performance of a number of different systems within the modeling framework of multi-grained models. The second is to analyze performance as a function of population size. We show that the most complex models within the framework perform the best and demonstrate that, in approximation, the identification error rate scales linearly with the log of the population size for the described system. Further, we develop a candidate rejection technique based on our analysis of the system performance which indicates a low confidence in the identity chosen. Upendra V. Chaudhari, Jirí Navrátil 0001, Ganesh N. Ramaswamy, Stéphane H. Maes |
ICASSP | 1 |
| 2001 | Speaker verification using target and background dependent linear transforms and multi-system fusionabstractThis paper describes a GMM-based speaker verification system that uses speaker-dependent background models transformed by speaker-specific maximum likelihood linear transforms to achieve a sharper separation between the target and the nontarget acoustic region. The effect of tying, or coupling, Gaussian components between the target and the background model is studied and shown to be a relevant factor with respect to the desired operating point. A fusion of scores from multiple systems built on different acoustic features via a neural network with performance gains over linear combination is also presented. The methods are experimentally studied on the 1999 NIST speaker recognition evaluation data. Jirí Navrátil 0001, Upendra V. Chaudhari, Ganesh N. Ramaswamy |
INTERSPEECH | 2 |
| 2000 | Transformation enhanced multi-grained modeling for text-independent speaker recognitionabstractWe describe our formulation of transformation enhanced data modeling used to develop a multi-grained data analysis approach to text independent speaker recognition. The broad goal is to address difficulties caused by sparse training and test data. First, our development of maximum likelihood transformation based recognition with diagonally constrained Gaussian mixture models is detailed. We give results to show its robustness to decreasing training data. Then using the these models as building blocks, a multigrained model structure is developed. For this, the training data must be labeled, e.g. with an HMM based phone labeler. A graduated phone class structure is then used to train the speaker model at various levels of detail. This structure is a tree with the root node containing all the phones. Subsequent levels partition the phones into increasingly finer grained linguistic classes. We demonstrate the effectiveness of the modeling with identification and verification experiments. 1. Upendra V. Chaudhari, Jirí Navrátil 0001, Stéphane H. Maes, Ramesh A. Gopinath |
INTERSPEECH | 1 |
| 1999 | A hierarchical approach to large-scale speaker recognitionabstractCORRECTIVE TRAINING FOR SPEAKER ADAPTATIONXiuyang Yu and Wayne WardCenter for Spoken Language UnderstandingUniversity of Colorado, Boulder, ColoradoABSTRACTThis paper reports results on an experiment to use correctivetraining techniques for rapid acoustic speaker adaptation in asemi-continuous speech recognition system. Decoder outputis used to adjust HMM acoustic models to improvediscrimination between correct words and near misses.Twenty sentences are used as an adaptation set. A speechrecognizer is run on each utterance to generate a wordlattice. The lattice is pruned relative to the correct path. Theforward-backward algorithm is used to align each path in thelattice against the speech input and compute observationcounts. For each input frame, counts in correct models areadjusted upward, and counts in incorrect models are adjusteddownward. The adjusted counts are normalized to generatenew observation probabilities for the models. Theparameters being adjusted are the mixture weights for thesemi-continuous HMMs. The technique reduced word errorfor a test subject by 37% relative.INTRODUCTIONSpeech recognition systems based on Hidden MarkovModels typically experience significant performancedegradation when a speaker is not represented well in thedata that the system was trained on. Rapid speakeradaptation techniques can be very effective in improvingperformance for a novel speaker. These techniques use asmall number of sentences (20-40), whose transcript isknown, to quickly adapt HMM acoustic models to a newspeaker. Such techniques are important, since in order forlonger term adaptation to be able to work, the system mustfunction well enough for the user to be productive. Thesetechniques attempt to correct very poor models quickly toget the system to a usable point for the user. This paperreports results on an experiment to use corrective trainingtechniques to rapidly adapt HMM acoustic models to apoorly recognized speaker.CORRECTIVE TRAININGCorrective training was introduced for speaker-dependentisolated word recognition by [1] and extended to speaker-independent continuous speech recognition by [2]. HiddenMarkov Models are normally trained according to amaximum likelihood criterion; parameters are adjusted tomaximize the probability of the training set. This processdoes not directly minimize the word errors, it does thisindirectly by attempting to assign a high probability to thecorrect utterance. It makes no representation of whatprobability is assigned to near misses. Corrective trainingseeks to directly minimize the number of word errors byadjusting parameters so as to improve discriminationbetween correct words and near misses. The general processis:1. Generate a set of near misses, words which areconfusable with the correct words.2. Align the correct words against the input.3. Align a near miss against the input.4. Modify model such that correct words are more likelyand incorrect ones less likely.5. Repeat steps 2-3 for other near misses.Speech recognizers are used to generate the near misses. In[1] and [2], a sp eech recognizer was used to generate an n-best list for an isolated word or sentence. This list was usedas the set of near misses. Both the correct utterance and nearmiss are aligned against the input. Model parameters arethen adjusted to make the correct words more likely and theincorrect ones less likely.SPHINX-II OBSERVATION ESTIMATESIn order to describe how corrective training is used to adaptour model, it is first necessary to describe the basic model.Our experimental system uses the Carnegie MellonUniversity Sphinx-II system for speech recognition[3],[4],[5]. Sphinx-II uses semi-continuous Hidden MarkovModels [3] to model context dependent phones. Likecontinuous HMM systems, semi-continuous systems use aweighted sum of points on Gaussian Probability DensityFunctions to estimate observation probabilities. Thedifference is that, while continuous systems estimate a set ofdistributions for each HMM state in the system, semi- Homayoon Beigi, Stéphane H. Maes, Upendra V. Chaudhari, Jeffrey S. Sorensen |
EUROSPEECH | 3 |
| 1995 | Reducing word error rate on conversational speech from the Switchboard corpusabstractSpeech recognition of conversational speech is a difficult task. The performance levels on the Switchboard corpus had been in the vicinity of 70% word error rate. In this paper, we describe the results of applying a variety of modifications to our speech recognition system and we show their impact on improving the performance on conversational speech. These modifications include the use of more complex models, trigram language models, and cross-word triphone models. We also show the effect of using additional acoustic training on the recognition performance. Finally, we present an approach to dealing with the abundance of short words, and examine how the variable speaking rate found in conversational speech impacts on the performance. Currently, the level of performance is at the vicinity of 50% error, a significant improvement over recent levels. Philippe Jeanrenaud, Ellen Eide, Upendra V. Chaudhari, John W. McDonough, Kenney Ng, Man-Hung Siu, Herbert Gish |
ICASSP | 3 |