EDBT 2026 Demo / reviewers in the wild / expert
Kartik Audhkhasi
dblp:04/8055
· DBLP profile ↗
72ranked-venue papers
25as first author
15since 2021 · last 2025
0000-0002-2340-1144ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 62 · 20 first-author · 14 since 2021Artificial intelligence and machine learning · 37 · 14 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Audio Diffusion with Large Language ModelsabstractIn this paper, we explore an alternate approach to the popular method of using large language models (LLMs) as a second decoder for Automated Speech Recognition (ASR) and speech understanding tasks. We propose to employ diffusion networks to generate a correction signal that can be applied on the original input audio features to improve performance. Specifically, the diffusion network is trained to predict the gradient of any ASR objective with respect to the input audio features conditioned on LLM embeddings. Our experiments are conducted on public corpora, namely, Librispeech and Common Voice. We show that the diffusion model is able to improve ASR performance on noisy and accented speech, with the addition of knowledge from the LLM, and also helps improve generalization to out-of-domain test sets. Kyle Kastner, Kartik Audhkhasi, Bhuvana Ramabhadran, Andrew Rosenberg |
ICASSP | 3 |
| 2025 | Weak-to-Strong Generalization in Speech RecognitionabstractTo surpass human-level accuracy, speech recognition models must go beyond relying solely on human labels. To this end, we must build stronger models from weaker supervisors and this is the main goal in weak-to-strong generalization (WSG). WSG methods normally incorporate additional information into weak teacher models to improve their performance for example reliability of teacher-generated labels. In this research, we investigate two sources of additional information to implement WSG for speech recognition: unsupervised data from the target language and supervised data from other similar languages. We study scenarios where unsupervised data boosts performance, and propose a new clustering method leveraging supervised data from similar languages for further gains. Our clustering method yields an average 8% reduction in word error rate compared to a universal speech model trained on 182 languages. By incorporating both supervised data from similar languages and unsupervised data from the target language, we further enhance the USM model by 10%. This improvement reaches 15% for the top-performing languages with a WER below 50%. Soheil Khorram, Rohit Prabhavalkar, Kartik Audhkhasi, Bhuvana Ramabhadran |
ICASSP | 4 |
| 2025 | Identifying and Mitigating Mismatched Language Code in Multilingual ASRabstractMultilingual speech recognition systems often use an input language code in order to prompt the transcription in the target language. However, the spoken language in the input audio may not always match the language code, as often prevalent in multilingual societies. This language mismatch can significantly reduce ASR quality. We present a technique to identify and mitigate this issue. We combine off-the-shelf language-ID and language verification models to determine the language code input to the ASR model. The language verification model acts as a gate that decides when to trust the provided language code or use the output of the language-ID model. We compare these approaches with baselines that include vanilla language-ID based and language-independent ASR models. Our experiments on YouTube, SPRING-INX and FLEURS datasets shows the efficacy of the proposed model especially in the mismatched language code setting. Sepand Mavandadi, Kartik Audhkhasi, Shikhar Bharadwaj, Brian Farris, Tongzhou Chen, Bhuvana Ramabhadran, Sriram Ganapathy |
ICASSP | 3 |
| 2024 | Task Vector Algebra for ASR ModelsabstractVector representations of text and speech signals such as word2vec and wav2vec are used commonly in automatic speech recognition (ASR) and spoken language understanding systems. Recent results in natural language processing have proposed a task vector, defined as the difference vector between a model’s converged parameters and its initial parameters. Task vector algebra provides a simple and computationally-efficient way to solve several modeling problems, including model editing to reduce undesirable behavior, multi-tasking, and improving domain generalization. We apply task vectors to ASR models for the first time. Our experiments with Conformer-RNNT models trained on the SpeechStew corpora show that task vectors retain their scaling and multi-tasking applications. We propose two novel applications of task vectors to ASR. First, we show that task vectors can perform zero-shot adaptation of ASR models to unseen domains without using any supervised data. Second, we present a novel "task analogy" formulation that enables us to use models trained on high-resource tasks to improve performance on low-resource tasks. We also explore a technique to improve the performance of task vector arithmetic for ASR models. Gowtham Ramesh, Kartik Audhkhasi, Bhuvana Ramabhadran |
ICASSP | 2 |
| 2023 | Modular Conformer Training for Flexible End-to-End ASRabstractThe state-of-the-art conformer used in automatic speech recognition combines feed-forward, convolution and multi-headed self-attention layers in a single model that is trained end-to-end with a decoder network. While this end-to-end training is simple and beneficial for word error rate, it restricts the ability to perform inference with the model at different operating points of word error rate and latency. Existing approaches to overcome this limitation include cascaded encoders and variable attention context models. We propose an alternative approach, called Modular Conformer training, which splits the Conformer model into a backbone convolutional model and attention submodels, which are added at each layer. We conduct experiments with a few training techniques on the Librispeech and Librilight corpus. We show that dropping-out the attention layers during the training of the backbone model allows for the largest WER improvements upon adding fine-tuned attention submodels, without impacting the WER of the backbone model itself. Kartik Audhkhasi, Brian Farris, Bhuvana Ramabhadran, Pedro J. Moreno 0001 |
ICASSP | 1 |
| 2023 | Large-Scale Language Model Rescoring on Long-Form DataabstractIn this work, we study the impact of Large-scale Language Models (LLM) on Automated Speech Recognition (ASR) of YouTube videos, which we use as a source for long-form ASR. We demonstrate up to 8% relative reduction in Word Error Eate (WER) on US English (en-us) and code-switched Indian English (en-in) long-form ASR test sets and a reduction of up to 30% relative on Salient Term Error Rate (STER) over a strong first-pass baseline that uses a maximum-entropy based language model. Improved lattice processing that results in a lattice with a proper (non-tree) digraph topology and carrying context from the 1-best hypothesis of the previous segment(s) results in significant wins in rescoring with LLMs. We also find that the gains in performance from the combination of LLMs trained on vast quantities of available data (such as C4 [1]) and conventional neural LMs is additive and significantly outperforms a strong first-pass baseline with a maximum entropy LM. Tongzhou Chen, Cyril Allauzen, Daniel S. Park, David Rybach, W. Ronny Huang, Rodrigo Cabrera, Kartik Audhkhasi, Bhuvana Ramabhadran, Pedro J. Moreno 0001, Michael Riley 0001 |
ICASSP | 8 |
| 2023 | Robust Knowledge Distillation from RNN-T Models with Noisy Training Labels Using Full-Sum LossabstractThis work studies knowledge distillation (KD) and addresses its constraints for recurrent neural network transducer (RNN-T) models. In hard distillation, a teacher model transcribes large amounts of unlabelled speech to train a student model. Soft distillation is another popular KD method that distills the output logits of the teacher model. Due to the nature of RNN-T alignments, applying soft distillation between RNNT architectures having different posterior distributions is challenging. In addition, bad teachers having high word-error-rate (WER) reduce the efficacy of KD. We investigate how to effectively distill knowledge from variable quality ASR teachers, which has not been studied before to the best of our knowledge. We show that a sequence-level KD, full-sum distillation, outperforms other distillation methods for RNN-T models, especially for bad teachers. We also propose a variant of full-sum distillation that distills the sequence discriminative knowledge of the teacher leading to further improvement in WER. We conduct experiments on public datasets namely SpeechStew and LibriSpeech, and on in-house production data. Mohammad Zeineldeen, Kartik Audhkhasi, Murali Karthick Baskar, Bhuvana Ramabhadran |
ICASSP | 2 |
| 2023 | O-1: Self-training with Oracle and 1-best Hypothesis
Murali Karthick Baskar, Andrew Rosenberg, Bhuvana Ramabhadran, Kartik Audhkhasi |
INTERSPEECH | 4 |
| 2022 | Federated Learning for Affective Computing TasksabstractFederated learning mitigates the need to store user data in a central datastore for machine learning tasks, and is particularly beneficial when working with sensitive user data or tasks. Although successfully used for applications such as improving keyboard query suggestions, it is not studied systematically for modeling affective computing tasks which are often laden with subjective labels and high variability across individuals/raters or even by the same participant. In this paper, we study the federated averaging algorithm FedAvg to model self-reported emotional experience and perception labels on a variety of speech, video and text datasets. We identify two learning paradigms that commonly arise in affective computing tasks: modeling of self-reports (user-as-client), and modeling perceptual judgments such as labeling sentiment of online comments (rater-as-client). In the user-as-client setting, we show that FedAvg generally performs on-par with a non-federated model in classifying self-reports. In the rater-as-client setting, FedAvg consistently performed poorer than its non-federated counterpart. We found that the performance of FedAvg degraded for classes where the inter-rater agreement was moderate to low. To address this finding, we propose an algorithm FedRater that learns client-specific label distributions in federated settings. Our experimental results show that FedRater not only improves the overall classification performance compared to FedAvg but also provides insights for estimating proxies of inter-rater agreement in distributed settings. Krishna Somandepalli, Brian Eoff, Alan Cowen, Kartik Audhkhasi, Josh Belanich, Brendan Jou |
ACII | 5 |
| 2022 | Analysis of Self-Attention Head Diversity for Conformer-based Automatic Speech RecognitionabstractAttention layers are an integral part of modern end-to-end automatic speech recognition systems, for instance as part of the Transformer or Conformer architecture.Attention is typically multi-headed, where each head has an independent set of learned parameters and operates on the same input feature sequence.The output of multi-headed attention is a fusion of the outputs from the individual heads.We empirically analyze the diversity between representations produced by the different attention heads and demonstrate that the heads become highly correlated during the course of training.We investigate a few approaches to increasing attention head diversity, including using different attention mechanisms for each head and auxiliary training loss functions to promote head diversity.We show that introducing diversity-promoting auxiliary loss functions during training is a more effective approach, and obtain WER improvements of up to 6% relative on the Librispeech corpus.Finally, we draw a connection between the diversity of attention heads and the similarity of the gradients of head parameters. Kartik Audhkhasi, Bhuvana Ramabhadran, Pedro J. Moreno 0001 |
INTERSPEECH | 1 |
| 2022 | Modular Hybrid Autoregressive TransducerabstractText-only adaptation of a transducer model remains challenging for end-to-end speech recognition since the transducer has no clearly separated acoustic model (AM), language model (LM) or blank model. In this work, we propose a modular hybrid autoregressive transducer (MHAT) that has structurally separated label and blank decoders to predict label and blank distributions, respectively, along with a shared acoustic encoder. The encoder and label decoder outputs are directly projected to AM and internal LM scores and then added to compute label posteriors. We train MHAT with an internal LM loss and a HAT loss to ensure that its internal LM becomes a standalone neural LM that can be effectively adapted to text. Moreover, text adaptation of MHAT fosters a much better LM fusion than internal LM subtraction-based methods. On Google's large-scale production data, a multi-domain MHAT adapted with 100B sentences achieves relative WER reductions of up to 12.4% without LM fusion and 21.5% with LM fusion from 400K-hour trained HAT. Zhong Meng, Tongzhou Chen, Rohit Prabhavalkar, Yu Zhang 0033, Gary Wang, Kartik Audhkhasi, Jesse Emond, Trevor Strohman, Bhuvana Ramabhadran, W. Ronny Huang, Ehsan Variani, Pedro J. Moreno 0001 |
SLT | 6 |
| 2021 | Convolutional Dropout and Wordpiece Augmentation for End-to-End Speech RecognitionabstractRegularization and data augmentation are crucial to training end-to-end automatic speech recognition systems. Dropout is a popular regularization technique, which operates on each neuron independently by multiplying it with a Bernoulli random variable. We propose a generalization of dropout, called "convolutional dropout", where each neuron’s activation is replaced with a randomly-weighted linear combination of neuron values in its neighborhood. We believe that this formulation combines the regularizing effect of dropout with the smoothing effects of the convolution operation. In addition to convolutional dropout, this paper also proposes using random word-piece segmentations as a data augmentation scheme during training, inspired by results in neural machine translation. We adopt both these methods during the training of transformer-transducer speech recognition models, and show consistent WER improvements on Librispeech as well as across different languages. Hainan Xu, Kartik Audhkhasi, Bhuvana Ramabhadran |
ICASSP | 4 |
| 2021 | Mixture Model Attention: Flexible Streaming and Non-Streaming Automatic Speech Recognition
Kartik Audhkhasi, Tongzhou Chen, Bhuvana Ramabhadran, Pedro J. Moreno 0001 |
Interspeech | 1 |
| 2021 | AVLnet: Learning Audio-Visual Language Representations from Instructional VideosabstractCurrent methods for learning visually grounded language from videos often rely on text annotation, such as human generated captions or machine generated automatic speech recognition (ASR) transcripts. In this work, we introduce the Audio-Video Language Network (AVLnet), a self-supervised network that learns a shared audio-visual embedding space directly from raw video inputs. To circumvent the need for text annotation, we learn audio-visual representations from randomly segmented video clips and their raw audio waveforms. We train AVLnet on HowTo100M, a large corpus of publicly available instructional videos, and evaluate on image retrieval and video retrieval tasks, achieving state-of-the-art performance. We perform analysis of AVLnet's learned representations, showing our model utilizes speech and natural sounds to learn audio-visual concepts. Further, we propose a tri-modal model that jointly processes raw audio, video, and text captions from videos to learn a multi-modal semantic embedding space useful for text-video retrieval. Our code, data, and trained models will be released at avlnet.csail.mit.edu Andrew Rouditchenko, Angie W. Boggust, David F. Harwath, Brian Chen 0001, Dhiraj Joshi, Samuel Thomas 0001, Kartik Audhkhasi, Hilde Kuehne, Rameswar Panda, Rogério Feris, Brian Kingsbury, Michael Picheny, Antonio Torralba 0001, James R. Glass |
Interspeech | 7 |
| 2021 | Regularizing Word Segmentation by Creating Misspellings
Hainan Xu, Kartik Audhkhasi, Jesse Emond, Bhuvana Ramabhadran |
Interspeech | 2 |
| 2020 | Leveraging Unpaired Text Data for Training End-To-End Speech-to-Intent SystemsabstractTraining an end-to-end (E2E) neural network speech-to-intent (S2I) system that directly extracts intents from speech requires large amounts of intent-labeled speech data, which is time consuming and expensive to collect. Initializing the S2I model with an ASR model trained on copious speech data can alleviate data sparsity. In this paper, we attempt to leverage NLU text resources. We implemented a CTC-based S2I system that matches the performance of a state-of-the-art, traditional cascaded SLU system. We performed controlled experiments with varying amounts of speech and text training data. When only a tenth of the original data is available, intent classification accuracy degrades by 7.6% absolute. Assuming we have additional text-to-intent data (without speech) available, we investigated two techniques to improve the S2I system: (1) transfer learning, in which acoustic embeddings for intent classification are tied to fine-tuned BERT text embeddings; and (2) data augmentation, in which the text-to-intent data is converted into speech-to-intent data using a multi-speaker text-to-speech system. The proposed approaches recover 80% of performance lost due to using limited intent-labeled speech. Hong-Kwang Jeff Kuo, Samuel Thomas 0001, Zvi Kons, Kartik Audhkhasi, Brian Kingsbury, Ron Hoory, Michael Picheny |
ICASSP | 5 |
| 2020 | Alignment-Length Synchronous Decoding for RNN TransducerabstractWe present a beam decoding strategy for recurrent neural network transducers which has the characteristic that all competing hypotheses within the beam have the same alignment length (number of output symbols plus BLANK symbols). We contrast the proposed technique with time-synchronous decoding where the competing hypotheses within the beam correspond to the same input frames (but can have different length output sequences). Experiments on the Switchboard 2000 hours corpus show that alignment-length synchronous decoding (ALSD) is 25% faster than time-synchronous decoding (TSD) for the same accuracy because ALSD performs 42% fewer joint network evaluations and hypothesis expansions during the search. Additionally, we discuss the benefit of caching and batching the prediction and joint network evaluations, of using prefix trees instead of full output vocabulary expansions, and of performing hypothesis recombination after pruning. With open beam decoding, we reach a 6.2% / 10.9% word error rate on the Switchboard and CallHome Hub5 2000 evaluation testsets which compares favorably to other published single-model results on this corpus. George Saon, Zoltán Tüske, Kartik Audhkhasi |
ICASSP | 3 |
| 2020 | Transliteration Based Data Augmentation for Training Multilingual ASR Acoustic Models in Low Resource Settings
Samuel Thomas 0001, Kartik Audhkhasi, Brian Kingsbury |
INTERSPEECH | 2 |
| 2020 | End-to-End Spoken Language Understanding Without Full TranscriptsabstractAn essential component of spoken language understanding (SLU) is slot filling: representing the meaning of a spoken utterance using semantic entity labels. In this paper, we develop end-to-end (E2E) spoken language understanding systems that directly convert speech input to semantic entities and investigate if these E2E SLU models can be trained solely on semantic entity annotations without word-for-word transcripts. Training such models is very useful as they can drastically reduce the cost of data collection. We created two types of such speech-to-entities models, a CTC model and an attention-based encoder-decoder model, by adapting models trained originally for speech recognition. Given that our experiments involve speech input, these systems need to recognize both the entity label and words representing the entity value correctly. For our speech-to-entities experiments on the ATIS corpus, both the CTC and attention models showed impressive ability to skip non-entity words: there was little degradation when trained on just entities versus full transcripts. We also explored the scenario where the entities are in an order not necessarily related to spoken order in the utterance. With its ability to do re-ordering, the attention model did remarkably well, achieving only about 2% degradation in speech-to-bag-of-entities F1 score. Hong-Kwang Jeff Kuo, Zoltán Tüske, Samuel Thomas 0001, Kartik Audhkhasi, Brian Kingsbury, Gakuto Kurata, Zvi Kons, Ron Hoory, Luis A. Lastras |
INTERSPEECH | 5 |
| 2020 | Single Headed Attention Based Sequence-to-Sequence Model for State-of-the-Art Results on SwitchboardabstractIt is generally believed that direct sequence-to-sequence (seq2seq) speech recognition models are competitive with hybrid models only when a large amount of data, at least a thousand hours, is available for training. In this paper, we show that state-of-the-art recognition performance can be achieved on the Switchboard-300 database using a single headed attention, LSTM based model. Using a cross-utterance language model, our single-pass speaker independent system reaches 6.4% and 12.5% word error rate (WER) on the Switchboard and CallHome subsets of Hub5'00, without a pronunciation lexicon. While careful regularization and data augmentation are crucial in achieving this level of performance, experiments on Switchboard-2000 show that nothing is more useful than more data. Overall, the combination of various regularizations and a simple but fairly large model results in a new state of the art, 4.7% and 7.8% WER on the Switchboard and CallHome sets, using SWB-2000 without any external data resources. Zoltán Tüske, George Saon, Kartik Audhkhasi, Brian Kingsbury |
INTERSPEECH | 3 |
| 2020 | Noise can speed backpropagation learning and deep bidirectional pretraining
Bart Kosko, Kartik Audhkhasi, Osonde Osoba |
Neural Networks | 2 |
| 2019 | Simplified LSTMS for Speech RecognitionabstractIn this paper we explore new variants of Long Short-Term Memory (LSTM) networks for sequential modeling of acoustic features. In particular, we show that: (i) removing the output gate, (ii) replacing the hyperbolic tangent nonlinearity at the cell output with hard tanh, and (iii) collapsing the cell and hidden state vectors leads to a model that is conceptually simpler than and comparable in effectiveness to a regular LSTM for speech recognition. The proposed model has 25% fewer parameters than an LSTM with the same number of cells, trains faster because it has larger gradients leading to larger steps in weight space, and reaches a better optimum because there are fewer nonlinearities to traverse across layers. We report experimental results for both hybrid and CTC acoustic models on three publicly available English datasets: Switchboard 300 hours telephone conversations, 400 hours broadcast news transcription, and the MALACH 176 hours corpus of Holocaust survivor testimonies. In all cases the proposed models achieve similar or better accuracy than regular LSTMs while being conceptually simpler. George Saon, Zoltán Tüske, Kartik Audhkhasi, Brian Kingsbury, Michael Picheny, Samuel Thomas 0001 |
ASRU | 3 |
| 2019 | Sequence Noise Injected Training for End-to-end Speech RecognitionabstractWe present a simple noise injection algorithm for training end-to-end ASR models which consists in adding to the spectra of training utterances the scaled spectra of random utterances of comparable length. We conjecture that the sequence information of the "noise" utterances is important and verify this via a contrast experiment where the frames of the utterances to be added are randomly shuffled. Experiments for both CTC and attention-based models show that the pro-posed scheme results in up to 9% relative word error rate improvements (depending on the model and test set) on the Switchboard 300 hours English conversational telephony database. Additionally, we set a new benchmark for attention-based encoder-decoder models on this corpus. George Saon, Zoltán Tüske, Kartik Audhkhasi, Brian Kingsbury |
ICASSP | 3 |
| 2019 | Acoustically Grounded Word Embeddings for Improved Acoustics-to-word Speech RecognitionabstractDirect acoustics-to-word (A2W) systems for end-to-end automatic speech recognition are simpler to train, and more efficient to decode with, than sub-word systems. However, A2W systems can have difficulties at training time when data is limited, and at decoding time when recognizing words outside the training vocabulary. To address these shortcomings, we investigate the use of recently proposed acoustic and acoustically grounded word embedding techniques in A2W systems. The idea is based on treating the final pre-softmax weight matrix of an AWE recognizer as a matrix of word embedding vectors, and using an externally trained set of word embeddings to improve the quality of this matrix. In particular we introduce two ideas: (1) Enforcing similarity at training time between the external embeddings and the recognizer weights, and (2) using the word embeddings at test time for predicting out-of-vocabulary words. Our word embedding model is acoustically grounded, that is it is learned jointly with acoustic embeddings so as to encode the words' acoustic-phonetic content; and it is parametric, so that it can embed any arbitrary (potentially out-of-vocabulary) sequence of characters. We find that both techniques improve the performance of an A2W recognizer on conversational telephone speech. Shane Settle, Kartik Audhkhasi, Karen Livescu, Michael Picheny |
ICASSP | 2 |
| 2019 | Forget a Bit to Learn Better: Soft Forgetting for CTC-Based Automatic Speech Recognition
Kartik Audhkhasi, George Saon, Zoltán Tüske, Brian Kingsbury, Michael Picheny |
INTERSPEECH | 1 |
| 2019 | Guiding CTC Posterior Spike Timings for Improved Posterior Fusion and Knowledge DistillationabstractConventional automatic speech recognition (ASR) systems trained from frame-level alignments can easily leverage posterior fusion to improve ASR accuracy and build a better single model with knowledge distillation.End-to-end ASR systems trained using the Connectionist Temporal Classification (CTC) loss do not require frame-level alignment and hence simplify model training.However, sparse and arbitrary posterior spike timings from CTC models pose a new set of challenges in posterior fusion from multiple models and knowledge distillation between CTC models.We propose a method to train a CTC model so that its spike timings are guided to align with those of a pre-trained guiding CTC model.As a result, all models that share the same guiding model have aligned spike timings.We show the advantage of our method in various scenarios including posterior fusion of CTC models and knowledge distillation between CTC models with different architectures.With the 300-hour Switchboard training data, the single word CTC model distilled from multiple models improved the word error rates to 13.7%/23.1% from 14.9%/24.1% on the Hub5 2000 Switchboard/CallHome test sets without using any data augmentation, language model, or complex decoder. Gakuto Kurata, Kartik Audhkhasi |
INTERSPEECH | 2 |
| 2019 | Multi-Task CTC Training with Auxiliary Feature Reconstruction for End-to-End Speech Recognition
Gakuto Kurata, Kartik Audhkhasi |
INTERSPEECH | 2 |
| 2019 | Challenging the Boundaries of Speech Recognition: The MALACH CorpusabstractThere has been huge progress in speech recognition over the last several years. Tasks once thought extremely difficult, such as SWITCHBOARD, now approach levels of human performance. The MALACH corpus (LDC catalog LDC2012S05), a 375-Hour subset of a large archive of Holocaust testimonies collected by the Survivors of the Shoah Visual History Foundation, presents significant challenges to the speech community. The collection consists of unconstrained, natural speech filled with disfluencies, heavy accents, age-related coarticulations, un-cued speaker and language switching, and emotional speech - all still open problems for speech recognition systems. Transcription is challenging even for skilled human annotators. This paper proposes that the community place focus on the MALACH corpus to develop speech recognition systems that are more robust with respect to accents, disfluencies and emotional speech. To reduce the barrier for entry, a lexicon and training and testing setups have been created and baseline results using current deep learning technologies are presented. The metadata has just been released by LDC (LDC2019S11). It is hoped that this resource will enable the community to build on top of these baselines so that the extremely important information in these and related oral histories becomes accessible to a wider audience. Michael Picheny, Zoltán Tüske, Brian Kingsbury, Kartik Audhkhasi, George Saon |
INTERSPEECH | 4 |
| 2019 | Detection and Recovery of OOVs for Improved English Broadcast News Captioning
Samuel Thomas 0001, Kartik Audhkhasi, Zoltán Tüske, Michael Picheny |
INTERSPEECH | 2 |
| 2019 | Advancing Sequence-to-Sequence Based Speech Recognition
Zoltán Tüske, Kartik Audhkhasi, George Saon |
INTERSPEECH | 2 |
| 2018 | Building Competitive Direct Acoustics-to-Word Models for English Conversational Speech RecognitionabstractDirect acoustics-to-word (A2W) models in the end-to-end paradigm have received increasing attention compared to conventional subword based automatic speech recognition models using phones, characters, or context-dependent hidden Markov model states. This is because A2W models recognize words from speech without any decoder, pronunciation lexicon, or externally-trained language model, making training and decoding with such models simple. Prior work has shown that A2W models require orders of magnitude more training data in order to perform comparably to conventional models. Our work also showed this accuracy gap when using the English Switchboard-Fisher data set. This paper describes a recipe to train an A2W model that closes this gap and is at-par with state-of-the-art sub-word based models. We achieve a word error rate of 8.8.8%/13.9% on the Hub5-2000 Switchboard/CallHome test sets without any decoder or language model. We find that model initialization, training data order, and regularization have the most impact on the A2W model performance. Next, we present a joint word-character A2W model that learns to first spell the word and then recognize it. This model provides a rich output to the user instead of simple word hypotheses, making it especially useful in the case of words unseen or rarely-seen during training. Kartik Audhkhasi, Brian Kingsbury, Bhuvana Ramabhadran, George Saon, Michael Picheny |
ICASSP | 1 |
| 2018 | Whole Sentence Neural Language ModelsabstractRecurrent neural networks have become increasingly popular for the task of language modeling achieving impressive gains in state-of-the-art speech recognition and natural language processing (NLP) tasks. Recurrent models exploit word dependencies over a much longer context window (as retained by the history states) than what is feasible with n-gram language models. However the training criterion of choice for recurrent language models continues to be the local conditional likelihood of generating the current word given the (pos-sibly long) word context, thus making local decisions at each word. This locally-conditional design fundamentally limits the ability of the model in exploiting whole sentence structures. In this paper, we present our initial results at whole sentence neural language models which assign a probability to the entire word sequence. We extend the previous work on whole sentence maximum entropy models to recurrent language models while using Noise Contrastive Estimation (NCE) for training, as these sentence models are fundamentally un-normalizable. We present results on a range of tasks: from sequence identification tasks such as, palindrome detection to large vocabulary automatic speech recognition (LVCSR) and demonstrate the modeling power of this approach. Abhinav Sethy, Kartik Audhkhasi, Bhuvana Ramabhadran |
ICASSP | 3 |
| 2018 | Joint Modeling of Accents and Acoustics for Multi-Accent Speech RecognitionabstractThe performance of automatic speech recognition systems degrades with increasing mismatch between the training and testing scenarios. Differences in speaker accents are a significant source of such mismatch. The traditional approach to deal with multiple accents involves pooling data from several accents during training and building a single model in multi-task fashion, where tasks correspond to individual accents. In this paper, we explore an alternate model where we jointly learn an accent classifier and a multi-task acoustic model. Experiments on the American English Wall Street Journal and British English Cambridge corpora demonstrate that our joint model outperforms the strong multi-task acoustic model baseline. We obtain a 5.94% relative improvement in word error rate on British English, and 9.47% relative improvement on American English. This illustrates that jointly modeling with accent information improves acoustic model performance. Xuesong Yang, Kartik Audhkhasi, Andrew Rosenberg, Samuel Thomas 0001, Bhuvana Ramabhadran, Mark Hasegawa-Johnson |
ICASSP | 2 |
| 2018 | Improved Knowledge Distillation from Bi-Directional to Uni-Directional LSTM CTC for End-to-End Speech RecognitionabstractEnd-to-end automatic speech recognition (ASR) promises to simplify model training and deployment. Most end-to-end ASR systems utilize a bi-directional Long Short-Term Memory (BiLSTM) acoustic model due to its ability to capture acoustic context from the entire utterance. However, BiLSTM models have high latency and cannot be used in streaming applications. Leveraging knowledge distillation to train a low-latency end-to-end uni-directional LSTM (UniLSTM) model from a BiLSTM model can be an option. However, it makes the strict assumption of shared frame-wise time alignments between the two models. We propose an improved knowledge distillation algorithm that relaxes this assumption and improves the accuracy of the UniLSTM model. We confirmed the advantage of the proposed method on a standard English conversational telephone speech recognition task. Gakuto Kurata, Kartik Audhkhasi |
SLT | 2 |
| 2018 | Modeling Multiple Time Series Annotations as Noisy Distortions of the Ground Truth: An Expectation-Maximization ApproachabstractStudies of time-continuous human behavioral phenomena often rely on ratings from multiple annotators. Since the ground truth of the target construct is often latent, the standard practice is to use ad-hoc metrics (such as averaging annotator ratings). Despite being easy to compute, such metrics may not provide accurate representations of the underlying construct. In this paper, we present a novel method for modeling multiple time series annotations over a continuous variable that computes the ground truth by modeling annotator specific distortions. We condition the ground truth on a set of features extracted from the data and further assume that the annotators provide their ratings as modification of the ground truth, with each annotator having specific distortion tendencies. We train the model using an Expectation-Maximization based algorithm and evaluate it on a study involving natural interaction between a child and a psychologist, to predict confidence ratings of the children's smiles. We compare and analyze the model against two baselines where: (i) the ground truth in considered to be framewise mean of ratings from various annotators and, (ii) each annotator is assumed to bear a distinct time delay in annotation and their annotations are aligned before computing the framewise mean. Rahul Gupta 0001, Kartik Audhkhasi, Zach Jacokes, Agata Rozga, Shri Narayanan |
IEEE Trans. Affect. Comput. | 2 |
| 2017 | End-to-end ASR-free keyword search from speechabstractEnd-to-end (E2E) systems have achieved competitive results compared to conventional hybrid hidden Markov model (HMM)-deep neural network based automatic speech recognition (ASR) systems. Such E2E systems are attractive due to the lack of dependence on alignments between input acoustic and output grapheme or HMM state sequence during training. This paper explores the design of an ASR-free end-to-end system for text query-based keyword search (KWS) from speech trained with minimal supervision. Our E2E KWS system consists of three sub-systems. The first sub-system is a recurrent neural network (RNN)-based acoustic auto-encoder trained to reconstruct the audio through a finite-dimensional representation. The second sub-system is a character-level RNN language model using embeddings learned from a convolutional neural network. Since the acoustic and text query embeddings occupy different representation spaces, they are input to a third feed-forward neural network that predicts whether the query occurs in the acoustic utterance or not. This E2E ASR-free KWS system performs respectably despite lacking a conventional ASR system and trains much faster. Kartik Audhkhasi, Andrew Rosenberg, Abhinav Sethy, Bhuvana Ramabhadran, Brian Kingsbury |
ICASSP | 1 |
| 2017 | Knowledge distillation across ensembles of multilingual models for low-resource languagesabstractThis paper investigates the effectiveness of knowledge distillation in the context of multilingual models. We show that with knowledge distillation, Long Short-Term Memory(LSTM) models can be used to train standard feed-forward Deep Neural Network (DNN) models for a variety of low-resource languages. We then examine how the agreement between the teacher's best labels and the original labels affects the student model's performance. Next, we show that knowledge distillation can be easily applied to semi-supervised learning to improve model performance. We also propose a promising data selection method to filter un-transcribed data. Then we focus on knowledge transfer among DNN models with multilingual features derived from CNN+DNN, LSTM, VGG, CTC and attention models. We show that a student model equipped with better input features not only learns better from the teacher's labels, but also outperforms the teacher. Further experiments suggest that by learning from each other, the original ensemble of various models is able to evolve into a new ensemble with even better combined performance. Jia Cui, Brian Kingsbury, Bhuvana Ramabhadran, George Saon, Tom Sercu, Kartik Audhkhasi, Abhinav Sethy, Markus Nußbaum-Thom, Andrew Rosenberg |
ICASSP | 6 |
| 2017 | End-to-end speech recognition and keyword search on low-resource languagesabstractIn recent years, so-called, “end-to-end” speech recognition systems have emerged as viable alternatives to traditional ASR frameworks. Keyword search, localizing an orthographic query in a speech corpus, is typically performed by using automatic speech recognition (ASR) to generate an index. Previous work has evaluated the use of end-to-end systems for ASR on well known corpora (WSJ, Switchboard, TIMIT, etc.) in high-resource languages like English and Mandarin. In this work, we investigate the use of Connectionist Temporal Classification (CTC) networks, recurrent encoder-decoders with attention, two end-to-end ASR systems for keyword search and speech recognition on low resource languages. We find end-to-end systems can generate high quality 1-best transcripts on low-resource languages, but, because they generate very sharp posteriors, their utility is limited for KWS. We explore a number of ways to address this limitation with modest success. Experimental results reported are based on the IARPA BABEL OP3 languages and evaluation framework. This paper represents the first results using “end-to-end” techniques for speech recognition and keyword search on low-resource languages. Andrew Rosenberg, Kartik Audhkhasi, Abhinav Sethy, Bhuvana Ramabhadran, Michael Picheny |
ICASSP | 2 |
| 2017 | Direct Acoustics-to-Word Models for English Conversational Speech RecognitionabstractRecent work on end-to-end automatic speech recognition (ASR) has shown that the connectionist temporal classification (CTC) loss can be used to convert acoustics to phone or character sequences.Such systems are used with a dictionary and separately-trained Language Model (LM) to produce word sequences.However, they are not truly end-to-end in the sense of mapping acoustics directly to words without an intermediate phone representation.In this paper, we present the first results employing direct acoustics-to-word CTC models on two well-known public benchmark tasks: Switchboard and Call-Home.These models do not require an LM or even a decoder at run-time and hence recognize speech with minimal complexity.However, due to the large number of word output units, CTC word models require orders of magnitude more data to train reliably compared to traditional systems.We present some techniques to mitigate this issue.Our CTC word model achieves a word error rate of 13.0%/18.8%on the Hub5-2000 Switchboard/CallHome test sets without any LM or decoder compared with 9.6%/16.0%for phone-based CTC with a 4-gram LM.We also present rescoring results on CTC word model lattices to quantify the performance benefits of a LM, and contrast the performance of word and phone CTC models. Kartik Audhkhasi, Bhuvana Ramabhadran, George Saon, Michael Picheny, David Nahamoo |
INTERSPEECH | 1 |
| 2017 | English Conversational Telephone Speech Recognition by Humans and MachinesabstractOne of the most difficult speech recognition tasks is accurate recognition of human to human communication. Advances in deep learning over the last few years have produced major speech recognition improvements on the representative Switchboard conversational corpus. Word error rates that just a few years ago were 14% have dropped to 8.0%, then 6.6% and most recently 5.8%, and are now believed to be within striking range of human performance. This then raises two issues - what IS human performance, and how far down can we still drive speech recognition error rates? A recent paper by Microsoft suggests that we have already achieved human performance. In trying to verify this statement, we performed an independent set of human performance measurements on two conversational tasks and found that human performance may be considerably better than what was earlier reported, giving the community a significantly harder goal to achieve. We also report on our own efforts in this area, presenting a set of acoustic and language modeling techniques that lowered the word error rate of our own English conversational telephone LVCSR system to the level of 5.5%/10.3% on the Switchboard/CallHome subsets of the Hub5 2000 evaluation, which - at least at the writing of this paper - is a new performance milestone (albeit not at what we measure to be human performance!). On the acoustic side, we use a score fusion of three models: one LSTM with multiple feature inputs, a second LSTM trained with speaker-adversarial multi-task learning and a third residual net (ResNet) with 25 convolutional layers and time-dilated convolutions. On the language modeling side, we use word and character LSTMs and convolutional WaveNet-style language models. George Saon, Gakuto Kurata, Tom Sercu, Kartik Audhkhasi, Samuel Thomas 0001, Dimitrios Dimitriadis, Bhuvana Ramabhadran, Michael Picheny, Lynn-Li Lim, Bergul Roomi, Phil Hall |
INTERSPEECH | 4 |
| 2016 | Semantic word embedding neural network language models for automatic speech recognitionabstractSemantic word embeddings have become increasingly important in natural language processing tasks over the last few years. This popularity is due to their ability to easily capture rich semantic information through a distributed representation and the availability of fast and scalable algorithms for learning them from large text corpora. State-of-the-art neural network language models (NNLMs) used in automatic speech recognition (ASR) and natural language processing also learn word embeddings optimized to model local N-gram dependencies given training text but are not optimized to capture semantic information. We hypothesize that semantic word embeddings provide diverse information compared to the word embeddings learned by NNLMs. We propose novel feedforward NNLM architectures that incorporate semantic word embeddings. We apply the resulting NNLMs to ASR on broadcast news and show improvements in both perplexity and word error rate. Kartik Audhkhasi, Abhinav Sethy, Bhuvana Ramabhadran |
ICASSP | 1 |
| 2016 | Efficient one-vs-one kernel ridge regression for speech recognitionabstractRecent evidences suggest that the performance of kernel methods may match that of deep neural networks (DNNs), which have been the state-of-the-art approach for speech recognition. In this work, we present an improvement of the kernel ridge regression studied in Huang et al., ICASSP 2014, and show that our proposal is computationally advantageous. Our approach performs classifications by using the one-vs-one scheme, which, under certain assumptions, reduces the costs of the one-vs-rest scheme by asymptotically a factor of c2 in training time and c in memory consumption. Here, c is the number of classes and it is typically on the order of hundreds and thousands for speech recognition. We demonstrate empirical results on the benchmark corpus TIMIT. In particular, the classification accuracy is one to two percentages higher (in the absolute term) than the best of the kernel methods and of the DNNs reported by Huang et al, and the speech recognition accuracy is highly comparable. Jie Chen 0007, Lingfei Wu 0001, Kartik Audhkhasi, Brian Kingsbury, Bhuvana Ramabhadran |
ICASSP | 3 |
| 2016 | Multilingual Data Selection for Low Resource Speech RecognitionabstractAbstract : Feature representations extracted from deep neural network-based multilingual frontends provide significant improvements to speech recognition systems in low resource settings. To effectively train these frontends, we introduce a data selection technique that discovers language groups from an available set of training languages. This data selection method reduces the required amount of training data and training time by approximately 40 , with minimal performance degradation. We present speech recognition results on 7 very limited language pack (VLLP) languages from the second option period of the IARPA Babel program using multilingual features trained on up to 10 languages. The proposed multilingual features provide up to 15 relative improvement over baseline acoustic features on the VLLP languages. Samuel Thomas 0001, Kartik Audhkhasi, Jia Cui, Brian Kingsbury, Bhuvana Ramabhadran |
INTERSPEECH | 2 |
| 2016 | Detecting paralinguistic events in audio stream using context in features and probabilistic decisions
Rahul Gupta 0001, Kartik Audhkhasi, Sungbok Lee, Shri Narayanan |
Comput. Speech Lang. | 2 |
| 2016 | Noise-enhanced convolutional neural networks
Kartik Audhkhasi, Osonde Osoba, Bart Kosko |
Neural Networks | 1 |
| 2015 | Multilingual representations for low resource speech recognition and keyword searchabstractThis paper examines the impact of multilingual (ML) acoustic representations on Automatic Speech Recognition (ASR) and keyword search (KWS) for low resource languages in the context of the OpenKWS15 evaluation of the IARPA Babel program. The task is to develop Swahili ASR and KWS systems within two weeks using as little as 3 hours of transcribed data. Multilingual acoustic representations proved to be crucial for building these systems under strict time constraints. The paper discusses several key insights on how these representations are derived and used. First, we present a data sampling strategy that can speed up the training of multilingual representations without appreciable loss in ASR performance. Second, we show that fusion of diverse multilingual representations developed at different LORELEI sites yields substantial ASR and KWS gains. Speaker adaptation and data augmentation of these representations improves both ASR and KWS performance (up to 8.7% relative). Third, incorporating un-transcribed data through semi-supervised learning, improves WER and KWS performance. Finally, we show that these multilingual representations significantly improve ASR and KWS performance (relative 9% for WER and 5% for MTWV) even when forty hours of transcribed audio in the target language is available. Multilingual representations significantly contributed to the LORELEI KWS systems winning the OpenKWS15 evaluation. Jia Cui, Brian Kingsbury, Bhuvana Ramabhadran, Abhinav Sethy, Kartik Audhkhasi, Ellen Eide, Lidia Mangu, Markus Nußbaum-Thom, Michael Picheny, Zoltán Tüske, Pavel Golik, Ralf Schlüter, Hermann Ney, Mark J. F. Gales, Kate M. Knill, Anton Ragni, Philip C. Woodland |
ASRU | 5 |
| 2015 | A mixture of experts approach towards intelligibility classification of pathological speechabstractPathological speech involves atypical speech production which may result from several factors including oral diseases, physical disabilities in the voice production system and atypical anatomy. Automatic evaluation of intelligibility in patients with pathological speech can assist accurate diagnosis of pathological conditions. Loss of intelligibility may be associated with one of the several pathological conditions, making automatic evaluation a challenging computational problem. A Mixture of Experts (MoE) models class boundaries using a weighted combination of several experts and can characterize the complex class boundaries arising due to pathological variability. We train an MoE for intelligibility evaluation using a modified Expectation Maximization (EM) algorithm based on joint simulated annealing-gradient ascent procedure. Our algorithm optimizes the expert parameters and simultaneously obtains the feature subsets for each expert. We observe that the MoE trained using the new EM algorithm not only outperforms a single classifier baseline but also the vanilla MoE. We perform further data analysis and interpret the weights assigned to each expert during inference. Also, we obtain a different feature subset per expert in the mixture. This illustrates feature use based on location of the data point in the feature space. Rahul Gupta 0001, Kartik Audhkhasi, Shri Narayanan |
ICASSP | 2 |
| 2014 | Fusion of diverse denoising systems for robust automatic speech recognitionabstractWe present a framework for combining different denoising front-ends for robust speech enhancement for recognition in noisy conditions. This is contrasted against results of optimally fusing diverse parameter settings for a single denoising algorithm. All frontends in the latter case exploit the same denoising algorithm, which combines harmonic decomposition, with noise estimation and spectral subtraction. The set of associated parameters involved in these steps are dependent on the noise conditions. Rather than explicitly tuning them, we suggest a strategy that tries to account for the trade-off between average word error rate and diversity to find an optimal subset of these parameter settings. We present the results on Aurora4 database and also compare against traditional speech enhancement methods e.g. Wiener filtering and spectral subtraction. Naveen Kumar 0004, Maarten Van Segbroeck, Kartik Audhkhasi, Peter Drotár, Shri Narayanan |
ICASSP | 3 |
| 2014 | Semi-supervised term-weighted value rescoring for keyword searchabstractWe present a semi-supervised algorithm for rescoring the output of a speech keyword search (KWS) system. Conventional loss functions such as squared-error and logistic loss are not suitable for optimizing the commonly-used KWS term-weighted value (TWV) performance metric. We derive a novel concave modified logistic log-likelihood function which lower-bounds TWV. We then use a manifold-regularized kernel classifier that maximizes this lower-bound. A manifold regularization term in our objective function uses available unlabeled speech data and makes our approach semi-supervised. This term is particularly useful for KWS in low-resource languages and ensures that the predicted keyword confidence scores are smooth on a low-dimensional manifold in the feature space. We conduct KWS experiments on the IARPA Babel Vietnamese task and show performance improvements in terms of the maximum TWV (MTWV). Our estimated confidence score is complementary with respect to the ASR posterior score and gives MTWV improvement upon interpolation with it. Kartik Audhkhasi, Abhinav Sethy, Bhuvana Ramabhadran, Shri Narayanan |
ICASSP | 1 |
| 2014 | Training ensemble of diverse classifiers on feature subsetsabstractEnsembles of diverse classifiers often out-perform single classifiers as has been well-demonstrated across several applications. Existing training algorithms either learn a classifier ensemble on pre-defined feature sets or independently perform classifier training and feature selection. Neither of these schemes is optimal. We pose feature subset selection and training of diverse classifiers on selected subsets as a joint optimization problem. We propose a novel greedy algorithm to solve this problem. We sequentially learn an ensemble of classifiers where each subsequent classifier is encouraged to learn data instances misclassified by previous classifiers on a concurrently selected feature set. Our experiments on synthetic and real-world data sets show the effectiveness of our algorithm. We observe that ensembles trained by our algorithm performs better than both a single classifier and an ensemble of classifiers learnt on pre-defined feature sets. We also test our algorithm as a feature selector on a synthetic dataset to filter out irrelevant features. Rahul Gupta 0001, Kartik Audhkhasi, Shri Narayanan |
ICASSP | 2 |
| 2014 | Theoretical Analysis of Diversity in an Ensemble of Automatic Speech Recognition SystemsabstractDiversity or complementarity of automatic speech recognition (ASR) systems is crucial for achieving a reduction in word error rate (WER) upon fusion using the ROVER algorithm. We present a theoretical proof explaining this often-observed link between ASR system diversity and ROVER performance. This is in contrast to many previous works that have only presented empirical evidence for this link or have focused on designing diverse ASR systems using intuitive algorithmic modifications. We prove that the WER of the ROVER output approximately decomposes into a difference of the average WER of the individual ASR systems and the average WER of the ASR systems with respect to the ROVER output. We refer to the latter quantity as the diversity of the ASR system ensemble because it measures the spread of the ASR hypotheses about the ROVER hypothesis. This result explains the trade-off between the WER of the individual systems and the diversity of the ensemble. We support this result through ROVER experiments using multiple ASR systems trained on standard data sets with the Kaldi toolkit. We use the proposed theorem to explain the lower WERs obtained by ASR confidence-weighted ROVER as compared to word frequency-based ROVER. We also quantify the reduction in ROVER WER with increasing diversity of the N-best list. We finally present a simple discriminative framework for jointly training multiple diverse acoustic models (AMs) based on the proposed theorem. Our framework generalizes and provides a theoretical basis for some recent intuitive modifications to well-known discriminative training criterion for training diverse AMs. Kartik Audhkhasi, Andreas M. Zavou, Panayiotis G. Georgiou, Shri Narayanan |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2013 | Joint training of interpolated exponential n-gram modelsabstractFor many speech recognition tasks, the best language model performance is achieved by collecting text from multiple sources or domains, and interpolating language models built separately on each individual corpus. When multiple corpora are available, it has also been shown that when using a domain adaptation technique such as feature augmentation [1], the performance on each individual domain can be improved by training a joint model across all of the corpora. In this paper, we explore whether improving each domain model via joint training also improves performance when interpolating the models together. We show that the diversity of the individual models is an important consideration, and propose a method for adjusting diversity to optimize overall performance. We present results using word n-gram models and Model M, a class-based n-gram model, and demonstrate improvements in both perplexity and word-error rate relative to state-of-the-art results on a Broadcast News transcription task. Abhinav Sethy, Stanley F. Chen, Ebru Arisoy, Bhuvana Ramabhadran, Kartik Audhkhasi, Shri Narayanan, Paul Vozila |
ASRU | 5 |
| 2013 | Noise benefits in backpropagation and deep bidirectional pre-trainingabstractWe prove that noise can speed convergence in the backpropagation algorithm. The proof consists of two separate results. The first result proves that the backpropagation algorithm is a special case of the generalized Expectation-Maximization (EM) algorithm for iterative maximum likelihood estimation. The second result uses the recent EM noise benefit to derive a sufficient condition for backpropagation training. The noise adds directly to the training data. A noise benefit also applies to the deep bidirectional pre-training of the neural network as well as to the backpropagation training of the network. The geometry of the noise benefit depends on the probability structure of the neurons at each layer. Logistic sigmoidal neurons produce a forbidden noise region that lies below a hyperplane. Then all noise on or above the hyperplane can only speed convergence of the neural network. The forbidden noise region is a sphere if the neurons have a Gaussian signal or activation function. These noise benefits all follow from the general noise benefit of the EM algorithm. Monte Carlo sample means estimate the population expectations in the EM algorithm. We demonstrate the noise benefits using MNIST digit classification. Kartik Audhkhasi, Osonde Osoba, Bart Kosko |
IJCNN | 1 |
| 2013 | Noisy hidden Markov models for speech recognitionabstractWe show that noise can speed training in hidden Markov models (HMMs). The new Noisy Expectation-Maximization (NEM) algorithm shows how to inject noise when learning the maximum-likelihood estimate of the HMM parameters because the underlying Baum-Welch training algorithm is a special case of the Expectation-Maximization (EM) algorithm. The NEM theorem gives a sufficient condition for such an average noise boost. The condition is a simple quadratic constraint on the noise when the HMM uses a Gaussian mixture model at each state. Simulations show that a noisy HMM converges faster than a noiseless HMM on the TIMIT data set. Kartik Audhkhasi, Osonde Osoba, Bart Kosko |
IJCNN | 1 |
| 2013 | Empirical link between hypothesis diversity and fusion performance in an ensemble of automatic speech recognition systems
Kartik Audhkhasi, Andreas M. Zavou, Panayiotis G. Georgiou, Shri Narayanan |
INTERSPEECH | 1 |
| 2013 | Classifying language-related developmental disorders from speech cues: the promise and the potential confoundsabstractSpeech and spoken language cues offer a valuable means to measure and model human behavior. Computational models of speech behavior have the potential to support health care through assistive technologies, informed intervention, and effi-cient long-term monitoring. The Interspeech 2013 Autism Sub-Challenge addresses two developmental disorders that manifest in speech: autism spectrum disorders and specific language im-pairment. We present classification results with an analysis on the development set including a discussion of potential con-founds in the data such as recording condition differences. We hence propose study of features within these domains that may inform realistic separability between groups as well as have the potential to be used for behavioral intervention and monitoring. We investigate template-based prosodic and formant modeling as well as goodness of pronunciation modeling, reporting above chance classification accuracies. Index Terms: autism spectrum disorders, intonation, specific language impairment, goodness of pronunciation Daniel Bone, Theodora Chaspari, Kartik Audhkhasi, James Gibson, Andreas Tsiartas, Maarten Van Segbroeck, Ming Li 0026, Sungbok Lee, Shri Narayanan |
INTERSPEECH | 3 |
| 2013 | Paralinguistic event detection from speech using probabilistic time-series smoothing and masking
Rahul Gupta 0001, Kartik Audhkhasi, Sungbok Lee, Shri Narayanan |
INTERSPEECH | 2 |
| 2013 | Which ASR should I choose for my dialogue system?
Fabrizio Morbini, Kartik Audhkhasi, Kenji Sagae, Ron Artstein, Dogan Can, Panayiotis G. Georgiou, Shri Narayanan, Anton Leuski, David R. Traum |
SIGDIAL Conference | 2 |
| 2013 | A Globally-Variant Locally-Constant Model for Fusion of Labels from Multiple Diverse Experts without Using Reference LabelsabstractResearchers have shown that fusion of categorical labels from multiple experts—humans or machine classifiers—improves the accuracy and generalizability of the overall classification system. Simple plurality is a popular technique for performing this fusion, but it gives equal importance to labels from all experts, who may not be equally reliable or consistent across the dataset. Estimation of expert reliability without knowing the reference labels is, however, a challenging problem. Most previous works deal with these challenges by modeling expert reliability as constant over the entire data (feature) space. This paper presents a model based on the consideration that in dealing with real-world data, expert reliability is variable over the complete feature space but constant over local clusters of homogeneous instances. This model jointly learns a classifier and expert reliability parameters without assuming knowledge of the reference labels using the Expectation-Maximization (EM) algorithm. Classification experiments on simulated data, data from the UCI Machine Learning Repository, and two emotional speech classification datasets show the benefits of the proposed model. Using a metric based on the Jensen-Shannon divergence, we empirically show that the proposed model gives greater benefit for datasets where expert reliability is highly variable over the feature space. Kartik Audhkhasi, Shri Narayanan |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2012 | Analyzing quality of crowd-sourced speech transcriptions of noisy audio for acoustic model adaptationabstractThe accuracy of crowd-sourced speech transcriptions varies depending on a variety of factors. This paper studies the impact of one such factor, namely, the quality of audio. We employed a speech database with babble noise at three SNR levels (clean, 2 dB and -2 dB) and asked workers on Amazon Mechanical Turk to transcribe it. Two interesting observations emerge. First, as expected, the quality of transcripts combined by word frequency based ROVER decreases with decreasing SNR. Further, we demonstrate that the use of some unsupervised reliability scores can improve the transcription quality, with increasing benefits at lower SNR. Second, we do not observe a significant drop in the performance of acoustic models adapted with increasing transcription noise. This highlights the surprising robustness of crowd-sourced transcripts for acoustic model adaptation. Kartik Audhkhasi, Panayiotis G. Georgiou, Shri Narayanan |
ICASSP | 1 |
| 2012 | Creating ensemble of diverse maximum entropy modelsabstractDiversity of a classifier ensemble has been shown to benefit overall classification performance. But most conventional methods of training ensembles offer no control on the extent of diversity and are meta-learners. We present a method for creating an ensemble of diverse maximum entropy (∂MaxEnt) models, which are popular in speech and language processing. We modify the objective function for conventional training of a MaxEnt model such that its output posterior distribution is diverse with respect to a reference model. Two diversity scores are explored - KL divergence and posterior cross-correlation. Experiments on the CoNLL-2003 Named Entity Recognition task and the IEMOCAP emotion recognition database show the benefits of a ∂MaxEnt ensemble. Kartik Audhkhasi, Abhinav Sethy, Bhuvana Ramabhadran, Shri Narayanan |
ICASSP | 1 |
| 2012 | Speaker Personality Classification Using Systems Based on Acoustic-Lexical Cues and an Optimal Tree-Structured Bayesian NetworkabstractAutomatic classification of human personality along the Big Five dimensions is an interesting problem with several prac-tical applications. This paper makes some contributions in this regard. First, we propose a few automatically-derived personality-discriminating lexical features which provide infor-mation complementary to the conventional acoustic-prosodic cues. We also design a frame-level Gaussian mixture model based system which adds complimentary information to the sys-tems trained on global statistical functionals. Next, we note that the Big Five dimensions are correlated and thus model the de-pendency between these dimensions in the form of an optimal tree-structured Bayesian network. Our final sub-system con-sists of within class covariance normalization followed by L1-regularized logistic regression. Fusion of all these sub-systems achieves better classification performance than independently trained classifiers using just acoustic features. Kartik Audhkhasi, Angeliki Metallinou, Ming Li 0026, Shri Narayanan |
INTERSPEECH | 1 |
| 2012 | A reranking approach for recognition and classification of speech input in conversational dialogue systemsabstractWe address the challenge of interpreting spoken input in a conversational dialogue system with an approach that aims to exploit the close relationship between the tasks of speech recognition and language understanding through joint modeling of these two tasks. Instead of using a standard pipeline approach where the output of a speech recognizer is the input of a language understanding module, we merge multiple speech recognition and utterance classification hypotheses into one list to be processed by a joint reranking model. We obtain substantially improved performance in language understanding in experiments with thousands of user utterances collected from a deployed spoken dialogue system. Fabrizio Morbini, Kartik Audhkhasi, Ron Artstein, Maarten Van Segbroeck, Kenji Sagae, Panayiotis G. Georgiou, David R. Traum, Shri Narayanan |
SLT | 2 |
| 2011 | Accurate transcription of broadcast news speech using multiple noisy transcribers and unsupervised reliability metricsabstractProfessional manual transcription of speech is an expensive and time consuming process. This paper focuses on the problem of combining noisy transcriptions from multiple non-expert transcribers, where the quality of work from each worker varies. Computing transcriber reliability is a difficult task in the absence of gold standard reference transcripts. Three simple metrics for quantifying this reliability without using a gold standard are proposed. We create a database of 1000 Mexican Spanish broadcast news audio clips transcribed by five transcribers each through Amazon Mechanical Turk. Combination of multiple noisy transcripts using these reliability scores improves the word error rate of the combined transcript with respect to the LDC gold standard by 8% relative, and the sentence error rate by 4.1% relative, when compared with a combination without any reliability information. Kartik Audhkhasi, Panayiotis G. Georgiou, Shri Narayanan |
ICASSP | 1 |
| 2011 | Emotion classification from speech using evaluator reliability-weighted combination of ranked listsabstractIn emotion recognition, a widely-used method to reconciliate disagreement between multiple human evaluators is to perform majority-voting on their assigned class labels. Instead, we propose asking evaluators to rank emotional categories given an audio clip, followed by a combination of these ranked lists. We compare two well-known ranked list voting methods Borda count and Schulze's method, with majority-voting and an evaluator model-based combination of the top ranked-labels. When tested on an emotional speech database with ground truth labels available, two interesting observations emerge. First, majority-voting performs significantly worse than the other three methods in the estimation of the given ground truth labels. Second, when performing classification using the combined labels, the two ranked list voting methods perform the best. We then propose evaluator reliability-weighted versions of these two methods, which improve the classification accuracy even further. Kartik Audhkhasi, Shri Narayanan |
ICASSP | 1 |
| 2011 | Reliability-Weighted Acoustic Model Adaptation Using Crowd Sourced TranscriptionsabstractThis paper focuses on adaptation of acoustic models using speech transcribed by multiple noisy experts. A simple approach involves combining multiple transcripts using word frequency based Recognizer Output Voting Error Reduction (ROVER) followed by adaptation using the combined transcripts. But this assumes that the transcripts being combined are equally reliable. To overcome this assumption, we use two sets of scores to estimate this reliability. The first set is based on answers to some questions given by the transcribers. The second set is derived in an unsupervised way using the word frequency based ROVER transcripts and baseline acoustic models. The overall confidence is a convex combination of these scores and is used to perform a confidence weighted fusion. We adapt the baseline acoustic models using these combined transcripts. Recognition results for a Mexican Spanish ASR system show an absolute improvement of 0.5% in word error rate and 0.9% in sentence error rate. Kartik Audhkhasi, Panayiotis G. Georgiou, Shri Narayanan |
INTERSPEECH | 1 |
| 2010 | Data-dependent evaluator modeling and its application to emotional valence classification from speechabstractPractical supervised learning scenarios involving subjectively evaluated data have multiple evaluators, each giving their noisy version of the hidden ground truth. Majority logic combination of labels assumes equally skilled evaluators, and is generally suboptimal. Previously proposed models have assumed data independent evaluator behavior. This paper presents a data dependent evaluator model, and an algorithm to jointly learn evaluator behavior and a classifier. This model is based on the intuition that real world evaluators have varying performance depending on the data. Experiments on an emotional valence classification task show modest performance improvements of the proposed algorithm as compared to the majority logic baseline and a data independent evaluator model. But more critically, the algorithm also provides accurate estimates of individual evaluator performance, thus paving the way for incorporating active learning, evaluator feedback and unreliable data detection. Kartik Audhkhasi, Shri Narayanan |
INTERSPEECH | 1 |
| 2010 | Automatic speech recognition system channel modelingabstractIn this paper, we present a systems approach for channel mod-eling of an Automatic Speech Recognition (ASR) system. This can have implications in improving speech recognition com-ponents, such as through discriminative language modeling. We simulate the ASR corruption using a phrase-based machine translation system trained between the reference phoneme and output phoneme sequences of a real ASR. We demonstrate that local optimization on the quality of phoneme-to-phoneme map-pings does not directly translate to overall improvement of the entire model. However, we are still able to capitalize on contex-tual information of the phonemes which a simple acoustic dis-tance model is not able to accomplish. Hence we show that the use of longer context results in a significantly improved model of the ASR channel. Qun Feng Tan, Kartik Audhkhasi, Panayiotis G. Georgiou, Emil Ettelaie, Shri Narayanan |
INTERSPEECH | 2 |
| 2009 | Lattice-based lexical cues for word fragment detection in conversational speechabstractPrevious approaches to the problem of word fragment detection in speech have focussed primarily on acoustic-prosodic features. This paper proposes that the output of a continuous automatic speech recognition (ASR) system can also be used to derive robust lexical features for the task. We hypothesize that the confusion in the word lattice generated by the ASR system can be exploited for detecting word fragments. Two sets of lexical features are proposed -one which is based on the word confusion, and the other based on the pronunciation confusion between the word hypotheses in the lattice. Classification experiments with a support vector machine (SVM) classifier show that these lexical features perform better than the previously proposed acoustic-prosodic features by around 5.20% (relative) on a corpus chosen from the DARPA Transtac Iraqi-English (San Diego) corpus. A combination of both these feature sets improves the word fragment detection accuracy by 11.50% relative to using just the acoustic-prosodic features. Kartik Audhkhasi, Panayiotis G. Georgiou, Shri Narayanan |
ASRU | 1 |
| 2009 | Formant-based technique for automatic filled-pause detection in spontaneous spoken englishabstractDetection of filled pauses is a challenging research problem which has several practical applications. It can be used to evaluate the spoken fluency skills of the speaker, to improve the performance of automatic speech recognition systems or to predict the mental state of the speaker. This paper presents an algorithm for filled pause detection that is based on the premise that the vocal tract characteristics, and hence the formants, are stable during the production of a filled pause. The performance of the proposed algorithm is evaluated on real-life recordings of call center agents where the locations of the filled pauses are hand labeled. The proposed algorithm outperforms a standard cepstral stability based filled pause detection algorithm and a standard pitch-based detection technique. Kartik Audhkhasi, Kundan Kandhway, Om Deshmukh, Ashish Verma 0001 |
ICASSP | 1 |
| 2009 | Automatic evaluation of spoken english fluencyabstractThis paper presents a method to automatically quantify the spoken English fluency skills of speakers. The focus of this work is to automatically compute a numeric score of spoken fluency that is correlated with the numerical score the human assessors would assign. The proposed method combines several novel prosodic and lexical features to compute the fluency score. It is shown that the prosodic and the lexical features provide complementary information for fluency evaluation. Extensive evaluation on human-labeled utterances shows that the proposed technique exhibits similar trends in performance and confusions as shown by human assessors. The proposed technique leads to 84.2% classification accuracy when the two extreme classes of fluency are considered. Om Deshmukh, Kundan Kandhway, Ashish Verma 0001, Kartik Audhkhasi |
ICASSP | 4 |
| 2007 | Keyword Search using Modified Minimum Edit Distance MeasureabstractA popular approach for keyword search in speech files is the phone lattice search. Recently minimum edit distance (MED) has been used as a measure of similarity between strings rather than using simple string matching while searching the phone lattice for the keyword. In this paper, we propose a variation of the MED, where the substitution penalties are automatically derived from the phone confusion matrix of the recognizer, as compared to heuristic or class based penalties used earlier. The results show that the substitution penalties derived from the phone confusion matrix lead to a considerable improvement in the accuracy of the keyword search algorithm. Kartik Audhkhasi, Ashish Verma 0001 |
ICASSP (4) | 1 |