EDBT 2026 Demo / reviewers in the wild / expert
Hong-Kwang Jeff Kuo
dblp:82/5117 · also Hong-Kwang Kuo
· DBLP profile ↗
82ranked-venue papers
21as first author
16since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 78 · 19 first-author · 16 since 2021Artificial intelligence and machine learning · 46 · 14 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Granite-speech: open-source speech-aware LLMs with strong English ASR capabilitiesabstractGranite-speech LLMs are compact and efficient speech language models specifically designed for English ASR1and automatic speech translation (AST). The models were trained by modality aligning granite-3.3-instruct to speech on publicly available open-source corpora. Comprehensive benchmarking on English ASR shows that they outperform several competitors’ models that were trained on orders of magnitude more proprietary data, and they keep pace on English-to-X AST for major European languages, Japanese, and Mandarin. The speech-specific components are: a conformer acoustic encoder using block attention and self-conditioning trained with connectionist temporal classification, a windowed query-transformer speech modality adapter used to do temporal downsampling of the acoustic embeddings and map them to the LLM text embedding space, and LoRA adapters to further fine-tune the text LLM. The models are freely available on HuggingFace2under a permissive Apache 2.0 license.1The latest models (revision 3.3.2) support multilingual ASR in English, French, German, Spanish and Portuguese and bidirectional speech translation to and from English. This paper covers the initial English-only release.2https://huggingface.co/ibm-granite/granite-speech-3.3-2b (and…-8b). George Saon, Avihu Dekel, Alexi Brooks, Tohru Nagano, Abraham Daniels, Aharon Satt, Ashish R. Mittal, Brian Kingsbury, David Haws, Edmilson da Silva Morais, Gakuto Kurata, Hagai Aronowitz, Ibrahim Ibrahim, Hong-Kwang Jeff Kuo, Kate Soule, Luis A. Lastras, Masayuki Suzuki, Ron Hoory, Samuel Thomas 0001, Sashi Novitasari, Takashi Fukuda, Vishal Sunder, Zvi Kons |
ASRU | 14 |
| 2025 | LLM based Text Generation for Improved Low-resource Speech Recognition ModelsabstractLimited transcribed spoken style data is a critical bottleneck in building automatic speech recognition (ASR) systems for low-resource languages. Prompting a large language model (LLM) to paraphrase input text can generate novel text data that is constrained to be semantically similar to the source data. We leverage this capability of LLMs to improve the performance of low-resource ASR systems by increasing the limited text training data while keeping the same spoken style. Since word sequences in the training data are now more diverse and the vocabulary of the ASR model is also expanded, this approach allows for building general purpose ASR without prior knowledge of various domains in the low-resource language. In our experiments with Brazilian Portuguese as a low-resource language, paraphrased data enhanced the n-gram language model (LM) used to build the weighted finite state transducer (WFST) for decoding with a Conformer-CTC speech recognition model, resulting in improvement of word error rate (WER) by 15.6% over the baseline model. Synthesizing the paraphrased text into speech and using it to fine-tune the acoustic model (AM) component helped to further improve the WER by 2.9%, achieving a combined improvement of 18.5%. We also demonstrate the usefulness of our proposed approach for high-resource languages like English. Tohru Nagano, Gakuto Kurata, Samuel Thomas 0001, Hong-Kwang Jeff Kuo, Daniel Bolaños, Hyun Jung, George Saon |
ICASSP | 4 |
| 2023 | Fine-Grained Textual Knowledge Transfer to Improve RNN Transducers for Speech Recognition and UnderstandingabstractRNN Tranducer (RNN-T) technology is very popular for building deployable models for end-to-end (E2E) automatic speech recognition (ASR) and spoken language understanding (SLU). Since these are E2E models operating on speech directly, there remains a potential to improve their performance using purely text based models like BERT, which have strong language understanding capabilities. In this paper, we propose a new training criteria for RNN-T based E2E ASR and SLU to transfer BERT’s knowledge into these systems. In the first stage of our proposed mechanism, we improve ASR performance by using a fine-grained, tokenwise knowledge transfer from BERT. In the second stage, we fine-tune the ASR model for SLU such that the above knowledge is explicitly utilized by the RNN-T model for improved performance. Our techniques improve ASR performance on the Switchboard and CallHome test sets of the NIST Hub5 2000 evaluation and on the recently released SLURP dataset on which we achieve a new state-of-the-art performance. For SLU, we show significant improvements on the SLURP slot filling task, outperforming HuBERT-base and reaching a performance close to HuBERTlarge. Compared to large transformer based speech models like HuBERT, our model is significantly more compact and uses only 300 hours of speech pretraining data. Vishal Sunder, Samuel Thomas 0001, Hong-Kwang Jeff Kuo, Brian Kingsbury, Eric Fosler-Lussier |
ICASSP | 3 |
| 2023 | Multi-Speaker Data Augmentation for Improved end-to-end Automatic Speech RecognitionabstractPublicly available datasets traditionally used to train E2E ASR models for conversational telephone speech recognition are based on clean, short duration, single speaker utterances collected on separate channels. While E2E ASR models achieve state-of-the-art performance on recognition tasks that match well with such training data, they are observed to fail on test recordings that contain multiple speakers, significant channel or background noise or span longer durations than training data utterances. To mitigate these issues, we propose an on-the-fly data augmentation strategy that transforms single speaker training data into multiple speaker data by appending together multiple single speaker utterances. The proposed technique encourages the E2E model to become robust to speaker changes and also process longer utterances effectively. During training, the model is also guided by a teacher model trained on single speaker utterances to map its multi-speaker encoder embeddings to better performing single speaker representations. With the proposed technique we obtain 7-14% relative improvement on various single speaker and multiple speaker test sets. We also show how this technique is able to improve recognition performance by up to 14% by capturing useful information from preceding spoken utterances used as dialog history. Samuel Thomas 0001, Hong-Kwang Jeff Kuo, George Saon, Brian Kingsbury |
ICASSP | 2 |
| 2023 | ConvKT: Conversation-Level Knowledge Transfer for Context Aware End-to-End Spoken Language Understanding
Vishal Sunder, Eric Fosler-Lussier, Samuel Thomas 0001, Hong-Kwang Jeff Kuo, Brian Kingsbury |
INTERSPEECH | 4 |
| 2022 | A New Data Augmentation Method for Intent Classification Enhancement and its Application on Spoken Conversation Datasets
Zvi Kons, Aharon Satt, Hong-Kwang Jeff Kuo, Samuel Thomas 0001, Boaz Carmeli, Ron Hoory, Brian Kingsbury |
ICASSP | 3 |
| 2022 | Improving End-to-end Models for Set Prediction in Spoken Language UnderstandingabstractThe goal of spoken language understanding (SLU) systems is to determine the meaning of the input speech signal, unlike speech recognition which aims to produce verbatim transcripts. Advances in end-to-end (E2E) speech modeling have made it possible to train solely on semantic entities, which are far cheaper to collect than verbatim transcripts. We focus on this set prediction problem, where entity order is unspecified. Using two classes of E2E models, RNN transducers and attention based encoder-decoders, we show that these models work best when the training entity sequence is arranged in spoken order. To improve E2E SLU models when entity spoken order is unknown, we propose a novel data augmentation technique along with an implicit attention based alignment method to infer the spoken order. F1 scores significantly increased by more than 11% for RNN-T and about 2% for attention based encoder-decoder SLU models, outperforming previously reported results. Hong-Kwang Jeff Kuo, Zoltán Tüske, Samuel Thomas 0001, Brian Kingsbury, George Saon |
ICASSP | 1 |
| 2022 | Towards End-to-End Integration of Dialog History for Improved Spoken Language UnderstandingabstractDialog history plays an important role in spoken language understanding (SLU) performance in a dialog system. For end-to-end (E2E) SLU, previous work has used dialog history in text form, which makes the model dependent on a cascaded automatic speech recognizer (ASR). This rescinds the benefits of an E2E system which is intended to be compact and robust to ASR errors. In this paper, we propose a hierarchical conversation model that is capable of directly using dialog history in speech form, making it fully E2E. We also distill semantic knowledge from the available gold conversation transcripts by jointly training a similar text-based conversation model with an explicit tying of acoustic and semantic embeddings. We also propose a novel technique that we call DropFrame to deal with the long training time incurred by adding dialog history in an E2E manner. On the HarperValleyBank dialog dataset, our E2E history integration outperforms a history independent baseline by 7.7% absolute F1 score on the task of dialog action recognition. Our model performs competitively with the state-of-the-art history based cascaded baseline, but uses 48% fewer parameters. In the absence of gold transcripts to fine-tune an ASR model, our model outperforms this baseline by a significant margin of 10% absolute F1 score. Vishal Sunder, Samuel Thomas 0001, Hong-Kwang Jeff Kuo, Jatin Ganhotra, Brian Kingsbury, Eric Fosler-Lussier |
ICASSP | 3 |
| 2022 | Towards Reducing the Need for Speech Training Data to Build Spoken Language Understanding SystemsabstractThe lack of speech data annotated with labels required for spoken language understanding (SLU) is often a major hurdle in building end-to-end (E2E) systems that can directly process speech inputs. In contrast, large amounts of text data with suitable labels are usually available. In this paper, we propose a novel text representation and training methodology that allows E2E SLU systems to be effectively constructed using these text resources. With very limited amounts of additional speech, we show that these models can be further improved to perform at levels close to similar systems built on the full speech datasets. The efficacy of our proposed approach is demonstrated on both intent and entity tasks using three different SLU datasets. With text-only training, the proposed system achieves up to 90% of the performance possible with full speech training. With just an additional 10% of speech data, these models significantly improve further to 97% of full performance. Samuel Thomas 0001, Hong-Kwang Jeff Kuo, Brian Kingsbury, George Saon |
ICASSP | 2 |
| 2022 | Integrating Text Inputs for Training and Adapting RNN Transducer ASR ModelsabstractCompared to hybrid automatic speech recognition (ASR) systems that use a modular architecture in which each component can be in-dependently adapted to a new domain, recent end-to-end (E2E) ASR system are harder to customize due to their all-neural monolithic construction. In this paper, we propose a novel text representation and training framework for E2E ASR models. With this approach, we show that a trained RNN Transducer (RNN-T) model’s internal LM component can be effectively adapted with text-only data. An RNN-T model trained using both speech and text inputs improves over a baseline model trained on just speech with close to 13% word error rate (WER) reduction on the Switchboard and CallHome test sets of the NIST Hub5 2000 evaluation. The usefulness of the proposed approach is further demonstrated by customizing this general purpose RNN-T model to three separate datasets. We observe 20-45% relative word error rate (WER) reduction in these settings with this novel LM style customization technique using only unpaired text data from the new domains. Samuel Thomas 0001, Brian Kingsbury, George Saon, Hong-Kwang Jeff Kuo |
ICASSP | 4 |
| 2022 | Extending RNN-T-based speech recognition systems with emotion and language classificationabstractSpeech transcription, emotion recognition, and language identification are usually considered to be three different tasks.Each one requires a different model with a different architecture and training process.We propose using a recurrent neural network transducer (RNN-T)-based speech-to-text (STT) system as a common component that can be used for emotion recognition and language identification as well as for speech recognition.Our work extends the STT system for emotion classification through minimal changes, and shows successful results on the IEMOCAP and MELD datasets.In addition, we demonstrate that by adding a lightweight component to the RNN-T module, it can also be used for language identification.In our evaluations, this new classifier demonstrates state-of-the-art accuracy for the NIST-LRE-07 dataset. Zvi Kons, Hagai Aronowitz, Edmilson da Silva Morais, Matheus Damasceno, Hong-Kwang Jeff Kuo, Samuel Thomas 0001, George Saon |
INTERSPEECH | 5 |
| 2022 | Tokenwise Contrastive Pretraining for Finer Speech-to-BERT Alignment in End-to-End Speech-to-Intent SystemsabstractRecent advances in End-to-End (E2E) Spoken Language Understanding (SLU) have been primarily due to effective pretraining of speech representations.One such pretraining paradigm is the distillation of semantic knowledge from state-of-the-art text-based models like BERT to speech encoder neural networks.This work is a step towards doing the same in a much more efficient and fine-grained manner where we align speech embeddings and BERT embeddings on a token-by-token basis.We introduce a simple yet novel technique that uses a cross-modal attention mechanism to extract token-level contextual embeddings from a speech encoder such that these can be directly compared and aligned with BERT based contextual embeddings.This alignment is performed using a novel tokenwise contrastive loss.Fine-tuning such a pretrained model to perform intent recognition using speech directly yields state-of-the-art performance on two widely used SLU datasets.Our model improves further when fine-tuned with additional regularization using SpecAugment especially when speech is noisy, giving an absolute improvement as high as 8% over previous results. Vishal Sunder, Eric Fosler-Lussier, Samuel Thomas 0001, Hong-Kwang Jeff Kuo, Brian Kingsbury |
INTERSPEECH | 4 |
| 2021 | RNN Transducer Models for Spoken Language UnderstandingabstractWe present a comprehensive study on building and adapting RNN transducer (RNN-T) models for spoken language understanding (SLU). These end-to-end (E2E) models are constructed in three practical settings: a case where verbatim transcripts are available, a constrained case where the only available annotations are SLU labels and their values, and a more restrictive case where transcripts are available but not corresponding audio. We show how RNN-T SLU models can be developed starting from pre-trained automatic speech recognition (ASR) systems, followed by an SLU adaptation step. In settings where real audio data is not available, artificially synthesized speech is used to successfully adapt various SLU models. When evaluated on two SLU data sets, the ATIS corpus and a customer call center data set, the proposed models closely track the performance of other E2E models and achieve state-of-the-art results. Samuel Thomas 0001, Hong-Kwang Jeff Kuo, George Saon, Zoltán Tüske, Brian Kingsbury, Gakuto Kurata, Zvi Kons, Ron Hoory |
ICASSP | 2 |
| 2021 | End-to-End Spoken Language Understanding Using Transformer Networks and Self-Supervised Pre-Trained FeaturesabstractTransformer networks and self-supervised pre-training have consistently delivered state-of-art results in the field of natural language processing (NLP); however, their merits in the field of spoken language understanding (SLU) still need further investigation. In this paper we introduce a modular End-to-End (E2E) SLU transformer network based architecture which allows the use of self-supervised pre- trained acoustic features, pre-trained model initialization and multi-task training. Several SLU experiments for predicting intent and entity labels/values using the ATIS dataset are performed. These experiments investigate the interaction of pre-trained model initialization and multi-task training with either traditional filterbank or self-supervised pre-trained acoustic features. Results show not only that self-supervised pre-trained acoustic features outperform filterbank features in almost all the experiments, but also that when these features are used in combination with multi-task training, they almost eliminate the necessity of pre-trained model initialization. Edmilson da Silva Morais, Hong-Kwang Jeff Kuo, Samuel Thomas 0001, Zoltán Tüske, Brian Kingsbury |
ICASSP | 2 |
| 2021 | Speak or Chat with Me: End-to-End Spoken Language Understanding System with Flexible InputsabstractA major focus of recent research in spoken language understanding (SLU) has been on the end-to-end approach where a single model can predict intents directly from speech inputs without intermediate transcripts.However, this approach presents some challenges.First, since speech can be considered as personally identifiable information, in some cases only automatic speech recognition (ASR) transcripts are accessible.Second, intent-labeled speech data is scarce.To address the first challenge, we propose a novel system that can predict intents from flexible types of inputs: speech, ASR transcripts, or both.We demonstrate strong performance for either modality separately, and when both speech and ASR transcripts are available, through system combination, we achieve better results than using a single input modality.To address the second challenge, we leverage a semantically robust pre-trained BERT model and adopt a cross-modal system that co-trains text embeddings and acoustic embeddings in a shared latent space.We further enhance this system by utilizing an acoustic module pre-trained on LibriSpeech and domain-adapting the text module on our target datasets.Our experiments show significant advantages for these pre-training and fine-tuning strategies, resulting in a system that achieves competitive intent-classification performance on Snips SLU and Fluent Speech Commands datasets. Sujeong Cha, Wangrui Hou, Hyun Jung, My Phung, Michael Picheny, Hong-Kwang Jeff Kuo, Samuel Thomas 0001, Edmilson da Silva Morais |
Interspeech | 6 |
| 2021 | Integrating Dialog History into End-to-End Spoken Language Understanding SystemsabstractEnd-to-end spoken language understanding (SLU) systems that process human-human or human-computer interactions are often context independent and process each turn of a conversation independently. Spoken conversations on the other hand, are very much context dependent, and dialog history contains useful information that can improve the processing of each conversational turn. In this paper, we investigate the importance of dialog history and how it can be effectively integrated into end-to-end SLU systems. While processing a spoken utterance, our proposed RNN transducer (RNN-T) based SLU model has access to its dialog history in the form of decoded transcripts and SLU labels of previous turns. We encode the dialog history as BERT embeddings, and use them as an additional input to the SLU model along with the speech features for the current utterance. We evaluate our approach on a recently released spoken dialog data set, the HarperValleyBank corpus. We observe significant improvements: 8% for dialog action and 30% for caller intent recognition tasks, in comparison to a competitive context independent end-to-end baseline system. Jatin Ganhotra, Samuel Thomas 0001, Hong-Kwang Jeff Kuo, Sachindra Joshi, George Saon, Zoltán Tüske, Brian Kingsbury |
Interspeech | 3 |
| 2020 | Leveraging Unpaired Text Data for Training End-To-End Speech-to-Intent SystemsabstractTraining an end-to-end (E2E) neural network speech-to-intent (S2I) system that directly extracts intents from speech requires large amounts of intent-labeled speech data, which is time consuming and expensive to collect. Initializing the S2I model with an ASR model trained on copious speech data can alleviate data sparsity. In this paper, we attempt to leverage NLU text resources. We implemented a CTC-based S2I system that matches the performance of a state-of-the-art, traditional cascaded SLU system. We performed controlled experiments with varying amounts of speech and text training data. When only a tenth of the original data is available, intent classification accuracy degrades by 7.6% absolute. Assuming we have additional text-to-intent data (without speech) available, we investigated two techniques to improve the S2I system: (1) transfer learning, in which acoustic embeddings for intent classification are tied to fine-tuned BERT text embeddings; and (2) data augmentation, in which the text-to-intent data is converted into speech-to-intent data using a multi-speaker text-to-speech system. The proposed approaches recover 80% of performance lost due to using limited intent-labeled speech. Hong-Kwang Jeff Kuo, Samuel Thomas 0001, Zvi Kons, Kartik Audhkhasi, Brian Kingsbury, Ron Hoory, Michael Picheny |
ICASSP | 2 |
| 2020 | End-to-End Spoken Language Understanding Without Full TranscriptsabstractAn essential component of spoken language understanding (SLU) is slot filling: representing the meaning of a spoken utterance using semantic entity labels. In this paper, we develop end-to-end (E2E) spoken language understanding systems that directly convert speech input to semantic entities and investigate if these E2E SLU models can be trained solely on semantic entity annotations without word-for-word transcripts. Training such models is very useful as they can drastically reduce the cost of data collection. We created two types of such speech-to-entities models, a CTC model and an attention-based encoder-decoder model, by adapting models trained originally for speech recognition. Given that our experiments involve speech input, these systems need to recognize both the entity label and words representing the entity value correctly. For our speech-to-entities experiments on the ATIS corpus, both the CTC and attention models showed impressive ability to skip non-entity words: there was little degradation when trained on just entities versus full transcripts. We also explored the scenario where the entities are in an order not necessarily related to spoken order in the utterance. With its ability to do re-ordering, the attention model did remarkably well, achieving only about 2% degradation in speech-to-bag-of-entities F1 score. Hong-Kwang Jeff Kuo, Zoltán Tüske, Samuel Thomas 0001, Kartik Audhkhasi, Brian Kingsbury, Gakuto Kurata, Zvi Kons, Ron Hoory, Luis A. Lastras |
INTERSPEECH | 1 |
| 2018 | A Recorded Debating Dataset
Shachar Mirkin, Michal Jacovi, Tamar Lavee, Hong-Kwang Jeff Kuo, Samuel Thomas 0001, Leslie Sager, Lili Kotlerman, Elad Venezian, Noam Slonim |
LREC | 4 |
| 2016 | The IBM 2016 English Conversational Telephone Speech Recognition SystemabstractWe describe the latest improvements to the IBM English conversational telephone speech recognition system.Some of the techniques that were found beneficial are: maxout networks with annealed dropout rates; networks with a very large number of outputs trained on 2000 hours of data; joint modeling of partially unfolded recurrent neural networks and convolutional nets by combining the bottleneck and output layers and retraining the resulting model; and lastly, sophisticated language model rescoring with exponential and neural network LMs.These techniques result in an 8.0% word error rate on the Switchboard part of the Hub5-2000 evaluation test set which is 23% relative better than our previous best published result. George Saon, Tom Sercu, Steven J. Rennie, Hong-Kwang Jeff Kuo |
INTERSPEECH | 4 |
| 2015 | The IBM 2015 English conversational telephone speech recognition systemabstractWe describe the latest improvements to the IBM English conversational telephone speech recognition system. Some of the techniques that were found beneficial are: maxout networks with annealed dropout rates; networks with a very large number of outputs trained on 2000 hours of data; joint modeling of partially unfolded recurrent neural networks and convolutional nets by combining the bottleneck and output layers and retraining the resulting model; and lastly, sophisticated language model rescoring with exponential and neural network LMs. These techniques result in an 8.0% word error rate on the Switchboard part of the Hub5-2000 evaluation test set which is 23% relative better than our previous best published result. George Saon, Hong-Kwang Jeff Kuo, Steven J. Rennie, Michael Picheny |
INTERSPEECH | 2 |
| 2015 | The IBM BOLT speech transcription systemabstractWe describe the IBM automatic speech recognition (ASR) sys-tem for the DARPA Broad Operational Language Translation (BOLT) program. The system is used to transcribe conversa-tional telephone speech (CTS) prior to machine translation for Phase 3 of the program’s Activity A. The ASR system is a com-bination of novel sequence trained ensemble deep neural net-work acoustic models on speaker adapted features and convolu-tional neural network models on two kinds of spectro-temporal representations of speech, in conjunction with a variety of class, neural network and n-gram based language models. Acoustic and language models for the recognition system are built on transcribed audio released under the program and further opti-mized for the final machine translation task as well. The evalua-tion system has a word error rate of 32.7 % on a 2 hour Egyptian Arabic development set for this task. Index Terms: Automatic speech recognition, conversational telephone speech, deep neural networks, machine translation Samuel Thomas 0001, George Saon, Hong-Kwang Jeff Kuo, Lidia Mangu |
INTERSPEECH | 3 |
| 2014 | Out-of-vocabulary word detection in a speech-to-speech translation systemabstractIn this paper we describe progress we have made in detecting out-of-vocabulary words (OOVs) for a speech-to-speech translation system for the purpose of playing back audio to the user for clarification and correction. Our OOV detector follows a strategy of first identifying a rough location of the OOV and then merging adjacent decoded words to cover the true OOV word. We show the advantage of our OOV detection strategy and report on improvements using a real-time implementation of a new Convolutional Neural Network acoustic model. We discuss why commonly used metrics for OOV detection do not meet our needs and explore an overlap metric as well as a Jaccard metric for evaluating our ability to detect the OOVs and localize them accurately in time. We have found different metrics to be useful at different stages of development. Hong-Kwang Jeff Kuo, Ellen Eide, Lidia Mangu, Hagen Soltau, Tomás Beran |
ICASSP | 1 |
| 2014 | Efficient spoken term detection using confusion networksabstractIn this paper, we present a fast, vocabulary independent algorithm for spoken term detection (STD) that demonstrates a word-based index is sufficient to achieve good performance for both in-vocabulary (IV) and out-of-vocabulary (OOV) terms. Previous approaches have required that a separate index be built at the sub-word level and then expanded to allow for matching OOV terms. Such a process, while accurate, is expensive in both time and memory. In the proposed architecture, a word-level confusion network (CN) based index is used for both IV and OOV search. This is implemented using a flexible WFST framework. Comparisons on 3 Babel languages (Tagalog, Pashto and Turkish) show that CN-based indexing results in better performance compared with the lattice approach while being orders of magnitude faster and having a much smaller footprint. Lidia Mangu, Brian Kingsbury, Hagen Soltau, Hong-Kwang Jeff Kuo, Michael Picheny |
ICASSP | 4 |
| 2013 | The IBM keyword search system for the DARPA RATS programabstractThe paper describes a state-of-the-art keyword search (KWS) system in which significant improvements are obtained by using Convolutional Neural Network acoustic models, a two-step speech segmentation approach and a simplified ASR architecture optimized for KWS. The system described in this paper had the best performance in the 2013 DARPA RATS evaluation for both Levantine and Farsi. Lidia Mangu, Hagen Soltau, Hong-Kwang Jeff Kuo, George Saon |
ASRU | 3 |
| 2013 | Exploiting diversity for spoken term detectionabstractThe paper describes a state-of-the-art spoken term detection system in which significant improvements are obtained by diversifying the ASR engines used for indexing and combining the search results. First, we describe the design factors that, when varied, produce complementary STD systems and show that the performance of the combined system is 3 times better than the best individual component. Next, we describe different strategies for system combination and show that significant improvements can be achieved by normalizing the combined scores. We propose a classifier-based system combination strategy which outperforms a highly optimized baseline. The system described in this paper had the highest accuracy in the 2012 DARPA RATS evaluation. Lidia Mangu, Hagen Soltau, Hong-Kwang Jeff Kuo, Brian Kingsbury, George Saon |
ICASSP | 3 |
| 2013 | Morpheme-based feature-rich language models using Deep Neural Networks for LVCSR of Egyptian ArabicabstractEgyptian Arabic (EA) is a colloquial version of Arabic. It is a low-resource morphologically rich language that causes problems in Large Vocabulary Continuous Speech Recognition (LVCSR). Building LMs on morpheme level is considered a better choice to achieve higher lexical coverage and better LM probabilities. Another approach is to utilize information from additional features such as morphological tags. On the other hand, LMs based on Neural Networks (NNs) with a single hidden layer have shown superiority over the conventional n-gram LMs. Recently, Deep Neural Networks (DNNs) with multiple hidden layers have achieved better performance in various tasks. In this paper, we explore the use of feature-rich DNN-LMs, where the inputs to the network are a mixture of words and morphemes along with their features. Significant Word Error Rate (WER) reductions are achieved compared to the traditional word-based LMs. Amr El-Desoky Mousa, Hong-Kwang Jeff Kuo, Lidia Mangu, Hagen Soltau |
ICASSP | 2 |
| 2013 | Neural network acoustic models for the DARPA RATS program
Hagen Soltau, Hong-Kwang Jeff Kuo, Lidia Mangu, George Saon, Tomás Beran |
INTERSPEECH | 2 |
| 2012 | Large Scale Hierarchical Neural Network Language Models
Hong-Kwang Jeff Kuo, Ebru Arisoy, Ahmad Emami, Paul Vozila |
INTERSPEECH | 1 |
| 2011 | Minimum Bayes risk discriminative language models for Arabic speech recognitionabstractIn this paper we explore discriminative language modeling (DLM) on highly optimized state-of-the-art large vocabulary Arabic broadcast speech recognition systems used for the Phase 5 DARPA GALE Evaluation. In particular, we study in detail a minimum Bayes risk (MBR) criterion for DLM. MBR training outperforms perceptron training. Interestingly, we found that our DLMs generalized to mismatched conditions, such as using a different acoustic model during testing. We also examine the interesting problem of unsupervised DLM training using a Bayes risk metric as a surrogate for word error rate (WER). In some experiments, we were able to obtain about half of the gain of the supervised DLM. Hong-Kwang Jeff Kuo, Ebru Arisoy, Lidia Mangu, George Saon |
ASRU | 1 |
| 2011 | The IBM 2011 GALE Arabic speech transcription systemabstractWe describe the Arabic broadcast transcription system fielded by IBM in the GALE Phase 5 machine translation evaluation. Key advances over our Phase 4 system include a new Bayesian Sensing HMM acoustic model; multistream neural network features; a MADA vowelized acoustic model; and the use of a variety of language model techniques with significant additive gains. These advances were instrumental in achieving a word error rate of 7.4% on the Phase 5 evaluation set, and an absolute improvement of 0.9% word error rate over our 2009 system on the unsequestered Phase 4 evaluation data. Lidia Mangu, Hong-Kwang Jeff Kuo, Stephen M. Chu, Brian Kingsbury, George Saon, Hagen Soltau, Fadi Biadsy |
ASRU | 2 |
| 2011 | The IBM 2009 GALE Arabic speech transcription systemabstractWe describe the Arabic broadcast transcription system fielded by IBM in the GALE Phase 4 machine translation evaluation. Key advances over our Phase 3.5 system include improvements to context-dependent modeling in vowelized Arabic acoustic models; the use of neural-network features provided by the International Computer Science Institute; Model M language models; a neural network language model that uses syntactic and morphological features; and improvements to our system combination strategy. These advances were instrumental in achieving a word error rate of 8.9% on the Phase 4 evaluation set, and an absolute improvement of 1.6% word error rate over our 2008 system on the unsequestered Phase 3.5 evaluation data. Brian Kingsbury, Hagen Soltau, George Saon, Stephen M. Chu, Hong-Kwang Jeff Kuo, Lidia Mangu, Suman V. Ravuri, Nelson Morgan, Adam Janin |
ICASSP | 5 |
| 2011 | Feature Combination Approaches for Discriminative Language Models
Ebru Arisoy, Bhuvana Ramabhadran, Hong-Kwang Jeff Kuo |
INTERSPEECH | 3 |
| 2010 | The 2009 IBM GALE Mandarin broadcast transcription systemabstractThis paper gives an up-to-date description of the IBM Mandarin broadcast transcription system developed under the DARPA GALE program. Technical advances over our previous system include a novel acoustic modeling approach using subspace Gaussian mixture models, a speaking rate adaptation method using frame rate normalization, and an effective recipe for lattice combination. We present results on three consortium-defined test sets. It is shown that with these advances, the new system attains a 9% relative reduction in character error rate compared to our previous GALE evaluation system. The reported 9.1% error rate on the phase three evaluation set represents the state of the art in Mandarin broadcast speech transcription. Stephen M. Chu, Daniel Povey, Hong-Kwang Jeff Kuo, Lidia Mangu, Shilei Zhang, Qin Shi 0001, Yong Qin 0001 |
ICASSP | 3 |
| 2010 | Morphological and syntactic features for Arabic speech recognitionabstractIn this paper, we study the use of morphological and syntactic context features to improve speech recognition of a morphologically rich language like Arabic. We examine a variety of syntactic features, including part-of-speech tags, shallow parse tags, and exposed head words and their non-terminal labels both before and after the word to be predicted. Neural network LMs are used to model these features since they generalize better to unseen events by modeling words and other context features in continuous space. Using morphological and syntactic features, we can improve the word error rate (WER) significantly on various test sets, including EVAL'08U, the unsequestered portion of the DARPA GALE Phase 3 evaluation test set. Hong-Kwang Jeff Kuo, Lidia Mangu, Ahmad Emami, Imed Zitouni |
ICASSP | 1 |
| 2010 | A comparative study on system combination schemes for LVCSRabstractWe present a comparative study on combination schemes for large vocabulary continuous speech recognition by incorporating long-span class posterior probability features into conventional short-time cepstral features. System combination can improve the overall speech recognition performance when multiple systems exhibit different error patterns and multiple knowledge sources encode complementary information. A variety of combination approaches are investigated in this paper, e.g., feature concatenation single stream system, model combination multi-stream system, lattice rescoring and ROVER. These techniques work at different levels of a LVCSR system and have different computational cost. We compared their performance and analyzed their advantages and disadvantages on large vocabulary English broadcast news transcription tasks. Experimental results showed that model combination with independent tree consistently outperforms ROVER, feature concatenation and lattice rescoring. In addition, the phoneme posterior probability features do provide complementary information to short-time cepstral features. Chengyuan Ma, Hong-Kwang Jeff Kuo, Hagen Soltau, Upendra V. Chaudhari, Lidia Mangu |
ICASSP | 2 |
| 2010 | The IBM 2008 GALE Arabic speech transcription systemabstractThis paper describes the Arabic broadcast transcription system fielded by IBM in the GALE Phase 3.5 machine translation evaluation. Key advances compared to our Phase 2.5 system include improved discriminative training, the use of Subspace Gaussian Mixture Models (SGMM), neural network acoustic features, variable frame rate decoding, training data partitioning experiments, unpruned n-gram language models and neural network language models. These advances were instrumental in achieving a word error rate of 8.9% on the evaluation test set. George Saon, Hagen Soltau, Upendra V. Chaudhari, Stephen M. Chu, Brian Kingsbury, Hong-Kwang Jeff Kuo, Lidia Mangu, Daniel Povey |
ICASSP | 6 |
| 2010 | Augmented context features for Arabic speech recognition
Ahmad Emami, Hong-Kwang Jeff Kuo, Imed Zitouni, Lidia Mangu |
INTERSPEECH | 2 |
| 2009 | Syntactic features for Arabic speech recognitionabstractWe report word error rate improvements with syntactic features using a neural probabilistic language model through N-best re-scoring. The syntactic features we use include exposed head words and their non-terminal labels both before and after the predicted word. Neural network LMs generalize better to unseen events by modeling words and other context features in continuous space. They are suitable for incorporating many different types of features, including syntactic features, where there is no pre-defined back-off order. We choose an N-best re-scoring framework to be able to take full advantage of the complete parse tree of the entire sentence. Using syntactic features, along with morphological features, improves the word error rate (WER) by up to 5.5% relative, from 9.4% to 8.6%, on the latest GALE evaluation test set. Hong-Kwang Jeff Kuo, Lidia Mangu, Ahmad Emami, Imed Zitouni, Young-Suk Lee 0001 |
ASRU | 1 |
| 2009 | A framework for rapid development of conversational natural language call routing systems for call centers
Ea-Ee Jan, Hong-Kwang Jeff Kuo, Osamuyimen Stewart, David M. Lubensky |
INTERSPEECH | 2 |
| 2009 | Advances in Arabic Speech Transcription at IBM Under the DARPA GALE ProgramabstractThis paper describes the Arabic broadcast transcription system fielded by IBM in the GALE Phase 2.5 machine translation evaluation. Key advances include the use of additional training data from the Linguistic Data Consortium (LDC), use of a very large vocabulary comprising 737 K words and 2.5 M pronunciation variants, automatic vowelization using flat-start training, cross-adaptation between unvowelized and vowelized acoustic models, and rescoring with a neural-network language model. The resulting system achieves word error rates below 10% on Arabic broadcasts. Very large scale experiments with unsupervised training demonstrate that the utility of unsupervised data depends on the amount of supervised data available. While unsupervised training improves system performance when a limited amount (135 h) of supervised data is available, these gains disappear when a greater amount (848 h) of supervised data is used, even with a very large (7069 h) corpus of unsupervised data. We also describe a method for modeling Arabic dialects that avoids the problem of data sparseness entailed by dialect-specific acoustic models via the use of non-phonetic, dialect questions in the decision trees. We show how this method can be used with a statically compiled decoding graph by partitioning the decision trees into a static component and a dynamic component, with the dynamic component being replaced by a mapping that is evaluated at run-time. Hagen Soltau, George Saon, Brian Kingsbury, Hong-Kwang Jeff Kuo, Lidia Mangu, Daniel Povey, Ahmad Emami |
IEEE Trans. Speech Audio Process. | 4 |
| 2008 | Recent advances in the IBM GALE Mandarin transcription systemabstractThis paper describes the system and algorithmic developments in the automatic transcription of Mandarin broadcast speech made at IBM in the second year of the DARPA GALE program. Technical advances over our previous system include improved acoustic models using embedded tone modeling, and a new topic-adaptive language model (LM) rescoring technique based on dynamically generated LMs. We present results on three community-defined test sets designed to cover both the broadcast news and the broadcast conversation domain. It is shown that our new baseline system attains a 15.4% relative reduction in character error rate compared with our previous GALE evaluation system. And a further 13.6% improvement over the baseline is achieved with the two described techniques. Selina M. Chu, Hong-Kwang Jeff Kuo, Lidia Mangu, Yi Y. Liu 0002, Yong Qin 0001, Qin Shi 0001, Shilei Zhang, Hagai Aronowitz |
ICASSP | 2 |
| 2008 | Discriminative graph training for ultra-fast low-footprint speech indexing
Upendra V. Chaudhari, Hong-Kwang Jeff Kuo, Brian Kingsbury |
INTERSPEECH | 2 |
| 2008 | A study of unsupervised clustering techniques for language modeling
Sangyun Hahn, Abhinav Sethy, Hong-Kwang Jeff Kuo, Bhuvana Ramabhadran |
INTERSPEECH | 3 |
| 2008 | XMLLR for improved speaker adaptation in speech recognitionabstractIn this paper we describe a novel technique for adaptation of Gaussian means. The technique is related to Maximum Likelihood Linear Regression (MLLR), but we regress not on the mean itself but on a vector associated with each mean. These associated vectors are initialized by an ingenious technique based on eigen decomposition. As the only form of adaptation this technique outperforms MLLR, even with multiple regression classes and Speaker Adaptive Training (SAT). However, when combined with Constrained MLLR (CMLLR) and Vocal Tract Length Normalization (VTLN) the improvements disappear. The combination of two forms of SAT (CMLLR-SAT and MLLR-SAT) which we performed as a baseline is itself a useful result; we describe it more fully in a companion paper. XMLLR is an interesting approach which we hope may have utility in other contexts, for example in speaker identification. Daniel Povey, Hong-Kwang Jeff Kuo |
INTERSPEECH | 2 |
| 2008 | Fast speaker adaptive training for speech recognition
Daniel Povey, Hong-Kwang Jeff Kuo, Hagen Soltau |
INTERSPEECH | 2 |
| 2008 | Search and classification based language model adaptation
Qin Shi 0001, Stephen M. Chu, Hong-Kwang Jeff Kuo, Yi Y. Liu 0002, Yong Qin 0001 |
INTERSPEECH | 4 |
| 2007 | The IBM Mandarin Broadcast Speech Transcription SystemabstractThis paper describes the technical and system building advances in the automatic transcription of Mandarin broadcast speech made at IBM in the first year of the DARPA GALE program. In particular, we discuss the application of minimum phone error (MPE) discriminative training and a new topic-adaptive language modeling technique. We present results on both the RT04 evaluation data and two larger community-defined test sets designed to cover both the broadcast news and the broadcast conversation domain. It is shown that with the described advances, the new transcription system achieves a 26.3% relative reduction in character error rate over our previous best-performing system, and is competitive with published numbers on these datasets. Stephen M. Chu, Hong-Kwang Jeff Kuo, Yi Y. Liu 0002, Yong Qin 0001, Qin Shi 0001, Geoffrey Zweig |
ICASSP (2) | 2 |
| 2007 | Discriminative Training of Decoding Graphs for Large Vocabulary Continuous Speech RecognitionabstractFinite-state decoding graphs integrate the decision trees, pronunciation model and language model for speech recognition into a unified representation of the search space. We explore discriminative training of the transition weights in the decoding graph in the context of large vocabulary speech recognition. In preliminary experiments on the RT-03 English Broadcast News evaluation set, the word error rate was reduced by about 5.7% relative, from 23.0% to 21.7%. We discuss how this method is particularly applicable to low-latency and low-resource applications such as real-time closed captioning of broadcast news and interactive speech-to-speech translation. Hong-Kwang Jeff Kuo, Brian Kingsbury, Geoffrey Zweig |
ICASSP (4) | 1 |
| 2007 | The IBM 2006 Gale Arabic ASR SystemabstractThis paper describes the advances made in IBM's Arabic broadcast news transcription system which was fielded in the 2006 GALE ASR and machine translation evaluation. These advances were instrumental in lowering the word error rate by 42% relative over the course of one year and include: training on additional LDC data, large-scale discriminative training on 1800 hours of unsupervised data, automatic vowelization using a flat-start approach, use of a large vocabulary with 617K words and 2 million pronunciations and lastly, a system architecture based on cross-adaptation between unvowelized and vowelized acoustic models. Hagen Soltau, George Saon, Brian Kingsbury, Hong-Kwang Jeff Kuo, Lidia Mangu, Daniel Povey, Geoffrey Zweig |
ICASSP (4) | 4 |
| 2007 | A data visualization and analysis method for natural language call routing system design
Hong-Kwang Jeff Kuo, Vaibhava Goel |
INTERSPEECH | 1 |
| 2006 | IBM Mastor: Multilingual Automatic Speech-To-Speech TranslatorabstractIn this paper, we describe the IBM MASTOR systems which handle spontaneous free-form speech-to-speech translation on both laptop and hand-held PDAs. Challenges include speech recognition and machine translation in adverse environments, lack of data and linguistic resources for under-studied languages, and the need to rapidly develop capabilities for new languages. Importantly, the code and models must fit within the limited memory and computational resources of hand-held devices. We describe our approaches, experience, and success in building working free-form S2S systems that can handle two language pairs (including a low-resource language). Bowen Zhou 0006, Liang Gu, Ruhi Sarikaya, Hong-Kwang Jeff Kuo, Antti-Veikko I. Rosti, Mohamed Afify, Weizhong Zhu |
ICASSP (5) | 5 |
| 2006 | On the use of morphological analysis for dialectal Arabic speech recognitionabstractArabic has a large number of affixes that can modify a stem to form words. In automatic speech recognition (ASR) this leads to a high out-of-vocabulary (OOV) rate for typical lexicon size, and hence a potential increase in WER. This is even more pronounced for dialects of Arabic where additional affixes are often introduced and the available data is typically sparse. To address this problem we introduce a simple word decomposition algorithm which only requires a text corpus and a predefined list of affixes. Using this al-gorithm to create the lexicon for Iraqi Arabic ASR results in about 10 % relative improvement in word error rate (WER). Also using the union of the segmented and unsegmented vocabularies and in-terpolating the corresponding language models results in further WER reduction. The net WER improvement is about 13%. 1. Mohamed Afify, Ruhi Sarikaya, Hong-Kwang Jeff Kuo, Laurent Besacier |
INTERSPEECH | 3 |
| 2006 | Maximum entropy direct models for speech recognitionabstractTraditional statistical models for speech recognition have mostly been based on a Bayesian framework using generative models such as hidden Markov models (HMMs). This paper focuses on a new framework for speech recognition using maximum entropy direct modeling, where the probability of a state or word sequence given an observation sequence is computed directly from the model. In contrast to HMMs, features can be asynchronous and overlapping. This model therefore allows for the potential combination of many different types of features, which need not be statistically independent of each other. In this paper, a specific kind of direct model, the maximum entropy Markov model (MEMM), is studied. Even with conventional acoustic features, the approach already shows promising results for phone level decoding. The MEMM significantly outperforms traditional HMMs in word error rate when used as stand-alone acoustic models. Preliminary results combining the MEMM scores with HMM and language model scores show modest improvements over the best HMM speech recognizer. Hong-Kwang Jeff Kuo |
IEEE Trans. Speech Audio Process. | 1 |
| 2005 | Portability challenges in developing interactive dialogue systemsabstractStatistical methods commonly used in developing interactive dialogue systems require large amounts of training data to achieve high accuracy and robustness. This becomes a major bottleneck in building free-style dialogue systems in a new domain or for a new language. Portability challenges hence arise regarding how to build statistical models rapidly and with low cost in terms of data collection, transcription and annotation. In this paper, we discuss challenges as well as potential solutions in several critical issues of efficient language modeling, utilization of untranscribed speech data, automatic annotation, and cross-lingual modeling. We believe that current approaches in these areas are far from mature and call for serious efforts from the research community. Liang Gu, Hong-Kwang Jeff Kuo |
ICASSP (5) | 3 |
| 2005 | Language Model Estimation for Optimizing End-to-end Performance of a Natural Language Call Routing SystemabstractConventional methods for training statistical models for automatic speech recognition, such as acoustic and language models, have focused on criteria such as maximum likelihood and sentence or word error rate (WER). However, unlike dictation systems, the goal for spoken dialogue systems is to understand the meaning of what a person says, not to get every word correctly transcribed. For such systems, we propose to optimize the statistical models under end-to-end system performance criteria. We illustrate this principle by focusing on the estimation of the language model (LM) component of a natural language call routing system. This estimation, carried out under a conditional maximum likelihood objective, aims at optimizing the call routing (classification) accuracy, which is often the criterion of interest in these systems. LM updates are derived using the extended Baum-Welch procedure (Gopalakrishnan et al. (1991)). In our experiments, we find that our estimation procedure leads to a small but promising gain in classification accuracy. Interestingly, the estimated language models also lead to an increase in the word error rate while improving the classification accuracy, showing that the system with the best classification accuracy is not necessarily the one with the lowest WER. Significantly, our LM estimation procedure does not require the correct transcription of the training data, and can therefore be applied to unsupervised learning from untranscribed speech data. Vaibhava Goel, Hong-Kwang Jeff Kuo, Sabine Deligne |
ICASSP (1) | 2 |
| 2005 | Rapid transition to new spoken dialogue domains: language model training using knowledge from previous domain applications and web text resources
Murat Akbacak, Liang Gu, Hong-Kwang Jeff Kuo |
INTERSPEECH | 4 |
| 2005 | Active learning with minimum expected error for spoken language understanding
Hong-Kwang Jeff Kuo, Vaibhava Goel |
INTERSPEECH | 1 |
| 2005 | Exploiting unlabeled data using multiple classifiers for improved natural language call-routing
Ruhi Sarikaya, Hong-Kwang Jeff Kuo, Vaibhava Goel |
INTERSPEECH | 2 |
| 2005 | Improving end-to-end performance of call classification through data confusion reduction and model tolerance enhancement
Xiang Li 0071, Hong-Kwang Jeff Kuo, E. E. Jan, Vaibhava Goel, David M. Lubensky |
INTERSPEECH | 3 |
| 2005 | A framework for predicting speech recognition errors
Eric Fosler-Lussier, Ingunn Amdal, Hong-Kwang Jeff Kuo |
Speech Commun. | 3 |
| 2004 | Maximum entropy direct model as a unified model for acoustic modeling in speech recognition
Hong-Kwang Jeff Kuo |
INTERSPEECH | 1 |
| 2004 | An automatic dialogue generation platform for personalized dialogue applications
Andrew N. Pargellis, Hong-Kwang Jeff Kuo |
Speech Commun. | 2 |
| 2003 | Minimum verification error training for topic verificationabstractWe propose a new formulation of minimum verification error training and apply it to the problem of topic verification as an example. In topic verification, a decision is made as to whether a document truly belongs to a particular topic of interest. Such a decision typically depends on a comparison between a model for the desired topic and a model for background topics, using a decision threshold. We propose modeling the background topics as a cohort model consisting of a weighted combination of the M closest topics discovered from the training data. The weights and the decision threshold are optimized using the generalized probabilistic descent algorithm to explicitly minimize the verification error rate, which is defined to be a weighted sum of the Type I (false rejection) and Type II (false acceptance) errors. Hong-Kwang Jeff Kuo, Imed Zitouni, Eric Fosler-Lussier |
ICASSP (1) | 1 |
| 2003 | Boosting and combination of classifiers for natural language call routing systems
Imed Zitouni, Hong-Kwang Jeff Kuo |
Speech Commun. | 2 |
| 2003 | Discriminative training of natural language call routersabstractThis paper shows how discriminative training can significantly improve classifiers used in natural language processing, using as an example the task of natural language call routing, where callers are transferred to desired departments based on natural spoken responses to an open-ended "How may I direct your call?" prompt. With vector-based natural language call routing, callers are transferred using a routing matrix trained on statistics of occurrence of words and word sequences in a training corpus. By re-training the routing matrix parameters using a minimum classification error criterion, a relative error rate reduction of 10-30% was achieved on a banking task. Increased robustness was demonstrated in that with 10% rejection, the error rate was reduced by 40%. Discriminative training also improves portability; we were able to train call routers with the highest known performance using as input only text transcription of routed calls, without any human intervention or knowledge about what terms are important or irrelevant for the routing task. This strategy was validated with both the banking task and a more difficult task involving calls to operators in the UK. The proposed formulation is applicable to algorithms addressing a broad range of speech understanding, information retrieval, and topic identification problems. Hong-Kwang Jeff Kuo |
IEEE Trans. Speech Audio Process. | 1 |
| 2002 | Discriminative training of language models for speech recognitionabstractIn this paper we describe how discriminative training can be applied to language models for speech recognition. Language models are important to guide the speech recognition search, particularly in compensating for mistakes in acoustic decoding. A frequently used measure of the quality of language models is the perplexity; however, what is more important for accurate decoding is not necessarily having the maximum likelihood hypothesis, but rather the best separation of the correct string from the competing, acoustically confusible hypotheses. Discriminative training can help to improve language models for the purpose of speech recognition by improving the separation of the correct hypothesis from the competing hypotheses. We describe the algorithm and demonstrate modest improvements in word and sentence error rates on the DARPA Communicator task without any increase in language model complexity. Hong-Kwang Jeff Kuo, Eric Fosler-Lussier, Hui Jiang 0001 |
ICASSP | 1 |
| 2002 | Adaptive language models for spoken dialogue systemsabstractIn this paper, we investigate both generative and statistical approaches for language modeling in spoken dialogue systems. Semantic class-based finite state and n-gram grammars are used for improving coverage and modeling accuracy when little training data is available. We have implemented dialogue-state specific language model adaptation to reduce perplexity and improve the efficiency of grammars for spoken dialogue systems. A novel algorithm for combining state-independent n-gram and state-dependent finite state grammars using acoustic confidence scores is proposed. Using this combination strategy, a relative word error reduction of 12% is achieved for certain dialogue states within a travel reservation task. Finally, semantic class multigrams are proposed and briefly evaluated for language modeling in dialogue systems. Roger Argiles Solsona, Eric Fosler-Lussier, Hong-Kwang Jeff Kuo, Alexandros Potamianos, Imed Zitouni |
ICASSP | 3 |
| 2002 | Combination of boosting and discriminative training for natural language call steering systemsabstractIn this paper, we describe the combination of two different techniques to improve natural language call routing: boosting and discriminative training. The goal of boosting is to re-weight the data in order to train a set of classifiers whose errors may be uncorrelated so that when combined, the classification error rate (CER) can be reduced. We propose using discriminative training to improve the individual classifier accuracy at each iteration of the boosting algorithm. Compared to the baseline classifiers, an improvement in the CER of 41–50% was observed on call routing for a banking task. More importantly, synergistic effects of discriminative training on the boosting algorithm were demonstrated: more iterations were possible because discriminative training reduced the CER of individual classifiers trained on re-weighted data by an average of 72%. Imed Zitouni, Hong-Kwang Jeff Kuo |
ICASSP | 2 |
| 2002 | Discriminative training for call classification and routing
Hong-Kwang Jeff Kuo, Imed Zitouni, Eric Fosler-Lussier, Egbert Ammicht |
INTERSPEECH | 1 |
| 2002 | Backoff hierarchical class n-gram language modelling for automatic speech recognition systems
Imed Zitouni, Olivier Siohan, Hong-Kwang Jeff Kuo |
INTERSPEECH | 3 |
| 2001 | Using semantic class information for rapid development of language models within ASR dialogue systemsabstractWhen dialogue system developers tackle a new domain, much effort is required; the development of different parts of the system usually proceeds independently. Yet it may be profitable to coordinate development efforts between different modules. We focus our efforts on extending small amounts of language model training data by integrating semantic classes that were created for a natural language understanding module. By converting finite state parses of a training corpus into a probabilistic context free grammar and subsequently generating artificial data from the context free grammar, we can significantly reduce perplexity and automatic speech recognition (ASR) word error for situations with little training data. Experiments are presented using data from the ATIS and DARPA Communicator travel corpora. Eric Fosler-Lussier, Hong-Kwang Jeff Kuo |
ICASSP | 2 |
| 2001 | Simplifying design specification for automatic training of robust natural language call routerabstractWe study techniques that allow us to relax some constraints imposed by expert knowledge in task specifications of a natural language call router design. We intend to fully automate the training of the routing matrix while still maintaining the same level of performance (over 90% accuracy) as that in an optimized system. Two specific issues are investigated: (1) reducing the matrix size by removing word pairs and triplets in key term definition while using only single word terms; and (2) increasing the matrix size by removing the need for defining stop words and performing stop word filtering. Since simplification of design often implies a degradation of performance, discriminative training of routing matrix parameters becomes an essential procedure. We show in our experiments that the performance degradation caused by relaxing the design constraints can be compensated entirely by minimum error classification (MCE) training even with the above two simplifications. We believe the procedure is applicable to algorithms addressing a broad range of speech understanding, topic identification, and information retrieval problems. Hong-Kwang Jeff Kuo |
ICASSP | 1 |
| 2001 | OASIS natural language call steering trialabstractA recent trial of natural language call steering on live UK calls to the operator is described along with its results. The characteristics of the problem are described along with the acoustic, language, semantic and dialogue modelling approaches employed. Natural language call steering is found to be viable, with recognition and semantic accuracy the current limiting factors. Peter J. Durston, Mark Farrell, David Attwater, James Allen, Hong-Kwang Jeff Kuo, Mohamed Afify, Eric Fosler-Lussier |
INTERSPEECH | 5 |
| 2001 | A portability study on natural language call steeringabstractIn this paper we examine the portability of the vector-based call router to a new task involving calls to the operator in the UK. One component of the router was shown to require expert knowledge and hand-tuning: the stop word list. Stop word filtering involves replacing certain words with place markers and is necessary to reduce the number of features and parameters used by the classifier. Two specific approaches that eliminates the need for stop word filtering were investigated that led to comparable classification performance: (1) using trigram, bigram, and unigram features and using SVD to reduce the number of parameters, and (2) using only unigram features and applying discriminative training to boost the performance. After discriminative training, the classification error rate was reduced by 18-30% over the baseline unigram results. Increased robustness is demonstrated by a 24-48% reduction in error rate at 20% false rejection rate. Hong-Kwang Jeff Kuo |
INTERSPEECH | 1 |
| 2000 | Natural language call steering for service applicationsabstractABSTRACT In this paper, a dialogue system for natural language based call steering is described and studied. The system is based on natural language speech recognition and understanding within a mixed initiative dialogue. The system is implemented on Bell Labs. Speech Technology Integration Platform (BLSTIP) using dialogue and natural language understanding components from BT laboratories. A prototype system in the operator service domain [2] is described. In order to improve the acoustic and language modeling for natural language based dialogue applications, various approaches are described and studied. The structure of the dialogue manager is also presented in which mixed-initiative dialogue can be supported with efficiency. Call classification and steering experiments were performed. The results confirm the efficacy of the proposed approach. 1. INTRODUCTION Natural language dialogue between human and machine is a challenge. In order to make a natural language based dialogue system successful, various efforts are made to improve the accuracy, flexibility and robustness of the system component technologies, such as speech recognition, speech understanding, dialogue generation and dialogue manager, text-to-speech synthesis, etc. Such a complex dialogue application imposes stringent requirements on the flexibility of the system platform. One of the drawbacks in systems deployed in the past is the limitation imposed by the finite state grammar on the language that a user can use to communicate with the machine. Although such constraint alleviates the complexity and problem in recognizing human speech, it becomes an obstacle to support more powerful, user friendly and flexible dialogue systems for mixed-initiative dialogues. In this paper, we study issues encountered in designing and implementing a natural language based call steering application for telephone service calls. This is a complicated application, and it performs a detailed diagnostic dialogue to identify the service problem, such as a troubled telephone line and etc., that the user is experiencing. It provides the desired service after receiving user’s consent and confirmation [2]. In the prototype system studied in this paper, the dialogue can go deep through many turns. The natural language based request and query from the user is recognized through natural language based automatic speech recognition. There is no constraint on the way that the user should communicate to the system. It allows the user to make direct requests as well as provide a description of the problem where the final action will be identified as the outcome of the dialogue. A call classifier provides natural language understanding based on the word string from the speech recognition output. The dialogue manager uses this understanding to determine the next appropriate system action. The organization of this paper is as follows. In Section 2, the dialogue system architecture and design are presented which support natural language based mixed-initiative dialogue applications such as call steering, movie locator, etc. Section 3 is devoted to natural language based speech recognition and statistical language modeling for dialogue applications. Section 4 is concentrated on the dialogue manager design and automatic query generation. Call classification and steering are studied in Section 5 and results are given based on a case study in a telephone service application. Wu Chou, Qiru Zhou, Hong-Kwang Jeff Kuo, Antoine Saad, David Attwater, Peter J. Durston, Mark Farrell, Frank Scahill |
INTERSPEECH | 3 |
| 2000 | Discriminative training in natural language call routing
Hong-Kwang Jeff Kuo |
INTERSPEECH | 1 |
| 2000 | Dialogue management in the Bell Labs communicator system
Alexandros Potamianos, Egbert Ammicht, Hong-Kwang Jeff Kuo |
INTERSPEECH | 3 |
| 2000 | Statistical recursive finite state machine parsing for speech understanding
Alexandros Potamianos, Hong-Kwang Jeff Kuo |
INTERSPEECH | 2 |
| 1999 | Phrase-based language models for speech recognitionabstractIncluding phrases in the vocabulary list can improve n-gram language models used in speech recognition. In this paper, we report results of automatic extraction of phrases from the training text using frequency, likelihood, and correlation criteria. We show how a language model built from a vocabulary that includes useful phrases can systematically improve language model perplexity in a natural language call-routing task and the 20K-Nov92 Wall Street Journal evaluation. We also discuss the impact of such phrase-based language models on recognition word error rate. Hong-Kwang Jeff Kuo, Wolfgang Reichl |
EUROSPEECH | 1 |
| 1999 | Automatic dialogue generator creates user defined applications
Andrew N. Pargellis, Hong-Kwang Jeff Kuo |
EUROSPEECH | 2 |
| 1992 | Speaker set identification through speaker group modeling
Hong-Kwang Jeff Kuo, Aaron E. Rosenberg |
ICSLP | 1 |