VLDB 2026 Research / reviewers in the wild / expert
Gakuto Kurata
dblp:20/4496
· DBLP profile ↗
60ranked-venue papers
20as first author
14since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 56 · 18 first-author · 13 since 2021Artificial intelligence and machine learning · 39 · 12 first-author · 9 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Granite-speech: open-source speech-aware LLMs with strong English ASR capabilitiesabstractGranite-speech LLMs are compact and efficient speech language models specifically designed for English ASR1and automatic speech translation (AST). The models were trained by modality aligning granite-3.3-instruct to speech on publicly available open-source corpora. Comprehensive benchmarking on English ASR shows that they outperform several competitors’ models that were trained on orders of magnitude more proprietary data, and they keep pace on English-to-X AST for major European languages, Japanese, and Mandarin. The speech-specific components are: a conformer acoustic encoder using block attention and self-conditioning trained with connectionist temporal classification, a windowed query-transformer speech modality adapter used to do temporal downsampling of the acoustic embeddings and map them to the LLM text embedding space, and LoRA adapters to further fine-tune the text LLM. The models are freely available on HuggingFace2under a permissive Apache 2.0 license.1The latest models (revision 3.3.2) support multilingual ASR in English, French, German, Spanish and Portuguese and bidirectional speech translation to and from English. This paper covers the initial English-only release.2https://huggingface.co/ibm-granite/granite-speech-3.3-2b (and…-8b). George Saon, Avihu Dekel, Alexi Brooks, Tohru Nagano, Abraham Daniels, Aharon Satt, Ashish R. Mittal, Brian Kingsbury, David Haws, Edmilson da Silva Morais, Gakuto Kurata, Hagai Aronowitz, Ibrahim Ibrahim, Hong-Kwang Jeff Kuo, Kate Soule, Luis A. Lastras, Masayuki Suzuki, Ron Hoory, Samuel Thomas 0001, Sashi Novitasari, Takashi Fukuda, Vishal Sunder, Zvi Kons |
ASRU | 11 |
| 2025 | Knowledge Distillation Based Training of Unified Conformer CTC Models for Multi-form ASRabstractThere is an on-going body of research on training separate dedicated models for either short-form or long-form utterances. Multi-form acoustic models that are simply trained on combined data from long-form and short-form utterances often suffer from various negative impacts due to the diversity of a speaking style, an accent, and a recording condition. In addition, a linguistic mismatch that comes from an utterance length is also another factor of the degradation. In this paper we investigate novel techniques for training unified Conformer-based models on multi-form speech data obtained from diverse domains and sources to serve multiple downstream applications with a single model. Our approach incorporates chunk-wise short-term discriminative knowledge distillation with an encoder embedding masking and mitigates the aforementioned problems that appear for single unified models. We show the benefit of our proposed technique on long and short-form ASR test sets by comparing our models against several variants trained by mixing utterances with various audio lengths. The proposed technique provides a significant improvement of up to 8.5% relative WER reduction over baseline systems that operate at a similar decoding cost. Takashi Fukuda, Gakuto Kurata, George Saon |
ICASSP | 2 |
| 2025 | LLM based Text Generation for Improved Low-resource Speech Recognition ModelsabstractLimited transcribed spoken style data is a critical bottleneck in building automatic speech recognition (ASR) systems for low-resource languages. Prompting a large language model (LLM) to paraphrase input text can generate novel text data that is constrained to be semantically similar to the source data. We leverage this capability of LLMs to improve the performance of low-resource ASR systems by increasing the limited text training data while keeping the same spoken style. Since word sequences in the training data are now more diverse and the vocabulary of the ASR model is also expanded, this approach allows for building general purpose ASR without prior knowledge of various domains in the low-resource language. In our experiments with Brazilian Portuguese as a low-resource language, paraphrased data enhanced the n-gram language model (LM) used to build the weighted finite state transducer (WFST) for decoding with a Conformer-CTC speech recognition model, resulting in improvement of word error rate (WER) by 15.6% over the baseline model. Synthesizing the paraphrased text into speech and using it to fine-tune the acoustic model (AM) component helped to further improve the WER by 2.9%, achieving a combined improvement of 18.5%. We also demonstrate the usefulness of our proposed approach for high-resource languages like English. Tohru Nagano, Gakuto Kurata, Samuel Thomas 0001, Hong-Kwang Jeff Kuo, Daniel Bolaños, Hyun Jung, George Saon |
ICASSP | 2 |
| 2025 | Improving End-to-end Mixed-case ASR with Knowledge Distillation and Integration of Voice Activity Cues
Sashi Novitasari, Takashi Fukuda, Gakuto Kurata |
INTERSPEECH | 3 |
| 2025 | Voice Activity-based Text Segmentation for ASR Text Denormalization
Sashi Novitasari, Takashi Fukuda, Gakuto Kurata |
INTERSPEECH | 3 |
| 2024 | Multiple Representation Transfer from Large Language Models to End-to-End ASR SystemsabstractTransferring the knowledge of large language models (LLMs) is a promising technique to incorporate linguistic knowledge into end-to-end automatic speech recognition (ASR) systems. However, existing works only transfer a single representation of LLM (e.g. the last layer of pretrained BERT), while the representation of a text is inherently non-unique and can be obtained variously from different layers, contexts and models. In this work, we explore a wide range of techniques to obtain and transfer multiple representations of LLMs into a transducer-based ASR system. While being conceptually simple, we show that transferring multiple representations of LLMs can be an effective alternative to transferring only a single LLM representation. Takuma Udagawa, Masayuki Suzuki, Gakuto Kurata, Masayasu Muraoka, George Saon |
ICASSP | 3 |
| 2023 | Speech-enriched Memory for Inference-time Adaptation of ASR Models to Word DictionariesabstractDespite the impressive performance of ASR models on mainstream benchmarks, their performance on rare words is unsatisfactory.In enterprise settings, often a focused list of entities (such as locations, names, etc) are available which can be used to adapt the model to the terminology of specific domains.In this paper, we present a novel inference algorithm that improves the prediction of state-of-the-art ASR models using nearest-neighbor-based matching on an inference-time word list.We consider both the Transducer architecture that is useful in the streaming setting, and state-of-the-art encoder-decoder models such as Whisper.In our approach, a list of rare entities is indexed in a memory by synthesizing speech for each entry, and then storing the internal acoustic and language model states obtained from the best possible alignment on the ASR model.The memory is organized as a trie which we harness to perform a stateful lookup during inference.A key property of our extension is that we prevent spurious matches by restricting to only word-level matches.In our experiments on publicly available datasets and private benchmarks, we show that our method is effective in significantly improving rare word recognition. Ashish R. Mittal, Sunita Sarawagi, Preethi Jyothi, George Saon, Gakuto Kurata |
EMNLP | 5 |
| 2022 | Improving Generalization of Deep Neural Network Acoustic Models with Length Perturbation and N-best Based Label SmoothingabstractWe introduce two techniques, length perturbation and n-best based label smoothing, to improve generalization of deep neural network (DNN) acoustic models for automatic speech recognition (ASR).Length perturbation is a data augmentation algorithm that randomly drops and inserts frames of an utterance to alter the length of the speech feature sequence.N-best based label smoothing randomly injects noise to ground truth labels during training in order to avoid overfitting, where the noisy labels are generated from n-best hypotheses.We evaluate these two techniques extensively on the 300-hour Switchboard (SWB300) dataset and an in-house 500-hour Japanese (JPN500) dataset using recurrent neural network transducer (RNNT) acoustic models for ASR.We show that both techniques improve the generalization of RNNT models individually and they can also be complementary.In particular, they yield good improvements over a strong SWB300 baseline and give state-of-art performance on SWB300 using RNNT models. George Saon, Tohru Nagano, Masayuki Suzuki, Takashi Fukuda, Brian Kingsbury, Gakuto Kurata |
INTERSPEECH | 7 |
| 2022 | Global RNN Transducer Models For Multi-dialect Speech Recognition
Takashi Fukuda, Samuel Thomas 0001, Masayuki Suzuki, Gakuto Kurata, George Saon, Brian Kingsbury |
INTERSPEECH | 4 |
| 2022 | Improving ASR Robustness in Noisy Condition Through VAD Integration
Sashi Novitasari, Takashi Fukuda, Gakuto Kurata |
INTERSPEECH | 3 |
| 2022 | Effect and Analysis of Large-scale Language Model Rescoring on Competitive ASR Systems
Takuma Udagawa, Masayuki Suzuki, Gakuto Kurata, Nobuyasu Itoh, George Saon |
INTERSPEECH | 3 |
| 2021 | RNN Transducer Models for Spoken Language UnderstandingabstractWe present a comprehensive study on building and adapting RNN transducer (RNN-T) models for spoken language understanding (SLU). These end-to-end (E2E) models are constructed in three practical settings: a case where verbatim transcripts are available, a constrained case where the only available annotations are SLU labels and their values, and a more restrictive case where transcripts are available but not corresponding audio. We show how RNN-T SLU models can be developed starting from pre-trained automatic speech recognition (ASR) systems, followed by an SLU adaptation step. In settings where real audio data is not available, artificially synthesized speech is used to successfully adapt various SLU models. When evaluated on two SLU data sets, the ATIS corpus and a customer call center data set, the proposed models closely track the performance of other E2E models and achieve state-of-the-art results. Samuel Thomas 0001, Hong-Kwang Jeff Kuo, George Saon, Zoltán Tüske, Brian Kingsbury, Gakuto Kurata, Zvi Kons, Ron Hoory |
ICASSP | 6 |
| 2021 | Generalized Knowledge Distillation from an Ensemble of Specialized Teachers Leveraging Unsupervised Neural ClusteringabstractThis paper proposes an improved generalized knowledge distillation framework with multiple dissimilar teacher networks, each of which is specialized for a specific domain, to make a deployable student network more robust to challenging acoustic environments. In this paper, we first address a method to partition the training data for constructing ensembles of the teachers from unsupervised neural clustering with features based on context-dependent phonemes representing each acoustic domain. Second, we illustrate how a single student network designed from partitioned data is effectively trained with multiple specialized teachers. During the training step, the weights of the student network are updated using a composite two-part cross entropy loss obtained from a pair consisting of a specialized teacher corresponding to input speech and a generalized teacher trained with a balanced data set. Unlike system combination methods, we aim to incorporate the benefits from multiple models into a single student network via knowledge distillation that does not increase any computational costs during the decoding time. The improvement of the proposed technique is shown on acoustically diverse signals contaminated by challenging practical noises. Takashi Fukuda, Gakuto Kurata |
ICASSP | 2 |
| 2021 | Improving Customization of Neural Transducers by Mitigating Acoustic Mismatch of Synthesized Audio
Gakuto Kurata, George Saon, Brian Kingsbury, David Haws, Zoltán Tüske |
Interspeech | 1 |
| 2020 | Converting Written Language to Spoken Language with Neural Machine Translation for Language ModelingabstractWhen building a language model (LM) for spontaneous speech, the ideal situation is to have a large amount of spoken, in-domain training data. Having such abundant data, however, is not realistic. We address this problem by generating texts in spoken language from those in written language by using a neural machine translation (NMT) model. We collected faithful transcripts of fully spontaneous speech and corresponding written versions and used them as a parallel corpus to train the NMT model. We used top-k random sampling, which generates a large variety of texts of higher quality as compared to other generation methods for NMT. We indicate that the NMT model is capable of converting written texts in a certain domain to spoken texts, and that the converted texts are effective for training LMs. Our experimental results show significant improvement of speech recognition accuracy with the LMs. Shintaro Ando, Masayuki Suzuki, Nobuyasu Itoh, Gakuto Kurata, Nobuaki Minematsu |
ICASSP | 4 |
| 2020 | Speaker Embeddings Incorporating Acoustic Conditions for DiarizationabstractWe present our work on training speaker embeddings, especially effective for speaker diarization. For various speaker recognition tasks, extracting speaker embeddings using Deep Neural Networks (DNNs) has become major methods. These embeddings are generally trained to be discriminate speakers and be robust with respect to different acoustic conditions. In speaker diarization, however, the acoustic conditions can be used as consistent information for discriminating speakers. Such information can include the distances to a microphone in a meeting, or the channels for each speaker in telephone conversation recorded in monaural. Hence, the proposed speaker-embedding network leverages differences in acoustic conditions to train effective speaker embeddings for speaker diarization. The information on acoustic conditions can be anything that contributes to distinguishing between recording environments; for example, we explore using i-vectors. Experiments conducted on a practical diarization system demonstrated that the proposed embeddings significantly improve performance over embeddings without information on acoustic conditions. Yosuke Higuchi, Masayuki Suzuki, Gakuto Kurata |
ICASSP | 3 |
| 2020 | New Advances in Speaker Diarization
Hagai Aronowitz, Weizhong Zhu, Masayuki Suzuki, Gakuto Kurata, Ron Hoory |
INTERSPEECH | 4 |
| 2020 | End-to-End Spoken Language Understanding Without Full TranscriptsabstractAn essential component of spoken language understanding (SLU) is slot filling: representing the meaning of a spoken utterance using semantic entity labels. In this paper, we develop end-to-end (E2E) spoken language understanding systems that directly convert speech input to semantic entities and investigate if these E2E SLU models can be trained solely on semantic entity annotations without word-for-word transcripts. Training such models is very useful as they can drastically reduce the cost of data collection. We created two types of such speech-to-entities models, a CTC model and an attention-based encoder-decoder model, by adapting models trained originally for speech recognition. Given that our experiments involve speech input, these systems need to recognize both the entity label and words representing the entity value correctly. For our speech-to-entities experiments on the ATIS corpus, both the CTC and attention models showed impressive ability to skip non-entity words: there was little degradation when trained on just entities versus full transcripts. We also explored the scenario where the entities are in an order not necessarily related to spoken order in the utterance. With its ability to do re-ordering, the attention model did remarkably well, achieving only about 2% degradation in speech-to-bag-of-entities F1 score. Hong-Kwang Jeff Kuo, Zoltán Tüske, Samuel Thomas 0001, Kartik Audhkhasi, Brian Kingsbury, Gakuto Kurata, Zvi Kons, Ron Hoory, Luis A. Lastras |
INTERSPEECH | 7 |
| 2020 | Knowledge Distillation from Offline to Streaming RNN Transducer for End-to-End Speech Recognition
Gakuto Kurata, George Saon |
INTERSPEECH | 1 |
| 2019 | Data Augmentation Based on Vowel Stretch for Improving Children's Speech RecognitionabstractProlongation is a speech disfluency that lengthens some portions of speech utterances. It is frequently observed in children's spontaneous speech, while it is rare in read speech. To make acoustic models more robust to children's spontaneous speech, collecting a large amount of children's speech data containing prolongation is usually required, which is very impractical in many cases. To tackle this problem, we propose a novel data augmentation method that virtually generates additional data by simulating prolongation. The method inserts pseudo frames into specific positions of speech utterances to simulate prolongation. The acoustic features of the inserted frames are calculated from the original frames on both sides. This is based on our analysis that many of vowels are actually stretched in children's spontaneous speech. Our proposed procedure can generate partially stretched utterances with low computational costs, unlike a conventional speed or tempo perturbation method that extends and shrinks entire utterances at a uniform rate. The effectiveness of the proposed method were confirmed with the experiments of acoustic model adaptations, in which our proposed method focusing on vowel stretch showed consistent improvement compared with conventional speed and tempo perturbation approach. Tohru Nagano, Takashi Fukuda, Masayuki Suzuki, Gakuto Kurata |
ASRU | 4 |
| 2019 | Improvements to N-gram Language Model Using Text Generated from Neural Language ModelabstractAlthough neural language models have emerged, n-gram language models are still used for many speech recognition tasks. This paper proposes four methods to improve n-gram language models using text generated from a recurrent neural network language model (RNNLM). First, we use multiple RNNLMs from different domains instead of a single RNNLM. The final n-gram language model is obtained by interpolating generated n-gram models from each domain. Second, we use subwords instead of words for RNNLM to reduce the out-of-vocabulary rate. Third, we generate text templates using an RNNLM for template-based data augmentation for named entities. Fourth, we use both forward RNNLM and backward RNNLM to generate text. We found that these four methods improved performance of speech recognition up to 4% relative in various tasks. Masayuki Suzuki, Nobuyasu Itoh, Tohru Nagano, Gakuto Kurata, Samuel Thomas 0001 |
ICASSP | 4 |
| 2019 | English Broadcast News Speech Recognition by Humans and MachinesabstractWith recent advances in deep learning, considerable attention has been given to achieving automatic speech recognition performance close to human performance on tasks like conversational telephone speech (CTS) recognition. In this paper we evaluate the usefulness of these proposed techniques on broadcast news (BN), a similar challenging task. We also perform a set of recognition measurements to understand how close the achieved automatic speech recognition results are to human performance on this task. On two publicly available BN test sets, DEV04F and RT04, our speech recognition system using LSTM and residual network based acoustic models with a combination of n-gram and neural network language models performs at 6.5% and 5.9% word error rate. By achieving new performance milestones on these test sets, our experiments show that techniques developed on other related tasks, like CTS, can be transferred to achieve similar performance. In contrast, the best measured human recognition performance on these test sets is much lower, at 3.6% and 2.8% respectively, indicating that there is still room for new techniques and improvements in this space, to reach human performance levels. Samuel Thomas 0001, Masayuki Suzuki, Gakuto Kurata, Zoltán Tüske, George Saon, Brian Kingsbury, Michael Picheny, Tom Dibert, Alice Kaiser-Schatzlein, Bern Samko |
ICASSP | 4 |
| 2019 | Direct Neuron-Wise Fusion of Cognate Neural Networks
Takashi Fukuda, Masayuki Suzuki, Gakuto Kurata |
INTERSPEECH | 3 |
| 2019 | Guiding CTC Posterior Spike Timings for Improved Posterior Fusion and Knowledge DistillationabstractConventional automatic speech recognition (ASR) systems trained from frame-level alignments can easily leverage posterior fusion to improve ASR accuracy and build a better single model with knowledge distillation.End-to-end ASR systems trained using the Connectionist Temporal Classification (CTC) loss do not require frame-level alignment and hence simplify model training.However, sparse and arbitrary posterior spike timings from CTC models pose a new set of challenges in posterior fusion from multiple models and knowledge distillation between CTC models.We propose a method to train a CTC model so that its spike timings are guided to align with those of a pre-trained guiding CTC model.As a result, all models that share the same guiding model have aligned spike timings.We show the advantage of our method in various scenarios including posterior fusion of CTC models and knowledge distillation between CTC models with different architectures.With the 300-hour Switchboard training data, the single word CTC model distilled from multiple models improved the word error rates to 13.7%/23.1% from 14.9%/24.1% on the Hub5 2000 Switchboard/CallHome test sets without using any data augmentation, language model, or complex decoder. Gakuto Kurata, Kartik Audhkhasi |
INTERSPEECH | 1 |
| 2019 | Multi-Task CTC Training with Auxiliary Feature Reconstruction for End-to-End Speech Recognition
Gakuto Kurata, Kartik Audhkhasi |
INTERSPEECH | 1 |
| 2018 | Data Augmentation Improves Recognition of Foreign Accented Speech
Takashi Fukuda, Raul Fernandez, Andrew Rosenberg, Samuel Thomas 0001, Bhuvana Ramabhadran, Alexander Sorin, Gakuto Kurata |
INTERSPEECH | 7 |
| 2018 | Inference-Invariant Transformation of Batch Normalization for Domain Adaptation of Acoustic Models
Masayuki Suzuki, Tohru Nagano, Gakuto Kurata, Samuel Thomas 0001 |
INTERSPEECH | 3 |
| 2018 | Improved Knowledge Distillation from Bi-Directional to Uni-Directional LSTM CTC for End-to-End Speech RecognitionabstractEnd-to-end automatic speech recognition (ASR) promises to simplify model training and deployment. Most end-to-end ASR systems utilize a bi-directional Long Short-Term Memory (BiLSTM) acoustic model due to its ability to capture acoustic context from the entire utterance. However, BiLSTM models have high latency and cannot be used in streaming applications. Leveraging knowledge distillation to train a low-latency end-to-end uni-directional LSTM (UniLSTM) model from a BiLSTM model can be an option. However, it makes the strict assumption of shared frame-wise time alignments between the two models. We propose an improved knowledge distillation algorithm that relaxes this assumption and improves the accuracy of the UniLSTM model. We confirmed the advantage of the proposed method on a standard English conversational telephone speech recognition task. Gakuto Kurata, Kartik Audhkhasi |
SLT | 1 |
| 2017 | Language modeling with highway LSTMabstractLanguage models (LMs) based on Long Short Term Memory (LSTM) have shown good gains in many automatic speech recognition tasks. In this paper, we extend an LSTM by adding highway networks inside an LSTM and use the resulting Highway LSTM (HW-LSTM) model for language modeling. The added highway networks increase the depth in the time dimension. Since a typical LSTM has two internal states, a memory cell and a hidden state, we compare various types of HW-LSTM by adding highway networks onto the memory cell and/or the hidden state. Experimental results on English broadcast news and conversational telephone speech recognition show that the proposed HW-LSTM LM improves speech recognition accuracy on top of a strong LSTM LM baseline. We report 5.1% and 9.9% on the Switchboard and CallHome subsets of the Hub5 2000 evaluation, which reaches the best performance numbers reported on these tasks to date. Gakuto Kurata, Bhuvana Ramabhadran, George Saon, Abhinav Sethy |
ASRU | 1 |
| 2017 | Effective joint training of denoising feature space transforms and Neural Network based acoustic modelsabstractNeural Network (NN) based acoustic frontends, such as denoising autoencoders, are actively being investigated to improve the robustness of NN based acoustic models to various noise conditions. In recent work the joint training of such frontends with backend NNs has been shown to significantly improve speech recognition performance. In this paper, we propose an effective algorithm to jointly train such a denoising feature space transform and a NN based acoustic model with various kinds of data. Our proposed method first pretrains a Convolutional Neural Network (CNN) based denoising frontend and then jointly trains this frontend with a NN backend acoustic model. In the unsupervised pretraining stage, the frontend is designed to estimate clean log Mel-filterbank features from noisy log-power spectral input features. A subsequent multi-stage training of the proposed frontend, with the dropout technique applied only at the joint layer between the frontend and backend NNs, leads to significant improvements in the overall performance. On the Aurora-4 task, our proposed system achieves an average WER of 9.98%. This is a 9.0% relative improvement over one of the best reported speaker independent baseline system's performance. A final semi-supervised adaptation of the frontend NN, similar to feature space adaptation, reduces the average WER to 7.39%, a further relative WER improvement of 25%. Takashi Fukuda, Osamu Ichikawa, Gakuto Kurata, Ryuki Tachibana, Samuel Thomas 0001, Bhuvana Ramabhadran |
ICASSP | 3 |
| 2017 | Harmonic feature fusion for robust neural network-based acoustic modelingabstractAcoustic modeling with deep learning has drastically improved the performance of automatic speech recognition (ASR) where the main stream of the acoustic feature is still log-Mel filtered one. While the log-Mel filtered features lose harmonic-structure information, they still include useful information for ASR. Several attempts have been made to integrate higher-resolution information into the network. In order to improve the ASR accuracy in noisy conditions, we propose new features integrated into acoustic modeling to represent which parts in the time-frequency domain have a distinct harmonic structure, since it is partially observed in noisy environments. The new features are combined with the standard acoustic features, and the network is trained with them using various noisy data. Through these operations, it learns the acoustic features with a kind of quality tag describing which parts are clean or degraded. Our model reduced the word error rate in an Aurora-4 task by 10.3% in DNN compared with the strong baseline while retaining the high accuracy in clean test cases. Osamu Ichikawa, Takashi Fukuda, Masayuki Suzuki, Gakuto Kurata, Bhuvana Ramabhadran |
ICASSP | 4 |
| 2017 | Efficient Knowledge Distillation from an Ensemble of Teachers
Takashi Fukuda, Masayuki Suzuki, Gakuto Kurata, Samuel Thomas 0001, Jia Cui, Bhuvana Ramabhadran |
INTERSPEECH | 3 |
| 2017 | Ensembles of Multi-Scale VGG Acoustic Models
Michael Heck, Masayuki Suzuki, Takashi Fukuda, Gakuto Kurata, Satoshi Nakamura 0001 |
INTERSPEECH | 4 |
| 2017 | Factorial Modeling for Effective Suppression of Directional Noise
Osamu Ichikawa, Takashi Fukuda, Gakuto Kurata, Steven J. Rennie |
INTERSPEECH | 3 |
| 2017 | Empirical Exploration of Novel Architectures and Objectives for Language Models
Gakuto Kurata, Abhinav Sethy, Bhuvana Ramabhadran, George Saon |
INTERSPEECH | 1 |
| 2017 | English Conversational Telephone Speech Recognition by Humans and MachinesabstractOne of the most difficult speech recognition tasks is accurate recognition of human to human communication. Advances in deep learning over the last few years have produced major speech recognition improvements on the representative Switchboard conversational corpus. Word error rates that just a few years ago were 14% have dropped to 8.0%, then 6.6% and most recently 5.8%, and are now believed to be within striking range of human performance. This then raises two issues - what IS human performance, and how far down can we still drive speech recognition error rates? A recent paper by Microsoft suggests that we have already achieved human performance. In trying to verify this statement, we performed an independent set of human performance measurements on two conversational tasks and found that human performance may be considerably better than what was earlier reported, giving the community a significantly harder goal to achieve. We also report on our own efforts in this area, presenting a set of acoustic and language modeling techniques that lowered the word error rate of our own English conversational telephone LVCSR system to the level of 5.5%/10.3% on the Switchboard/CallHome subsets of the Hub5 2000 evaluation, which - at least at the writing of this paper - is a new performance milestone (albeit not at what we measure to be human performance!). On the acoustic side, we use a score fusion of three models: one LSTM with multiple feature inputs, a second LSTM trained with speaker-adversarial multi-task learning and a third residual net (ResNet) with 25 convolutional layers and time-dilated convolutions. On the language modeling side, we use word and character LSTMs and convolutional WaveNet-style language models. George Saon, Gakuto Kurata, Tom Sercu, Kartik Audhkhasi, Samuel Thomas 0001, Dimitrios Dimitriadis, Bhuvana Ramabhadran, Michael Picheny, Lynn-Li Lim, Bergul Roomi, Phil Hall |
INTERSPEECH | 2 |
| 2017 | Symbol Sequence Search from Telephone Conversation
Masayuki Suzuki, Gakuto Kurata, Abhinav Sethy, Bhuvana Ramabhadran, Kenneth Church 0001, Mark Drake |
INTERSPEECH | 2 |
| 2016 | Leveraging Sentence-level Information with Encoder LSTM for Semantic Slot FillingabstractRecurrent Neural Network (RNN) and one of its specific architectures, Long Short-Term Memory (LSTM), have been widely used for sequence labeling.Explicitly modeling output label dependencies on top of RNN/LSTM is a widely-studied and effective extension.We propose another extension to incorporate the global information spanning over the whole input sequence.The proposed method, encoder-labeler LSTM, first encodes the whole input sequence into a fixed length vector with the encoder LSTM, and then uses this encoded vector as the initial state of another LSTM for sequence labeling.With this method, we can predict the label sequence while taking the whole input sequence information into consideration.In the experiments of a slot filling task, which is an essential component of natural language understanding, with using the standard ATIS corpus, we achieved the state-of-the-art F 1 -score of 95.66%. Gakuto Kurata, Bing Xiang, Bowen Zhou 0006, Mo Yu |
EMNLP | 1 |
| 2016 | Speech recognition robust against speech overlapping in monaural recordings of telephone conversationsabstractMonaural (single-channel) recording is sometimes used for telephone conversations in call centers. Generally speaking, the accuracy of automatic speech recognition of a monaural recording is worse than that of the multi-channel recording of the same conversation where each speaker's voice is separately recorded. The major reason is that the recognition system fails not only at the overlapping segments where the voices of the multiple speakers overlap, but also at the neighboring segments surrounding the overlapping segments. In this paper, we tackle this problem by using a combination of garbage modeling and noise-robust monaural acoustic modeling. Our proposed method trains the models by making use of multi-channel recordings and transcripts, which are relatively easy to prepare than monaural recordings and transcripts. We present experimental results where the proposed methods reduced the error rates by approximately 3% relative to the baseline methods for both of GMM-HMM and CNN-HMM cases. Because the proposed method is quite simple, the proposed method is easy to deploy to wide range of ASR systems for monaural speech transcription. Masayuki Suzuki, Gakuto Kurata, Tohru Nagano, Ryuki Tachibana |
ICASSP | 2 |
| 2016 | Improved Neural Network Initialization by Grouping Context-Dependent Targets for Acoustic Modeling
Gakuto Kurata, Brian Kingsbury |
INTERSPEECH | 1 |
| 2016 | Labeled Data Generation with Encoder-Decoder LSTM for Semantic Slot Filling
Gakuto Kurata, Bing Xiang, Bowen Zhou 0006 |
INTERSPEECH | 1 |
| 2016 | Improved Neural Network-based Multi-label Classification with Better Initialization Leveraging Label Co-occurrence
Gakuto Kurata, Bing Xiang, Bowen Zhou 0006 |
HLT-NAACL | 1 |
| 2015 | A metric for evaluating speech recognizer output based on human-perception modelabstractWord error rate or character error rate are usually used as the metrics for evaluating the accuracy of speech recognition. These are naturally-defined objective metrics and are helpful for comparing recognition methods fairly. However the overall performance of the recognition systems and the usefulness of the results are not necessarily considered. To address this problem, we study and propose a metric which replicates human-annotated scores using their perception to the recognition results. The features that we use are the numbers of insertion errors, deletion errors, and substitution errors in the characters and the syllables. In addition we studied the numbers of consecutive errors, the misrecognized keywords, and the locations of errors. We created models using linear regression and random forest, predicted human-perceived scores, and compared them with the actual scores using Spearman’s rank-based correlation. According to our experiments the correlation of human perceived scores with character error rates is 0.456, while those with the predicted scores by using a random forest of 10 features is 0.715. The latter is close to the averaged correlation between the scores of the human subjects, 0.765, which suggests that we can predict the human-perceived scores using those features and that we can leverage human perception model for evaluating speech recognition performance. The important factors (features) for the prediction are the numbers of substitution errors and consecutive errors. Nobuyasu Itoh, Gakuto Kurata, Ryuki Tachibana, Masafumi Nishimura |
INTERSPEECH | 2 |
| 2015 | Deep neural network training emphasizing central frames
Gakuto Kurata, Daniel Willett |
INTERSPEECH | 1 |
| 2015 | Discriminative re-ranking for automatic speech recognition by leveraging invariant structures
Masayuki Suzuki, Gakuto Kurata, Masafumi Nishimura, Nobuaki Minematsu |
Speech Commun. | 2 |
| 2012 | Discriminative Reranking for LVCSR Leveraging Invariant StructureabstractAn invariant structure is one of the long-span acoustic represen-tations, where acoustic variations caused by non-linguistic fac-tors are effectively removed from speech. We present in this pa-per a new method to leverage the invariant structures as features of discriminative reranking for Large Vocabulary Continuous Speech Recognition (LVCSR). First we use a traditional HMM-based LVCSR system to get a list of N-best candidates with phone alignments and construct an invariant structure for each candidate using its phone alignment. Here, the invariant struc-ture is composed of lengths between every two phonemes in the candidate. Then we estimate a score of each phoneme-pair in the invariant structure, and rerank the N-best candidates using a weighted sum of the phoneme-pair scores, where the weights are trained discriminatively by averaged perceptron. Experi-mental results show a relative CER improvement of 6.69 % over the baseline HMM-based LVCSR system. Index Terms: Invariant Structure, LVCSR, Discriminative reranking 1. Masayuki Suzuki, Gakuto Kurata, Masafumi Nishimura, Nobuaki Minematsu |
INTERSPEECH | 2 |
| 2012 | Leveraging word confusion networks for named entity modeling and detection from conversational telephone speech
Gakuto Kurata, Nobuyasu Itoh, Masafumi Nishimura, Abhinav Sethy, Bhuvana Ramabhadran |
Speech Commun. | 1 |
| 2012 | Acoustically discriminative language model training with pseudo-hypothesis
Gakuto Kurata, Abhinav Sethy, Bhuvana Ramabhadran, Ariya Rastrow, Nobuyasu Itoh, Masafumi Nishimura |
Speech Commun. | 1 |
| 2011 | Training of error-corrective model for ASR without using audio dataabstractThis paper introduces a method to train an error-corrective model for Automatic Speech Recognition (ASR) without using audio data. In existing techniques, it is assumed that sufficient audio data of the tar get application is available and negative samples can be prepared by having ASR recognize this audio data. However, this assumption is not always true. We propose generating probable N-best lists, which the ASR may produce, directly from the text data of the target application by taking phoneme similarity into consideration. We call this process "Pseudo-ASR". We conduct discriminative reranking with the error-corrective model by regarding the text data as positive samples and the N-best lists from the Pseudo-ASR as negative samples. Experiments with Japanese call center data showed that discriminative reranking based on the Pseudo-ASR improved the accuracy of the ASR. Gakuto Kurata, Nobuyasu Itoh, Masafumi Nishimura |
ICASSP | 1 |
| 2011 | Named entity recognition from Conversational Telephone Speech leveraging Word Confusion Networks for training and recognitionabstractNamed Entity (NE) recognition from the results of Automatic Speech Recognition (ASR) is challenging because of ASR errors. To detect NEs, one of the options is to use a statistical NE model that is usually trained with ASR one-best results. In order to make NE recognition more robust to ASR errors, we propose using Word Confusion Networks (WCNs), sequences of bundled words, for both NE modeling and recognition by regarding the word bundles as units instead of the independent words. This is done by clustering similar word bundles that may originate from the same word. We trained the NE models with the maximum entropy principle and evaluated the performance using real-life call-center data. The results showed that by using the WCNs, the error of NE recognition was relatively reduced by up to 33.0%. Gakuto Kurata, Nobuyasu Itoh, Masafumi Nishimura, Abhinav Sethy, Bhuvana Ramabhadran |
ICASSP | 1 |
| 2011 | Acoustic Model Training with Detecting Transcription Errors in the Training Data
Gakuto Kurata, Nobuyasu Itoh, Masafumi Nishimura |
INTERSPEECH | 1 |
| 2011 | Continuous Digits Recognition Leveraging Invariant StructureabstractRecently, an invariant structure of speech was proposed, where the inevitable acoustic variations caused by non-linguistic fac-tors are effectively removed from speech. The invariant struc-ture was applied to isolated word recognition and the experi-mental results showed good performance. However, the pre-vious method can’t apply to continuous speech recognition di-rectly because there was no efficient decoding algorithm. In this paper, we propose a method to leverage the invariant structure in continuous digits recognition. We use a traditional HMM-based Automatic Speech Recognition (ASR) system to get N-best lists with phone alignments. Then we construct invariant structures using these phone alignments and re-rank the N-best lists by investigating which hypothesis is structurally more valid. Experimental results show a relative WER improvement of 17.4 % over the baseline HMM-based ASR system. Masayuki Suzuki, Gakuto Kurata, Masafumi Nishimura, Nobuaki Minematsu |
INTERSPEECH | 2 |
| 2009 | Acoustically discriminative training for language modelsabstractThis paper introduces a discriminative training for language models (LMs) by leveraging phoneme similarities estimated from an acoustic model. To train an LM discriminatively, we needed the correct word sequences and the recognized results that automatic speech recognition (ASR) produced by processing the utterances of those correct word sequences. But, sufficient utterances are not always available. We propose to generate the probable N-best lists, which the ASR may produce, directly from the correct word sequences by leveraging the phoneme similarities. We call this process the ldquoPseudo-ASRrdquo. We train the LM discriminatively by comparing the correct word sequences and the corresponding N-best lists from the Pseudo-ASR. Experiments with real-life data from a Japanese call center showed that the LM trained with the proposed method improved the accuracy of the ASR. Gakuto Kurata, Nobuyasu Itoh, Masafumi Nishimura |
ICASSP | 1 |
| 2007 | Unsupervised Lexicon Acquisition from Speech and TextabstractWhen introducing a large vocabulary continuous speech recognition (LVCSR) system into a specific domain, it is preferable to add the necessary domain-specific words and their correct pronunciations selectively to the lexicon, especially in the areas where the LVCSR system should be updated frequently by adding new words. In this paper, we propose an unsupervised method of word acquisition in Japanese, where no spaces exist between words. In our method, by taking advantage of the speech of the target domain, we selected the domain-specific words among an enormous number of word candidates extracted from the raw corpora. The experiments showed that the acquired lexicon was of good quality and that it contributed to the performance of the LVCSR system for the target domain. Gakuto Kurata, Shinsuke Mori, Nobuyasu Itoh, Masafumi Nishimura |
ICASSP (4) | 1 |
| 2007 | Preliminary experiments toward automatic generation of new TTS voices from recorded speech alone
Ryuki Tachibana, Tohru Nagano, Gakuto Kurata, Masafumi Nishimura, Noboru Babaguchi |
INTERSPEECH | 3 |
| 2006 | Phoneme-to-Text Transcription System with an Infinite VocabularyabstractThe noisy channel model approach is successfully applied to various natural language processing tasks. Currently the main research focus of this approach is adaptation methods, how to capture characteristics of words and expressions in a target domain given example sentences in that domain. As a solution we describe a method enlarging the vocabulary of a language model to an almost infinite size and capturing their context information. Especially the new method is suitable for languages in which words are not delimited by whitespace. We applied our method to a phoneme-to-text transcription task in Japanese and reduced about 10% of the errors in the results of an existing method. Shinsuke Mori, Daisuke Takuma, Gakuto Kurata |
ACL | 3 |
| 2006 | Unsupervised Adaptation of a Stochastic Language Model Using a Japanese Raw CorpusabstractThe target uses of large vocabulary continuous speech recognition (LVCSR) systems are spreading. It takes a lot of time to build a good LVCSR system specialized for the target domain because experts need to manually segment the corpus of the target domain, which is a labor-intensive task. In this paper, we propose a new method to adapt an LVCSR system to a new domain. In our method, we stochastically segment a Japanese raw corpus of the target domain. Then a domain-specific language model (LM) is built based on this corpus. All of the domain-specific words can be added to the lexicon for LVCSR. Most importantly, the proposed method is fully automatic. Therefore, we can reduce the time for introducing an LVCSR system drastically. In addition, the proposed method yielded a comparable or even superior performance to use of expensive manual segmentation Gakuto Kurata, Shinsuke Mori, Masafumi Nishimura |
ICASSP (1) | 1 |
| 2005 | Class-based variable memory length Markov modelabstractIn this paper, we present a class-based variable memory length Markov model and its learning algorithm. This is an extension of a variable memory length Markov model. Our model is based on a class-based probabilistic suffix tree, whose nodes have an automatically acquired wordclass relation. We experimentally compared our new model with a word-based bi-gram model, a word-based tri-gram model, a class-based bi-gram model, and a word-based variable memory length Markov model. The results show that a class-based variable memory length Markov model outperforms the other models in perplexity and model size. 1. Shinsuke Mori, Gakuto Kurata |
INTERSPEECH | 2 |
| 2002 | Integration of MLLR adaptation with pronunciation proficiency adaptation for non-native speech recognitionabstractTo recognize non-native speech, larger acoustic/linguistic distortions must be handled adequately in acoustic modeling, language modeling, lexical modeling, and/or decoding strategy. In this paper, a novel method to enhance MLLR adaptation of acoustic models for non-native speech recognition is proposed. In the case of native speech recognition, MLLR speaker adaptation was successfully introduced because it enables efficient adaptation with a small number of adaptation data by using a regression tree of Gaussian mixtures of HMMs. However, as for non-native speech, most of the cases, the regression tree built from the baseline HMMs does not match with pronunciation proficiency of a speaker. This paper provides a solution for this problem, where the speaker’s proficiency is automatically estimated and the tree suited for the proficiency is built, which can be viewed as proficiency adaptation. Recognition experiments show that MLLR with the new tree raises the averaged error reduction rate up to about 30 % from the baseline MLLR performance of approximately 20 %. Nobuaki Minematsu, Gakuto Kurata, Keikichi Hirose |
INTERSPEECH | 2 |
| 2002 | Corpus-based analysis of English spoken by Japanese students in view of the entire phonemic system of English
Nobuaki Minematsu, Gakuto Kurata, Keikichi Hirose |
INTERSPEECH | 2 |