Shigeki Karita

dblp:166/2740 · DBLP profile ↗
← Back
21ranked-venue papers
7as first author
6since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 21 · 7 first-author · 6 since 2021Artificial intelligence and machine learning · 15 · 5 first-author · 5 since 2021
YearPublicationVenuePosition
2024 FLEURS-R: A Restored Multilingual Speech Corpus for Generation Tasks
Yuma Koizumi, Shigeki Karita, Heiga Zen, Jason Riesa, Haruko Ishikawa, Michiel Bacchiani
INTERSPEECH3
2023 LibriTTS-R: A Restored Multi-Speaker Text-to-Speech Corpus
Yuma Koizumi, Heiga Zen, Shigeki Karita, Yifan Ding 0004, Kohei Yatabe, Nobuyuki Morioka, Michiel Bacchiani, Yu Zhang 0033, Wei Han 0002, Ankur Bapna
INTERSPEECH3
2022 Knowledge Transfer from Large-Scale Pretrained Language Models to End-To-End Speech Recognizers
abstract
End-to-end speech recognition is a promising technology for enabling compact automatic speech recognition (ASR) systems since it can unify the acoustic and language model into a single neural network. However, as a drawback, training of end-to-end speech recognizers always requires transcribed utterances. Since end-to-end models are also known to be severely data hungry, this constraint is crucial especially because obtaining transcribed utterances is costly and can possibly be impractical or impossible. This paper proposes a method for alleviating this issue by transferring knowledge from a language model neural network that can be pretrained with text-only data. Specifically, this paper attempts to transfer semantic knowledge acquired in embedding vectors of large-scale language models. Since embedding vectors can be assumed as implicit representations of linguistic information such as part-of-speech, intent, and so on, those are also expected to be useful modeling cues for ASR decoders. This paper extends two types of ASR decoders, attention-based decoders and neural transducers, by modifying training loss functions to include embedding prediction terms. The proposed systems were shown to be effective for error rate reduction without incurring extra computational costs in the decoding phase.
Yotaro Kubo, Shigeki Karita, Michiel Bacchiani
ICASSP2
2022 SNRi Target Training for Joint Speech Enhancement and Recognition
Yuma Koizumi, Shigeki Karita, Arun Narayanan, Sankaran Panchapagesan, Michiel Bacchiani
INTERSPEECH2
2021 A Comparative Study on Neural Architectures and Training Methods for Japanese Speech Recognition
abstract
End-to-end (E2E) modeling is advantageous for automatic speech recognition (ASR) especially for Japanese since word-based tokenization of Japanese is not trivial, and E2E modeling is able to model character sequences directly. This paper focuses on the latest E2E modeling techniques, and investigates their performances on character-based Japanese ASR by conducting comparative experiments. The results are analyzed and discussed in order to understand the relative advantages of long short-term memory (LSTM), and Conformer models in combination with connectionist temporal classification, transducer, and attention-based loss functions. Furthermore, the paper investigates on effectivity of the recent training techniques such as data augmentation (SpecAugment), variational noise injection, and exponential moving average. The best configuration found in the paper achieved the state-of-the-art character error rates of 4.1%, 3.2%, and 3.5% for Corpus of Spontaneous Japanese (CSJ) eval1, eval2, and eval3 tasks, respectively. The system is also shown to be computationally efficient thanks to the efficiency of Conformer transducers.
Shigeki Karita, Yotaro Kubo, Michiel Bacchiani, Llion Jones
Interspeech1
2021 Unsupervised Learning of Disentangled Speech Content and Style Representation
abstract
We present an approach for unsupervised learning of speech representation disentangling contents and styles. Our model consists of: (1) a local encoder that captures per-frame information; (2) a global encoder that captures per-utterance information; and (3) a conditional decoder that reconstructs speech given local and global latent variables. Our experiments show that (1) the local latent variables encode speech contents, as reconstructed speech can be recognized by ASR with low word error rates (WER), even with a different global encoding; (2) the global latent variables encode speaker style, as reconstructed speech shares speaker identity with the source utterance of the global encoding. Additionally, we demonstrate an useful application from our pre-trained model, where we can train a speaker recognition model from the global latent variables and achieve high accuracy by fine-tuning with as few data as one label per speaker.
Andros Tjandra, Ruoming Pang, Yu Zhang 0033, Shigeki Karita
Interspeech4
2020 Self-Distillation for Improving CTC-Transformer-Based ASR Systems
Takafumi Moriya, Tsubasa Ochiai, Shigeki Karita, Hiroshi Sato 0002, Tomohiro Tanaka, Takanori Ashihara, Ryo Masumura, Yusuke Shinohara, Marc Delcroix
INTERSPEECH3
2019 A Comparative Study on Transformer vs RNN in Speech Applications
abstract
Sequence-to-sequence models have been widely used in end-to-end speech processing, for example, automatic speech recognition (ASR), speech translation (ST), and text-to-speech (TTS). This paper focuses on an emergent sequence-to-sequence model called Transformer, which achieves state-of-the-art performance in neural machine translation and other natural language processing applications. We undertook intensive studies in which we experimentally compared and analyzed Transformer and conventional recurrent neural networks (RNN) in a total of 15 ASR, one multilingual ASR, one ST, and two TTS benchmarks. Our experiments revealed various training tips and significant performance benefits obtained with Transformer for each task including the surprising superiority of Transformer in 13/15 ASR benchmarks in comparison with RNN. We are preparing to release Kaldi-style reproducible recipes using open source and publicly available datasets for all the ASR, ST, and TTS tasks for the community to succeed our exciting outcomes.
Shigeki Karita, Xiaofei Wang 0007, Shinji Watanabe 0001, Takenori Yoshimura, Wangyou Zhang, Nanxin Chen, Tomoki Hayashi, Takaaki Hori, Hirofumi Inaguma, Ziyan Jiang, Masao Someki, Nelson Enrique Yalta Soplin, Ryuichi Yamamoto
ASRU1
2019 Semi-supervised End-to-end Speech Recognition Using Text-to-speech and Autoencoders
abstract
We introduce speech and text autoencoders that share encoders and decoders with an automatic speech recognition (ASR) model to improve ASR performance with large speech only and text only training datasets. To build the speech and text autoencoders, we leverage state-of-the-art ASR and text-to-speech (TTS) encoder decoder architectures. These autoencoders learn features from speech only and text only datasets by switching the encoders and decoders used in the ASR and TTS models. Simultaneously, they aim to encode features to be compatible with ASR and TTS models by a multi-task loss. Additionally, we anticipate that TTS joint training can also improve the ASR performance because both ASR and TTS models learn transformations between speech and text. The experimental result we obtained with our semi-supervised end-to-end ASR/TTS training revealed reductions from a model initially trained with a small paired subset of the LibriSpeech corpus in the character error rate from 10.4% to 8.4% and word error rate from 20.6% to 18.0% by retraining the model with a large unpaired subset of the corpus.
Shigeki Karita, Shinji Watanabe 0001, Tomoharu Iwata, Marc Delcroix, Atsunori Ogawa, Tomohiro Nakatani
ICASSP1
2019 End-to-End SpeakerBeam for Single Channel Target Speech Recognition
Marc Delcroix, Shinji Watanabe 0001, Tsubasa Ochiai, Keisuke Kinoshita, Shigeki Karita, Atsunori Ogawa, Tomohiro Nakatani
INTERSPEECH5
2019 Improving Transformer-Based End-to-End Speech Recognition with Connectionist Temporal Classification and Language Model Integration
Shigeki Karita, Nelson Enrique Yalta Soplin, Shinji Watanabe 0001, Marc Delcroix, Atsunori Ogawa, Tomohiro Nakatani
INTERSPEECH1
2019 Improved Deep Duel Model for Rescoring N-Best Speech Recognition List Using Backward LSTMLM and Ensemble Encoders
Atsunori Ogawa, Marc Delcroix, Shigeki Karita, Tomohiro Nakatani
INTERSPEECH3
2018 Frame-by-Frame Closed-Form Update for Mask-Based Adaptive MVDR Beamforming
abstract
Beamforming approaches using time-frequency masks have recently been investigated and have shown promising results for noise robust automatic speech recognition (ASR) in many tasks. The time-frequency masks are estimated to compute the spatial statistics of target speech and noise signals, and then the statistics are used to derive a beamformer. Although its effectiveness has been clearly shown in batch and blockwise processing, it has not been well extended to frame-by-frame processing, which is a very important procedure for many actual applications. In this paper, we derive a frame-by-frame update rule for a mask-based minimum variance distortion-less response (MVDR) beamformer, which enables us to obtain enhanced signals without a long delay by combining it with uni-directional recurrent neural network-based mask estimation. Based on the Woodbury matrix identity, our algorithm achieves a closed-form solution of the mask-based MVDR beamformer at every time frame without any matrix inversion. Experimental results show that our frame-by-frame beamformer outperforms baseline block-wise beamforming on the CHiME-3 simulation dataset even with a shorter time delay.
Takuya Higuchi, Keisuke Kinoshita, Nobutaka Ito, Shigeki Karita, Tomohiro Nakatani
ICASSP4
2018 Sequence Training of Encoder-Decoder Model Using Policy Gradient for End-to-End Speech Recognition
abstract
The standard evaluation metric of automatic speech recognition (ASR) is the word error rate (WER), which measures the dissimilarity between recognized word sequences and their ground truth. Many training algorithms designed to reduce sequence-level errors such as WER have been proposed for hidden Markov model (HMM)-based ASR, e.g., state-level minimum Bayes risk (sMBR). However, these approaches cannot be used directly for encoder-decoder model based end-to-end ASR, because the encoder-decoder model employs very different mechanisms from HMM-based approaches. In this paper, we propose a new method for optimizing the encoder-decoder model based on a sequence-level evaluation metric. Since the WER is not directly differentiable, we adopt a policy gradient objective function to train the encoder-decoder model, which enables us to minimize the expected WER of the model predictions. This training method employs the scoring of multiple hypotheses as in the decoding stage while usual cross entropy training uses only the ground truth. Therefore, we can expect it to improve the decoding results of the encoder-decoder model. We perform experiments using the Tedlium corpus to demonstrate the potential of our proposed method for improving the recognition performance of the encoder-decoder model.
Shigeki Karita, Atsunori Ogawa, Marc Delcroix, Tomohiro Nakatani
ICASSP1
2018 Rescoring N-Best Speech Recognition List Based on One-on-One Hypothesis Comparison Using Encoder-Classifier Model
abstract
This paper proposes a new model for accurately rescoring (reranking) N-best speech recognition hypothesis lists. The model is based on state-of-the-art neural networks (NNs) and provides the minimum necessary functionality to perform N-best rescoring, i.e. one-on-one hypothesis comparison on a given N-best list in terms of word error rates (WERs). The model is composed of a long short-term memory (LSTM)-based encoder network followed by a fully-connected feedforward NN-based binary-class classifier network. Given the feature vector sequences of two hypotheses to be compared, this encoder-classifier (EC) model encodes these features and outputs binary-class probabilities that indicate which hypothesis has the lower WER. Then, depending on the output, the ranks of these hypotheses can be swapped. By repeating this one-on-one hypothesis comparison, a reranked N-best list can be obtained. In N-best rescoring experiments using a large scale speech corpus, the proposed EC model steadily outperforms an LSTM-based language model (LSTMLM), which is a strong and widely-used competitor. In addition, by incorporating the LSTMLM scores as an additional feature vector dimension, the N-best rescoring performance of the EC model is further improved. The improved EC model achieves a 10% relative WER reduction from the LSTMLM baseline.
Atsunori Ogawa, Marc Delcroix, Shigeki Karita, Tomohiro Nakatani
ICASSP3
2018 Auxiliary Feature Based Adaptation of End-to-end ASR Systems
Marc Delcroix, Shinji Watanabe 0001, Atsunori Ogawa, Shigeki Karita, Tomohiro Nakatani
INTERSPEECH4
2018 Semi-Supervised End-to-End Speech Recognition
Shigeki Karita, Shinji Watanabe 0001, Tomoharu Iwata, Atsunori Ogawa, Marc Delcroix
INTERSPEECH1
2018 ESPnet: End-to-End Speech Processing Toolkit
abstract
This paper introduces a new open source platform for end-to-end speech processing named ESPnet. ESPnet mainly focuses on end-to-end automatic speech recognition (ASR), and adopts widely-used dynamic neural network toolkits, Chainer and PyTorch, as a main deep learning engine. ESPnet also follows the Kaldi ASR toolkit style for data processing, feature extraction/format, and recipes to provide a complete setup for speech recognition and other speech processing experiments. This paper explains a major architecture of this software platform, several important functionalities, which differentiate ESPnet from other open source ASR toolkits, and experimental results with major ASR benchmarks.
Shinji Watanabe 0001, Takaaki Hori, Shigeki Karita, Tomoki Hayashi, Jiro Nishitoba, Yuya Unno, Nelson Enrique Yalta Soplin, Jahn Heymann, Matthew Wiesner, Nanxin Chen, Adithya Renduchintala, Tsubasa Ochiai
INTERSPEECH3
2017 Forward-Backward Convolutional LSTM for Acoustic Modeling
Shigeki Karita, Atsunori Ogawa, Marc Delcroix, Tomohiro Nakatani
INTERSPEECH1
2017 Unfolded Deep Recurrent Convolutional Neural Network with Jump Ahead Connections for Acoustic Modeling
Dung T. Tran, Marc Delcroix, Shigeki Karita, Michael Hentschel, Atsunori Ogawa, Tomohiro Nakatani
INTERSPEECH3
2015 Far-field speech recognition using CNN-DNN-HMM with convolution in time
abstract
Recent studies in speech recognition have shown that the performance of convolutional neural networks (CNNs) is superior to that of fully connected deep neural networks (DNNs). In this paper, we explore the use of CNNs in far-field speech recognition for dealing with reverberation, which blurs spectral energies along the time axis. Unlike most previous CNN applications to speech recognition, we consider convolution in time to examine whether it provides an improved reverberation modelling capability. Experimental results show that a CNN coupled with a fully connected DNN can model short time correlations in feature vectors with fewer parameters than a DNN and thus generalise better to unseen test environments. Combining this approach with signal-space dereverberation, which copes with long-term correlations, is shown to result in further improvement, where the gains from both approaches are almost additive. An initial investigation of the use of restricted convolution forms is also undertaken.
Takuya Yoshioka, Shigeki Karita, Tomohiro Nakatani
ICASSP2