EDBT 2026 Demo / reviewers in the wild / expert
Takaaki Hori
dblp:46/3941
· DBLP profile ↗
121ranked-venue papers
26as first author
15since 2021 · last 2025
0000-0003-4560-8039ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 105 · 21 first-author · 14 since 2021Artificial intelligence and machine learning · 57 · 15 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Delayed Fusion: Integrating Large Language Models into First-Pass Decoding in End-to-end Speech RecognitionabstractThis paper presents an efficient decoding approach for end-to-end automatic speech recognition (E2E-ASR) with large language models (LLMs). Although shallow fusion is the most common approach to incorporate language models into E2E-ASR decoding, we face two practical problems with LLMs. (1) LLM inference is computationally costly. (2) There may be a vocabulary mismatch between the ASR model and the LLM. To resolve this mismatch, we need to retrain the ASR model and/or the LLM, which is at best time-consuming and in many cases not feasible. We propose delayed fusion, which applies LLM scores to ASR hypotheses with a delay during decoding and enables easier use of pre-trained LLMs in ASR tasks. This method can reduce not only the number of hypotheses scored by the LLM but also the number of LLM inference calls. It also allows re-tokenizion of ASR hypotheses during decoding if ASR and LLM employ different tokenizations. We demonstrate that delayed fusion provides improved decoding speed and accuracy compared to shallow fusion and N-best rescoring using the LibriHeavy ASR corpus and three public LLMs, OpenLLaMA 3B & 7B and Mistral 7B. Takaaki Hori, Martin Kocour, Adnan Haider, Erik McDermott, Xiaodan Zhuang |
ICASSP | 1 |
| 2024 | End-to-End Speech Recognition: A SurveyabstractIn the last decade of automatic speech recognition (ASR) research, the introduction of deep learning has brought considerable reductions in word error rate of more than 50% relative, compared to modeling without deep learning. In the wake of this transition, a number of all-neural ASR architectures have been introduced. These so-calledend-to-end(E2E) models provide highly integrated, completely neural ASR models, which rely strongly on general machine learning knowledge, learn more consistently from data, with lower dependence on ASR domain-specific experience. The success and enthusiastic adoption of deep learning, accompanied by more generic model architectures has led to E2E models now becoming the prominent ASR approach. The goal of this survey is to provide a taxonomy of E2E ASR models and corresponding improvements, and to discuss their properties and their relationship to classical hidden Markov model (HMM) based ASR architectures. All relevant aspects of E2E ASR are covered in this work: modeling, training, decoding, and external language model integration, discussions of performance and deployment opportunities, as well as an outlook into potential future developments. Rohit Prabhavalkar, Takaaki Hori, Tara N. Sainath, Ralf Schlüter, Shinji Watanabe 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | Variable Attention Masking for Configurable Transformer Transducer Speech RecognitionabstractThis work studies the use of attention masking in transformer transducer based speech recognition for building a single configurable model for different deployment scenarios. We present a comprehensive set of experiments comparing fixed masking, where the same attention mask is applied at every frame, with chunked masking, where the attention mask for each frame is determined by chunk boundaries, in terms of recognition accuracy and latency. We then explore the use of variable masking, where the attention masks are sampled from a target distribution at training time, to build models that can work in different configurations. Finally, we investigate how a single configurable model can be used to perform both first pass streaming recognition and second pass acoustic rescoring. Experiments show that chunked masking achieves a better accuracy vs latency trade-off compared to fixed masking, both with and without FastEmit. We also show that variable masking improves the accuracy by up to 8% relative in the acoustic re-scoring scenario. Pawel Swietojanski, Dogan Can, Thiago Fraga da Silva, Arnab Ghoshal, Takaaki Hori, Roger Hsiao, Henry Mason, Erik McDermott, Honza Silovsky, Ruchir Travadi, Xiaodan Zhuang |
ICASSP | 6 |
| 2022 | Extended Graph Temporal Classification for Multi-Speaker End-to-End ASRabstractGraph-based temporal classification (GTC), a generalized form of the connectionist temporal classification loss, was recently proposed to improve automatic speech recognition (ASR) systems using graph-based supervision. For example, GTC was first used to encode an N-best list of pseudo-label sequences into a graph for semi-supervised learning. In this paper, we propose an extension of GTC to model the posteriors of both labels and label transitions by a neural network, which can be applied to a wider range of tasks. As an example application, we use the extended GTC (GTC-e) for the multi-speaker speech recognition task. The transcriptions and speaker information of multi-speaker speech are represented by a graph, where the speaker information is associated with the transitions and ASR outputs with the nodes. Using GTC-e, multi-speaker ASR modelling becomes very similar to single-speaker ASR modeling, in that tokens by multiple speakers are recognized as a single merged sequence in chronological order. For evaluation, we perform experiments on a simulated multi-speaker speech dataset derived from LibriSpeech, obtaining promising results with performance close to classical benchmarks for the task. Xuankai Chang, Niko Moritz, Takaaki Hori, Shinji Watanabe 0001, Jonathan Le Roux |
ICASSP | 3 |
| 2022 | Advancing Momentum Pseudo-Labeling with Conformer and Initialization StrategyabstractPseudo-labeling (PL), a semi-supervised learning (SSL) method where a seed model performs self-training using pseudo-labels generated from untranscribed speech, has been shown to enhance the performance of end-to-end automatic speech recognition (ASR). Our prior work proposed momentum pseudo-labeling (MPL), which performs PL-based SSL via an interaction between online and offline models, inspired by the mean teacher framework. MPL achieves remarkable results on various semi-supervised settings, showing robustness to variations in the amount of data and domain mismatch severity. However, there is further room for improving the seed model used to initialize the MPL training, as it is in general critical for a PL-based method to start training from high-quality pseudo-labels. To this end, we propose to enhance MPL by (1) introducing the Conformer architecture to boost the overall recognition accuracy and (2) exploiting iterative pseudo-labeling with a language model to improve the seed model before applying MPL. The experimental results demonstrate that the proposed approaches effectively improve MPL performance, outperforming other PL-based methods. We also present in-depth investigations to make our improvements effective, e.g., with regard to batch normalization typically used in Conformer and LM quality. Yosuke Higuchi, Niko Moritz, Jonathan Le Roux, Takaaki Hori |
ICASSP | 4 |
| 2022 | Sequence Transduction with Graph-Based SupervisionabstractThe recurrent neural network transducer (RNN-T) objective plays a major role in building today’s best automatic speech recognition (ASR) systems for production. Similarly to the connectionist temporal classification (CTC) objective, the RNN-T loss uses specific rules that define how a set of alignments is generated to form a lattice for the full-sum training. However, it is yet largely unknown if these rules are optimal and do lead to the best possible ASR results. In this work, we present a new transducer objective function that generalizes the RNN-T loss to accept a graph representation of the labels, thus providing a flexible and efficient framework to manipulate training lattices, e.g., for studying different transition rules, implementing different transducer losses, or restricting alignments. We demonstrate that transducer-based ASR with CTC-like lattice achieves better results compared to standard RNN-T, while also ensuring a strictly monotonic alignment, which will allow better optimization of the decoding procedure. For example, the proposed CTC-like transducer achieves an improvement of 4.8% on the test-other condition of LibriSpeech relative to an equivalent RNN-T based system. Niko Moritz, Takaaki Hori, Shinji Watanabe 0001, Jonathan Le Roux |
ICASSP | 2 |
| 2022 | Audio-Visual Scene-Aware Dialog and Reasoning Using Audio-Visual Transformers with Joint Student-Teacher LearningabstractIn previous work, we have proposed the Audio-Visual Scene-Aware Dialog (AVSD) task, collected an AVSD dataset, developed AVSD technologies, and hosted an AVSD challenge track at both the 7th and 8th Dialog System Technology Challenges (DSTC7, DSTC8). In these challenges, the best-performing systems relied heavily on human-generated descriptions of the video content, which were available in the datasets but would be unavailable in real-world applications. To promote further advancements for real-world applications, we proposed a third AVSD challenge, at DSTC10, with two modifications: 1) the human-created description is unavailable at inference time, and 2) systems must demonstrate temporal reasoning by finding evidence from the video to support each answer. This paper introduces the new task that includes temporal reasoning and our new extension of the AVSD dataset for DSTC10, for which we collected human-generated temporal reasoning data. We also introduce a baseline system built using an AV-transformer, which we released along with the new dataset. Finally, this paper introduces a new system that extends our baseline system with attentional multimodal fusion, joint student-teacher learning (JSTL), and model combination techniques, achieving state-of-the-art performances on the AVSD datasets for DSTC7, DSTC8, and DSTC10. We also propose two temporal reasoning methods for AVSD: one attention-based, and one based on a time-domain region proposal network. Ankit P. Shah, Shijie Geng, Peng Gao 0007, Anoop Cherian, Takaaki Hori, Tim K. Marks, Jonathan Le Roux, Chiori Hori |
ICASSP | 5 |
| 2022 | Low-Latency Online Streaming VideoQA Using Audio-Visual Transformers
Chiori Hori, Takaaki Hori, Jonathan Le Roux |
INTERSPEECH | 2 |
| 2021 | Unsupervised Domain Adaptation for Speech Recognition via Uncertainty Driven Self-TrainingabstractThe performance of automatic speech recognition (ASR) systems typically degrades significantly when the training and test data domains are mismatched. In this paper, we show that self-training (ST) combined with an uncertainty-based pseudo-label filtering approach can be effectively used for domain adaptation. We propose DUST, a dropout-based uncertainty-driven self-training technique which uses agreement between multiple predictions of an ASR system obtained for different dropout settings to measure the model’s uncertainty about its prediction. DUST excludes pseudo-labeled data with high uncertainties from the training, which leads to substantially improved ASR results compared to ST without filtering, and accelerates the training time due to a reduced training data set. Domain adaptation experiments using WSJ as a source domain and TED-LIUM 3 as well as SWITCHBOARD as the target domains show that up to 80% of the performance of a system trained on ground-truth data can be recovered. Sameer Khurana, Niko Moritz, Takaaki Hori, Jonathan Le Roux |
ICASSP | 3 |
| 2021 | Capturing Multi-Resolution Context by Dilated Self-AttentionabstractSelf-attention has become an important and widely used neural network component that helped to establish new state-of-the-art results for various applications, such as machine translation and automatic speech recognition (ASR). However, the computational complexity of self-attention grows quadratically with the input sequence length. This can be particularly problematic for applications such as ASR, where an input sequence generated from an utterance can be relatively long. In this work, we propose a combination of restricted self-attention and a dilation mechanism, which we refer to as dilated self-attention. The restricted self-attention allows attention to neighboring frames of the query at a high resolution, and the dilation mechanism summarizes distant information to allow attending to it with a lower resolution. Different methods for summarizing distant frames are studied, such as subsampling, mean-pooling, and attention-based pooling. ASR results demonstrate substantial improvements compared to restricted self-attention alone, achieving similar results compared to full-sequence based self-attention with a fraction of the computational costs. Niko Moritz, Takaaki Hori, Jonathan Le Roux |
ICASSP | 2 |
| 2021 | Semi-Supervised Speech Recognition Via Graph-Based Temporal ClassificationabstractSemi-supervised learning has demonstrated promising results in automatic speech recognition (ASR) by self-training using a seed ASR model with pseudo-labels generated for unlabeled data. The effectiveness of this approach largely relies on the pseudo-label accuracy, for which typically only the 1-best ASR hypothesis is used. However, alternative ASR hypotheses of an N-best list can provide more accurate labels for an unlabeled speech utterance and also reflect uncertainties of the seed ASR model. In this paper, we propose a generalized form of the connectionist temporal classification (CTC) objective that accepts a graph representation of the training labels. The newly proposed graph-based temporal classification (GTC) objective is applied for self-training with WFST-based supervision, which is generated from an N-best list of pseudo-labels. In this setup, GTC is used to learn not only a temporal alignment, similarly to CTC, but also a label alignment to obtain the optimal pseudo-label sequence from the weighted graph. Results show that this approach can effectively exploit an N-best list of pseudo-labels with associated scores, considerably outperforming standard pseudo-labeling, with ASR results approaching an oracle experiment in which the best hypotheses of the N-best lists are selected manually. Niko Moritz, Takaaki Hori, Jonathan Le Roux |
ICASSP | 2 |
| 2021 | Momentum Pseudo-Labeling for Semi-Supervised Speech RecognitionabstractPseudo-labeling (PL) has been shown to be effective in semisupervised automatic speech recognition (ASR), where a base model is self-trained with pseudo-labels generated from unlabeled data.While PL can be further improved by iteratively updating pseudo-labels as the model evolves, most of the previous approaches involve inefficient retraining of the model or intricate control of the label update.We present momentum pseudo-labeling (MPL), a simple yet effective strategy for semisupervised ASR.MPL consists of a pair of online and offline models that interact and learn from each other, inspired by the mean teacher method.The online model is trained to predict pseudo-labels generated on the fly by the offline model.The offline model maintains a momentum-based moving average of the online model.MPL is performed in a single training process and the interaction between the two models effectively helps them reinforce each other to improve the ASR performance.We apply MPL to an end-to-end ASR model based on the connectionist temporal classification.The experimental results demonstrate that MPL effectively improves over the base model and is scalable to different semi-supervised scenarios with varying amounts of data or domain mismatch. Yosuke Higuchi, Niko Moritz, Jonathan Le Roux, Takaaki Hori |
Interspeech | 4 |
| 2021 | Optimizing Latency for Online Video Captioning Using Audio-Visual TransformersabstractVideo captioning is an essential technology to understand scenes and describe events in natural language.To apply it to real-time monitoring, a system needs not only to describe events accurately but also to produce the captions as soon as possible.Low-latency captioning is needed to realize such functionality, but this research area for online video captioning has not been pursued yet.This paper proposes a novel approach to optimize each caption's output timing based on a trade-off between latency and caption quality.An audio-visual Transformer is trained to generate ground-truth captions using only a small portion of all video frames, and to mimic outputs of a pre-trained Transformer to which all the frames are given.A CNN-based timing detector is also trained to detect a proper output timing, where the captions generated by the two Transformers become sufficiently close to each other.With the jointly trained Transformer and timing detector, a caption can be generated in the early stages of an event-triggered video clip, as soon as an event happens or when it can be forecasted.Experiments with the ActivityNet Captions dataset show that our approach achieves 94% of the caption quality of the upper bound given by the pre-trained Transformer using the entire video clips, using only 28% of frames from the beginning. Chiori Hori, Takaaki Hori, Jonathan Le Roux |
Interspeech | 2 |
| 2021 | Advanced Long-Context End-to-End Speech Recognition Using Context-Expanded TransformersabstractThis paper addresses end-to-end automatic speech recognition (ASR) for long audio recordings such as lecture and conversational speeches.Most end-to-end ASR models are designed to recognize independent utterances, but contextual information (e.g., speaker or topic) over multiple utterances is known to be useful for ASR.In our prior work, we proposed a contextexpanded Transformer that accepts multiple consecutive utterances at the same time and predicts an output sequence for the last utterance, achieving 5-15% relative error reduction from utterance-based baselines in lecture and conversational ASR benchmarks.Although the results have shown remarkable performance gain, there is still potential to further improve the model architecture and the decoding process.In this paper, we extend our prior work by (1) introducing the Conformer architecture to further improve the accuracy, (2) accelerating the decoding process with a novel activation recycling technique, and (3) enabling streaming decoding with triggered attention.We demonstrate that the extended Transformer provides state-of-the-art end-to-end ASR performance, obtaining a 17.3% character error rate for the HKUST dataset and 12.0%/6.3%word error rates for the Switchboard-300 Eval2000 CallHome/Switchboard test sets.The new decoding method reduces decoding time by more than 50% and further enables streaming ASR with limited accuracy degradation. Takaaki Hori, Niko Moritz, Chiori Hori, Jonathan Le Roux |
Interspeech | 1 |
| 2021 | Dual Causal/Non-Causal Self-Attention for Streaming End-to-End Speech RecognitionabstractAttention-based end-to-end automatic speech recognition (ASR) systems have recently demonstrated state-of-the-art results for numerous tasks. However, the application of self-attention and attention-based encoder-decoder models remains challenging for streaming ASR, where each word must be recognized shortly after it was spoken. In this work, we present the dual causal/non-causal self-attention (DCN) architecture, which in contrast to restricted self-attention prevents the overall context to grow beyond the look-ahead of a single layer when used in a deep architecture. DCN is compared to chunk-based and restricted self-attention using streaming transformer and conformer architectures, showing improved ASR performance over restricted self-attention and competitive ASR results compared to chunk-based self-attention, while providing the advantage of frame-synchronous processing. Combined with triggered attention, the proposed streaming end-to-end ASR systems obtained state-of-the-art results on the LibriSpeech, HKUST, and Switchboard ASR tasks. Niko Moritz, Takaaki Hori, Jonathan Le Roux |
Interspeech | 2 |
| 2020 | Streaming Automatic Speech Recognition with the Transformer ModelabstractEncoder-decoder based sequence-to-sequence models have demonstrated state-of-the-art results in end-to-end automatic speech recognition (ASR). Recently, the transformer architecture, which uses self-attention to model temporal context information, has been shown to achieve significantly lower word error rates (WERs) compared to recurrent neural network (RNN) based system architectures. Despite its success, the practical usage is limited to offline ASR tasks, since encoder-decoder architectures typically require an entire speech utterance as input. In this work, we propose a transformer based end-to-end ASR system for streaming ASR, where an output must be generated shortly after each spoken word. To achieve this, we apply time-restricted self-attention for the encoder and triggered attention for the encoder-decoder attention mechanism. Our proposed streaming transformer architecture achieves 2.8% and 7.3% WER for the "clean" and "other" test data of LibriSpeech, which to our knowledge is the best published streaming end-to-end ASR result for this task. Niko Moritz, Takaaki Hori, Jonathan Le Roux |
ICASSP | 2 |
| 2020 | Unsupervised Speaker Adaptation Using Attention-Based Speaker Memory for End-to-End ASRabstractWe propose an unsupervised speaker adaptation method inspired by the neural Turing machine for end-to-end (E2E) automatic speech recognition (ASR). The proposed model contains a memory block that holds speaker i-vectors extracted from the training data and reads relevant i-vectors from the memory through an attention mechanism. The resulting memory vector (M-vector) is concatenated to the acoustic features or to the hidden layer activations of an E2E neural network model. The E2E ASR system is based on the joint connectionist temporal classification and attention-based encoder-decoder architecture. M-vector and i-vector results are compared for inserting them at different layers of the encoder neural network using the WSJ and TED-LIUM2 ASR benchmarks. We show that M-vectors, which do not require an auxiliary speaker embedding extraction system at test time, achieve similar word error rates (WERs) compared to i-vectors for single speaker utterances and significantly lower WERs for utterances in which there are speaker changes. Leda Sari, Niko Moritz, Takaaki Hori, Jonathan Le Roux |
ICASSP | 3 |
| 2020 | Transformer-Based Long-Context End-to-End Speech Recognition
Takaaki Hori, Niko Moritz, Chiori Hori, Jonathan Le Roux |
INTERSPEECH | 1 |
| 2020 | All-in-One Transformer: Unifying Speech Recognition, Audio Tagging, and Event Detection
Niko Moritz, Gordon Wichern, Takaaki Hori, Jonathan Le Roux |
INTERSPEECH | 3 |
| 2020 | Multi-Stream End-to-End Speech RecognitionabstractAttention-based methods and Connectionist Temporal Classification (CTC) network have been promising research directions for end-to-end (E2E) Automatic Speech Recognition (ASR). The joint CTC/Attention model has achieved great success by utilizing both architectures during multi-task training and joint decoding. In this article, we present a multi-stream framework based on joint CTC/Attention E2E ASR with parallel streams represented by separate encoders aiming to capture diverse information. On top of the regular attention networks, the Hierarchical Attention Network (HAN) is introduced to steer the decoder toward the most informative encoders. A separate CTC network is assigned to each stream to force monotonic alignments. Two representative framework have been proposed and discussed, which are Multi-Encoder Multi-Resolution (MEM-Res) framework and Multi-Encoder Multi-Array (MEM-Array) framework, respectively. In MEM-Res framework, two heterogeneous encoders with different architectures, temporal resolutions and separate CTC networks work in parallel to extract complementary information from same acoustics. Experiments are conducted on Wall Street Journal (WSJ) and CHiME-4, resulting in relative Word Error Rate (WER) reduction of 18.0-32.1% and the best WER of 3.6% in the WSJ eval92 test set. The MEM-Array framework aims at improving the far-field ASR robustness using multiple microphone arrays which are activated by separate encoders. Compared with the best single-array results, the proposed framework has achieved relative WER reduction of 3.7% and 9.7% in AMI and DIRHA multi-array corpora, respectively, which also outperforms conventional fusion strategies. Xiaofei Wang 0007, Sri Harish Reddy Mallidi, Shinji Watanabe 0001, Takaaki Hori, Hynek Hermansky |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2019 | A Comparative Study on Transformer vs RNN in Speech ApplicationsabstractSequence-to-sequence models have been widely used in end-to-end speech processing, for example, automatic speech recognition (ASR), speech translation (ST), and text-to-speech (TTS). This paper focuses on an emergent sequence-to-sequence model called Transformer, which achieves state-of-the-art performance in neural machine translation and other natural language processing applications. We undertook intensive studies in which we experimentally compared and analyzed Transformer and conventional recurrent neural networks (RNN) in a total of 15 ASR, one multilingual ASR, one ST, and two TTS benchmarks. Our experiments revealed various training tips and significant performance benefits obtained with Transformer for each task including the surprising superiority of Transformer in 13/15 ASR benchmarks in comparison with RNN. We are preparing to release Kaldi-style reproducible recipes using open source and publicly available datasets for all the ASR, ST, and TTS tasks for the community to succeed our exciting outcomes. Shigeki Karita, Xiaofei Wang 0007, Shinji Watanabe 0001, Takenori Yoshimura, Wangyou Zhang, Nanxin Chen, Tomoki Hayashi, Takaaki Hori, Hirofumi Inaguma, Ziyan Jiang, Masao Someki, Nelson Enrique Yalta Soplin, Ryuichi Yamamoto |
ASRU | 8 |
| 2019 | Streaming End-to-End Speech Recognition with Joint CTC-Attention Based ModelsabstractIn this paper, we present a one-pass decoding algorithm for streaming recognition with joint connectionist temporal classification (CTC) and attention-based end-to-end automatic speech recognition (ASR) models. The decoding scheme is based on a frame-synchronous CTC prefix beam search algorithm and the recently proposed triggered attention concept. To achieve a fully streaming end-to-end ASR system, the CTC-triggered attention decoder is combined with a unidirectional encoder neural network based on parallel time-delayed long short-term memory (PTDLSTM) streams, which has demonstrated superior performance compared to various other streaming encoder architectures in earlier work. A new type of pre-training method is studied to further improve our streaming ASR models by adding residual connections to the encoder neural network and layer-wise removing them during the training process. The proposed joint CTC-triggered attention decoding algorithm, which enables streaming recognition of attention-based ASR systems, achieves similar ASR results compared to offline CTC-attention decoding and significantly better results compared to CTC prefix beam search decoding alone. Niko Moritz, Takaaki Hori, Jonathan Le Roux |
ASRU | 2 |
| 2019 | Promising Accurate Prefix Boosting for Sequence-to-sequence ASRabstractIn this paper, we present promising accurate prefix boosting (PAPB), a discriminative training technique for attention based sequence-to-sequence (seq2seq) ASR. PAPB is devised to unify the training and testing scheme effectively. The training procedure involves maximizing the score of each partial correct sequence obtained during beam search compared to other hypotheses. The training objective also includes minimization of token (character) error rate. PAPB shows its efficacy by achieving 10.8% and 3.8% WER with and without external RNNLM respectively on Wall Street Journal dataset. Murali Karthick Baskar, Lukás Burget, Shinji Watanabe 0001, Martin Karafiát, Takaaki Hori, Jan Cernocký |
ICASSP | 5 |
| 2019 | Language Model Integration Based on Memory Control for Sequence to Sequence Speech RecognitionabstractIn this paper, we explore several new schemes to train a seq2seq model to integrate a pre-trained language model (LM). Our proposed fusion methods focus on the memory cell state and the hidden state in the seq2seq decoder long short-term memory (LSTM), and the memory cell state is updated by the LM unlike the prior studies. This means the memory retained by the main seq2seq would be adjusted by the external LM. These fusion methods have several variants depending on the architecture of this memory cell update and the use of memory cell and hidden states which directly affects the final label inference. We performed the experiments to show the effectiveness of the proposed methods in a mono-lingual ASR setup on the Librispeech corpus and in a transfer learning setup from a multilingual ASR (MLASR) base model to a low-resourced language. In Librispeech, our best model improved WER by 3.7%, 2.4% for test clean, test other relatively to the shallow fusion baseline, with multilevel decoding. In transfer learning from an MLASR base model to the IARPA Babel Swahili model, the best scheme improved the transferred model on eval set by 9.9%, 9.8% in CER, WER relatively to the 2-stage transfer baseline. Shinji Watanabe 0001, Takaaki Hori, Murali Karthick Baskar, Hirofumi Inaguma, Jesús Villalba 0001, Najim Dehak |
ICASSP | 3 |
| 2019 | End-to-end Audio Visual Scene-aware Dialog Using Multimodal Attention-based Video FeaturesabstractIn order for machines interacting with the real world to have conversations with users about the objects and events around them, they need to understand dynamic audiovisual scenes. The recent revolution of neural network models allows us to combine various modules into a single end-to-end differentiable network. As a result, Audio Visual Scene-Aware Dialog (AVSD) systems for real-world applications can be developed by integrating state-of-the-art technologies from multiple research areas, including end-to-end dialog technologies, visual question answering (VQA) technologies, and video description technologies. In this paper, we introduce a new data set of dialogs about videos of human behaviors, as well as an end-to-end Audio Visual Scene-Aware Dialog (AVSD) model, trained using this new data set, that generates responses in a dialog about a video. By using features that were developed for multimodal attention-based video description, our system improves the quality of generated dialog about dynamic video scenes. Chiori Hori, Huda AlAmri, Jue Wang 0010, Gordon Wichern, Takaaki Hori, Anoop Cherian, Tim K. Marks, Vincent Cartillier, Raphael Gontijo Lopes, Abhishek Das 0002, Irfan A. Essa, Dhruv Batra, Devi Parikh |
ICASSP | 5 |
| 2019 | Cycle-consistency Training for End-to-end Speech RecognitionabstractThis paper presents a method to train end-to-end automatic speech recognition (ASR) models using unpaired data. Although the end-to-end approach can eliminate the need for expert knowledge such as pronunciation dictionaries to build ASR systems, it still requires a large amount of paired data, i.e., speech utterances and their transcriptions. Cycle-consistency losses have been recently proposed as a way to mitigate the problem of limited paired data. These approaches compose a reverse operation with a given transformation, e.g., text-to-speech (TTS) with ASR, to build a loss that only requires unsupervised data, speech in this example. Applying cycle consistency to ASR models is not trivial since fundamental information, such as speaker traits, are lost in the intermediate text bottleneck. To solve this problem, this work presents a loss that is based on the speech encoder state sequence instead of the raw speech signal. This is achieved by training a Text-To-Encoder model and defining a loss based on the encoder reconstruction error. Experimental results on the LibriSpeech corpus show that the proposed cycle-consistency training reduced the word error rate by 14.7% from an initial model trained with 100-hour paired data, using an additional 360 hours of audio data without transcriptions. We also investigate the use of text-only data mainly for language modeling to further improve the performance in the unpaired data training scenario. Takaaki Hori, Ramón Fernandez Astudillo, Tomoki Hayashi, Yu Zhang 0033, Shinji Watanabe 0001, Jonathan Le Roux |
ICASSP | 1 |
| 2019 | Triggered Attention for End-to-end Speech RecognitionabstractA new system architecture for end-to-end automatic speech recognition (ASR) is proposed that combines the alignment capabilities of the connectionist temporal classification (CTC) approach and the modeling strength of the attention mechanism. The proposed system architecture, named triggered attention (TA), uses a CTC-based classifier to control the activation of an attention-based decoder neural network. This allows for a frame-synchronous decoding scheme with an adjustable look-ahead parameter to control the induced delay and opens the door to streaming recognition with attention-based end-to-end ASR systems. We present ASR results of the TA model on three data sets of different size and language and compare the scores to a well-tuned attention-based end-to-end ASR baseline system, which consumes input frames in the traditional full-sequence manner. The proposed triggered attention (TA) decoder concept achieves similar or better ASR results in all experiments compared to the full-sequence attention model, while also limiting the decoding delay to two look-ahead frames, which in our setup corresponds to an output delay of 80 ms. Niko Moritz, Takaaki Hori, Jonathan Le Roux |
ICASSP | 2 |
| 2019 | Stream Attention-based Multi-array End-to-end Speech RecognitionabstractAutomatic Speech Recognition (ASR) using multiple microphone arrays has achieved great success in the far-field robustness. Taking advantage of all the information that each array shares and contributes is crucial in this task. Motivated by the advances of joint Connectionist Temporal Classification (CTC)/attention mechanism in the End-to-End (E2E) ASR, a stream attention-based multi-array framework is proposed in this work. Microphone arrays, acting as information streams, are activated by separate encoders and decoded under the instruction of both CTC and attention networks. In terms of attention, a hierarchical structure is adopted. On top of the regular attention networks, stream attention is introduced to steer the decoder toward the most informative encoders. Experiments have been conducted on AMI and DIRHA multi-array corpora using the encoder-decoder architecture. Compared with the best single-array results, the proposed framework has achieved relative Word Error Rates (WERs) reduction of 3.7% and 9.7% in the two datasets, respectively, which is better than conventional strategies as well. Xiaofei Wang 0007, Sri Harish Reddy Mallidi, Takaaki Hori, Shinji Watanabe 0001, Hynek Hermansky |
ICASSP | 4 |
| 2019 | Semi-Supervised Sequence-to-Sequence ASR Using Unpaired Speech and TextabstractSequence-to-sequence automatic speech recognition (ASR) models require large quantities of data to attain high performance. For this reason, there has been a recent surge in interest for unsupervised and semi-supervised training in such models. This work builds upon recent results showing notable improvements in semi-supervised training using cycle-consistency and related techniques. Such techniques derive training procedures and losses able to leverage unpaired speech and/or text data by combining ASR with Text-to-Speech (TTS) models. In particular, this work proposes a new semi-supervised loss combining an end-to-end differentiable ASR$\rightarrow$TTS loss with TTS$\rightarrow$ASR loss. The method is able to leverage both unpaired speech and text data to outperform recently proposed related techniques in terms of \%WER. We provide extensive results analyzing the impact of data quantity and speech and text modalities and show consistent gains across WSJ and Librispeech corpora. Our code is provided in ESPnet to reproduce the experiments. Murali Karthick Baskar, Shinji Watanabe 0001, Ramón Fernandez Astudillo, Takaaki Hori, Lukás Burget, Jan Cernocký |
INTERSPEECH | 4 |
| 2019 | Joint Student-Teacher Learning for Audio-Visual Scene-Aware Dialog
Chiori Hori, Anoop Cherian, Tim K. Marks, Takaaki Hori |
INTERSPEECH | 4 |
| 2019 | Analysis of Multilingual Sequence-to-Sequence Speech Recognition SystemsabstractThis paper investigates the applications of various multilingual approaches developed in conventional hidden Markov model (HMM) systems to sequence-to-sequence (seq2seq) automatic speech recognition (ASR). On a set composed of Babel data, we first show the effectiveness of multi-lingual training with stacked bottle-neck (SBN) features. Then we explore various architectures and training strategies of multi-lingual seq2seq models based on CTC-attention networks including combinations of output layer, CTC and/or attention component re-training. We also investigate the effectiveness of language-transfer learning in a very low resource scenario when the target language is not included in the original multi-lingual training data. Interestingly, we found multilingual features superior to multilingual models, and this finding suggests that we can efficiently combine the benefits of the HMM system with the seq2seq system through these multilingual feature techniques. Martin Karafiát, Murali Karthick Baskar, Shinji Watanabe 0001, Takaaki Hori, Matthew Wiesner, Jan Cernocký |
INTERSPEECH | 4 |
| 2019 | Unidirectional Neural Network Architectures for End-to-End Automatic Speech Recognition
Niko Moritz, Takaaki Hori, Jonathan Le Roux |
INTERSPEECH | 2 |
| 2019 | Vectorized Beam Search for CTC-Attention-Based Speech Recognition
Hiroshi Seki, Takaaki Hori, Shinji Watanabe 0001, Niko Moritz, Jonathan Le Roux |
INTERSPEECH | 2 |
| 2019 | End-to-End Multilingual Multi-Speaker Speech Recognition
Hiroshi Seki, Takaaki Hori, Shinji Watanabe 0001, Jonathan Le Roux, John R. Hershey |
INTERSPEECH | 2 |
| 2019 | Overview of the sixth dialog system technology challenge: DSTC6
Chiori Hori, Julien Perez, Ryuichiro Higashinaka, Takaaki Hori, Y-Lan Boureau, Michimasa Inaba, Yuiko Tsunomori, Tetsuro Takahashi, Koichiro Yoshino, Seokhwan Kim |
Comput. Speech Lang. | 4 |
| 2019 | Adversarial training and decoding strategies for end-to-end neural conversation models
Takaaki Hori, Yusuke Koji, Chiori Hori, Bret Harsham, John R. Hershey |
Comput. Speech Lang. | 1 |
| 2018 | A Purely End-to-End System for Multi-speaker Speech RecognitionabstractRecently, there has been growing interest in multi-speaker speech recognition, where the utterances of multiple speakers are recognized from their mixture.Promising techniques have been proposed for this task, but earlier works have required additional training data such as isolated source signals or senone alignments for effective learning.In this paper, we propose a new sequence-to-sequence framework to directly decode multiple label sequences from a single speech sequence by unifying source separation and speech recognition functions in an end-toend manner.We further propose a new objective function to improve the contrast between the hidden vectors to avoid generating similar hypotheses.Experimental results show that the model is directly able to learn a mapping from a speech mixture to multiple label sequences, achieving 83.1% relative improvement compared to a model trained without the proposed objective.Interestingly, the results are comparable to those produced by previous endto-end works featuring explicit separation and recognition modules. Hiroshi Seki, Takaaki Hori, Shinji Watanabe 0001, Jonathan Le Roux, John R. Hershey |
ACL (1) | 2 |
| 2018 | Speaker Adaptation for Multichannel End-to-End Speech RecognitionabstractRecent work on multichannel end-to-end automatic speech recognition (ASR) has shown that multichannel speech enhancement and speech recognition functions can be integrated into a deep neural network (DNN)-based system, and promising experimental results have been shown using the CHiME-4 and AMI corpora. In other recent DNN-based hidden Markov model (DNN-HMM) hybrid architectures, the effectiveness of speaker adaptation has been well established. Motivated by these results, we propose a multi-path adaptation scheme for end-to-end multichannel ASR, which combines the unprocessed noisy speech features with a speech-enhanced pathway to improve upon previous end-to-end ASR approaches. Experimental results using CHiME-4 show that (1) our proposed multi-path adaptation scheme improves ASR performance and (2) adapting the encoder network is more effective than adapting the neural beamformer, attention mechanism, or decoder network. Tsubasa Ochiai, Shinji Watanabe 0001, Shigeru Katagiri, Takaaki Hori, John R. Hershey |
ICASSP | 4 |
| 2018 | An End-to-End Language-Tracking Speech Recognizer for Mixed-Language SpeechabstractEnd-to-end automatic speech recognition (ASR) can significantly reduce the burden of developing ASR systems for new languages, by eliminating the need for linguistic information such as pronunciation dictionaries. This also creates an opportunity to build a monolithic multilingual ASR system with a language-independent neural network architecture. In our previous work, we proposed a monolithic neural network architecture that can recognize multiple languages, and showed its effectiveness compared with conventional language-dependent models. However, the model is not guaranteed to properly handle switches in language within an utterance, thus lacking the flexibility to recognize mixed-language speech such as code-switching. In this paper, we extend our model to enable dynamic tracking of the language within an utterance, and propose a training procedure that takes advantage of a newly created mixed-language speech corpus. Experimental results show that the extended model outperforms both language-dependent models and our previous model without suffering from performance degradation that could be associated with language switching. Hiroshi Seki, Shinji Watanabe 0001, Takaaki Hori, Jonathan Le Roux, John R. Hershey |
ICASSP | 3 |
| 2018 | End-to-End Multi-Speaker Speech RecognitionabstractCurrent advances in deep learning have resulted in a convergence of methods across a wide range of tasks, opening the door for tighter integration of modules that were previously developed and optimized in isolation. Recent ground-breaking works have produced end-to-end deep network methods for both speech separation and end-to-end automatic speech recognition (ASR). Speech separation methods such as deep clustering address the challenging cocktail-party problem of distinguishing multiple simultaneous speech signals. This is an enabling technology for real-world human machine interaction (HMI). However, speech separation requires ASR to interpret the speech for any HMI task. Likewise, ASR requires speech separation to work in an unconstrained environment. Although these two components can be trained in isolation and connected after the fact, this paradigm is likely to be sub-optimal, since it relies on artificially mixed data. In this paper, we develop the first fully end-to-end, jointly trained deep learning system for separation and recognition of overlapping speech signals. The joint training framework synergistically adapts the separation and recognition to each other. As an additional benefit, it enables training on more realistic data that contains only mixed signals and their transcriptions, and thus is suited to large scale training on existing transcribed data. Shane Settle, Jonathan Le Roux, Takaaki Hori, Shinji Watanabe 0001, John R. Hershey |
ICASSP | 3 |
| 2018 | ESPnet: End-to-End Speech Processing ToolkitabstractThis paper introduces a new open source platform for end-to-end speech processing named ESPnet. ESPnet mainly focuses on end-to-end automatic speech recognition (ASR), and adopts widely-used dynamic neural network toolkits, Chainer and PyTorch, as a main deep learning engine. ESPnet also follows the Kaldi ASR toolkit style for data processing, feature extraction/format, and recipes to provide a complete setup for speech recognition and other speech processing experiments. This paper explains a major architecture of this software platform, several important functionalities, which differentiate ESPnet from other open source ASR toolkits, and experimental results with major ASR benchmarks. Shinji Watanabe 0001, Takaaki Hori, Shigeki Karita, Tomoki Hayashi, Jiro Nishitoba, Yuya Unno, Nelson Enrique Yalta Soplin, Jahn Heymann, Matthew Wiesner, Nanxin Chen, Adithya Renduchintala, Tsubasa Ochiai |
INTERSPEECH | 2 |
| 2018 | Multilingual Sequence-to-Sequence Speech Recognition: Architecture, Transfer Learning, and Language ModelingabstractSequence-to-sequence (seq2seq) approach for low-resource ASR is a relatively new direction in speech research. The approach benefits by performing model training without using lexicon and alignments. However, this poses a new problem of requiring more data compared to conventional DNN-HMM systems. In this work, we attempt to use data from 10 BABEL languages to build a multilingual seq2seq model as a prior model, and then port them towards 4 other BABEL languages using transfer learning approach. We also explore different architectures for improving the prior multilingual seq2seq model. The paper also discusses the effect of integrating a recurrent neural network language model (RNNLM) with a seq2seq model during decoding. Experimental results show that the transfer learning approach from the multilingual model shows substantial gains over monolingual models across all 4 BABEL languages. Incorporating an RNNLM also brings significant improvements in terms of %WER, and achieves recognition performance comparable to the models trained with twice more training data. Murali Karthick Baskar, Matthew Wiesner, Sri Harish Reddy Mallidi, Nelson Enrique Yalta Soplin, Martin Karafiát, Shinji Watanabe 0001, Takaaki Hori |
SLT | 9 |
| 2018 | Back-Translation-Style Data Augmentation for end-to-end ASRabstractIn this paper we propose a novel data augmentation method for attention-based end-to-end automatic speech recognition (E2E-ASR), utilizing a large amount of text which is not paired with speech signals. Inspired by the back-translation technique proposed in the field of machine translation, we build a neural text-to-encoder model which predicts a sequence of hidden states extracted by a pre-trained E2E-ASR encoder from a sequence of characters. By using hidden states as a target instead of acoustic features, it is possible to achieve faster attention learning and reduce computational cost, thanks to sub-sampling in E2E-ASR encoder, also the use of the hidden states can avoid to model speaker dependencies unlike acoustic features. After training, the text-to-encoder model generates the hidden states from a large amount of unpaired text, then E2E-ASR decoder is retrained using the generated hidden states as additional training data. Experimental evaluation using LibriSpeech dataset demonstrates that our proposed method achieves improvement of ASR performance and reduces the number of unknown words without the need for paired data. Tomoki Hayashi, Shinji Watanabe 0001, Yu Zhang 0033, Tomoki Toda, Takaaki Hori, Ramón Fernandez Astudillo, Kazuya Takeda |
SLT | 5 |
| 2018 | End-to-end Speech Recognition With Word-Based Rnn Language ModelsabstractThis paper investigates the impact of word-based RNN language models (RNN-LMs) on the performance of end-to-end automatic speech recognition (ASR). In our prior work, we have proposed a multi-level LM, in which character-based and word-based RNN-LMs are combined in hybrid CTC/attention-based ASR. Although this multi-level approach achieves significant error reduction in the Wall Street Journal (WSJ) task, two different LMs need to be trained and used for decoding, which increase the computational cost and memory usage. In this paper, we further propose a novel word-based RNN-LM, which allows us to decode with only the word-based LM, where it provides look-ahead word probabilities to predict next characters instead of the character-based LM, leading competitive accuracy with less computation compared to the multi-level LM. We demonstrate the efficacy of the word-based RNN-LMs using a larger corpus, LibriSpeech, in addition to WSJ we used in the prior work. Furthermore, we show that the proposed model achieves 5.1 %WER for WSJ Eval'92 test set when the vocabulary size is increased, which is the best WER reported for end-to-end ASR systems on this benchmark. Takaaki Hori, Shinji Watanabe 0001 |
SLT | 1 |
| 2017 | Joint CTC/attention decoding for end-to-end speech recognitionabstractEnd-to-end automatic speech recognition (ASR) has become a popular alternative to conventional DNN/HMM systems because it avoids the need for linguistic resources such as pronunciation dictionary, tokenization, and contextdependency trees, leading to a greatly simplified model-building process.There are two major types of end-to-end architectures for ASR: attention-based methods use an attention mechanism to perform alignment between acoustic frames and recognized symbols, and connectionist temporal classification (CTC), uses Markov assumptions to efficiently solve sequential problems by dynamic programming.This paper proposes a joint decoding algorithm for end-to-end ASR with a hybrid CTC/attention architecture, which effectively utilizes both advantages in decoding.We have applied the proposed method to two ASR benchmarks (spontaneous Japanese and Mandarin Chinese), and showing the comparable performance to conventional state-of-the-art DNN/HMM ASR systems without linguistic resources. Takaaki Hori, Shinji Watanabe 0001, John R. Hershey |
ACL (1) | 1 |
| 2017 | Early and late integration of audio features for automatic video descriptionabstractThis paper presents our approach to improve video captioning by integrating audio and video features. Video captioning is the task of generating a textual description to describe the content of a video. State-of-the-art approaches to video captioning are based on sequence-to-sequence models, in which a single neural network accepts sequential images and audio data, and outputs a sequence of words that best describe the input data in natural language. The network thus learns to encode the video input into an intermediate semantic representation, which can be useful in applications such as multimedia indexing, automatic narration, and audio-visual question answering. In our prior work, we proposed an attention-based multi-modal fusion mechanism to integrate image, motion, and audio features, where the multiple features are integrated in the network. Here, we apply hypothesis-level integration based on minimum Bayes-risk (MBR) decoding to further improve the caption quality, focusing on well-known evaluation metrics (BLEU and METEOR scores). Experiments with the YouTube2Text and MSR-VTT datasets demonstrate that combinations of early and late integration of multimodal features significantly improve the audio-visual semantic representation, as measured by the resulting caption quality. In addition, we compared the performance of our method using two different types of audio features: MFCC features, and the audio features extracted using SoundNet, which was trained to recognize objects and scenes from videos using only the audio signals. Chiori Hori, Takaaki Hori, Tim K. Marks, John R. Hershey |
ASRU | 2 |
| 2017 | Multi-level language modeling and decoding for open vocabulary end-to-end speech recognitionabstractWe propose a combination of character-based and word-based language models in an end-to-end automatic speech recognition (ASR) architecture. In our prior work, we combined a character-based LSTM RNN-LM with a hybrid attention/connectionist temporal classification (CTC) architecture. The character LMs improved recognition accuracy to rival state-of-the-art DNN/HMM systems in Japanese and Mandarin Chinese tasks. Although a character-based architecture can provide for open vocabulary recognition, the character-based LMs generally under-perform relative to word LMs for languages such as English with a small alphabet, because of the difficulty of modeling Linguistic constraints across long sequences of characters. This paper presents a novel method for end-to-end ASR decoding with LMs at both the character and word level. Hypotheses are first scored with the character-based LM until a word boundary is encountered. Known words are then re-scored using the word-based LM, while the character-based LM provides for out-of-vocabulary scores. In a standard Wall Street Journal (WSJ) task, we achieved 5.6 % WER for the Eval'92 test set using only the SI284 training set and WSJ text data, which is the best score reported for end-to-end ASR systems on this benchmark. Takaaki Hori, Shinji Watanabe 0001, John R. Hershey |
ASRU | 1 |
| 2017 | Language independent end-to-end architecture for joint language identification and speech recognitionabstractEnd-to-end automatic speech recognition (ASR) can significantly reduce the burden of developing ASR systems for new languages, by eliminating the need for linguistic information such as pronunciation dictionaries. This also creates an opportunity, which we fully exploit in this paper, to build a monolithic multilingual ASR system with a language-independent neural network architecture. We present a model that can recognize speech in 10 different languages, by directly performing grapheme (character/chunked-character) based speech recognition. The model is based on our hybrid attention/connectionist temporal classification (CTC) architecture which has previously been shown to achieve the state-of-the-art performance in several ASR benchmarks. Here we augment its set of output symbols to include the union of character sets appearing in all the target languages. These include Roman and Cyrillic Alphabets, Arabic numbers, simplified Chinese, and Japanese Kanji/Hiragana/Katakana characters (5,500 characters in all). This allows training of a single multilingual model, whose parameters are shared across all the languages. The model can jointly identify the language and recognize the speech, automatically formatting the recognized text in the appropriate character set. The experiments, which used speech databases composed of Wall Street Journal (English), Corpus of Spontaneous Japanese, HKUST Mandarin CTS, and Voxforge (German, Spanish, French, Italian, Dutch, Portuguese, Russian), demonstrate comparable/superior performance relative to language-dependent end-to-end ASR systems. Shinji Watanabe 0001, Takaaki Hori, John R. Hershey |
ASRU | 2 |
| 2017 | BLSTM-HMM hybrid system combined with sound activity detection network for polyphonic Sound Event DetectionabstractThis paper presents a new hybrid approach for polyphonic Sound Event Detection (SED) which incorporates a temporal structure modeling technique based on a hidden Markov model (HMM) with a frame-by-frame detection method based on a bidirectional long short-term memory (BLSTM) recurrent neural network (RNN). The proposed BLSTM-HMM hybrid system makes it possible to model sound event-dependent temporal structures and also to perform sequence-by-sequence detection without having to resort to thresholding such as in the conventional frame-by-frame methods. Furthermore, to effectively reduce insertion errors of sound events, which often occurs under noisy conditions, we additionally implement a binary mask post-processing using a sound activity detection (SAD) network to identify segments with any sound event activity. We conduct an experiment using the DCASE 2016 task 2 dataset to compare our proposed method with typical conventional methods, such as non-negative matrix factorization (NMF) and a standard BLSTM-RNN. Our proposed method outperforms the conventional methods and achieves an F1-score 74.9 % (error rate of 44.7 %) on the event-based evaluation, and an F1-score of 80.5 % (error rate of 33.8 %) on the segment-based evaluation, most of which also outperforms the best reported result in the DCASE 2016 task 2 challenge. Tomoki Hayashi, Shinji Watanabe 0001, Tomoki Toda, Takaaki Hori, Jonathan Le Roux, Kazuya Takeda |
ICASSP | 4 |
| 2017 | Joint CTC-attention based end-to-end speech recognition using multi-task learningabstractRecently, there has been an increasing interest in end-to-end speech recognition that directly transcribes speech to text without any predefined alignments. One approach is the attention-based encoder-decoder framework that learns a mapping between variable-length input and output sequences in one step using a purely data-driven method. The attention model has often been shown to improve the performance over another end-to-end approach, the Connectionist Temporal Classification (CTC), mainly because it explicitly uses the history of the target character without any conditional independence assumptions. However, we observed that the performance of the attention has shown poor results in noisy condition and is hard to learn in the initial training stage with long input sequences. This is because the attention model is too flexible to predict proper alignments in such cases due to the lack of left-to-right constraints as used in CTC. This paper presents a novel method for end-to-end speech recognition to improve robustness and achieve fast convergence by using a joint CTC-attention model within the multi-task learning framework, thereby mitigating the alignment issue. An experiment on the WSJ and CHiME-4 tasks demonstrates its advantages over both the CTC and attention-based encoder-decoder baselines, showing 5.4-14.6% relative improvements in Character Error Rate (CER). Suyoun Kim, Takaaki Hori, Shinji Watanabe 0001 |
ICASSP | 2 |
| 2017 | Student-teacher network learning with enhanced featuresabstractRecent advances in distant-talking ASR research have confirmed that speech enhancement is an essential technique for improving the ASR performance, especially in the multichannel scenario. However, speech enhancement inevitably distorts speech signals, which can cause significant degradation when enhanced signals are used as training data. Thus, distant-talking ASR systems often resort to using the original noisy signals as training data and the enhanced signals only at test time, and give up on taking advantage of enhancement techniques in the training stage. This paper proposes to make use of enhanced features in the student-teacher learning paradigm. The enhanced features are used as input to a teacher network to obtain soft targets, while a student network tries to mimic the teacher network's outputs using the original noisy features as input, so that speech enhancement is implicitly performed within the student network. Compared with conventional student-teacher learning, which uses a better network as teacher, the proposed self-supervised method uses better (enhanced) inputs to a teacher. This setup matches the above scenario of making use of enhanced features in network training. Experiments with the CHiME-4 challenge real dataset show significant ASR improvements with an error reduction rate of 12% in the single-channel track and 15% in the 2-channel track, respectively, by using 6-channel beamformed features for the teacher model. Shinji Watanabe 0001, Takaaki Hori, Jonathan Le Roux, John R. Hershey |
ICASSP | 2 |
| 2017 | Attention-Based Multimodal Fusion for Video Description
Chiori Hori, Takaaki Hori, Teng-Yok Lee, Bret Harsham, John R. Hershey, Tim K. Marks, Kazuhiro Sumi |
ICCV | 2 |
| 2017 | Multichannel End-to-end Speech RecognitionabstractThe field of speech recognition is in the midst of a paradigm shift: end-to-end neural networks are challenging the dominance of hidden Markov models as a core technology. Using an attention mechanism in a recurrent encoder-decoder architecture solves the dynamic time alignment problem, allowing joint end-to-end training of the acoustic and language modeling components. In this paper we extend the end-to-end framework to encompass microphone array signal processing for noise suppression and speech enhancement within the acoustic encoding network. This allows the beamforming components to be optimized jointly within the recognition architecture to improve the end-to-end speech recognition objective. Experiments on the noisy speech benchmarks (CHiME-4 and AMI) show that our multichannel end-to-end system outperformed the attention-based baseline with input from a conventional adaptive beamformer. Tsubasa Ochiai, Shinji Watanabe 0001, Takaaki Hori, John R. Hershey |
ICML | 3 |
| 2017 | Advances in Joint CTC-Attention Based End-to-End Speech Recognition with a Deep CNN Encoder and RNN-LMabstractWe present a state-of-the-art end-to-end Automatic Speech Recognition (ASR) model. We learn to listen and write characters with a joint Connectionist Temporal Classification (CTC) and attention-based encoder-decoder network. The encoder is a deep Convolutional Neural Network (CNN) based on the VGG network. The CTC network sits on top of the encoder and is jointly trained with the attention-based decoder. During the beam search process, we combine the CTC predictions, the attention-based decoder predictions and a separately trained LSTM language model. We achieve a 5-10\% error reduction compared to prior systems on spontaneous Japanese and Chinese speech, and our end-to-end model beats out traditional hybrid ASR systems. Takaaki Hori, Shinji Watanabe 0001, Yu Zhang 0033 |
INTERSPEECH | 1 |
| 2017 | Multi-microphone speech recognition integrating beamforming, robust feature extraction, and advanced DNN/RNN backend
Takaaki Hori, Zhuo Chen 0006, Hakan Erdogan, John R. Hershey, Jonathan Le Roux, Vikramjit Mitra, Shinji Watanabe 0001 |
Comput. Speech Lang. | 1 |
| 2017 | Error detection and accuracy estimation in automatic speech recognition using deep bidirectional recurrent neural networks
Atsunori Ogawa, Takaaki Hori |
Speech Commun. | 2 |
| 2017 | Duration-Controlled LSTM for Polyphonic Sound Event DetectionabstractThis paper presents a new hybrid approach called duration-controlled long short-term memory (LSTM) for polyphonic sound event detection (SED). It builds upon a state-of-the-art SED method that performs frame-by-frame detection using a bidirectional LSTM recurrent neural network (BLSTM), and incorporates a duration-controlled modeling technique based on a hidden semi-Markov model. The proposed approach makes it possible to model the duration of each sound event precisely and to perform sequence-by-sequence detection without having to resort to thresholding, as in conventional frame-by-frame methods. Furthermore, to effectively reduce sound event insertion errors, which often occur under noisy conditions, we also introduce a binary-mask-based postprocessing that relies on a sound activity detection network to identify segments with any sound event activity, an approach inspired by the well-known benefits of voice activity detection in speech recognition systems. We conduct an experiment using the DCASE2016 task 2 dataset to compare our proposed method with typical conventional methods, such as nonnegative matrix factorization and standard BLSTM. Our proposed method outperforms the conventional methods both in an event-based evaluation, achieving a 75.3% F1 score and a 44.2% error rate, and in a segment-based evaluation, achieving an 81.1% F1 score, and a 32.9% error rate, outperforming the best results reported in the DCASE2016 task 2 Challenge. Tomoki Hayashi, Shinji Watanabe 0001, Tomoki Toda, Takaaki Hori, Jonathan Le Roux, Kazuya Takeda |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2016 | Minimum word error training of long short-term memory recurrent neural network language models for speech recognitionabstractThis paper describes minimum word error (MWE) training of recurrent neural network language models (RNNLMs) for speech recognition. RNNLMs are usually trained to minimize a cross entropy of estimated word probabilities against the correct word sequence, which corresponds to maximum likelihood criterion. However, this training does not necessarily maximize a performance measure in a target task, i.e. it does not minimize word error rate (WER) explicitly in speech recognition. To solve such a problem, several discriminative training methods have already been proposed for n-gram language models, but those for RNNLMs have not sufficiently investigated. In this paper, we propose a MWE training method for RNNLMs, and report significant WER reductions when we applied the MWE method to a standard Elman-type RNNLM and a more advanced model, a Long Short-Term Memory (LSTM) RNNLM. We also present efficient MWE training with N-best lists on Graphics Processing Units (GPUs). Takaaki Hori, Chiori Hori, Shinji Watanabe 0001, John R. Hershey |
ICASSP | 1 |
| 2016 | Driver confusion status detection using recurrent neural networksabstractIn this paper, we present a method for estimating the confusion level of a driver using a classifier trained on multimodal sensor data. Using the driver confusion status detector, a car navigation system can proactively support the driver when he/she is confused. A corpus of data was collected during on-road driving in traffic using a navigation system and a car instrumented with a variety of sensors. The data was manually annotated with the driver's confusion status and with multiple features representing driver's behavior and the traffic conditions. We compared different types of classifiers trained from the data: logistic regression, a feed-forward neural network, a recurrent neural networks, and a long short-term memory (LSTM)-based recurrent neural network. The accuracy was evaluated using F-max as well as precision/recall. We found that the LSTM outperformed the other models. Chiori Hori, Shinji Watanabe 0001, Takaaki Hori, Bret Harsham, John R. Hershey, Yusuke Koji, Youichi Fujii, Yuki Furumoto |
ICME | 3 |
| 2016 | Context-Sensitive and Role-Dependent Spoken Language Understanding Using Bidirectional and Attention LSTMs
Chiori Hori, Takaaki Hori, Shinji Watanabe 0001, John R. Hershey |
INTERSPEECH | 2 |
| 2016 | Dialog state tracking with attention-based sequence-to-sequence learningabstractWe present an advanced dialog state tracking system designed for the 5th Dialog State Tracking Challenge (DSTC5). The main task of DSTC5 is to track the dialog state in a human-human dialog. For each utterance, the tracker emits a frame of slot-value pairs considering the full history of the dialog up to the current turn. Our system includes an encoder-decoder architecture with an attention mechanism to map an input word sequence to a set of semantic labels, i.e., slot-value pairs. This handles the problem of the unknown alignment between the utterances and the labels. By combining the attention-based tracker with rule-based trackers elaborated for English and Chinese, the F-score for the development set improved from 0.475 to 0.507 compared to the rule-only trackers. Moreover, we achieved 0.517 F-score by refining the combination strategy based on the topic and slot level performance of each tracker. In this paper, we also validate the efficacy of each technique and report the test set results submitted to the challenge. Takaaki Hori, Chiori Hori, Shinji Watanabe 0001, Bret Harsham, Jonathan Le Roux, John R. Hershey, Yusuke Koji, Yi Jing, Zhaocheng Zhu, Takeyuki Aikawa |
SLT | 1 |
| 2016 | Automated structure discovery and parameter tuning of neural network language model based on evolution strategyabstractLong short-term memory (LSTM) recurrent neural network based language models are known to improve speech recognition performance. However, significant effort is required to optimize network structures and training configurations. In this study, we automate the development process using evolutionary algorithms. In particular, we apply the covariance matrix adaptation-evolution strategy (CMA-ES), which has demonstrated robustness in other black box hyper-parameter optimization problems. By flexibly allowing optimization of various meta-parameters including layer wise unit types, our method automatically finds a configuration that gives improved recognition performance. Further, by using a Pareto based multi-objective CMA-ES, both WER and computational time were reduced jointly: after 10 generations, relative WER and computational time reductions for decoding were 4.1% and 22.7% respectively, compared to an initial baseline system whose WER was 8.7%. Tomohiro Tanaka, Takafumi Moriya, Takahiro Shinozaki, Shinji Watanabe 0001, Takaaki Hori, Kevin Duh |
SLT | 5 |
| 2016 | Estimating Speech Recognition Accuracy Based on Error Type ClassificationabstractMethods for estimating the speech recognition accuracy without using manually transcribed references are beneficial to the research and development of automatic speech recognition technology. This paper proposes recognition accuracy estimation methods based on error type classification (ETC). ETC is an extension of confidence estimation. In ETC, each word in the recognition results (recognized word sequences) for the target speech data is probabilistically classified into three categories: the correct recognition (C), substitution error (S), and insertion error (I). Deletion errors (D) that can occur at interword positions in the recognition results are also probabilistically detected. By summing these CSID probabilities individually, the numbers of CSIDs and, as a result, the two standard recognition accuracy measures, i.e., the percent correct and word accuracy (WAcc), for the speech data can be estimated without using the reference transcriptions. Two recognition accuracy estimation methods based on ETC are proposed. In the first easy-to-use method, ETC is performed by converting the recognition results represented as word confusion networks into word alignment networks (WANs). In the second and more accurate method, the WAN-based ETC results are refined with conditional random fields (CRFs) using various types of additional features extracted for each of the recognized words. Experiments using English and Japanese lecture speech corpora show that the recognition accuracy can be accurately estimated with the CRF-based method. The correlation coefficient and root mean square error between the lecture-level true WAccs calculated using the reference transcriptions and those estimated with the CRF-based method are 0.97 and lower than 2%, respectively. A series of additional experiments and analyses are also conducted to better understand the effectiveness of the CRF-based method. Atsunori Ogawa, Takaaki Hori, Atsushi Nakamura |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2015 | The MERL/SRI system for the 3RD CHiME challenge using beamforming, robust feature extraction, and advanced speech recognitionabstractThis paper introduces the MERL/SRI system designed for the 3rd CHiME speech separation and recognition challenge (CHiME-3). Our proposed system takes advantage of recurrent neural networks (RNNs) throughout the model from the front speech enhancement to the language modeling. Two different types of beamforming are used to combine multi-microphone signals to obtain a single higher quality signal. Beamformed signal is further processed by a single-channel bi-directional long short-term memory (LSTM) enhancement network which is used to extract stacked mel-frequency cepstral coefficients (MFCC) features. In addition, two proposed noise-robust feature extraction methods are used with the beamformed signal. The features are used for decoding in speech recognition systems with deep neural network (DNN) based acoustic models and large-scale RNN language models to achieve high recognition accuracy in noisy environments. Our training methodology includes data augmentation and speaker adaptive training, whereas at test time model combination is used to improve generalization. Results on the CHiME-3 benchmark show that the full cadre of techniques substantially reduced the word error rate (WER). Combining hypotheses from different robust-feature systems ultimately achieved 9.10% WER for the real test data, a 72.4% reduction relative to the baseline of 32.99% WER. Takaaki Hori, Zhuo Chen 0006, Hakan Erdogan, John R. Hershey, Jonathan Le Roux, Vikramjit Mitra, Shinji Watanabe 0001 |
ASRU | 1 |
| 2015 | Double-layer neighborhood graph based similarity search for fast query-by-example spoken term detectionabstractThis paper presents a novel double-layer neighborhood graph index for acceleration of similarity search that accomplishes fast querybyexample spoken term detection (STD). When a query segment is given, our proposed STD method finds similar segments to the query from an utterance data set by efficient similarity search that traverses the double-layer neighborhood graph (DLG) with a low computational cost. The segment is a sequence of Gaussian mixture model posteriorgram frames and corresponds to a vertex in the DLG. A dissimilarity between vertices is measured by dynamic time warping. The DLG consists of two distinct degree-reduced k-nearest neighbor graphs in a base and an upper layer. The base layer's graph has all the vertices in the data set while the upper layer's graph includes only representatives extracted from the vertices in the base layer. By way of analogy, search in the DLG resembles driving on general roads and express highways appropriately for travel-time saving. Experimental results on the MIT lecture corpus demonstrate that the proposed method achieves CPU time reduction by 40% and more than 60% compared to the most recent method and the ordinary graphbased method, keeping almost the same precision. Kazuo Aoyama, Atsunori Ogawa, Takashi Hattori, Takaaki Hori |
ICASSP | 4 |
| 2015 | Context adaptive deep neural networks for fast acoustic model adaptationabstractDeep neural networks (DNNs) are widely used for acoustic modeling in automatic speech recognition (ASR), since they greatly outperform legacy Gaussian mixture model-based systems. However, the levels of performance achieved by current DNN-based systems remain far too low in many tasks, e.g. when the training and testing acoustic contexts differ due to ambient noise, reverberation or speaker variability. Consequently, research on DNN adaptation has recently attracted much interest. In this paper, we present a novel approach for the fast adaptation of a DNN-based acoustic model to the acoustic context. We introduce a context adaptive DNN with one or several layers depending on external factors that represent the acoustic conditions. This is realized by introducing a factorized layer that uses a different set of parameters to process each class of factors. The output of the factorized layer is then obtained by weighted averaging over the contribution of the different factor classes, given posteriors over the factor classes. This paper introduces the concept of context adaptive DNN and describes preliminary experiments with the TIMIT phoneme recognition task showing consistent improvement with the proposed approach. Marc Delcroix, Keisuke Kinoshita, Takaaki Hori, Tomohiro Nakatani |
ICASSP | 3 |
| 2015 | WFST-based structural classification integrating dnn acoustic features and RNN language features for speech recognitionabstractThis paper proposes a method to train Weighted Finite State Transducer (WFST) based structural classifiers using deep neural network (DNN) acoustic features and recurrent neural network (RNN) language features for speech recognition. Structural classification is an effective approach to achieve highly accurate recognition of structured data in which the classifier is optimized to maximize the discriminative performance using different kinds of features. A WFST-based classifier, which can integrate acoustic, pronunciation, and language features embedded in a composed WFST, was recently extended to incorporate DNN bottleneck (DNNBN) features. In this paper, we further investigate the integration of a RNN language model (RNNLM) with the WFST classifier. To this end, we introduce a lattice rescoring method using a RNNLM for efficient classifier training. In a lecture transcription task, we reduced the word error rate from 19.2% to 18.6% by optimizing the WFST parameters for the DNNBN acoustic and RNNLM language features. Quoc Truong Do, Satoshi Nakamura 0001, Marc Delcroix, Takaaki Hori |
ICASSP | 4 |
| 2015 | ASR error detection and recognition rate estimation using deep bidirectional recurrent neural networksabstractRecurrent neural networks (RNNs) have recently been applied as the classifiers for sequential labeling problems. In this paper, deep bidirectional RNNs (DBRNNs) are applied for the first time to error detection in automatic speech recognition (ASR), which is a sequential labeling problem. We investigate three types of ASR error detection tasks, i.e. confidence estimation, out-of-vocabulary word detection and error type classification. We also estimate recognition rates from the error type classification results. Experimental results show that the DBRNNs greatly outperform conditional random fields (CRFs), especially for the detection of infrequent error labels. The DBRNNs also slightly outperform the CRFs in recognition rate estimation. In addition, experiments using a reduced size of training data suggest that the DBRNNs have a better generalization ability than the CRFs owing to their word vector representation in a low-dimensional continuous space. As a result, the DBRNNs trained using only 20% of the training data show higher error detection performance than the CRFs trained using the full training data. Atsunori Ogawa, Takaaki Hori |
ICASSP | 2 |
| 2015 | Multiscale recurrent neural network based language model
Tsuyoshi Morioka, Tomoharu Iwata, Takaaki Hori, Tetsunori Kobayashi |
INTERSPEECH | 3 |
| 2014 | Zero-resource spoken term detection using hierarchical graph-based similarity searchabstractThis paper presents fast zero-resource spoken term detection (STD) in a large-scale data set, by using a hierarchical graph-based similarity search method (HGSS). HGSS is an improved graph-based similarity search method (GSS) in terms of a search space for high-speed performance. Instead of a degree-reduced k-nearest neighbor (k-DR) graph for GSS, a hierarchical k-DR graph, which is constructed based on a cluster structure in the corresponding k-DR graph, is used as an index for HGSS. A search algorithm for the hierarchical k-DR graph effectively utilizes the cluster structure, resulting in the reduction of the search space. HGSS inherits the useful property of GSS; it is available for any data sets without limits on a data type nor a defined dissimilarity since a graph is a general expression of a relationship between objects. A vertex and an edge in the hierarchical graph correspond to a Gaussian mixture model (GMM) posterior-gram segment and the relationship between a pair of GMM poste-riorgram segments, which is measured by dynamic time warping, respectively. Experimental results demonstrate that HGSS successfully reduces the computational cost by more than 40 % at nearly the same accuracy, compared to GSS. Kazuo Aoyama, Atsunori Ogawa, Takashi Hattori, Takaaki Hori, Atsushi Nakamura |
ICASSP | 4 |
| 2014 | Real-time one-pass decoding with recurrent neural network language model for speech recognitionabstractThis paper proposes an efficient one-pass decoding method for realtime speech recognition employing a recurrent neural network language model (RNNLM). An RNNLM is an effective language model that yields a large gain in recognition accuracy when it is combined with a standard n-gram model. However, since every word probability distribution based on an RNNLM is dependent on the entire history from the beginning of the speech, the search space in Viterbi decoding grows exponentially with the length of the recognition hypotheses and makes computation prohibitively expensive. Therefore, an RNNLM is usually used by N-best rescoring or by approximating it to a back-off n-gram model. In this paper, we present another approach that enables one-pass Viterbi decoding with an RNNLM without approximation, where the RNNLM is represented as a prefix tree of possible word sequences, and only the part needed for decoding is generated on-the-fly and used to rescore each hypothesis using an on-the-fly composition technique we previously proposed. Experimental results on the MIT lecture transcription task show that our proposed method enables one-pass decoding with small overhead for the RNNLM and achieves a slightly higher accuracy than 1000-best rescoring. Furthermore, it reduces the latency from the end of each utterance in two-pass decoding by a factor of 10. Takaaki Hori, Yotaro Kubo, Atsushi Nakamura |
ICASSP | 1 |
| 2014 | Fast segment search for corpus-based speech enhancement based on speech recognition technologyabstractCorpus-based speech enhancement has received increasing attention recently since it shows high enhancement performance in highly non-stationary noisy environments by precisely modeling the long-term temporal dynamics of speech. However, it has a disadvantage in that the cost is very high for searching the longest matching clean speech segments from a multi-condition parallel speech corpus. This paper proposes a fast segment search method for corpus-based speech enhancement. It is mainly based on two techniques derived from speech recognition technology. The first is an A* search like segment evaluation function for accurately finding the longest matching segments. The second is a tree and linear connected search space for efficiently sharing the segment likelihood calculations. In the experiments for non-stationary noisy observations using the 26 multi-condition TIMIT parallel speech corpus, the proposed search method found the segments almost in real-time without degrading the quality of the enhanced speech. Our method was about 7 to 13 times faster than the conventional segment search method. Atsunori Ogawa, Keisuke Kinoshita, Takaaki Hori, Tomohiro Nakatani, Atsushi Nakamura |
ICASSP | 3 |
| 2014 | Restructuring output layers of deep neural networks using minimum risk parameter clustering
Yotaro Kubo, Jun Suzuki 0001, Takaaki Hori, Atsushi Nakamura |
INTERSPEECH | 3 |
| 2013 | Graph index based query-by-example search on a large speech data setabstractThis paper presents a neighborhood graph index approach for query-by-example search using dynamic time warping (DTW) on Gaussian mixture model (GMM) posteriorgram sequences. The approach is intended to achieve a significant speed-up of a spoken term detection (STD) task for resource-limited situations. The proposed method employs a degree-reduced k-nearest neighbor (k-DR) graph as an index. A set of k-DR graphs is pre-constructed off-line from a large number of GMM posteriorgram sequences. Given a query posteriorgram sequence, one k-DR graph is selected from the set as the index. By applying a newly introduced combination of greedy-search (GS) and breadth-first search (BFS) algorithms to the selected k-DR graph index, the proposed method efficiently achieves query-by-example STD. Experimental results on the MIT lecture corpus demonstrate that the proposed method works much faster than a state-of-art method by more than one order magnitude, keeping almost the same precision. Kazuo Aoyama, Atsunori Ogawa, Takashi Hattori, Takaaki Hori, Atsushi Nakamura |
ICASSP | 4 |
| 2013 | Feature space variational Bayesian linear regression and its combination with model space VBLRabstractIn this paper, we propose a tuning-free Bayesian linear regression approach for speaker adaptation. We first formulate feature space variational Bayesian linear regression (fVBLR). Using a lower bound as the objective function, we can optimize a binary tree structure and control parameters for prior density scaling. We experimentally verified the proposed fVBLR could achieve performance comparable to that of the conventional fine-tuned fSMAPLR and SMAPLR. For further performance improvement regardless of the amount of adaptation data, we combine fVBLR with model space VBLR (fVBLR+VBLR). Therefore, feature space normalization and model space adaptation are consistently performed based on a variational Bayesian approach without any tuning parameters. In the experiment, the proposed fVBLR+VBLR showed performance improvement compared with both fVBLR and VBLR. Seong-Jun Hahm, Atsunori Ogawa, Marc Delcroix, Masakiyo Fujimoto, Takaaki Hori, Atsushi Nakamura |
ICASSP | 5 |
| 2013 | Large vocabulary continuous speech recognition based on WFST structured classifiers and deep bottleneck featuresabstractRecently, structured classification approaches have been considered important with a view to achieving unified modeling of the acoustic and linguistic aspects of speech recognizers. With these approaches, unified representation is achieved by directly optimizing a score function that measures the correspondence between the input and output of the system. Since structured classifiers typically employ a linear function as a score function, extracting expressive features from the input and output of the system is very important. On the other hand, the effectiveness of deep neural networks has been verified by several experiments, and it has been suggested that the outputs of hidden layers in deep neural networks (DNNs) are essential speech features that purely express phonetic information. In this paper, we propose a method for structured classification with DNN features. The proposed method expands conventional DNN- based acoustic models so that they optimizes the weight terms of the arcs in a decoding WFST, which is constructed with the on-the-fly composition method. Since DNN-based features can be considered enhancements in the input representation, the enhancements in the output representation based on the WFST arcs are expected to complement the DNN-based features. The proposed method achieved an 8 % relative error reduction even compared with a strong acoustic model based on DNNs. Yotaro Kubo, Takaaki Hori, Atsushi Nakamura |
ICASSP | 2 |
| 2013 | Coupling beamforming with spatial and spectral feature based spectral enhancement and its application to meeting recognitionabstractThis paper discusses microphone array based interference reduction approaches for robust automatic speech recognition. A model based multichannel spectral enhancement approach has recently been proposed for effectively reducing interference by exploiting both the spatial and spectral features of the signals. With the goal of further improving the effectiveness of this approach, we propose a new framework that combines this approach with a microphone-array based beamforming approach. Because the two approaches can work in a complementary manner in the proposed framework, they can greatly improve the interference reduction performance. We apply the proposed framework to the recognition of actual meetings, and show that it is superior to the use of beamforming or spectral enhancement alone in terms of the word error rates. Tomohiro Nakatani, Mehrez Souden, Shoko Araki, Takuya Yoshioka, Takaaki Hori, Atsunori Ogawa |
ICASSP | 5 |
| 2013 | Discriminative recognition rate estimation for N-best list and its application to N-best rescoringabstractTechniques for estimating recognition rates without using reference transcriptions are essential if we are to judge whether or not speech recognition technology is applicable to a new task. We have proposed a discriminative recognition rate estimation (DRRE) method for 1-best recognition hypotheses and shown its good estimation performance experimentally. In this paper, we extend our DRRE to N-best lists of recognition hypotheses by modifying its feature extraction procedures and efficiently selecting N-best hypotheses for its discriminative model training. In addition, we apply our extended DRRE to N-best rescoring. In the experiments, the extended DRRE also showed good estimation performance for the N-best lists. And using the estimated recognition rates, the 1-best word accuracy was significantly improved by N-best rescoring from the baseline. Atsunori Ogawa, Takaaki Hori, Atsushi Nakamura |
ICASSP | 2 |
| 2013 | A method for structure estimation of weighted finite-state transducers and its application to grapheme-to-phoneme conversion
Yotaro Kubo, Takaaki Hori, Atsushi Nakamura |
INTERSPEECH | 2 |
| 2013 | Unsupervised discriminative language modeling using error rate estimator
Takanobu Oba, Atsunori Ogawa, Takaaki Hori, Hirokazu Masataki, Atsushi Nakamura |
INTERSPEECH | 3 |
| 2013 | Speech recognition in living rooms: Integrated speech enhancement and recognition system based on spatial, spectral and temporal modeling of sounds
Marc Delcroix, Keisuke Kinoshita, Tomohiro Nakatani, Shoko Araki, Atsunori Ogawa, Takaaki Hori, Shinji Watanabe 0001, Masakiyo Fujimoto, Takuya Yoshioka, Takanobu Oba, Yotaro Kubo, Mehrez Souden, Seong-Jun Hahm, Atsushi Nakamura |
Comput. Speech Lang. | 6 |
| 2013 | Prior-shared feature and model space speaker adaptation by consistently employing map estimation
Seong-Jun Hahm, Shinji Watanabe 0001, Atsunori Ogawa, Masakiyo Fujimoto, Takaaki Hori, Atsushi Nakamura |
Speech Commun. | 5 |
| 2012 | Handling uncertain observations in unsupervised topic-mixture language model adaptationabstractWe propose an extension to the recent approaches in topic-mixture modeling such as Latent Dirichlet Allocation and Topic Tracking Model for the purpose of unsupervised adaptation in speech recognition. Instead of using the 1-best input given by the speech recognizer, the proposed model takes confusion network as an input to alleviate recognition errors. We incorporate a selection variable which helps reweight the recognition output, thus creating a more accurate latent topic estimate. Compared to adapting based on just one recognition hypothesis, the proposed model show WER improvements on two different tasks. Ekapol Chuangsuwanich, Shinji Watanabe 0001, Takaaki Hori, Tomoharu Iwata, James R. Glass |
ICASSP | 3 |
| 2012 | Spoken document retrieval by discriminative modeling in a high dimensional feature spaceabstractThis paper proposes discriminative modeling in a high dimensional feature space for spoken document retrieval (SDR). To estimate the parameters of a high dimensional model properly, a large quantity of data is necessary, but there is no such large corpus for document retrieval. This paper employs two approaches to overcome this problem. One is a reranking approach. A baseline system first gives each document a score and then the score is compensated by employing a high dimensional model. The other approach is automatic query generation. A large number of queries are automatically generated and used for parameter estimation. Our experimental result shows that our proposed method can greatly improve SDR performance. Takanobu Oba, Takaaki Hori, Atsushi Nakamura, Akinori Ito |
ICASSP | 2 |
| 2012 | Error type classification and word accuracy estimation using alignment features from word confusion networkabstractThis paper addresses error type classification in continuous speech recognition (CSR). In CSR, errors are classified into three types, namely, the substitution, insertion and deletion errors, by making an alignment between a recognized word sequence and its reference transcription with a dynamic programming (DP) procedure. We propose a method for deriving such alignment features from a word confusion network (WCN) without using the reference transcription. We show experimentally that the WCN-based alignment features steadily improve the performance of error type classification. They also improve the performance of out-of-vocabulary (OOV) word detection, since OOV word utterances are highly correlated with a particular alignment pattern. In addition, we show that the word accuracy can be estimated from the WCN-based alignment features and more accurately from the error type classification result without using the reference transcription. Atsunori Ogawa, Takaaki Hori, Atsushi Nakamura |
ICASSP | 2 |
| 2012 | Bag Of ARCS: New representation of speech segment features based on finite state machinesabstractThis paper proposes a new feature representation, Bag Of Arcs (BOA) for speech segments. A speech segment in BOA is simply represented as a set of counts for unique arcs in a finite state machine. Similar to the Bag Of Words model (BOW), BOA disregards the order of arcs, and thus, efficiently models speech segments. A strong motivation to use BOA is provided by a fact that the BOA representation is tightly connected to the output of a Weighted Finite State Transducer (WFST) based ASR decoder. Thus, BOA directly represents elements in the search network of a WFST-based ASR decoder, and can include information about context-dependent HMM topologies, lexicons, and back-off smoothed n-gram networks. In addition, the counts of BOA are accumulated by using the WFST decoder output directly, and we do not require an additional overhead and a change of decoding algorithms to extract the features. Consequently, we can combine the ASR decoder and post-processing without a process to extract word features from the decoder outputs or re-compiling WFST networks. We show the effectiveness of the proposed approach for some ASR post-processing applications in utterance classification experiments, and in speaker adaptation experiments by achieving absolute 1% improvement in WER from baseline results. We also show examples of latent semantic analysis for BOA by using latent Dirichlet allocation. Shinji Watanabe 0001, Yotaro Kubo, Takanobu Oba, Takaaki Hori, Atsushi Nakamura |
ICASSP | 4 |
| 2012 | Speaker Adaptation Using Variational Bayesian Linear Regression in Normalized Feature Space
Seong-Jun Hahm, Atsunori Ogawa, Masakiyo Fujimoto, Takaaki Hori, Atsushi Nakamura |
INTERSPEECH | 4 |
| 2012 | Efficient Beam Width Control to Suppress Excessive Speech Recognition Computation Time Based on Prior Score Range Normalization
Satoshi Kobashikawa, Takaaki Hori, Yoshikazu Yamaguchi, Taichi Asami, Hirokazu Masataki, Satoshi Takahashi |
INTERSPEECH | 2 |
| 2012 | Integrating Deep Neural Networks into Structural Classification Approach based on Weighted Finite-State Transducers
Yotaro Kubo, Takaaki Hori, Atsushi Nakamura |
INTERSPEECH | 2 |
| 2012 | Recognition rate estimation based on word alignment network and discriminative error type classificationabstractTechniques for estimating recognition rates without using reference transcriptions are essential if we are to judge whether or not speech recognition technology is applicable to a new task. This paper proposes two recognition rate estimation methods for continuous speech recognition. The first is an easy-to-use method based on a word alignment network (WAN) obtained from a word confusion network through simple conversion procedures. A WAN contains the correct (C), substitution error (S), insertion error (I) and deletion error (D) probabilities word-by-word for a recognition result. By summing these CSID probabilities individually, the percent correct and word accuracy (WACC) can be estimated without using a reference transcription. The second more advanced method refines the CSID probabilities provided by a WAN based on discriminative error type classification (ETC) and estimates the recognition rates more accurately. In the experiments on the MIT lecture speech corpus, we obtained 0.97 of correlation coefficient between the true WACCs calculated by a scoring tool using reference transcriptions and the WACCs estimated from the discriminative ETC results. Atsunori Ogawa, Takaaki Hori, Atsushi Nakamura |
SLT | 2 |
| 2012 | Efficient prior and incremental beam width control to suppress excessive speech recognition time based on score range estimationabstractThis paper proposes a technique that efficiently controls the beam width to yield practical computation times when auto-transcribing massive volumes of speeches. We focus on the fact that a lot of time is wasted by recognizing poor quality speeches that will yield, with inordinate slowness, erroneous transcriptions and provide no useful results. To stabilize the time regardless of quality, our proposal controls the beam width based on prolonged score spread against the target speech; it formulates the score range within the width and maximizes computation efficiency by regulating the range relevant to the hypotheses' survival rate. The proposed technique can control the width rapidly by using just monophones prior to decoding. It also restricts the width in decoding by using the processing speed and remaining data time to better handle stubborn speeches. Experiments with several SNRs and actual call-center speeches confirm a reduction in computation time while matching the accuracy of existing techniques. Satoshi Kobashikawa, Takaaki Hori, Yoshikazu Yamaguchi, Taichi Asami, Hirokazu Masataki, Satoshi Takahashi |
SLT | 2 |
| 2012 | Efficient training of discriminative language models by sample selection
Takanobu Oba, Takaaki Hori, Atsushi Nakamura |
Speech Commun. | 2 |
| 2012 | Low-Latency Real-Time Meeting Recognition and Understanding Using Distant Microphones and Omni-Directional CameraabstractThis paper presents our real-time meeting analyzer for monitoring conversations in an ongoing group meeting. The goal of the system is to recognize automatically “who is speaking what” in an online manner for meeting assistance. Our system continuously captures the utterances and face poses of each speaker using a microphone array and an omni-directional camera positioned at the center of the meeting table. Through a series of advanced audio processing operations, an overlapping speech signal is enhanced and the components are separated into individual speaker's channels. Then the utterances are sequentially transcribed by our speech recognizer with low latency. In parallel with speech recognition, the activity of each participant (e.g., speaking, laughing, watching someone) and the circumstances of the meeting (e.g., topic, activeness, casualness) are detected and displayed on a browser together with the transcripts. In this paper, we describe our techniques and our attempt to achieve the low-latency monitoring of meetings, and we show our experimental results for real-time meeting transcription. Takaaki Hori, Shoko Araki, Takuya Yoshioka, Masakiyo Fujimoto, Shinji Watanabe 0001, Takanobu Oba, Atsunori Ogawa, Kazuhiro Otsuka, Dan Mikami, Keisuke Kinoshita, Tomohiro Nakatani, Atsushi Nakamura, Junji Yamato |
IEEE Trans. Speech Audio Process. | 1 |
| 2012 | Structural Classification Methods Based on Weighted Finite-State Transducers for Automatic Speech RecognitionabstractThe potential of structural classification methods for automatic speech recognition (ASR) has been attracting the speech community since they can realize the unified modeling of acoustic and linguistic aspects of recognizers. However, the structural classification approaches involve well-known tradeoffs between the richness of features and the computational efficiency of decoders. If we are to employ, for example, a frame-synchronous one-pass decoding technique, features considered to calculate the likelihood of each hypothesis must be restricted to the same form as the conventional acoustic and language models. This paper tackles this limitation directly by exploiting the structure of the weighted finite-state transducers (WFSTs) used for decoding. Although WFST arcs provide rich contextual information, close integration with a computationally efficient decoding technique is still possible since most decoding techniques only require that their likelihood functions are factorizable for each decoder arc and time frame. In this paper, we compare two methods for structural classification with the WFST-based features; the structured perceptron and conditional random field (CRF) techniques. To analyze the advantages of these two classifiers, we present experimental results for the TIMIT continuous phoneme recognition task, the WSJ transcription task, and the MIT lecture transcription task. We confirmed that the proposed approach improved the ASR performance without sacrificing the computational efficiency of the decoders, even though the baseline systems are already trained with discriminative training techniques (e.g., MPE). Yotaro Kubo, Shinji Watanabe 0001, Takaaki Hori, Atsushi Nakamura |
IEEE Trans. Speech Audio Process. | 3 |
| 2012 | Round-Robin Duel Discriminative Language ModelsabstractDiscriminative training has received a lot of attention from both the machine learning and speech recognition communities. The idea behind the discriminative approach is to construct a model that distinguishes correct samples from incorrect samples, while the conventional generative approach estimates the distributions of correct samples. We propose a novel discriminative training method and apply it to a language model for reranking speech recognition hypotheses. Our proposed method has round-robin duel discrimination (R2D2) criteria in which all the pairs of sentence hypotheses including pairs of incorrect sentences are distinguished from each other, taking their error rate into account. Since the objective function is convex, the global optimum can be found through a normal parameter estimation method such as the quasi-Newton method. Furthermore, the proposed method is an expansion of the global conditional log-linear model whose objective function corresponds to the conditional random fields. Our experimental results show that R2D2 outperforms conventional methods in many situations, including different languages, different feature constructions and different difficulties. Takanobu Oba, Takaaki Hori, Atsushi Nakamura, Akinori Ito |
IEEE Trans. Speech Audio Process. | 2 |
| 2011 | Round-robin duel discriminative language models in one-pass decoding with on-the-fly error correctionabstractThis paper focuses on discriminative n-gram language models for large vocabulary speech recognition. We have proposed a novel training method called the round-robin duel discrimination (R2D2) method. Our previous report showed that R2D2 outperforms conventional methods on word n-gram based discriminative language models (DLMs). In this paper, we achieve additional error reduction and one-pass decoding at the same time. The keys to achieving this are the use of morphological features and the on-the-fly composition of weighted finite-state transducers (WFSTs) that represent both word and morphological discriminative features. Our experimental results show that R2D2 can reduce recognition errors more effectively than conventional methods in the reranking of n-best hypotheses and one-pass decoding can be accomplished with an equivalent accuracy. Takanobu Oba, Takaaki Hori, Akinori Ito, Atsushi Nakamura |
ICASSP | 2 |
| 2011 | Gibbs sampling based Multi-scale Mixture Model for speaker clusteringabstractThe aim of this work is to apply a sampling approach to speech modeling, and propose a Gibbs sampling based Multi-scale Mixture Model (M3). The proposed approach focuses on the multi-scale property of speech dynamics, i.e., dynamics in speech can be observed on, for instance, short-time acoustical, linguistic-segmental, and utterance-wise temporal scales. M3is an extension of the Gaussian mixture model and is considered a hierarchical mixture model, where mixture components in each time scale will change at intervals of the corresponding time unit. We derive a fully Bayesian treatment of the multi-scale mixture model based on Gibbs sampling. The advantage of the proposed model is that each speaker cluster can be precisely modeled based on the Gaussian mixture model unlike conventional single-Gaussian based speaker clustering (e.g., using the Bayesian Information Criterion (BIC)). In addition, Gibbs sampling offers the potential to avoid a serious local optimum problem. Speaker clustering experiments confirmed these advantages and obtained a significant improvement over the conventional BIC based approaches. Shinji Watanabe 0001, Daichi Mochihashi, Takaaki Hori, Atsushi Nakamura |
ICASSP | 3 |
| 2011 | Topic tracking language model for speech recognition
Shinji Watanabe 0001, Tomoharu Iwata, Takaaki Hori, Atsushi Sako, Yasuo Ariki |
Comput. Speech Lang. | 3 |
| 2010 | Search error risk minimization in Viterbi beam search for speech recognitionabstractThis paper proposes a method to optimize Viterbi beam search based on search error risk minimization in large vocabulary continuous speech recognition (LVCSR). Most speech recognizers employ beam search to speed up the decoding process, in which unpromising partial hypotheses are pruned during decoding. However, the pruning step involves the risk of missing the best complete hypothesis by discarding a partial hypothesis that might grow into the best. Missing the best hypothesis is called search error. Our purpose is to reduce search error by optimizing the pruning step. While conventional methods use heuristic criteria to prune each hypothesis based on its score, rank, and so on, our proposed method introduces a pruning function that makes a more precise decision using the rich features extracted from each hypothesis. The parameters of the function can be estimated efficiently to minimize the search error risk using recognition lattices at the training step. We implemented the new method in a WFST-based decoder and achieved a significant reduction of search errors in a 200K-word LVCSR task. Takaaki Hori, Shinji Watanabe 0001, Atsushi Nakamura |
ICASSP | 1 |
| 2010 | A comparative study on methods of Weighted language model training for reranking lvcsr N-best hypothesesabstractThis paper focuses on discriminative n-gram language models for a large vocabulary speech recognition task. Specifically we compare three training methods, Reranking Boosting (ReBst), Minimum Error Rate Training (MERT) and the Weighted Global Log-Linear Model (W-GCLM). They have a mechanism for handling sample weights, which are useful for providing an accurate model and work as impact factors of hypotheses for training. W-GCLM is proposed in this paper. We discuss the relationship between the three methods by comparing their loss functions. We also compare them experimentally by reranking N-best hypotheses under several conditions. We show that MERT and W-GCLM are different types of expansion of ReBst and have different respective advantages. Our experimental results reveal that W-GCLM outperforms ReBst and whether MERT or W-GCLM is superior depends on the training and test conditions. Takanobu Oba, Takaaki Hori, Atsushi Nakamura |
ICASSP | 2 |
| 2010 | A discriminative model for continuous speech recognition based on Weighted Finite State TransducersabstractThis paper proposes a discriminative model for speech recognition that directly optimizes the parameters of a speech model represented in the form of a decoding graph. In the process of recognition, a decoder, given an input speech signal, searches for an appropriate label sequence among possible combinations from separate knowledge sources of speech, e.g., acoustic, lexicon, and language models. It is more reasonable to use an integrated knowledge source, which is composed of these models and forms an overall space to be searched by a decoder, than to use separate ones. This paper aims to estimate a speech model composed in this way directly in the search network, unlike discriminative training approaches, which estimate parameters in acoustic or language model layers. Our approach is formulated as the weight parameter optimization of log-linear distributions in the decoding arcs of a Weighted Finite State Transducer (WFST) to efficiently handle a large network statically. The weight parameters are estimated by an averaged perceptron algorithm. The experimental results show that, especially when the model size is small, the proposed approach provided better recognition performance than the conventional maximum likelihood and comparable to or slightly better performance than discriminative training approaches. Shinji Watanabe 0001, Takaaki Hori, Erik McDermott, Atsushi Nakamura |
ICASSP | 2 |
| 2010 | Improvements of search error risk minimization in viterbi beam search for speech recognitionabstractThis paper describes improvements in a search error risk minimization approach to fast beam search for speech recognition. In our previous work, we proposed this approach to reduce search errors by optimizing the pruning criterion. While conventional methods use heuristic criteria to prune hypotheses, our proposed method employs a pruning function that makes a more precise decision using rich features extracted from each hypothesis. The parameters of the function can be estimated to minimize a loss function based on the search error risk. In this paper, we improve this method by introducing a modified loss function, arc-averaged risk, which potentially has a higher correlation with actual error rate than the original one. We also investigate various combinations of features. Experimental results show that further search error reduction over the original method is obtained in a 100K-word vocabulary lecture speech transcription task. Takaaki Hori, Shinji Watanabe 0001, Atsushi Nakamura |
INTERSPEECH | 1 |
| 2010 | Round-robin discrimination model for reranking ASR hypotheses
Takanobu Oba, Takaaki Hori, Atsushi Nakamura |
INTERSPEECH | 2 |
| 2010 | Large vocabulary continuous speech recognition using WFST-based linear classifier for structured dataabstractThis paper describes a discriminative approach that further advances the framework for Weighted Finite State Transducer (WFST) based decoding. The approach introduces additional linear models for adjusting the scores of a decoding graph composed of conventional information source models (e.g., hidden Markov models and N-gram models), and reviews the WFSTbased decoding process as a linear classifier for structured data (e.g., sequential multiclass data). The difficulty with the approach is that the number of dimensions of the additional linear models becomes very large in proportion to the number of arcs in a WFST, and our previous study only applied it to a small task (TIMIT phoneme recognition). This paper proposes a training method for a large-scale linear classifier employed in WFSTbased decoding by using a distributed perceptron algorithm. The experimental results show that the proposed approach was successfully applied to a large vocabulary continuous speech recognition task, and achieved an improvement compared with the performance of the minimum phone error based discriminative training of acoustic models. Index Terms: speech recognition, weighted finite state transducer, linear classifier, distributed perceptron, large vocabulary continuous speech recognition Shinji Watanabe 0001, Takaaki Hori, Atsushi Nakamura |
INTERSPEECH | 2 |
| 2010 | Real-time meeting recognition and understanding using distant microphones and omni-directional cameraabstractThis paper presents our newly developed real-time meeting analyzer for monitoring conversations in an ongoing group meeting. The goal of the system is to automatically recognize “who is speaking what” in an online manner for meeting assistance. Our system continuously captures the utterances and the face pose of each speaker using a distant microphone array and an omni-directional camera at the center of the meeting table. Through a series of advanced audio processing operations, an overlapping speech signal is enhanced and the components are separated into individual speaker's channels. Then the utterances are sequentially transcribed by our speech recognizer with low latency. In parallel with speech recognition, the activity of each participant (e.g. speaking, laughing, watching someone) and the situation of the meeting (e.g. topic, activeness, casualness) are detected and displayed on a browser together with the transcripts. In this paper, we describe our techniques and our attempt to achieve the low-latency monitoring of meetings, and we show our experimental results for real-time meeting transcription. Takaaki Hori, Shoko Araki, Takuya Yoshioka, Masakiyo Fujimoto, Shinji Watanabe 0001, Takanobu Oba, Atsunori Ogawa, Kazuhiro Otsuka, Dan Mikami, Keisuke Kinoshita, Tomohiro Nakatani, Atsushi Nakamura, Junji Yamato |
SLT | 1 |
| 2010 | Application of topic tracking model to language model adaptation and meeting analysisabstractIn a real environment, acoustic and language features often vary depending on the speakers, speaking styles and topic changes. This paper focuses on changes in the language environment, and applies a topic tracking model to language model adaptation for speech recognition and topic word extraction for meeting analysis. The topic tracking model can adaptively track changes in topics based on current text information and previously estimated topic models in an online manner. The effectiveness of the proposed method is shown experimentally by the improvement in speech recognition performance achieved with the Corpus of Spontaneous Japanese and by providing appropriate topic information in an automatic meeting analyzer. Shinji Watanabe 0001, Tomoharu Iwata, Takaaki Hori, Atsushi Sako, Yasuo Ariki |
SLT | 3 |
| 2008 | Sequential dependency analysis for online spontaneous speech processing
Takanobu Oba, Takaaki Hori, Atsushi Nakamura |
Speech Commun. | 2 |
| 2007 | Open-Vocabulary Spoken Utterance Retrieval using Confusion NetworksabstractThis paper presents a novel approach to open-vocabulary spoken utterance retrieval using confusion networks. If out-of-vocabulary (OOV) words are present in queries and the corpus, word-based indexing will not be sufficient. For this problem, we apply phone confusion networks and combine them with word confusion networks. With this approach, we can generate a more compact index table that enables robust keyword matching compared with typical lattice-based methods. In the retrieval experiments with speech recordings in MIT lecture corpus, our method using phone confusion networks outperformed lattice-based methods especially for OOV queries. Takaaki Hori, I. Lee Hetherington, Timothy J. Hazen, James R. Glass |
ICASSP (4) | 1 |
| 2007 | An approach to efficient generation of high-accuracy and compact error-corrective models for speech recognition
Takanobu Oba, Takaaki Hori, Atsushi Nakamura |
INTERSPEECH | 2 |
| 2007 | Efficient WFST-Based One-Pass Decoding With On-The-Fly Hypothesis Rescoring in Extremely Large Vocabulary Continuous Speech RecognitionabstractThis paper proposes a novel one-pass search algorithm with on-the-fly composition of weighted finite-state transducers (WFSTs) for large-vocabulary continuous-speech recognition. In the standard search method with on-the-fly composition, two or more WFSTs are composed during decoding, and a Viterbi search is performed based on the composed search space. With this new method, a Viterbi search is performed based on the first of the two WFSTs. The second WFST is only used to rescore the hypotheses generated during the search. Since this rescoring is very efficient, the total amount of computation required by the new method is almost the same as when using only the first WFST. In a 65k-word vocabulary spontaneous lecture speech transcription task, our proposed method significantly outperformed the standard search method. Furthermore, our method was faster than decoding with a single fully composed and optimized WFST, where our method used only 38% of the memory required for decoding with the single WFST. Finally, we have achieved high-accuracy one-pass real-time speech recognition with an extremely large vocabulary of 1.8 million words Takaaki Hori, Chiori Hori, Yasuhiro Minami, Atsushi Nakamura |
IEEE Trans. Speech Audio Process. | 1 |
| 2006 | An Extremely Large Vocabulary Approach to Named Entity Extraction from SpeechabstractThis paper describes an approach to named entity (NE) extraction from speech data, in which an extremely large vocabulary lexicon including all NEs occurring in a large text corpus is used for automatic speech recognition (ASR). Accordingly, NEs appear in the recognition results just as they are. Our approach is implemented by the following steps: (1) run an NE-tagger for a whole text corpus and make an NE-tagged corpus in which each NE is padded with its category, (2) construct a lexicon and a language model for ASR using the tagged corpus where each NE is considered as a regular word, and (3) run the speech recognizer in one pass. Although a very large vocabulary is necessary to ensure a high coverage of NEs, that is no longer a major problem since we recently achieved real-time extremely large vocabulary ASR using a WEST framework. In experiments on NE extraction from spoken queries for an open-domain question-answering system, our approach yielded higher F-measure values than a conventional approach Takaaki Hori, Atsushi Nakamura |
ICASSP (1) | 1 |
| 2006 | Sentence boundary detection using sequential dependency analysis combined with CRF-based chunking
Takanobu Oba, Takaaki Hori, Atsushi Nakamura |
INTERSPEECH | 2 |
| 2005 | Efficient Generation of high-order context-dependent Weighted Finite State Transducers for Speech RecognitionabstractThis paper describes an algorithm for efficient building of weighted finite state transducers for speech recognition when high-order context-dependent models of order K>3 (triphones) with tied states are used. We show how an algorithm to build a part of the needed composed transducers directly from the decision trees in combination with an improved compilation process can lead to much faster, simpler and more memory-efficient compilation. In our case, it also resulted in substantially smaller final networks. With the described algorithm, it is simple to use high-order full cross-word models with little overhead directly within a one-pass time-synchronous search, which we test comparing resulting final network sizes, recognition rates and speed on a large, spontaneous Japanese speech database. Using the proposed algorithm, it is possible to do real-time recognition using full crossword quinphones with a large acoustic model in about 125 MB of memory at about 9% search error. Mike Schuster, Takaaki Hori |
ICASSP (1) | 2 |
| 2005 | Generalized fast on-the-fly composition algorithm for WFST-based speech recognitionabstractThis paper describes a Generalized Fast On-the-fly Composition (GFOC) algorithm for Weighted Finite-State Transducers (WFSTs) in speech recognition. We already proposed the original version of GFOC, which yields fast and memory-efficient decoding using two WFSTs. GFOC enables fast on-the-fly composition of three or more WFSTs during decoding. In many cases, it is actually difficult or impossible to organize an entire transduction process of speech recognition using only one or two WFST(s) since some types of models considerably enlarge after written in WFST form. For example, a class n-gram model often results in a large WFST which is several times larger than a word n-gram model for the same vocabulary. GFOC makes it possible to use such a model after decomposing it into small multiple WFSTs. In a spontaneous speech transcription task, we evaluated the size of WFSTs, decoding speed, and word accuracy of several decoding approaches. The results show that GFOC with three or more WFSTs is an efficient algorithm when using a class-based language model. 1. Takaaki Hori, Atsushi Nakamura |
INTERSPEECH | 1 |
| 2005 | Experiments with probabilistic principal component analysis in LVCSR
Mike Schuster, Takaaki Hori, Atsushi Nakamura |
INTERSPEECH | 2 |
| 2004 | Fast on-the-fly composition for weighted finite-state transducers in 1.8 million-word vocabulary continuous speech recognitionabstractThis paper proposes a new on-the-fly composition algorithm for Weighted Finite-State Transducers (WFSTs) in large-vocabulary continuous-speech recognition. In general on-the-fly composition, two transducers are composed during decoding, and a Viterbi search is performed based on the composed search space. In this new method, a Viterbi search is performed based on the first of two transducers. The second transducer is only used to rescore the hypotheses generated during the search. Since this rescoring is very efficient, the total amount of computation in the new method is almost the same as when using only the first transducer. In a 30kword vocabulary spontaneous lecture speech transcription task, our proposed method significantly outperformed the general on-the-fly composition method. Furthermore the speed of our method was slightly faster than that of decoding with a single fully composed and optimized WFST, where our method consumed only 20 % of the memory usage required for decoding with the single WFST. Finally, we have achieved one-pass real-time speech recognition in an extremely large vocabulary of 1.8 million words. 1. Takaaki Hori, Chiori Hori, Yasuhiro Minami |
INTERSPEECH | 1 |
| 2003 | Deriving disambiguous queries in a spoken interactive ODQA systemabstractRecently, open-domain question answering (ODQA) systems that extract an exact answer from large text corpora based on text input are intensively being investigated. However, the information in the first question input by a user is not usually enough to yield the desired answer. Interactions for collecting additional information to accomplish QA is needed. This paper proposes an interactive approach for spoken interactive ODQA systems. When the reliabilities for answer hypotheses obtained by an ODQA system are low, the system automatically derives disambiguous queries (DQ) that draw out additional information. The additional information based on the DQ should contribute to distinguishing effectively an exact answer and to supplementing a lack of information by recognition errors. In our spoken interactive ODQA system, SPIQA, spoken questions are recognized by an ASR system, and DQ are automatically generated to disambiguate the transcribed questions. We confirmed the appropriateness of the derived DQ by comparing them with manually prepared ones. Chiori Hori, Takaaki Hori, Hideki Isozaki, Eisaku Maeda, Shigeru Katagiri, Sadaoki Furui |
ICASSP (1) | 2 |
| 2003 | Language model adaptation using WFST-based speaking-style translationabstractThis paper describes a new approach to language model adaptation for speech recognition based on the statistical framework of speech translation. The main idea of this approach is to compose a weighted finite-state transducer (WFST) that translates sentence styles from in-domain to out-of-domain. It enables to integrate language models of different styles of speaking or dialects and even of different vocabularies. The WFST is built by combining in-domain and out-of-domain models through the translation, while each model and the translation itself is expressed as a WFST. We apply this technique to building language models for spontaneous speech recognition using large written-style corpora. We conducted experiments on a 20k-word Japanese spontaneous speech recognition task. With a small in-domain corpus, a 2.9% absolute improvement in word error rate is achieved over the in-domain model. Takaaki Hori, Daniel Willett, Yasuhiro Minami |
ICASSP (1) | 1 |
| 2003 | Evaluation method for automatic speech summarizationabstractWe have proposed an automatic speech summarization approach that extracts words from transcription results obtained by automatic speech recognition (ASR) systems. To numerically evaluate this approach, the automatic summarization results are compared with manual summarization generated by humans through word extraction. We have proposed three metrics, weighted word precision, word strings precision and summarization accuracy (SumACCY), based on a word network created by merging manual summarization results. In this paper, we propose a new metric for automatic summarization results, weighted summarization accuracy (WSumACCY). This accuracy is weighted by the posterior probability of the manual summaries in the network to give the reliability of each answer extracted from the network. We clarify the goal of each metric and use these metrics to provide automatic evaluation results of the summarized speech. To compare the performance of each evaluation metric, correlations between the evaluation results using these metrics and subjective evaluation by hand are measured. It is confirmed that WSumACCY is an effective and robust measure for automatic summarization. Chiori Hori, Takaaki Hori, Sadaoki Furui |
INTERSPEECH | 2 |
| 2003 | Speech summarization using weighted finite-state transducersabstractThis paper proposes an integrated framework to summarize spontaneous speech into written-style compact sentences. Most current speech recognition systems attempt to transcribe whole spoken words correctly. However, recognition results of spontaneous speech are usually difficult to understand, even if the recognition is perfect, because spontaneous speech includes redundant information, and its style is different to that of written sentences. In particular, the style of spoken Japanese is very different to that of the written language. Therefore, techniques to summarize recognition results into readable and compact sentences are indispensable for generating captions or minutes from speech. Our speech summarization includes speech recognition, paraphrasing, and sentence compaction, which are integrated in a single Weighted Finite-State Transducer (WFST). This approach enables the decoder to employ all the knowledge sources in a one-pass search strategy and reduces the search errors, since all the constraints of the models are used from the beginning of the search. We conducted experiments on a 20kword Japanese lecture speech recognition and summarization task. Our approach yielded improvements in both recognition accuracy and summarization accuracy compared with other approaches that perform speech recognition and summarization separately. 1. Takaaki Hori, Chiori Hori, Yasuhiro Minami |
INTERSPEECH | 1 |
| 2001 | Improved phoneme-history-dependent search for large-vocabulary continuous-speech recognition
Takaaki Hori, Yoshiaki Noda, Shoichi Matsunaga |
INTERSPEECH | 1 |