Atsunori Ogawa

dblp:32/2157 · DBLP profile ↗
← Back
107ranked-venue papers
25as first author
31since 2021 · last 2026
0000-0002-2888-101XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 100 · 24 first-author · 30 since 2021Artificial intelligence and machine learning · 56 · 8 first-author · 19 since 2021
YearPublicationVenuePosition
2026 Microphone array geometry-independent multi-talker distant ASR: NTT system for DASR task of the CHiME-8 challenge
Naoyuki Kamo, Naohiro Tawara, Atsushi Ando, Takatomo Kano, Hiroshi Sato 0002, Rintaro Ikeshita, Takafumi Moriya, Shota Horiguchi, Kohei Matsuura, Atsunori Ogawa, Alexis Plaquet, Takanori Ashihara, Tsubasa Ochiai, Masato Mimura, Marc Delcroix, Tomohiro Nakatani, Taichi Asami, Shoko Araki
Comput. Speech Lang.10
2025 Predictive ASR and Turn-taking Prediction at Once: Towards More Responsive Spoken Dialog System
abstract
Spoken dialog systems usually wait for users to finish speaking before generating responses, resulting in response delays. A possible solution for reducing the response delay is to predict future words and/or turn-ends while the user is speaking. To realize this, we propose a method to jointly perform predictive automatic speech recognition and turn-taking prediction. Our model receives partial utterances as input and performs speech recognition, future word prediction, and turntaking prediction via autoregressive decoding. It enables turntaking prediction based on prosodic and linguistic cues of observed partial utterances and predicted future linguistic cues. We also incorporate dialogue contexts to improve the performance. Experiments on the Switchboard corpus showed that our multi-task model outperforms a single-task model in turn-taking prediction. We found that conditioning turn-taking prediction on predicted words improved performance when words were correctly predicted.
Ryo Fukuda, Takatomo Kano, Naohiro Tawara, Marc Delcroix, Atsunori Ogawa, Yuya Chiba, Atsushi Ando
ASRU5
2025 All-in-One ASR: Unifying Encoder-Decoder Models of CTC, Attention, and Transducer in Dual-Mode ASR
abstract
This paper proposes a unified framework, All-in-One ASR, that allows a single model to support multiple automatic speech recognition (ASR) paradigms, including connectionist temporal classification (CTC), attention-based encoder-decoder (AED), and Transducer, in both offline and streaming modes. While each ASR architecture offers distinct advantages and trade-offs depending on the application, maintaining separate models for each scenario incurs substantial development and deployment costs. To address this issue, we introduce a multi-mode joiner that enables seamless integration of various ASR modes within a single unified model. Experiments show that All-in-One ASR significantly reduces the total model footprint while matching or even surpassing the recognition performance of individually optimized ASR models. Furthermore, joint decoding leverages the complementary strengths of different ASR modes, yielding additional improvements in recognition accuracy.
Takafumi Moriya, Masato Mimura, Tomohiro Tanaka, Hiroshi Sato 0002, Ryo Masumura, Atsunori Ogawa
ASRU6
2025 Speech Emotion Recognition Based on Large-Scale Automatic Speech Recognizer
abstract
This paper proposes a novel speech emotion recognition (SER) method that fully leverages the architecture of Whisper, a large-scale automatic speech recognition (ASR) model. The conventional SER models using a pre-trained speech encoder may fail to capture linguistic content since their decoders are too simple. Our proposed method addresses this shortcoming by adopting the decoder of Whisper, which has been discarded in conventional SER, to leverage its language modeling capability. The proposed method introduces special tokens corresponding to the target emotions and then fine-tunes the entire Whisper model. Furthermore, we also propose a new training scheme suitable for Whisper, named serialized multi-task learning (SerialMTL), to consider various speech information as context for the objective SER task. In SerialMTL, the model initially predicts subtask tokens, such as transcription and gender tokens, and then estimates the emotion token. An advantage of the proposed method is the simplicity of the model structure, even when adding any new subtasks. Experimental results show that our model, based on the entire Whisper, achieves better SER performance than the conventional model and further improves with SerialMTL training via ASR and gender recognition subtasks.
Ryo Fukuda, Takatomo Kano, Atsushi Ando, Atsunori Ogawa
ICASSP4
2025 Bridging Speech and Text Foundation Models with ReShape Attention
abstract
This paper investigates cascade approaches bridging speech and text foundation models (FMs) for speech translation (ST). We address the limitations of cascade systems which suffer from the propagation of speech recognition errors and the lack of access to acoustic information. We propose a ReShape Attention (RSA) that bridges speech embeddings of Whisper, a speech FM, to LLaMA2, a text FM. Speech and text embeddings have temporal and dimensional gaps, which make merging them challenging. RSA reshapes the speech and text embeddings into a sequence of subvectors sharing the same feature dimension. RSA performs cross-attention in the LLaMA2 layers between these two sequences, which allows combining the two embeddings. The RSA allows text FM to directly access speech FM embeddings and optimize the entire ST system for input speech. RSA improves 8.5% relative BLEU score compared to the baseline ST system, which cascades Whisper and LLaMA2. Moreover, our analyses show that the proposed method could even improve performance with ground-truth transcriptions, which suggests that our bridging approach is not limited to mitigating the effect of recognition errors but can also exploit the benefit of acoustic information.
Takatomo Kano, Atsunori Ogawa, Marc Delcroix, Ryo Fukuda, Kohei Matsuura, Takanori Ashihara, Shinji Watanabe 0001
ICASSP2
2025 Why is children's ASR so difficult? Analyzing children's phonological error patterns using SSL-based phoneme recognizers
Koharu Horii, Naohiro Tawara, Atsunori Ogawa, Shoko Araki
INTERSPEECH3
2025 Pick and Summarize: Integrating Extractive and Abstractive Speech Summarization
Takatomo Kano, Atsunori Ogawa, Marc Delcroix, Ryo Fukuda, Shinji Watanabe 0001
INTERSPEECH2
2025 Domain adaptation using non-parallel target domain corpus for self-supervised learning-based automatic speech recognition
abstract
The recognition accuracy of conventional automatic speech recognition (ASR) systems depends heavily on the amount of speech and associated transcription data available in the target domain for model training. However, preparing parallel speech and text data each time a model is trained for a new domain is costly and time-consuming. To solve this problem, we propose a method of domain adaptation that does not require the use of a large amount of parallel target domain training data, as most of the data used for model training is not from the target domain. Instead, only target domain speech is used for model training, along with non-target domain speech and its parallel text data, i.e., the domains and contents of the two types of training data do not correspond to one another. Collecting this type of training data is relatively inexpensive. Domain adaptation is performed in two steps: (1) A pre-trained wav2vec 2.0 model is further pre-trained using a large amount of target domain speech data and is then fine-tuned using a large amount of non-target domain speech and its transcriptions. (2) The density ratio approach (DRA) is applied during inference to a language model (LM) trained using target domain text unrelated to, and independently from, the wav2vec 2.0 training. Experimental evaluation illustrated that the proposed domain adaptation obtained character error rate (CER) 10.4 pts lower than baseline with wav2vec 2.0 and 3.9 pts with XLS-R under the situation that the parallel target domain data is unavailable against the target domain test set, achieving 34.4% and 16.2% reductions in relative CER.
Takahiro Kinouchi, Atsunori Ogawa, Yukoh Wakabayashi, Kengo Ohta, Norihide Kitaoka
Speech Commun.2
2024 Train Long and Test Long: Leveraging Full Document Contexts in Speech Processing
abstract
The quadratic memory complexity of self-attention has generally restricted Transformer-based models to utterance-based speech processing, preventing models from leveraging long-form contexts. A common solution has been to formulate long-form speech processing into a streaming problem, only using limited prior context. We propose a new and simple paradigm, encoding entire documents at once, which has been unexplored in Automatic Speech Recognition (ASR) and Speech Translation (ST) due to its technical infeasibility. We exploit developments in efficient attention mechanisms, such as Flash Attention, and show that Transformer-based models can be easily adapted to document-level processing. We experiment with methods to address the quadratic complexity of attention by replacing it with simpler alternatives. As such, our models can handle up to 30 minutes of speech during both training and testing. We evaluate our models on ASR, ST, and Speech Summarization (SSUM) using How2, TEDLIUM3, and SLUE-TED. With document-level context, our ASR models achieve 33.3% and 6.5% relative improvements in WER on How2 and TEDLIUM3 over prior work. Finally, we use our findings to propose a new attention-free self-supervised model, LongHuBERT, capable of handling long inputs. In doing so, we achieve state-of-the-art performance on SLUE-TED SSUM, outperforming cascaded systems that have dominated the benchmark.
Takatomo Kano, Atsunori Ogawa, Marc Delcroix, Shinji Watanabe 0001
ICASSP3
2024 NTT Speaker Diarization System for Chime-7: Multi-Domain, Multi-Microphone end-to-end and Vector Clustering Diarization
abstract
This paper details our speaker diarization system designed for multi-domain, multi-microphone casual conversations. The proposed diarization pipeline uses weighted prediction error (WPE)based dereverberation as a front end, and separately applies end-to-end neural diarization with vector clustering (EEND-VC) to each channel. It integrates the diarization result obtained from each channel using diarization output voting error reduction plus overlap (DOVER-Lap). To harness the knowledge from the target domain and the results integrated across all channels, we apply self-supervised adaptation for each session by retraining the EEND-VC with pseudo-labels derived from DOVER-Lap. We incorporated our proposed system into NTT’s submission for a distant automatic speech recognition task in the CHiME-7 challenge. Our system obtained third place in the diarization performance by improving the development and evaluation sets by 65 % and 62 % compared to the organizer-provided, VC-based baseline diarization system.
Naohiro Tawara, Marc Delcroix, Atsushi Ando, Atsunori Ogawa
ICASSP4
2024 Boosting CTC-based ASR using inter-layer attention-based CTC loss
abstract
This paper addresses improving the performance of CTC-based models, which leverage the intermediate outputs of all encoder layers with an attention mechanism.Several previous studies have used the intermediate outputs of the encoder layer to modify CTC-based models.Here, we focus on the role of the Transformer encoder layer, and each encoder layer is computed for two CTC losses by weighting the intermediate outputs of its lower and upper layers using an attention mechanism.By dividing the layer into two groups, it is expected to be possible to calculate the loss, taking into account both acoustic and linguistic features.Experimental results showed that the proposed method improved the baseline recognition performance of TEDLIUM2 speech data, achieving a WER of 9.9% on the dev set and 11.8% on the test set.Our method outperformed the conventional methods for WER with only slightly increased inference speed measured by RTF.
Keigo Hojo, Yukoh Wakabayashi, Kengo Ohta, Atsunori Ogawa, Norihide Kitaoka
INTERSPEECH4
2024 Sentence-wise Speech Summarization: Task, Datasets, and End-to-End Modeling with LM Knowledge Distillation
Kohei Matsuura, Takanori Ashihara, Takafumi Moriya, Masato Mimura, Takatomo Kano, Atsunori Ogawa, Marc Delcroix
INTERSPEECH6
2024 Text-only Domain Adaptation for CTC-based Speech Recognition through Substitution of Implicit Linguistic Information in the Search Space
Tatsunari Takagi, Yukoh Wakabayashi, Atsunori Ogawa, Norihide Kitaoka
INTERSPEECH3
2023 Summarize While Translating: Universal Model With Parallel Decoding for Summarization and Translation
abstract
Recently, multi-decoder and universal models have attracted increased interest in speech and language processing as they allow learning common representations across tasks. These models learn a common representation by sharing a part of or all network parameters. Moreover, such a universal model can handle tasks unseen during training (zero-shot tasks). However, these models do not fully exploit inter-dependencies between tasks during decoding since they usually perform decoding for each task independently. In this paper, we propose to address this issue by extending the universal model to perform multi-task parallel decoding with a cross-attention module between decoders to capture task inter-dependencies explicitly. We also introduce a novel multi-stream beam search algorithm to allow such parallel decoding. We test our proposed model on multi-lingual (English and Portuguese) text/speech translation and summarization, confirming its potential, especially in zero-shot tasks.
Takatomo Kano, Atsunori Ogawa, Marc Delcroix, Kohei Matsuura, Takanori Ashihara, Shinji Watanabe 0001
ASRU2
2023 Espnet-Summ: Introducing a Novel Large Dataset, Toolkit, and a Cross-Corpora Evaluation of Speech Summarization Systems
abstract
Speech summarization has garnered significant interest and progressed rapidly over the past few years. In particular, end-to-end models have recently emerged as a competitive alternative to cascade systems for abstractive video summarization. This paper aims to establish progress in this rapidly evolving research field, by introducing ESPNet-SUMM, a new open-source toolkit that facilitates a comprehensive comparison of end-to-end and cascade speech summarization models on 4 different speech summarization tasks spanning diverse applications. Experiments demonstrate that end-to-end models perform better for larger corpora with shorter inputs. This work also introduces Interview, the largest public open-domain multiparty interview corpus with $4400 \mathrm{~h}$ of conversations between radio hosts and guests. Finally, this work explores the use of multiple datasets to improve end-to-end summarization, and experiments demonstrate the benefit of multi-style training over fine-tuning. 1
Roshan S. Sharma, Takatomo Kano, Ruchira Sharma, Siddhant Arora, Shinji Watanabe 0001, Atsunori Ogawa, Marc Delcroix, Rita Singh, Bhiksha Raj
ASRU7
2023 Speech Summarization of Long Spoken Document: Improving Memory Efficiency of Speech/Text Encoders
abstract
Speech summarization requires processing several minute-long speech sequences to allow exploiting the whole context of a spoken document. A conventional approach is a cascade of automatic speech recognition (ASR) and text summarization (TS). However, the cascade systems are sensitive to ASR errors. Moreover, the cascade system cannot be optimized for input speech and utilize para-linguistic information. Recently, there has been an increased interest in end-to-end (E2E) approaches optimized to output summaries directly from speech. Such systems can thus mitigate the ASR errors of cascade approaches. However, E2E speech summarization requires massive computational resources because it needs to encode long speech sequences. We propose a speech summarization system that enables E2E summarization from 100 seconds, which is the limit of the conventional method, to up to 10 minutes (i.e., the duration of typical instructional videos on YouTube). However, the modeling capability of this model for minute-long speech sequences is weaker than the conventional approach. We thus exploit auxiliary text information from ASR transcriptions to improve the modeling capabilities. The resultant system consists of a dual speech/text encoder decoder-based summarization system. We perform experiments on the How2 dataset showing the proposed system improved METEOR scores by up to 2.7 points by fully exploiting the long spoken documents.
Takatomo Kano, Atsunori Ogawa, Marc Delcroix, Roshan S. Sharma, Kohei Matsuura, Shinji Watanabe 0001
ICASSP2
2023 Leveraging Large Text Corpora For End-To-End Speech Summarization
abstract
End-to-end speech summarization (E2E SSum) is a technique to directly generate summary sentences from speech. Compared with the cascade approach, which combines automatic speech recognition (ASR) and text summarization models, the E2E approach is more promising because it mitigates ASR errors, incorporates nonverbal information, and simplifies the overall system. However, since collecting a large amount of paired data (i.e., speech and summary) is difficult, the training data is usually insufficient to train a robust E2E SSum system. In this paper, we present two novel methods that leverage a large amount of external text summarization data for E2E SSum training. The first technique is to utilize a text-to-speech (TTS) system to generate synthesized speech, which is used for E2E SSum training with the text summary. The second is a TTS-free method that directly inputs phoneme sequence instead of synthesized speech to the E2E SSum model. Experiments show that our proposed TTS- and phoneme-based methods improve several metrics on the How2 dataset. In particular, our best system outperforms a previous state-of-the-art one by a large margin (i.e., METEOR score improvements of more than 6 points). To the best of our knowledge, this is the first work to use external language resources for E2E SSum. Moreover, we report a detailed analysis of the How2 dataset to confirm the validity of our proposed E2E SSum system.
Kohei Matsuura, Takanori Ashihara, Takafumi Moriya, Tomohiro Tanaka, Atsunori Ogawa, Marc Delcroix, Ryo Masumura
ICASSP5
2023 Iterative Shallow Fusion of Backward Language Model for End-To-End Speech Recognition
abstract
We propose a new shallow fusion (SF) method to exploit an external backward language model (BLM) for end-to-end automatic speech recognition (ASR). The BLM has complementary characteristics with a forward language model (FLM), and the effectiveness of their combination has been confirmed by rescoring ASR hypotheses as post-processing. In the proposed SF, we iteratively apply the BLM to partial ASR hypotheses in the backward direction (i.e., from the possible next token to the start symbol) during decoding, substituting the newly calculated BLM scores for the scores calculated at the last iteration. To enhance the effectiveness of this iterative SF (ISF), we train a partial sentence-aware BLM (PBLM) using reversed text data including partial sentences, considering the framework of ISF. In experiments using an attention-based encoder-decoder ASR system, we confirmed that ISF using the PBLM shows comparable performance with SF using the FLM. By performing ISF, early pruning of prospective hypotheses can be prevented during decoding, and we can obtain a performance improvement compared to applying the PBLM as post-processing. Finally, we confirmed that, by combining SF and ISF, further performance improvement can be obtained thanks to the complementarity of the FLM and PBLM.
Atsunori Ogawa, Takafumi Moriya, Naoyuki Kamo, Naohiro Tawara, Marc Delcroix
ICASSP1
2023 Impact of Residual Noise and Artifacts in Speech Enhancement Errors on Intelligibility of Human and Machine
Shoko Araki, Ayako Yamamoto, Tsubasa Ochiai, Kenichi Arai, Atsunori Ogawa, Tomohiro Nakatani, Toshio Irino
INTERSPEECH5
2023 Multi-Stream Extension of Variational Bayesian HMM Clustering (MS-VBx) for Combined End-to-End and Vector Clustering-based Diarization
Marc Delcroix, Naohiro Tawara, Mireia Díez, Federico Landini, Anna Silnova, Atsunori Ogawa, Tomohiro Nakatani, Lukás Burget, Shoko Araki
INTERSPEECH6
2023 What are differences? Comparing DNN and Human by Their Performance and Characteristics in Speaker Age Estimation
Yuki Kitagishi, Naohiro Tawara, Atsunori Ogawa, Ryo Masumura, Taichi Asami
INTERSPEECH3
2023 Transfer Learning from Pre-trained Language Models Improves End-to-End Speech Summarization
Kohei Matsuura, Takanori Ashihara, Takafumi Moriya, Tomohiro Tanaka, Takatomo Kano, Atsunori Ogawa, Marc Delcroix
INTERSPEECH6
2023 Knowledge Distillation for Neural Transducer-based Target-Speaker ASR: Exploiting Parallel Mixture/Single-Talker Speech Data
Takafumi Moriya, Hiroshi Sato 0002, Tsubasa Ochiai, Marc Delcroix, Takanori Ashihara, Kohei Matsuura, Tomohiro Tanaka, Ryo Masumura, Atsunori Ogawa, Taichi Asami
INTERSPEECH9
2022 Integrating Multiple ASR Systems into NLP Backend with Attention Fusion
abstract
Spoken language processing (SLP) systems such as speech summarization and translation can be achieved by cascade models. It combines an automatic speech recognition (ASR) frontend and a natural language processing (NLP) backend including machine translation (MT) or text summarization (TS). With this cascade approach, we can exploit large non-paired datasets to independently train state-of-the-art models for each module. However, ASR errors directly affect the performance of the NLP backend in the cascade approach. In this paper, we reduce the impact of ASR errors on the NLP back-end by combining transcriptions from various ASR systems. Recognizer output voting error reduction (ROVER) is a widely used technique for system combination. Although ROVER improves ASR performance, the combination process is not optimized for backend tasks. We propose a system combination that resembles ROVER using attention fusion to achieve the alignment and the combination of multiple ASR hypotheses. This allows the combination process to be optimized for the backend NLP task without changing the ASR frontend. Our proposed technique is general and can be applied to various SLP tasks. We confirm its effectiveness on both speech summarization and translation experiments.
Takatomo Kano, Atsunori Ogawa, Marc Delcroix, Shinji Watanabe 0001
ICASSP2
2022 Lattice Rescoring Based on Large Ensemble of Complementary Neural Language Models
abstract
We investigate the effectiveness of using a large ensemble of advanced neural language models (NLMs) for lattice rescoring on automatic speech recognition (ASR) hypotheses. Previous studies have reported the effectiveness of combining a small number of NLMs. In contrast, in this study, we combine up to eight NLMs, i.e., forward/backward long short-term memory/Transformer-LMs that are trained with two different random initialization seeds. We combine these NLMs through iterative lattice generation. Since these NLMs work complementarily with each other, by combining them one by one at each rescoring iteration, language scores attached to given lattice arcs can be gradually refined. Consequently, errors of the ASR hypotheses can be gradually reduced. We also investigate the effectiveness of carrying over contextual information (previous rescoring results) across a lattice sequence of a long speech such as a lecture speech. In experiments using a lecture speech corpus, by combining the eight NLMs and using context carry-over, we obtained a 24.4% relative word error rate reduction from the ASR 1-best baseline. For further comparison, we performed simultaneous (i.e., non-iterative) NLM combination and 100-best rescoring using the large ensemble of NLMs, which confirmed the advantage of lattice rescoring with iterative NLM combination.
Atsunori Ogawa, Naohiro Tawara, Marc Delcroix, Shoko Araki
ICASSP1
2022 End-to-End Spontaneous Speech Recognition Using Disfluency Labeling
Koharu Horii, Meiko Fukuda, Kengo Ohta, Ryota Nishimura, Atsunori Ogawa, Norihide Kitaoka
INTERSPEECH5
2021 Attention-Based Multi-Hypothesis Fusion for Speech Summarization
abstract
Speech summarization, which generates a text summary from speech, can be achieved by combining automatic speech recognition (ASR) and text summarization (TS). With this cascade approach, we can exploit state-of-the-art models and large training datasets for both subtasks, i.e., Transformer for ASR and Bidirectional Encoder Representations from Transformers (BERT) for TS. However, ASR errors directly affect the quality of the output summary in the cascade approach. We propose a cascade speech summarization model that is robust to ASR errors and that exploits multiple hypotheses generated by ASR to attenuate the effect of ASR errors on the summary. We investigate several schemes to combine ASR hypotheses. First, we propose using the sum of sub-word embedding vectors weighted by their posterior values provided by an ASR system as an input to a BERT-based TS system. Then, we introduce a more general scheme that uses an attention-based fusion module added to a pre-trained BERT module to align and combine several ASR hypotheses. Finally, we perform speech summarization experiments on the How2 dataset and a newly assembled TED-based dataset that we will release with this paper11https://github.com/nttcslab-sp-admin/TEDSummary. These experiments show that retraining the BERT-based TS system with these schemes can improve summarization performance and that the attention-based fusion module is particularly effective.
Takatomo Kano, Atsunori Ogawa, Marc Delcroix, Shinji Watanabe 0001
ASRU2
2021 Robust Speech-Age Estimation Using Local Maximum Mean Discrepancy Under Mismatched Recording Conditions
abstract
A recently proposed time-delay neural network (TDNN)-based age estimation system has yielded state-of-the-art performance in speech-age estimation tasks. However, the performance of this TDNN-based system can seriously degrade when the recording conditions of each utterance are different in the training and testing phases. To tackle this problem, we examine the efficiencies of a series of unsupervised domain adaptation (UDA) methods to obtain the model invariance against the difference of these conditions. In particular, we propose using local maximum mean discrepancy (LMMD) with soft-target labels to consider an ordinal relationship between age labels. In most UDA methods, the model is trained to obtain domain invariant representations by minimizing the statistical difference of the distributions between labeled source and unlabeled target data without considering their age class labels. In contrast, our LMMD-based approach locally minimizes the differences in their distributions on each age class while considering adjacent age classes using soft-target labels. We conducted speech-age estimation experiments on in-house datasets under mismatched conditions including different background noise, reverberation, and microphones. The experimental comparison demonstrated that the LMMD-based method contributed to efficiently reducing the effect of mismatches of input data, yielding significant improvements over other UDA methods, such as MMD and reverse gradients.
Naohiro Tawara, Atsunori Ogawa, Yuki Kitagishi, Hosana Kamiyama, Yusuke Ijima
ASRU2
2021 BLSTM-Based Confidence Estimation for End-to-End Speech Recognition
abstract
Confidence estimation, in which we estimate the reliability of each recognized token (e.g., word, sub-word, and character) in automatic speech recognition (ASR) hypotheses and detect incorrectly recognized tokens, is an important function for developing ASR applications. In this study, we perform confidence estimation for end-to-end (E2E) ASR hypotheses. Recent E2E ASR systems show high performance (e.g., around 5% token error rates) for various ASR tasks. In such situations, confidence estimation becomes difficult since we need to detect infrequent incorrect tokens from mostly correct token sequences. To tackle this imbalanced dataset problem, we employ a bidirectional long short-term memory (BLSTM)-based model as a strong binary-class (correct/incorrect) sequence labeler that is trained with a class balancing objective. We experimentally confirmed that, by utilizing several types of ASR decoding scores as its auxiliary features, the model steadily shows high confidence estimation performance under highly imbalanced settings. We also confirmed that the BLSTM-based model outperforms Transformer-based confidence estimation models, which greatly underestimate incorrect tokens.
Atsunori Ogawa, Naohiro Tawara, Takatomo Kano, Marc Delcroix
ICASSP1
2021 Age-VOX-Celeb: Multi-Modal Corpus for Facial and Speech Estimation
abstract
Estimating a speaker’s age from their speech is more challenging than age estimation from their face because of insufficiently available public corpora. To tackle this problem, we construct a new audio-visual age corpus named AgeVoxCeleb by annotating age labels to VoxCeleb2 videos. AgeVoxCeleb is the first large-scale, balanced, and multi-modal age corpus that contains both video and speech of the same speakers from a wide age range. Using AgeVox-Celeb, our paper makes the following contributions: (i) A facial age estimation model can outperform a speech age estimation model by comparing the state-of-the-art models in each task. (ii) Facial age estimation is more robust against the difference between training and test sets. (iii) We developed cross-modal transfer learning from face to speech age estimation, showing that the estimated age with a facial age estimation model can be used to train a speech age estimation model. Proposed AgeVoxCeleb will be published in https://github.com/nttcslab-sp/agevoxceleb.
Naohiro Tawara, Atsunori Ogawa, Yuki Kitagishi, Hosana Kamiyama
ICASSP2
2021 Comparison of Remote Experiments Using Crowdsourcing and Laboratory Experiments on Speech Intelligibility
abstract
Many subjective experiments have been performed to develop objective speech intelligibility measures, but the novel coronavirus outbreak has made it very difficult to conduct experiments in a laboratory. One solution is to perform remote testing using crowdsourcing; however, because we cannot control the listening conditions, it is unclear whether the results are entirely reliable. In this study, we compared speech intelligibility scores obtained in remote and laboratory experiments. The results showed that the mean and standard deviation (SD) of the remote experiments' speech reception threshold (SRT) were higher than those of the laboratory experiments. However, the variance in the SRTs across the speech-enhancement conditions revealed similarities, implying that remote testing results may be as useful as laboratory experiments to develop an objective measure. We also show that the practice session scores correlate with the SRT values. This is a priori information before performing the main tests and would be useful for data screening to reduce the variability of the SRT distribution.
Ayako Yamamoto, Toshio Irino, Kenichi Arai, Shoko Araki, Atsunori Ogawa, Keisuke Kinoshita, Tomohiro Nakatani
Interspeech5
2020 Improving Speaker-Attribute Estimation by Voting Based on Speaker Cluster Information
abstract
This paper proposes a general post-processing method for improving speaker-attribute estimation. Estimating speaker-specific attributes such as age and gender is an important task with a wide range of applications. While the recent proposed deep neural network-based end-to-end approach achieves high performance, the model tends to over-fit to specific speakers when the amount of training data is limited or imbalanced. To solve this over-fitting problem, we propose a general framework for correcting unreliable results. The proposed algorithm first clusters the target utterances into speaker clusters by speaker similarity based on i-vectors. Then, for each of the speaker cluster, the speaker-attribute class of the cluster is determined by voting on the utterances assigned to the cluster. By then replacing the result of each utterance with the clusters' speaker-attribute class, we can correct the result of unreliable utterances. We used two tasks to evaluate the proposed algorithm including age estimation using the NIST-SRE10 and age-gender classification using an in-house read speech corpus, yielding significant improvements in mean absolute and classification errors.
Naohiro Tawara, Hosana Kamiyama, Satoshi Kobashikawa, Atsunori Ogawa
ICASSP4
2020 Frame-Level Phoneme-Invariant Speaker Embedding for Text-Independent Speaker Recognition on Extremely Short Utterances
abstract
This paper investigates a phoneme-invariant speaker embedding approach for speaker recognition on extremely short utterances. Intuitively, phonemes are nuisance information for text-independent speaker recognition task since the contents of the speech are usually mismatched between enrolling and testing time. However, many studies have shown that incorporating phoneme information is quite effective to improve the performance of the speaker recognition system. One reasonable explanation for this counter-intuitive result is that the pooling mechanism of segment-based speaker embedding can focus on the specific phonemes which contain rich speaker information, and phoneme information may help this. From this insight, we hypothesize that the pooling mechanism and phoneme-aware training are harmful to extract the speaker embeddings from extremely short utterances. To verify this hypothesis, an adversarial framework is introduced to remove phoneme-variability from the frame-wise speaker embeddings. The experimental results on the Librispeech corpus confirm that our frame-wise, phoneme-adversarial approach outperforms the conventional segment-wise, phoneme-aware approach for short utterances of less than about 1.4 seconds.
Naohiro Tawara, Atsunori Ogawa, Tomoharu Iwata, Marc Delcroix, Tetsuji Ogawa
ICASSP2
2020 Predicting Intelligibility of Enhanced Speech Using Posteriors Derived from DNN-Based ASR System
Kenichi Arai, Shoko Araki, Atsunori Ogawa, Keisuke Kinoshita, Tomohiro Nakatani, Toshio Irino
INTERSPEECH3
2020 Language Model Data Augmentation Based on Text Domain Transfer
Atsunori Ogawa, Naohiro Tawara, Marc Delcroix
INTERSPEECH1
2019 A Unified Framework for Feature-based Domain Adaptation of Neural Network Language Models
abstract
An important task for language models is the adaptation of general-domain models to specific target domains. For neural network-based language models, feature-based domain adaptation has been a popular method in previous research. Conventional methods use an adaptation feature providing context information that is calculated from a topic model. However, such a topic model needs to be trained separately from the language model. To unify the language and context model training, we present an approach that combines an extractor network and a domain adaptation layer. The extractor network learns a context representation from a fixed-size window of past words and provides the context information for the adaptation layer. The benefit of our method is that the extractor network can be trained jointly with the language model in a single training step. Our proposed method showed superior performance over conventional domain adaptation with topic features on a dataset of TED talks with respect to perplexity and word error rate after 100-best rescoring.
Michael Hentschel, Marc Delcroix, Atsunori Ogawa, Tomoharu Iwata, Tomohiro Nakatani
ICASSP3
2019 Semi-supervised End-to-end Speech Recognition Using Text-to-speech and Autoencoders
abstract
We introduce speech and text autoencoders that share encoders and decoders with an automatic speech recognition (ASR) model to improve ASR performance with large speech only and text only training datasets. To build the speech and text autoencoders, we leverage state-of-the-art ASR and text-to-speech (TTS) encoder decoder architectures. These autoencoders learn features from speech only and text only datasets by switching the encoders and decoders used in the ASR and TTS models. Simultaneously, they aim to encode features to be compatible with ASR and TTS models by a multi-task loss. Additionally, we anticipate that TTS joint training can also improve the ASR performance because both ASR and TTS models learn transformations between speech and text. The experimental result we obtained with our semi-supervised end-to-end ASR/TTS training revealed reductions from a model initially trained with a small paired subset of the LibriSpeech corpus in the character error rate from 10.4% to 8.4% and word error rate from 20.6% to 18.0% by retraining the model with a large unpaired subset of the corpus.
Shigeki Karita, Shinji Watanabe 0001, Tomoharu Iwata, Marc Delcroix, Atsunori Ogawa, Tomohiro Nakatani
ICASSP5
2019 A Unified Framework for Neural Speech Separation and Extraction
abstract
The development of deep learning techniques has triggered the active investigation of neural network-based speech enhancement approaches. In particular, single-channel blind (uninformed) speech separation and speaker-aware (informed) speech extraction have received increased interest. Blind speech separation separates a speech mixture into all source signals without requiring any auxiliary information about the speakers. In contrast, speaker-aware speech extraction focuses on extracting speech from a target speaker using prior knowledge, such as an utterance spoken by the target speaker. Speaker extraction is therefore not fully blind, but it can mitigate the source permutation problem faced by blind source separation, and potentially achieve better speech quality by exploiting the auxiliary information. In this paper, to take advantage of both approaches, we propose a unified framework for both speech separation and speech extraction using a single model. This is realized by incorporating a speaker attention mechanism within a generalized permutation invariant training (PIT)-based blind speech separation model, and introducing a multitask separation/extraction objective for training the model. Experiments on the WSJ0-2mix dataset show that our proposed framework realizes both uninformed separation and informed extraction, and achieves better separation/extraction performance than a baseline PIT-based model.
Tsubasa Ochiai, Marc Delcroix, Keisuke Kinoshita, Atsunori Ogawa, Tomohiro Nakatani
ICASSP4
2019 ILP-based Compressive Speech Summarization with Content Word Coverage Maximization and Its Oracle Performance Analysis
abstract
We propose an integer linear programming (ILP)-based compressive speech summarization method that maximizes the coverage of content words in a resultant summary. It is an unsupervised method and, under the designed constraints, it performs a single-step globally optimal summarization of a given long speech recording, which is decoded as a confusion network form of an automatic speech recognition (ASR) hypothesis sequence. It selects as many different content words as possible from the speech input that inevitably includes a high level of redundancy (e.g. the repetition of the same word) under a given length constraint. In experiments using a lecture speech corpus, we obtained higher summarization performance in terms of ROUGE scores than with a baseline extractive summarization method. We further conduct experimental analyses to obtain the oracle (upper bound) performance of the summarization methods. The analysis results show that the oracle performance is very high even though the ASR hypotheses include recognition errors. It is significantly higher than the system performance and, in addition, the oracle performance of the compressive method is significantly higher than that of the extractive method. These results confirm that our method is a promising approach.
Atsunori Ogawa, Tsutomu Hirao, Tomohiro Nakatani, Masaaki Nagata
ICASSP1
2019 Predicting Speech Intelligibility of Enhanced Speech Using Phone Accuracy of DNN-Based ASR System
Kenichi Arai, Shoko Araki, Atsunori Ogawa, Keisuke Kinoshita, Tomohiro Nakatani, Katsuhiko Yamamoto, Toshio Irino
INTERSPEECH3
2019 End-to-End SpeakerBeam for Single Channel Target Speech Recognition
Marc Delcroix, Shinji Watanabe 0001, Tsubasa Ochiai, Keisuke Kinoshita, Shigeki Karita, Atsunori Ogawa, Tomohiro Nakatani
INTERSPEECH6
2019 Improving Transformer-Based End-to-End Speech Recognition with Connectionist Temporal Classification and Language Model Integration
Shigeki Karita, Nelson Enrique Yalta Soplin, Shinji Watanabe 0001, Marc Delcroix, Atsunori Ogawa, Tomohiro Nakatani
INTERSPEECH5
2019 Multimodal SpeakerBeam: Single Channel Target Speech Extraction with Audio-Visual Speaker Clues
Tsubasa Ochiai, Marc Delcroix, Keisuke Kinoshita, Atsunori Ogawa, Tomohiro Nakatani
INTERSPEECH4
2019 Improved Deep Duel Model for Rescoring N-Best Speech Recognition List Using Backward LSTMLM and Ensemble Encoders
Atsunori Ogawa, Marc Delcroix, Shigeki Karita, Tomohiro Nakatani
INTERSPEECH1
2018 Single Channel Target Speaker Extraction and Recognition with Speaker Beam
abstract
This paper addresses the problem of single channel speech recognition of a target speaker in a mixture of speech signals. We propose to exploit auxiliary speaker information provided by an adaptation utterance from the target speaker to extract and recognize only that speaker. Using such auxiliary information, we can build a speaker extraction neural network (NN) that is independent of the number of sources in the mixture, and that can track speakers across different utterances, which are two challenging issues occurring with conventional approaches for speech recognition of mixtures. We call such an informed speaker extraction scheme “SpeakerBeam”. SpeakerBeam exploits a recently developed context adaptive deep NN (CADNN) that allows tracking speech from a target speaker using a speaker adaptation layer, whose parameters are adjusted depending on auxiliary features representing the target speaker characteristics. SpeakerBeam was previously investigated for speaker extraction using a microphone array. In this paper, we demonstrate that it is also efficient for single channel speaker extraction. The speaker adaptation layer can be employed either to build a speaker adaptive acoustic model that recognizes only the target speaker or a mask-based speaker extraction network that extracts the target speech from the speech mixture signal prior to recognition. We also show that the latter speaker extraction network can be optimized jointly with an acoustic model to further improve ASR performance.
Marc Delcroix, Katerina Zmolíková, Keisuke Kinoshita, Atsunori Ogawa, Tomohiro Nakatani
ICASSP4
2018 Sequence Training of Encoder-Decoder Model Using Policy Gradient for End-to-End Speech Recognition
abstract
The standard evaluation metric of automatic speech recognition (ASR) is the word error rate (WER), which measures the dissimilarity between recognized word sequences and their ground truth. Many training algorithms designed to reduce sequence-level errors such as WER have been proposed for hidden Markov model (HMM)-based ASR, e.g., state-level minimum Bayes risk (sMBR). However, these approaches cannot be used directly for encoder-decoder model based end-to-end ASR, because the encoder-decoder model employs very different mechanisms from HMM-based approaches. In this paper, we propose a new method for optimizing the encoder-decoder model based on a sequence-level evaluation metric. Since the WER is not directly differentiable, we adopt a policy gradient objective function to train the encoder-decoder model, which enables us to minimize the expected WER of the model predictions. This training method employs the scoring of multiple hypotheses as in the decoding stage while usual cross entropy training uses only the ground truth. Therefore, we can expect it to improve the decoding results of the encoder-decoder model. We perform experiments using the Tedlium corpus to demonstrate the potential of our proposed method for improving the recognition performance of the encoder-decoder model.
Shigeki Karita, Atsunori Ogawa, Marc Delcroix, Tomohiro Nakatani
ICASSP2
2018 Language Model Domain Adaptation Via Recurrent Neural Networks with Domain-Shared and Domain-Specific Representations
abstract
Training recurrent neural network language models (RNNLMs) requires a large amount of data, which is difficult to collect for specific domains such as multiparty conversations. Data augmentation using external resources and model adaptation, which adjusts a model trained on a large amount of data to a target domain, have been proposed for low-resource language modeling. While there are the commonalities and discrepancies between the source and target domains in terms of the statistics of words and their contexts, these methods for domain adaptation make the commonalities and discrepancies jumbled. We propose novel domain adaptation techniques for RNNLM by introducing domain-shared and domain-specific word embedding and contextual features. This explicit modeling of the commonalities and discrepancies would improve the language modeling performance. Experimental comparisons using multiparty conversation data as the target domain augmented by lecture data from the source domain demonstrate that the proposed domain adaptation method exhibits improvements in the perplexity and word error rate over the long short-term memory based language model (LSTMLM) trained using the source and target domain data.
Tsuyoshi Morioka, Naohiro Tawara, Tetsuji Ogawa, Atsunori Ogawa, Tomoharu Iwata, Tetsunori Kobayashi
ICASSP4
2018 Rescoring N-Best Speech Recognition List Based on One-on-One Hypothesis Comparison Using Encoder-Classifier Model
abstract
This paper proposes a new model for accurately rescoring (reranking) N-best speech recognition hypothesis lists. The model is based on state-of-the-art neural networks (NNs) and provides the minimum necessary functionality to perform N-best rescoring, i.e. one-on-one hypothesis comparison on a given N-best list in terms of word error rates (WERs). The model is composed of a long short-term memory (LSTM)-based encoder network followed by a fully-connected feedforward NN-based binary-class classifier network. Given the feature vector sequences of two hypotheses to be compared, this encoder-classifier (EC) model encodes these features and outputs binary-class probabilities that indicate which hypothesis has the lower WER. Then, depending on the output, the ranks of these hypotheses can be swapped. By repeating this one-on-one hypothesis comparison, a reranked N-best list can be obtained. In N-best rescoring experiments using a large scale speech corpus, the proposed EC model steadily outperforms an LSTM-based language model (LSTMLM), which is a strong and widely-used competitor. In addition, by incorporating the LSTMLM scores as an additional feature vector dimension, the N-best rescoring performance of the EC model is further improved. The improved EC model achieves a 10% relative WER reduction from the LSTMLM baseline.
Atsunori Ogawa, Marc Delcroix, Shigeki Karita, Tomohiro Nakatani
ICASSP1
2018 Auxiliary Feature Based Adaptation of End-to-end ASR Systems
Marc Delcroix, Shinji Watanabe 0001, Atsunori Ogawa, Shigeki Karita, Tomohiro Nakatani
INTERSPEECH3
2018 Semi-Supervised End-to-End Speech Recognition
Shigeki Karita, Shinji Watanabe 0001, Tomoharu Iwata, Atsunori Ogawa, Marc Delcroix
INTERSPEECH4
2018 Context Adaptive Neural Network Based Acoustic Models for Rapid Adaptation
abstract
The adaptation of automatic speech recognition systems to a speaker or an environment is important if we are to achieve high speech recognition performance ubiquitously. Recently, deep neural network (DNN) based acoustic models have been made adaptive to speakers or environments by the addition of an auxiliary feature representing the acoustic context information such as speaker or noise characteristics to the network input. The addition of such auxiliary features to the input realizes only the adaptation of the bias term of the input layer. In this paper, we introduce “context adaptive neural networks,” which are an alternative approach for exploiting auxiliary features that can achieve adaptation of all the parameters of a layer including the linear transformation matrices and the bias terms. A context adaptive neural network is a neural network with one of its layers factorized into sublayers, each associated with an acoustic context class representing a class of speakers or noise conditions. The output of the factorized layer is obtained as a weighted sum of the contributions of all of the sublayers. The weighting coefficients, or context class weights, are derived from the auxiliary features, by transforming them through an auxiliary network. The auxiliary network and the main network can be trained jointly, which enables the context classes that optimize the training criterion to be learned automatically. We perform experiments on three tasks, i.e., two speaker adaptation experiments using DNN models with medium-sized (Wall Street Journal) and large (Continuous Spontaneous Japanese) training datasets, and one environmental adaptation of a convolutional neural network based acoustic model with CHiME3 data. These experiments confirm the potential of the proposed approach in various settings.
Marc Delcroix, Keisuke Kinoshita, Atsunori Ogawa, Christian Huemmer 0001, Tomohiro Nakatani
IEEE ACM Trans. Audio Speech Lang. Process.3
2017 Learning speaker representation for neural network based multichannel speaker extraction
abstract
Recently, schemes employing deep neural networks (DNNs) for extracting speech from noisy observation have demonstrated great potential for noise robust automatic speech recognition. However, these schemes are not well suited when the interfering noise is another speaker. To enable extracting a target speaker from a mixture of speakers, we have recently proposed to inform the neural network using speaker information extracted from an adaptation utterance from the same speaker. In our previous work, we explored ways how to inform the network about the speaker and found a speaker adaptive layer approach to be suitable for this task. In our experiments, we used speaker features designed for speaker recognition tasks as the additional speaker information, which may not be optimal for the speaker extraction task. In this paper, we propose a usage of a sequence summarizing scheme enabling to learn the speaker representation jointly with the network. Furthermore, we extend the previous experiments to demonstrate the potential of our proposed method as a front-end for speech recognition and explore the effect of additional noise on the performance of the method.
Katerina Zmolíková, Marc Delcroix, Keisuke Kinoshita, Takuya Higuchi, Atsunori Ogawa, Tomohiro Nakatani
ASRU5
2017 Online environmental adaptation of CNN-based acoustic models using spatial diffuseness features
abstract
We propose a new concept for adapting CNN-based acoustic models using spatial diffuseness features as auxiliary information about the acoustic environment: the spatial diffuseness features are simultaneously employed as acoustic-model input features and to estimate environmental cues for context adaptation, where one convolutional layer is factorized into several sub-layers to represent different acoustic conditions. This context-adaptive CNN-based acoustic model facilitates an online environmental adaptation and is experimentally verified for the real-world recordings provided by the CHiME-3 task. The best performing setup reduces the average word error rate scores achieved by the baseline system (without using spatial diffuseness features) from 19.4% to 15.9% and 12.2% to 10.7% considering two experimental setups with and without front-end signal enhancement, respectively.
Christian Huemmer 0001, Marc Delcroix, Atsunori Ogawa, Keisuke Kinoshita, Tomohiro Nakatani, Walter Kellermann
ICASSP3
2017 Deep mixture density network for statistical model-based feature enhancement
abstract
We propose a novel framework designed to extend conventional deep neural network (DNN)-based feature enhancement approaches. In general, the conventional DNN-based feature enhancement framework aims to map input noisy observation to clean speech or a binary/ soft mask in a deterministic way, assuming that there is one-to-one mapping between the input and the output without any uncertainty. However, when we consider that the general feature enhancement problem to be an ill-posed inverse problem where the mapping cannot be uniquely determined given an input signal, the assumption in the conventional approaches is not theoretically correct and potentially limits the performance of DNN-based feature enhancement. To overcome this problem, this paper proposes utilizing a mixture density network (MDN), which is a neural network that maps an input feature to a set of Gaussian mixture model (GMM) parameters representing the distribution of a target variable. By estimating the distribution of clean speech feature based on MDN, we are now able to explicitly consider the uncertainty in the parameter estimation. Then, we further utilizes the estimated GMM to obtain a refined clean speech estimate in the framework of statistical model-based feature enhancement. In this paper, after detailing the proposed framework and the MDN, we show mathematically and experimentally how MDN appropriately models the uncertainty information. We also show that the proposed method can outperform a conventional DNN-based feature enhancement method.
Keisuke Kinoshita, Marc Delcroix, Atsunori Ogawa, Takuya Higuchi, Tomohiro Nakatani
ICASSP3
2017 Cumulative moving averaged bottleneck speaker vectors for online speaker adaptation of CNN-based acoustic models
abstract
Adapting acoustic models to speakers have shown to greatly improve performance for many tasks. Among the adaptation approaches, exploiting auxiliary features characterizing speakers or environments has received great attention because they allow rapid adaptation, i.e. adaptation with limited amount of speech data such as a single utterance. However, the auxiliary features are usually computed in batch mode, which causes some inevitable latency. In this paper we explore an extension of the auxiliary feature-based adaptation to online processing. We employ auxiliary features obtained from bottleneck speaker vectors and extend their computation to online processing using cumulative moving averaging. We test our proposed approach for deep CNN-based acoustic models, using context adaptive networks to exploit the auxiliary features. Experimental results on the CHiME-3 task demonstrate that the proposed approach can realize online speaker adaptation.
Tsubasa Ochiai, Marc Delcroix, Keisuke Kinoshita, Atsunori Ogawa, Taichi Asami, Shigeru Katagiri, Tomohiro Nakatani
ICASSP4
2017 Feedback connection for deep neural network-based acoustic modeling
abstract
The use of auxiliary features is an effective way to improve the performance of deep neural network (DNN)-based acoustic models. Most approaches use auxiliary features that represent the speaker or the environment. These auxiliary features are usually computed independently of the acoustic model. This paper investigates a types of auxiliary features obtained from the output of a hidden layer that feeds back to the input layer of the network. Since the auxiliary features are extracted from the hidden layer of the network no external information is required such as the speaker or the environment. Experimentally, by forcing the extraction of the auxiliary features from the same networks, we can further improve the performance of the overall network and reduce the total number of parameters used. We tested this approach with different deep neural network architectures including: deep neural networks, convolutional neural networks and unfolded recurrent convolutional networks. We confirmed the effectiveness of this approach on the CHiME3 dataset.
Dung T. Tran, Marc Delcroix, Atsunori Ogawa, Christian Huemmer 0001, Tomohiro Nakatani
ICASSP3
2017 Forward-Backward Convolutional LSTM for Acoustic Modeling
Shigeki Karita, Atsunori Ogawa, Marc Delcroix, Tomohiro Nakatani
INTERSPEECH2
2017 Improved Example-Based Speech Enhancement by Using Deep Neural Network Acoustic Model for Noise Robust Example Search
Atsunori Ogawa, Keisuke Kinoshita, Marc Delcroix, Tomohiro Nakatani
INTERSPEECH1
2017 Unfolded Deep Recurrent Convolutional Neural Network with Jump Ahead Connections for Acoustic Modeling
Dung T. Tran, Marc Delcroix, Shigeki Karita, Michael Hentschel, Atsunori Ogawa, Tomohiro Nakatani
INTERSPEECH5
2017 Uncertainty Decoding with Adaptive Sampling for Noise Robust DNN-Based Acoustic Modeling
Dung T. Tran, Marc Delcroix, Atsunori Ogawa, Tomohiro Nakatani
INTERSPEECH3
2017 Speaker-Aware Neural Network Based Beamformer for Speaker Extraction in Speech Mixtures
Katerina Zmolíková, Marc Delcroix, Keisuke Kinoshita, Takuya Higuchi, Atsunori Ogawa, Tomohiro Nakatani
INTERSPEECH5
2017 Error detection and accuracy estimation in automatic speech recognition using deep bidirectional recurrent neural networks
Atsunori Ogawa, Takaaki Hori
Speech Commun.1
2016 Spatial correlation model based observation vector clustering and MVDR beamforming for meeting recognition
abstract
This paper addresses a minimum variance distortionless response (MVDR) beamforming based speech enhancement approach for meeting speech recognition. In a meeting situation, speaker overlaps and noise signals are not negligible. To handle these issues, we employ MVDR beamforming, where accurate estimation of the steering vector is paramount. We recently found that steering vector estimation by clustering the time-frequency components of microphone observation vectors performs well as regards real-world noise reduction. The clustering is performed by taking a cue from the spatial correlation matrix of each speaker, which is realized by modeling the time-frequency components of the observation vectors with a complex Gaussian mixture model (CGMM). Experimental results with real recordings show that the proposed MVDR scheme outperforms conventional null-beamformer based speech enhancement in a meeting situation.
Shoko Araki, Masahiro Okada, Takuya Higuchi, Atsunori Ogawa, Tomohiro Nakatani
ICASSP4
2016 Context adaptive deep neural networks for fast acoustic model adaptation in noisy conditions
abstract
Deep neural network (DNN) based acoustic models have greatly improved the performance of automatic speech recognition (ASR) for various tasks. Further performance improvements have been reported when making DNNs aware of the acoustic context (e.g. speaker or environment) for example by adding auxiliary features to the input, such as noise estimates or speaker i-vectors. We have recently proposed a context adaptive DNN (CA-DNN), which is another approach to exploit the acoustic context information within a DNN. A CA-DNN is a DNN that has one or several factorized layers, i.e. layers that use a different set of parameters to process each acoustic context class. The output of a factorized layer is obtained by the weighted sum over the contribution of the different context classes, given weights over the context classes. In our previous work, the class weights were computed independently of the recognizer. In this paper, we extend our previous work by introducing the joint training of the CA-DNN parameters and the class weights computation. Consequently, the class weights and the associated class definitions can be optimized for ASR. We report experimental results on the AURORA4 noisy speech recognition task showing the potential of our approach for fast unsupervised adaptation.
Marc Delcroix, Keisuke Kinoshita, Chengzhu Yu, Atsunori Ogawa, Takuya Yoshioka, Tomohiro Nakatani
ICASSP4
2016 Context Adaptive Neural Network for Rapid Adaptation of Deep CNN Based Acoustic Models
Marc Delcroix, Keisuke Kinoshita, Atsunori Ogawa, Takuya Yoshioka, Dung T. Tran, Tomohiro Nakatani
INTERSPEECH3
2016 Robust Example Search Using Bottleneck Features for Example-Based Speech Enhancement
Atsunori Ogawa, Shogo Seki, Keisuke Kinoshita, Marc Delcroix, Takuya Yoshioka, Tomohiro Nakatani, Kazuya Takeda
INTERSPEECH1
2016 Factorized Linear Input Network for Acoustic Model Adaptation in Noisy Conditions
Dung T. Tran, Marc Delcroix, Atsunori Ogawa, Tomohiro Nakatani
INTERSPEECH3
2016 Differenced maximum mutual information criterion for robust unsupervised acoustic model adaptation
Marc Delcroix, Atsunori Ogawa, Seong-Jun Hahm, Tomohiro Nakatani, Atsushi Nakamura
Comput. Speech Lang.2
2016 Estimating Speech Recognition Accuracy Based on Error Type Classification
abstract
Methods for estimating the speech recognition accuracy without using manually transcribed references are beneficial to the research and development of automatic speech recognition technology. This paper proposes recognition accuracy estimation methods based on error type classification (ETC). ETC is an extension of confidence estimation. In ETC, each word in the recognition results (recognized word sequences) for the target speech data is probabilistically classified into three categories: the correct recognition (C), substitution error (S), and insertion error (I). Deletion errors (D) that can occur at interword positions in the recognition results are also probabilistically detected. By summing these CSID probabilities individually, the numbers of CSIDs and, as a result, the two standard recognition accuracy measures, i.e., the percent correct and word accuracy (WAcc), for the speech data can be estimated without using the reference transcriptions. Two recognition accuracy estimation methods based on ETC are proposed. In the first easy-to-use method, ETC is performed by converting the recognition results represented as word confusion networks into word alignment networks (WANs). In the second and more accurate method, the WAN-based ETC results are refined with conditional random fields (CRFs) using various types of additional features extracted for each of the recognized words. Experiments using English and Japanese lecture speech corpora show that the recognition accuracy can be accurately estimated with the CRF-based method. The correlation coefficient and root mean square error between the lecture-level true WAccs calculated using the reference transcriptions and those estimated with the CRF-based method are 0.97 and lower than 2%, respectively. A series of additional experiments and analyses are also conducted to better understand the effectiveness of the CRF-based method.
Atsunori Ogawa, Takaaki Hori, Atsushi Nakamura
IEEE ACM Trans. Audio Speech Lang. Process.1
2015 The NTT CHiME-3 system: Advances in speech enhancement and recognition for mobile multi-microphone devices
abstract
CHiME-3 is a research community challenge organised in 2015 to evaluate speech recognition systems for mobile multi-microphone devices used in noisy daily environments. This paper describes NTT's CHiME-3 system, which integrates advanced speech enhancement and recognition techniques. Newly developed techniques include the use of spectral masks for acoustic beam-steering vector estimation and acoustic modelling with deep convolutional neural networks based on the "network in network" concept. In addition to these improvements, our system has several key differences from the official baseline system. The differences include multi-microphone training, dereverberation, and cross adaptation of neural networks with different architectures. The impacts that these techniques have on recognition performance are investigated. By combining these advanced techniques, our system achieves a 3.45% development error rate and a 5.83% evaluation error rate. Three simpler systems are also developed to perform evaluations with constrained set-ups.
Takuya Yoshioka, Nobutaka Ito, Marc Delcroix, Atsunori Ogawa, Keisuke Kinoshita, Masakiyo Fujimoto, Chengzhu Yu, Wojciech J. Fabian, Miquel Espi, Takuya Higuchi, Shoko Araki, Tomohiro Nakatani
ASRU4
2015 Double-layer neighborhood graph based similarity search for fast query-by-example spoken term detection
abstract
This paper presents a novel double-layer neighborhood graph index for acceleration of similarity search that accomplishes fast querybyexample spoken term detection (STD). When a query segment is given, our proposed STD method finds similar segments to the query from an utterance data set by efficient similarity search that traverses the double-layer neighborhood graph (DLG) with a low computational cost. The segment is a sequence of Gaussian mixture model posteriorgram frames and corresponds to a vertex in the DLG. A dissimilarity between vertices is measured by dynamic time warping. The DLG consists of two distinct degree-reduced k-nearest neighbor graphs in a base and an upper layer. The base layer's graph has all the vertices in the data set while the upper layer's graph includes only representatives extracted from the vertices in the base layer. By way of analogy, search in the DLG resembles driving on general roads and express highways appropriately for travel-time saving. Experimental results on the MIT lecture corpus demonstrate that the proposed method achieves CPU time reduction by 40% and more than 60% compared to the most recent method and the ordinary graphbased method, keeping almost the same precision.
Kazuo Aoyama, Atsunori Ogawa, Takashi Hattori, Takaaki Hori
ICASSP2
2015 ASR error detection and recognition rate estimation using deep bidirectional recurrent neural networks
abstract
Recurrent neural networks (RNNs) have recently been applied as the classifiers for sequential labeling problems. In this paper, deep bidirectional RNNs (DBRNNs) are applied for the first time to error detection in automatic speech recognition (ASR), which is a sequential labeling problem. We investigate three types of ASR error detection tasks, i.e. confidence estimation, out-of-vocabulary word detection and error type classification. We also estimate recognition rates from the error type classification results. Experimental results show that the DBRNNs greatly outperform conditional random fields (CRFs), especially for the detection of infrequent error labels. The DBRNNs also slightly outperform the CRFs in recognition rate estimation. In addition, experiments using a reduced size of training data suggest that the DBRNNs have a better generalization ability than the CRFs owing to their word vector representation in a low-dimensional continuous space. As a result, the DBRNNs trained using only 20% of the training data show higher error detection performance than the CRFs trained using the full training data.
Atsunori Ogawa, Takaaki Hori
ICASSP1
2015 Text-informed speech enhancement with deep neural networks
abstract
A speech signal captured by a distant microphone is generally contaminated by background noise, which severely degrades the audible quality and intelligibility of the observed signal. To resolve this issue, speech enhancement has been intensively studied. In this paper, we consider a text-informed speech enhancement, where the enhancement process is guided by the corresponding text information, i.e., a correct transcription of the target utterance. The proposed deep neural network (DNN)based framework is motivated by the recent success in the textto-speech (TTS) research in employing DNN as well as high audible-quality output signal of the corpus-based speech enhancement which borrows knowledge from the TTS research field. Taking advantage of the nature of DNN that allows us to utilize disparate features in an inference stage, the proposed method infers the clean speech features by jointly using the observed signal and widely-used TTS features derived from the corresponding text. In this paper, we first introduce the background and the details of the proposed method. Then, we show how the text information can be naturally integrated into speech enhancement by utilizing DNN and improve the enhancement performance. Index Terms: speech enhancement, text-to-speech, deep neural network
Keisuke Kinoshita, Marc Delcroix, Atsunori Ogawa, Tomohiro Nakatani
INTERSPEECH3
2015 Robust i-vector extraction for neural network adaptation in noisy environment
Chengzhu Yu, Atsunori Ogawa, Marc Delcroix, Takuya Yoshioka, Tomohiro Nakatani, John H. L. Hansen
INTERSPEECH2
2014 Zero-resource spoken term detection using hierarchical graph-based similarity search
abstract
This paper presents fast zero-resource spoken term detection (STD) in a large-scale data set, by using a hierarchical graph-based similarity search method (HGSS). HGSS is an improved graph-based similarity search method (GSS) in terms of a search space for high-speed performance. Instead of a degree-reduced k-nearest neighbor (k-DR) graph for GSS, a hierarchical k-DR graph, which is constructed based on a cluster structure in the corresponding k-DR graph, is used as an index for HGSS. A search algorithm for the hierarchical k-DR graph effectively utilizes the cluster structure, resulting in the reduction of the search space. HGSS inherits the useful property of GSS; it is available for any data sets without limits on a data type nor a defined dissimilarity since a graph is a general expression of a relationship between objects. A vertex and an edge in the hierarchical graph correspond to a Gaussian mixture model (GMM) posterior-gram segment and the relationship between a pair of GMM poste-riorgram segments, which is measured by dynamic time warping, respectively. Experimental results demonstrate that HGSS successfully reduces the computational cost by more than 40 % at nearly the same accuracy, compared to GSS.
Kazuo Aoyama, Atsunori Ogawa, Takashi Hattori, Takaaki Hori, Atsushi Nakamura
ICASSP2
2014 Fast segment search for corpus-based speech enhancement based on speech recognition technology
abstract
Corpus-based speech enhancement has received increasing attention recently since it shows high enhancement performance in highly non-stationary noisy environments by precisely modeling the long-term temporal dynamics of speech. However, it has a disadvantage in that the cost is very high for searching the longest matching clean speech segments from a multi-condition parallel speech corpus. This paper proposes a fast segment search method for corpus-based speech enhancement. It is mainly based on two techniques derived from speech recognition technology. The first is an A* search like segment evaluation function for accurately finding the longest matching segments. The second is a tree and linear connected search space for efficiently sharing the segment likelihood calculations. In the experiments for non-stationary noisy observations using the 26 multi-condition TIMIT parallel speech corpus, the proposed search method found the segments almost in real-time without degrading the quality of the enhanced speech. Our method was about 7 to 13 times faster than the conventional segment search method.
Atsunori Ogawa, Keisuke Kinoshita, Takaaki Hori, Tomohiro Nakatani, Atsushi Nakamura
ICASSP1
2013 Graph index based query-by-example search on a large speech data set
abstract
This paper presents a neighborhood graph index approach for query-by-example search using dynamic time warping (DTW) on Gaussian mixture model (GMM) posteriorgram sequences. The approach is intended to achieve a significant speed-up of a spoken term detection (STD) task for resource-limited situations. The proposed method employs a degree-reduced k-nearest neighbor (k-DR) graph as an index. A set of k-DR graphs is pre-constructed off-line from a large number of GMM posteriorgram sequences. Given a query posteriorgram sequence, one k-DR graph is selected from the set as the index. By applying a newly introduced combination of greedy-search (GS) and breadth-first search (BFS) algorithms to the selected k-DR graph index, the proposed method efficiently achieves query-by-example STD. Experimental results on the MIT lecture corpus demonstrate that the proposed method works much faster than a state-of-art method by more than one order magnitude, keeping almost the same precision.
Kazuo Aoyama, Atsunori Ogawa, Takashi Hattori, Takaaki Hori, Atsushi Nakamura
ICASSP2
2013 Unsupervised discriminative adaptation using differenced maximum mutual information based linear regression
abstract
This paper proposes a new approach for unsupervised model adaptation using a discriminative criterion. Discriminative criteria for acoustic model training have been widely used and have provided significantly improved performance compared with models trained using maximum likelihood. However, discriminative criteria are sensitive to errors in reference transcriptions, which limits their applicability to unsupervised adaptation. In this paper, we apply the recently proposed differenced maximum mutual information (dMMI) criteria to unsupervised linear regression based adaptation because dMMI has an intrinsic mechanism that mitigates the influence of transcription errors. We report unsupervised adaptation results for a large vocabulary continuous speech recognition task showing a significant improvement over maximum likelihood based linear regression.
Marc Delcroix, Atsunori Ogawa, Seong-Jun Hahm, Tomohiro Nakatani, Atsushi Nakamura
ICASSP2
2013 Feature space variational Bayesian linear regression and its combination with model space VBLR
abstract
In this paper, we propose a tuning-free Bayesian linear regression approach for speaker adaptation. We first formulate feature space variational Bayesian linear regression (fVBLR). Using a lower bound as the objective function, we can optimize a binary tree structure and control parameters for prior density scaling. We experimentally verified the proposed fVBLR could achieve performance comparable to that of the conventional fine-tuned fSMAPLR and SMAPLR. For further performance improvement regardless of the amount of adaptation data, we combine fVBLR with model space VBLR (fVBLR+VBLR). Therefore, feature space normalization and model space adaptation are consistently performed based on a variational Bayesian approach without any tuning parameters. In the experiment, the proposed fVBLR+VBLR showed performance improvement compared with both fVBLR and VBLR.
Seong-Jun Hahm, Atsunori Ogawa, Marc Delcroix, Masakiyo Fujimoto, Takaaki Hori, Atsushi Nakamura
ICASSP2
2013 Coupling beamforming with spatial and spectral feature based spectral enhancement and its application to meeting recognition
abstract
This paper discusses microphone array based interference reduction approaches for robust automatic speech recognition. A model based multichannel spectral enhancement approach has recently been proposed for effectively reducing interference by exploiting both the spatial and spectral features of the signals. With the goal of further improving the effectiveness of this approach, we propose a new framework that combines this approach with a microphone-array based beamforming approach. Because the two approaches can work in a complementary manner in the proposed framework, they can greatly improve the interference reduction performance. We apply the proposed framework to the recognition of actual meetings, and show that it is superior to the use of beamforming or spectral enhancement alone in terms of the word error rates.
Tomohiro Nakatani, Mehrez Souden, Shoko Araki, Takuya Yoshioka, Takaaki Hori, Atsunori Ogawa
ICASSP6
2013 Discriminative recognition rate estimation for N-best list and its application to N-best rescoring
abstract
Techniques for estimating recognition rates without using reference transcriptions are essential if we are to judge whether or not speech recognition technology is applicable to a new task. We have proposed a discriminative recognition rate estimation (DRRE) method for 1-best recognition hypotheses and shown its good estimation performance experimentally. In this paper, we extend our DRRE to N-best lists of recognition hypotheses by modifying its feature extraction procedures and efficiently selecting N-best hypotheses for its discriminative model training. In addition, we apply our extended DRRE to N-best rescoring. In the experiments, the extended DRRE also showed good estimation performance for the N-best lists. And using the estimated recognition rates, the 1-best word accuracy was significantly improved by N-best rescoring from the baseline.
Atsunori Ogawa, Takaaki Hori, Atsushi Nakamura
ICASSP1
2013 Unsupervised discriminative language modeling using error rate estimator
Takanobu Oba, Atsunori Ogawa, Takaaki Hori, Hirokazu Masataki, Atsushi Nakamura
INTERSPEECH2
2013 Speech recognition in living rooms: Integrated speech enhancement and recognition system based on spatial, spectral and temporal modeling of sounds
Marc Delcroix, Keisuke Kinoshita, Tomohiro Nakatani, Shoko Araki, Atsunori Ogawa, Takaaki Hori, Shinji Watanabe 0001, Masakiyo Fujimoto, Takuya Yoshioka, Takanobu Oba, Yotaro Kubo, Mehrez Souden, Seong-Jun Hahm, Atsushi Nakamura
Comput. Speech Lang.5
2013 Fast unsupervised adaptation based on efficient statistics accumulation using frame independent confidence within monophone states
Satoshi Kobashikawa, Atsunori Ogawa, Taichi Asami, Yoshikazu Yamaguchi, Hirokazu Masataki, Satoshi Takahashi
Comput. Speech Lang.2
2013 Prior-shared feature and model space speaker adaptation by consistently employing map estimation
Seong-Jun Hahm, Shinji Watanabe 0001, Atsunori Ogawa, Masakiyo Fujimoto, Takaaki Hori, Atsushi Nakamura
Speech Commun.3
2012 Discriminative feature transforms using differenced maximum mutual information
abstract
Recently feature compensation techniques that train feature transforms using a discriminative criterion have attracted much interest in the speech recognition community. Typically, the acoustic feature space is modeled by a Gaussian mixture model (GMM), and a feature transform is assigned to each Gaussian of the GMM. Feature compensation is then performed by transforming features using the transformation associated with each Gaussian, then summing up the transformed features weighted by the posterior probability of each Gaussian. Several discriminative criteria have been investigated for estimating the feature transformation parameters including maximum mutual information (MMI) and minimum phone error (MPE). Recently, the differenced MMI (dMMI) criterion that generalizes MMI andMPE, has been shown to provide competitive performance for acoustic model training. In this paper, we investigate the use of the dMMI criterion for discriminative feature transforms and demonstrate in a noisy speech recognition experiment that dMMI achieves recognition performance superior to that of MMI or MPE.
Marc Delcroix, Atsunori Ogawa, Shinji Watanabe 0001, Tomohiro Nakatani, Atsushi Nakamura
ICASSP2
2012 Error type classification and word accuracy estimation using alignment features from word confusion network
abstract
This paper addresses error type classification in continuous speech recognition (CSR). In CSR, errors are classified into three types, namely, the substitution, insertion and deletion errors, by making an alignment between a recognized word sequence and its reference transcription with a dynamic programming (DP) procedure. We propose a method for deriving such alignment features from a word confusion network (WCN) without using the reference transcription. We show experimentally that the WCN-based alignment features steadily improve the performance of error type classification. They also improve the performance of out-of-vocabulary (OOV) word detection, since OOV word utterances are highly correlated with a particular alignment pattern. In addition, we show that the word accuracy can be estimated from the WCN-based alignment features and more accurately from the error type classification result without using the reference transcription.
Atsunori Ogawa, Takaaki Hori, Atsushi Nakamura
ICASSP1
2012 Speaker Adaptation Using Variational Bayesian Linear Regression in Normalized Feature Space
Seong-Jun Hahm, Atsunori Ogawa, Masakiyo Fujimoto, Takaaki Hori, Atsushi Nakamura
INTERSPEECH2
2012 Automatic Vocabulary Adaptation Based on Semantic Similarity and Speech Recognition Confidence Measure
Shoko Yamahata, Yoshikazu Yamaguchi, Atsunori Ogawa, Hirokazu Masataki, Osamu Yoshioka, Satoshi Takahashi
INTERSPEECH3
2012 Recognition rate estimation based on word alignment network and discriminative error type classification
abstract
Techniques for estimating recognition rates without using reference transcriptions are essential if we are to judge whether or not speech recognition technology is applicable to a new task. This paper proposes two recognition rate estimation methods for continuous speech recognition. The first is an easy-to-use method based on a word alignment network (WAN) obtained from a word confusion network through simple conversion procedures. A WAN contains the correct (C), substitution error (S), insertion error (I) and deletion error (D) probabilities word-by-word for a recognition result. By summing these CSID probabilities individually, the percent correct and word accuracy (WACC) can be estimated without using a reference transcription. The second more advanced method refines the CSID probabilities provided by a WAN based on discriminative error type classification (ETC) and estimates the recognition rates more accurately. In the experiments on the MIT lecture speech corpus, we obtained 0.97 of correlation coefficient between the true WACCs calculated by a scoring tool using reference transcriptions and the WACCs estimated from the discriminative ETC results.
Atsunori Ogawa, Takaaki Hori, Atsushi Nakamura
SLT1
2012 Joint estimation of confidence and error causes in speech recognition
Atsunori Ogawa, Atsushi Nakamura
Speech Commun.1
2012 Low-Latency Real-Time Meeting Recognition and Understanding Using Distant Microphones and Omni-Directional Camera
abstract
This paper presents our real-time meeting analyzer for monitoring conversations in an ongoing group meeting. The goal of the system is to recognize automatically “who is speaking what” in an online manner for meeting assistance. Our system continuously captures the utterances and face poses of each speaker using a microphone array and an omni-directional camera positioned at the center of the meeting table. Through a series of advanced audio processing operations, an overlapping speech signal is enhanced and the components are separated into individual speaker's channels. Then the utterances are sequentially transcribed by our speech recognizer with low latency. In parallel with speech recognition, the activity of each participant (e.g., speaking, laughing, watching someone) and the circumstances of the meeting (e.g., topic, activeness, casualness) are detected and displayed on a browser together with the transcripts. In this paper, we describe our techniques and our attempt to achieve the low-latency monitoring of meetings, and we show our experimental results for real-time meeting transcription.
Takaaki Hori, Shoko Araki, Takuya Yoshioka, Masakiyo Fujimoto, Shinji Watanabe 0001, Takanobu Oba, Atsunori Ogawa, Kazuhiro Otsuka, Dan Mikami, Keisuke Kinoshita, Tomohiro Nakatani, Atsushi Nakamura, Junji Yamato
IEEE Trans. Speech Audio Process.7
2011 Machine and acoustical condition dependency analyses for fast acoustic likelihood calculation techniques
abstract
The acceleration of acoustic likelihood calculation has been an important research issue for developing practical speech recognition systems. And there are various specification machines and various acoustical conditions in the fields to which speech recognition is applied. In this paper, we reveal the machine and acoustical condition dependencies of fast acoustic likelihood calculation techniques. We employed state likelihood recycling as an approximation technique, batch state likelihood calculation as a technique based on computer architecture, and their combinations with or without acoustic backing-off that were our previously proposed efficient techniques. We evaluated and analyzed these four techniques in large vocabulary continuous speech recognition experiments by using four machines with different types of CPUs (Intel Pentium 4, Xeon, Core 2 Duo and Xeon X5570) under two acoustical conditions (clean and noisy). The combined technique with acoustic backing-off exhibited the best acceleration performance while preventing word accuracy degradation under all of the experimental conditions. The experimental and analytical results obtained in this paper are informative especially for developing speech recognition systems that are used in the fields.
Atsunori Ogawa, Satoshi Takahashi, Atsushi Nakamura
ICASSP1
2010 Discriminative confidence and error cause estimation for extended speech recognition function
abstract
Errors are unavoidable in speech recognition, and so confidence estimation, which scores the reliability of recognition results, plays a critical role in this procedure. If we are to develop speech recognition systems capable of practical use, in addition to achieving accurate confidence estimation, we will need to extend the functions of speech recognition engines. As the first step towards extending these functions, we have proposed a method that estimates the causes of recognition errors while simultaneously estimating the confidence of recognition results using a discriminative model, and shown its potential experimentally. In this paper, we modify our previously proposed method by dividing its simultaneous confidence and error cause estimation procedure into two separate procedures. In the speech recognition experiments, the separate estimation methods achieved the same confidence estimation accuracy as the simultaneous method but their error cause estimation accuracies were superior.
Atsunori Ogawa, Atsushi Nakamura
ICASSP1
2010 A novel confidence measure based on marginalization of jointly estimated error cause probabilities
Atsunori Ogawa, Atsushi Nakamura
INTERSPEECH1
2010 Real-time meeting recognition and understanding using distant microphones and omni-directional camera
abstract
This paper presents our newly developed real-time meeting analyzer for monitoring conversations in an ongoing group meeting. The goal of the system is to automatically recognize “who is speaking what” in an online manner for meeting assistance. Our system continuously captures the utterances and the face pose of each speaker using a distant microphone array and an omni-directional camera at the center of the meeting table. Through a series of advanced audio processing operations, an overlapping speech signal is enhanced and the components are separated into individual speaker's channels. Then the utterances are sequentially transcribed by our speech recognizer with low latency. In parallel with speech recognition, the activity of each participant (e.g. speaking, laughing, watching someone) and the situation of the meeting (e.g. topic, activeness, casualness) are detected and displayed on a browser together with the transcripts. In this paper, we describe our techniques and our attempt to achieve the low-latency monitoring of meetings, and we show our experimental results for real-time meeting transcription.
Takaaki Hori, Shoko Araki, Takuya Yoshioka, Masakiyo Fujimoto, Shinji Watanabe 0001, Takanobu Oba, Atsunori Ogawa, Kazuhiro Otsuka, Dan Mikami, Keisuke Kinoshita, Tomohiro Nakatani, Atsushi Nakamura, Junji Yamato
SLT7
2009 Efficient combination of likelihood recycling and batch calculation based on conditional fast processing and acoustic back-off
abstract
This paper proposes an efficient combination of state likelihood recycling and batch state likelihood calculation for accelerating acoustic likelihood calculation in an HMM-based speech recognizer. Recycling and batch calculation are each based on different technical approaches, i.e. the former is a purely algorithmic technique while the latter fully exploits PC architecture, and their good acceleration performances are reported in the literatures, respectively. To accelerate the recognition process further by combining them efficiently, we introduce conditional fast processing and acoustic back-off strategies. Our combination algorithm employs the conditional fast processing strategy that is conditioned by two criteria. The first potential activity criterion is used to control not only the recycling of state likelihoods at the current frame but also the precalculation of state likelihoods for several succeeding frames. The second reliability criterion and acoustic back-off are used to control the choice of recycled or batch calculated state likelihoods when they are contradictory in the combination and to prevent word accuracies from degrading. Large vocabulary spontaneous speech recognition experiments using four PCs with different specifications showed that, despite the PC specification dependence, the combined acceleration technique further reduced the total recognition time on all of the PCs.
Atsunori Ogawa, Satoshi Takahashi, Atsushi Nakamura
ICASSP1
2009 Rapid unsupervised adaptation using frame independent output probabilities of gender and context independent phoneme models
Satoshi Kobashikawa, Atsunori Ogawa, Yoshikazu Yamaguchi, Satoshi Takahashi
INTERSPEECH2
2009 Simultaneous estimation of confidence and error cause in speech recognition using discriminative model
Atsunori Ogawa, Atsushi Nakamura
INTERSPEECH1
2008 Weighted distance measures for efficient reduction of Gaussian mixture components in HMM-based acoustic model
abstract
In this paper, two weighted distance measures; the weighted K-L divergence and the Bayesian criterion-based distance measure are proposed to efficiently reduce the Gaussian mixture components in the HMM-based acoustic model. Conventional distance measures such as the K-L divergence and the Bhattacharyya distance consider only distribution parameters (i.e. mean and variance vectors of Gaussian pdfs). Another example considers only mixture weights. In contrast to them, the two proposed distance measures consider both distribution parameters and mixture weights. Experimental results showed that the component-reduced acoustic models created using the proposed distance measures were more compact and computationally efficient than those created using conventional distance measures.
Atsunori Ogawa, Satoshi Takahashi
ICASSP1
2005 Rapid response and robust speech recognition by preliminary model adaptation for additive and convolutional noise
Satoshi Kobashikawa, Satoshi Takahashi, Yoshikazu Yamaguchi, Atsunori Ogawa
INTERSPEECH4
2003 Non-native English speech recognition using bilingual English lexicon and acoustic models
abstract
This paper proposes an English speech recognition system which can recognize both non-native (i.e. Japanese) and native English speakers' pronunciation of English speech. The system uses a bilingual pronunciation lexicon in which each word has both English and Japanese phoneme transcriptions. The Japanese transcription is constructed considering typical Japanese pronunciation of English. Japanese and English acoustic models are used in recognizing both transcriptions, and the highest-likelihood word sequence obtained in combining with native English- and Japanese-pronounced words is the recognition result. Continuous speech recognition experiments show that the proposed system greatly improves Japanese-English speech recognition performance while maintaining the same performance level as that of a purely native English recognition system.
Shoichi Matsunaga, Atsunori Ogawa, Yoshikazu Yamaguchi, Akihiro Imamura
ICASSP (1)2
2003 Non-native English speech recognition using bilingual English lexicon and acoustic models
abstract
This paper proposes an English speech recognition system which can recognize both non-native (i.e. Japanese) and native English speaker's pronunciation of English speech. The system uses a bilingual pronunciation lexicon in which each word has both English and Japanese phoneme transcriptions. The Japanese transcription is constructed considering typical Japanese pronunciation of English. Japanese and English acoustic models are used in recognizing both transcriptions, and the highest-likelihood word sequence obtained in combining with native English- and Japanese-pronounced words is the recognition results. Continuous speech recognition experiments show that the proposed system greatly improves Japanese-English speech recognition performance while maintaining the same performance levels as that of a purely native English recognition system.
Shoichi Matsunaga, Atsunori Ogawa, Yoshikazu Yamaguchi, Akihiro Imamura
ICME2
2003 Speaker adaptation for non-native speakers using bilingual English lexicon and acoustic models
Shoichi Matsunaga, Atsunori Ogawa, Yoshikazu Yamaguchi, Akihiro Imamura
INTERSPEECH2
2000 Novel two-pass search strategy using time-asynchronous shortest-first second-pass beam search
Atsunori Ogawa, Yoshiaki Noda, Shoichi Matsunaga
INTERSPEECH1
1998 Balancing acoustic and linguistic probabilities
abstract
The length of the word sequence is not taken into account under language modeling of n-gram local probability modeling. Due to this property the optimal values of the language weight and word insertion penalty for balancing acoustic and linguistic probabilities is affected by the length of word sequence. To deal with this problem, a new language model is developed based on the Bernoulli trial model taking the length of the word sequence into account. Not only better recognition accuracy but also more robust balancing with acoustic probability compared with the normal n-gram model of the proposed method is confirmed through recognition experiments.
Atsunori Ogawa, Kazuya Takeda, Fumitada Itakura
ICASSP1
1998 Estimating entropy of a language from optimal word insertion penalty
Kazuya Takeda, Atsunori Ogawa, Fumitada Itakura
ICSLP2