Naohiro Tawara

dblp:79/10649 · DBLP profile ↗
← Back
38ranked-venue papers
12as first author
24since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 37 · 12 first-author · 23 since 2021Artificial intelligence and machine learning · 17 · 4 first-author · 11 since 2021
YearPublicationVenuePosition
2026 Microphone array geometry-independent multi-talker distant ASR: NTT system for DASR task of the CHiME-8 challenge
Naoyuki Kamo, Naohiro Tawara, Atsushi Ando, Takatomo Kano, Hiroshi Sato 0002, Rintaro Ikeshita, Takafumi Moriya, Shota Horiguchi, Kohei Matsuura, Atsunori Ogawa, Alexis Plaquet, Takanori Ashihara, Tsubasa Ochiai, Masato Mimura, Marc Delcroix, Tomohiro Nakatani, Taichi Asami, Shoko Araki
Comput. Speech Lang.2
2026 Effect of individual characteristics on impressions of one's own recorded voice
abstract
This study aims to identify individual characteristics such as age, gender, personality traits, and values that influence the perception of one’s own recorded voice. While previous studies have shown that the perception of one’s own recorded voice is different from that of others, and that these differences are influenced by individual characteristics, only a limited number of individual characteristics were examined in past research. In our study, we conducted a large-scale subjective experiment with 141 Japanese participants using multiple individual characteristics. Participants evaluated impressions of their own recorded voices and the voices of others, and we analyzed the relationship between each of the individual characteristics and the voice impressions. Our findings showed that individual characteristics such as the frequency of listening to one’s own recorded voice (which had not been examined in the previous studies) influenced the perception of one’s own recorded voice. We further analyzed the use of combinations of multiple individual characteristics, including those that influenced impressions in a single use, to predict impressions of one’s own recorded voice and found that they were better predicted by the combination of multiple individual characteristics than by the use of a single individual characteristic. • We show that impressions, such as familiarity, of one’s own recorded voice are different from those of others. • We show that multiple individual characteristics, such as the frequency of listening to one’s own recorded voice, influence the impression of one’s own recorded voice. • We show that combinations of multiple individual characteristics predicts the impression of one’s own recorded voice better than the use of a single individual characteristic.
Hikaru Yanagida, Yusuke Ijima, Naohiro Tawara
Speech Commun.3
2025 Predictive ASR and Turn-taking Prediction at Once: Towards More Responsive Spoken Dialog System
abstract
Spoken dialog systems usually wait for users to finish speaking before generating responses, resulting in response delays. A possible solution for reducing the response delay is to predict future words and/or turn-ends while the user is speaking. To realize this, we propose a method to jointly perform predictive automatic speech recognition and turn-taking prediction. Our model receives partial utterances as input and performs speech recognition, future word prediction, and turntaking prediction via autoregressive decoding. It enables turntaking prediction based on prosodic and linguistic cues of observed partial utterances and predicted future linguistic cues. We also incorporate dialogue contexts to improve the performance. Experiments on the Switchboard corpus showed that our multi-task model outperforms a single-task model in turn-taking prediction. We found that conditioning turn-taking prediction on predicted words improved performance when words were correctly predicted.
Ryo Fukuda, Takatomo Kano, Naohiro Tawara, Marc Delcroix, Atsunori Ogawa, Yuya Chiba, Atsushi Ando
ASRU3
2025 Can We Really Repurpose Multi-Speaker ASR Corpus for Speaker Diarization?
abstract
Neural speaker diarization is widely used for overlapaware speaker diarization, but it requires large multi-speaker datasets for training. To meet this data requirement, large datasets are often constructed by combining multiple corpora, including those originally designed for multi-speaker automatic speech recognition (ASR). However, ASR datasets often feature loosely defined segment boundaries that do not align with the stricter conventions of diarization benchmarks. In this work, we show that such boundary looseness significantly impacts the diarization error rate, reducing evaluation reliability. We also reveal that models trained on data with varying boundary precision tend to learn dataset-specific looseness, leading to poor generalization across out-of-domain datasets. Training with standardized tight boundaries via forced alignment improves not only diarization performance, especially in streaming scenarios, but also ASR performance when combined with simple post-processing.
Shota Horiguchi, Naohiro Tawara, Takanori Ashihara, Atsushi Ando, Marc Delcroix
ASRU2
2025 SoundBeam meets M2D: Target Sound Extraction with Audio Foundation Model
abstract
Target sound extraction (TSE) consists of isolating a desired sound from a mixture of arbitrary sounds using clues to identify it. A TSE system requires solving two problems at once, identifying the target source and extracting the target signal from the mixture. For increased practicability, the same system should work with various types of sound. The duality of the problem and the wide variety of sounds make it challenging to train a powerful TSE system from scratch. In this paper, to tackle this problem, we explore using a pre-trained audio foundation model that can provide rich feature representations of sounds within a TSE system. We chose the masked-modeling duo (M2D) foundation model, which appears especially suited for the TSE task, as it is trained using a dual objective consisting of sound-label predictions and improved masked prediction. These objectives are related to sound identification and the signal extraction problems of TSE. We propose a new TSE system that integrates the feature representation from M2D into SoundBeam, which is a strong TSE system that can exploit both target sound class labels and pre-recorded enrollments (or audio queries) as clues. We show experimentally that using M2D can increase extraction performance, especially when employing enrollment clues.
Carlos Hernandez-Olivan, Marc Delcroix, Tsubasa Ochiai, Daisuke Niizumi, Naohiro Tawara, Tomohiro Nakatani, Shoko Araki
ICASSP5
2025 Guided Speaker Embedding
abstract
This paper proposes a guided speaker embedding extraction system, which extracts speaker embeddings of the target speaker using speech activities of target and interference speakers as clues. Several methods for long-form overlapped multi-speaker audio processing are typically two-staged: i) segment-level processing and ii) inter-segment speaker matching. Speaker embeddings are often used for the latter purpose. Typical speaker embedding extraction approaches only use single-speaker intervals to avoid corrupting the embeddings with speech from interference speakers. However, this often makes speaker embeddings impossible to extract because sufficiently long non-overlapping intervals are not always available. In this paper, we propose using speaker activities as clues to extract the embedding of the speaker-of-interest directly from overlapping speech. Specifically, we concatenate the activity of target and non-target speakers to acoustic features before being fed to the model. We also condition the attention weights used for pooling so that the attention weights of the intervals in which the target speaker is inactive are zero. The effectiveness of the proposed method is demonstrated in speaker verification and speaker diarization.
Shota Horiguchi, Takafumi Moriya, Atsushi Ando, Takanori Ashihara, Hiroshi Sato 0002, Naohiro Tawara, Marc Delcroix
ICASSP6
2025 Mamba-based Segmentation Model for Speaker Diarization
abstract
Mamba is a newly proposed architecture that behaves like a recurrent neural network (RNN) with attention-like capabilities. These properties are promising for speaker diarization, as attention-based models have unsuitable memory requirements for long-form audio, and traditional RNN capabilities are too limited. In this paper, we propose to assess the potential of Mamba for diarization by comparing the state-of-the-art neural segmentation of the pyannote pipeline with our proposed Mamba-based variant. Mamba’s stronger processing capabilities allow usage of longer local windows, which significantly improve diarization quality by making the speaker embedding extraction more reliable. We find Mamba to be a superior alternative to both traditional RNN and the tested attention-based model. Our proposed Mamba-based system achieves state-of-the-art performance on three widely used diarization datasets.
Alexis Plaquet, Naohiro Tawara, Marc Delcroix, Shota Horiguchi, Atsushi Ando, Shoko Araki
ICASSP2
2025 Multi-channel Speaker Counting for EEND-VC-based Speaker Diarization on Multi-domain Conversation
abstract
This paper proposes a speaker counting scheme using multichannel microphones for end-to-end neural diarization with a vector clustering (EEND-VC) speaker diarization pipeline. The EEND-VC-based system estimates the number of speakers by clustering speaker embeddings from small chunks. However, conventional speaker counting struggles in short sessions with limited available embeddings. We address this issue by leveraging the most possible embeddings from multichannel signals to increase the number of embeddings. One challenge in using embeddings across channels is the biases caused by channel differences. To mitigate this issue, we extend the EEND-VC pipeline with two modifications: (1) applying speech enhancement before extracting speaker embedding to capture the speaker characteristics even from short chunks and (2) grouping microphones based on inter-channel correlation to perform speaker counting within each group and then aggregating these channel-wise results. The proposed scheme was integrated into our CHiME-8 diarization pipeline, achieving superior speaker counting accuracy compared to the CHiME-8 baseline, with 54.2% and 61.4% improvements in the development and evaluation sets, respectively.
Naohiro Tawara, Atsushi Ando, Shota Horiguchi, Marc Delcroix
ICASSP1
2025 Mitigating Non-Target Speaker Bias in Guided Speaker Embedding
Shota Horiguchi, Takanori Ashihara, Marc Delcroix, Atsushi Ando, Naohiro Tawara
INTERSPEECH5
2025 Pretraining Multi-Speaker Identification for Neural Speaker Diarization
Shota Horiguchi, Atsushi Ando, Naohiro Tawara, Marc Delcroix
INTERSPEECH3
2025 Why is children's ASR so difficult? Analyzing children's phonological error patterns using SSL-based phoneme recognizers
Koharu Horii, Naohiro Tawara, Atsunori Ogawa, Shoko Araki
INTERSPEECH2
2024 Discriminative Training of VBx Diarization
abstract
Bayesian HMM clustering of x-vector sequences (VBx) has become a widely adopted diarization baseline model in publications and challenges. It uses an HMM to model speaker turns, a generatively trained probabilistic linear discriminant analysis (PLDA) for speaker distribution modeling, and Bayesian inference to estimate the assignment of x-vectors to speakers. This paper presents a new framework for updating the VBx parameters using discriminative training, which directly optimizes a predefined loss. We also propose a new loss that better correlates with the diarization error rate compared to binary cross-entropy — the default choice for diarization end-to-end systems. Proof-of-concept results across three datasets (AMI, CALLHOME, and DIHARD II) demonstrate the method’s capability of automatically finding hyperparameters, achieving comparable performance to those found by extensive grid search, which typically requires additional hyperparameter behavior knowledge. Moreover, we show that discriminative fine-tuning of PLDA can further improve the model’s performance. We release the source code with this publication.
Dominik Klement, Mireia Díez, Federico Landini, Lukás Burget, Anna Silnova, Marc Delcroix, Naohiro Tawara
ICASSP7
2024 NTT Speaker Diarization System for Chime-7: Multi-Domain, Multi-Microphone end-to-end and Vector Clustering Diarization
abstract
This paper details our speaker diarization system designed for multi-domain, multi-microphone casual conversations. The proposed diarization pipeline uses weighted prediction error (WPE)based dereverberation as a front end, and separately applies end-to-end neural diarization with vector clustering (EEND-VC) to each channel. It integrates the diarization result obtained from each channel using diarization output voting error reduction plus overlap (DOVER-Lap). To harness the knowledge from the target domain and the results integrated across all channels, we apply self-supervised adaptation for each session by retraining the EEND-VC with pseudo-labels derived from DOVER-Lap. We incorporated our proposed system into NTT’s submission for a distant automatic speech recognition task in the CHiME-7 challenge. Our system obtained third place in the diarization performance by improving the development and evaluation sets by 65 % and 62 % compared to the organizer-provided, VC-based baseline diarization system.
Naohiro Tawara, Marc Delcroix, Atsushi Ando, Atsunori Ogawa
ICASSP1
2024 Recursive Attentive Pooling For Extracting Speaker Embeddings From Multi-Speaker Recordings
abstract
This paper proposes a method for extracting speaker embedding for each speaker from a variable-length recording containing multiple speakers. Speaker embeddings are crucial not only for speaker recognition but also for various multi-speaker speech applications such as speaker diarization and target-speaker speech processing. Despite the challenges of obtaining a single speaker’s speech without pre-registration in multi-speaker scenarios, most studies on speaker embedding extraction focus on extracting embeddings only from single-speaker recordings. Some methods have been proposed for extracting speaker embeddings directly from multi-speaker recordings, but they typically require preparing a model for each possible number of speakers or involve complicated training procedures. The proposed method computes the embeddings of multiple speakers by focusing on different parts of the frame-wise embeddings extracted from the input multi-speaker audio. This is achieved by recursively computing attention weights for pooling the frame-wise embeddings. Additionally, we propose using the calculated attention weights to estimate the number of speakers in the recording, which allows the same model to be applied to various numbers of speakers. Experimental evaluations demonstrate the effectiveness of the proposed method in speaker verification and diarization tasks.
Shota Horiguchi, Atsushi Ando, Takafumi Moriya, Takanori Ashihara, Hiroshi Sato 0002, Naohiro Tawara, Marc Delcroix
SLT6
2023 Iterative Shallow Fusion of Backward Language Model for End-To-End Speech Recognition
abstract
We propose a new shallow fusion (SF) method to exploit an external backward language model (BLM) for end-to-end automatic speech recognition (ASR). The BLM has complementary characteristics with a forward language model (FLM), and the effectiveness of their combination has been confirmed by rescoring ASR hypotheses as post-processing. In the proposed SF, we iteratively apply the BLM to partial ASR hypotheses in the backward direction (i.e., from the possible next token to the start symbol) during decoding, substituting the newly calculated BLM scores for the scores calculated at the last iteration. To enhance the effectiveness of this iterative SF (ISF), we train a partial sentence-aware BLM (PBLM) using reversed text data including partial sentences, considering the framework of ISF. In experiments using an attention-based encoder-decoder ASR system, we confirmed that ISF using the PBLM shows comparable performance with SF using the FLM. By performing ISF, early pruning of prospective hypotheses can be prevented during decoding, and we can obtain a performance improvement compared to applying the PBLM as post-processing. Finally, we confirmed that, by combining SF and ISF, further performance improvement can be obtained thanks to the complementarity of the FLM and PBLM.
Atsunori Ogawa, Takafumi Moriya, Naoyuki Kamo, Naohiro Tawara, Marc Delcroix
ICASSP4
2023 Multi-Stream Extension of Variational Bayesian HMM Clustering (MS-VBx) for Combined End-to-End and Vector Clustering-based Diarization
Marc Delcroix, Naohiro Tawara, Mireia Díez, Federico Landini, Anna Silnova, Atsunori Ogawa, Tomohiro Nakatani, Lukás Burget, Shoko Araki
INTERSPEECH2
2023 What are differences? Comparing DNN and Human by Their Performance and Characteristics in Speaker Age Estimation
Yuki Kitagishi, Naohiro Tawara, Atsunori Ogawa, Ryo Masumura, Taichi Asami
INTERSPEECH2
2023 Influence of Personal Traits on Impressions of One's Own Voice
Hikaru Yanagida, Yusuke Ijima, Naohiro Tawara
INTERSPEECH3
2022 Lattice Rescoring Based on Large Ensemble of Complementary Neural Language Models
abstract
We investigate the effectiveness of using a large ensemble of advanced neural language models (NLMs) for lattice rescoring on automatic speech recognition (ASR) hypotheses. Previous studies have reported the effectiveness of combining a small number of NLMs. In contrast, in this study, we combine up to eight NLMs, i.e., forward/backward long short-term memory/Transformer-LMs that are trained with two different random initialization seeds. We combine these NLMs through iterative lattice generation. Since these NLMs work complementarily with each other, by combining them one by one at each rescoring iteration, language scores attached to given lattice arcs can be gradually refined. Consequently, errors of the ASR hypotheses can be gradually reduced. We also investigate the effectiveness of carrying over contextual information (previous rescoring results) across a lattice sequence of a long speech such as a lecture speech. In experiments using a lecture speech corpus, by combining the eight NLMs and using context carry-over, we obtained a 24.4% relative word error rate reduction from the ASR 1-best baseline. For further comparison, we performed simultaneous (i.e., non-iterative) NLM combination and 100-best rescoring using the large ensemble of NLMs, which confirmed the advantage of lattice rescoring with iterative NLM combination.
Atsunori Ogawa, Naohiro Tawara, Marc Delcroix, Shoko Araki
ICASSP2
2021 Robust Speech-Age Estimation Using Local Maximum Mean Discrepancy Under Mismatched Recording Conditions
abstract
A recently proposed time-delay neural network (TDNN)-based age estimation system has yielded state-of-the-art performance in speech-age estimation tasks. However, the performance of this TDNN-based system can seriously degrade when the recording conditions of each utterance are different in the training and testing phases. To tackle this problem, we examine the efficiencies of a series of unsupervised domain adaptation (UDA) methods to obtain the model invariance against the difference of these conditions. In particular, we propose using local maximum mean discrepancy (LMMD) with soft-target labels to consider an ordinal relationship between age labels. In most UDA methods, the model is trained to obtain domain invariant representations by minimizing the statistical difference of the distributions between labeled source and unlabeled target data without considering their age class labels. In contrast, our LMMD-based approach locally minimizes the differences in their distributions on each age class while considering adjacent age classes using soft-target labels. We conducted speech-age estimation experiments on in-house datasets under mismatched conditions including different background noise, reverberation, and microphones. The experimental comparison demonstrated that the LMMD-based method contributed to efficiently reducing the effect of mismatches of input data, yielding significant improvements over other UDA methods, such as MMD and reverse gradients.
Naohiro Tawara, Atsunori Ogawa, Yuki Kitagishi, Hosana Kamiyama, Yusuke Ijima
ASRU1
2021 Integrating End-to-End Neural and Clustering-Based Diarization: Getting the Best of Both Worlds
abstract
Recent diarization technologies can be categorized into two approaches, i.e., clustering and end-to-end neural approaches, which have different pros and cons. The clustering-based approaches assign speaker labels to speech regions by clustering speaker embeddings such as x-vectors. While it can be seen as a current state-of-the-art approach that works for various challenging data with reasonable robustness and accuracy, it has a critical disadvantage that it cannot handle overlapped speech that is inevitable in natural conversational data. In contrast, the end-to-end neural diarization (EEND), which directly predicts diarization labels using a neural network, was devised to handle the overlapped speech. While the EEND, which can easily incorporate emerging deep-learning technologies, has started outperforming the x-vector clustering approach in some realistic database, it is difficult to make it work for long recordings (e.g., recordings longer than 10 minutes) because of, e.g., its huge memory consumption. Block-wise independent processing is also difficult because it poses an inter-block label permutation problem, i.e., an ambiguity of the speaker label assignments between blocks. In this paper, we propose a simple but effective hybrid diarization framework that works with overlapped speech and for long recordings containing an arbitrary number of speakers. It modifies the conventional EEND framework to output global speaker embeddings so that speaker clustering can be performed across blocks based on a constrained clustering algorithm to solve the permutation problem. With experiments based on simulated noisy reverberant 2-speaker meeting-like data, we show that the proposed framework works significantly better than the original EEND especially when the input data is long.
Keisuke Kinoshita, Marc Delcroix, Naohiro Tawara
ICASSP3
2021 BLSTM-Based Confidence Estimation for End-to-End Speech Recognition
abstract
Confidence estimation, in which we estimate the reliability of each recognized token (e.g., word, sub-word, and character) in automatic speech recognition (ASR) hypotheses and detect incorrectly recognized tokens, is an important function for developing ASR applications. In this study, we perform confidence estimation for end-to-end (E2E) ASR hypotheses. Recent E2E ASR systems show high performance (e.g., around 5% token error rates) for various ASR tasks. In such situations, confidence estimation becomes difficult since we need to detect infrequent incorrect tokens from mostly correct token sequences. To tackle this imbalanced dataset problem, we employ a bidirectional long short-term memory (BLSTM)-based model as a strong binary-class (correct/incorrect) sequence labeler that is trained with a class balancing objective. We experimentally confirmed that, by utilizing several types of ASR decoding scores as its auxiliary features, the model steadily shows high confidence estimation performance under highly imbalanced settings. We also confirmed that the BLSTM-based model outperforms Transformer-based confidence estimation models, which greatly underestimate incorrect tokens.
Atsunori Ogawa, Naohiro Tawara, Takatomo Kano, Marc Delcroix
ICASSP2
2021 Age-VOX-Celeb: Multi-Modal Corpus for Facial and Speech Estimation
abstract
Estimating a speaker’s age from their speech is more challenging than age estimation from their face because of insufficiently available public corpora. To tackle this problem, we construct a new audio-visual age corpus named AgeVoxCeleb by annotating age labels to VoxCeleb2 videos. AgeVoxCeleb is the first large-scale, balanced, and multi-modal age corpus that contains both video and speech of the same speakers from a wide age range. Using AgeVox-Celeb, our paper makes the following contributions: (i) A facial age estimation model can outperform a speech age estimation model by comparing the state-of-the-art models in each task. (ii) Facial age estimation is more robust against the difference between training and test sets. (iii) We developed cross-modal transfer learning from face to speech age estimation, showing that the estimated age with a facial age estimation model can be used to train a speech age estimation model. Proposed AgeVoxCeleb will be published in https://github.com/nttcslab-sp/agevoxceleb.
Naohiro Tawara, Atsunori Ogawa, Yuki Kitagishi, Hosana Kamiyama
ICASSP1
2021 Advances in Integration of End-to-End Neural and Clustering-Based Diarization for Real Conversational Speech
abstract
Recently, we proposed a novel speaker diarization method called End-to-End-Neural-Diarization-vector clustering (EEND-vector clustering) that integrates clustering-based and end-to-end neural network-based diarization approaches into one framework. The proposed method combines advantages of both frameworks, i.e. high diarization performance and handling of overlapped speech based on EEND, and robust handling of long recordings with an arbitrary number of speakers based on clustering-based approaches. However, the method was only evaluated so far on simulated 2-speaker meeting-like data. This paper is to (1) report recent advances we made to this framework, including newly introduced robust constrained clustering algorithms, and (2) experimentally show that the method can now significantly outperform competitive diarization methods such as Encoder-Decoder Attractor (EDA)-EEND, on CALLHOME data which comprises real conversational speech data including overlapped speech and an arbitrary number of speakers. By further analyzing the experimental results, this paper also discusses pros and cons of the proposed method and reveals potential for further improvement.
Keisuke Kinoshita, Marc Delcroix, Naohiro Tawara
Interspeech3
2020 Improving Speaker Discrimination of Target Speech Extraction With Time-Domain Speakerbeam
abstract
Target speech extraction, which extracts a single target source in a mixture given clues about the target speaker, has attracted increasing attention. We have recently proposed SpeakerBeam, which exploits an adaptation utterance of the target speaker to extract his/her voice characteristics that are then used to guide a neural network towards extracting speech of that speaker. SpeakerBeam presents a practical alternative to speech separation as it enables tracking speech of a target speaker across utterances, and achieves promising speech extraction performance. However, it sometimes fails when speakers have similar voice characteristics, such as in same-gender mixtures, because it is difficult to discriminate the target speaker from the interfering speakers. In this paper, we investigate strategies for improving the speaker discrimination capability of SpeakerBeam. First, we propose a time-domain implementation of SpeakerBeam similar to that proposed for a time-domain audio separation network (TasNet), which has achieved state-of-the-art performance for speech separation. Besides, we investigate (1) the use of spatial features to better discriminate speakers when microphone array recordings are available, (2) adding an auxiliary speaker identification loss for helping to learn more discriminative voice characteristics. We show experimentally that these strategies greatly improve speech extraction performance, especially for same-gender mixtures, and outperform TasNet in terms of target speech extraction.
Marc Delcroix, Tsubasa Ochiai, Katerina Zmolíková, Keisuke Kinoshita, Naohiro Tawara, Tomohiro Nakatani, Shoko Araki
ICASSP5
2020 Improving Speaker-Attribute Estimation by Voting Based on Speaker Cluster Information
abstract
This paper proposes a general post-processing method for improving speaker-attribute estimation. Estimating speaker-specific attributes such as age and gender is an important task with a wide range of applications. While the recent proposed deep neural network-based end-to-end approach achieves high performance, the model tends to over-fit to specific speakers when the amount of training data is limited or imbalanced. To solve this over-fitting problem, we propose a general framework for correcting unreliable results. The proposed algorithm first clusters the target utterances into speaker clusters by speaker similarity based on i-vectors. Then, for each of the speaker cluster, the speaker-attribute class of the cluster is determined by voting on the utterances assigned to the cluster. By then replacing the result of each utterance with the clusters' speaker-attribute class, we can correct the result of unreliable utterances. We used two tasks to evaluate the proposed algorithm including age estimation using the NIST-SRE10 and age-gender classification using an in-house read speech corpus, yielding significant improvements in mean absolute and classification errors.
Naohiro Tawara, Hosana Kamiyama, Satoshi Kobashikawa, Atsunori Ogawa
ICASSP1
2020 Frame-Level Phoneme-Invariant Speaker Embedding for Text-Independent Speaker Recognition on Extremely Short Utterances
abstract
This paper investigates a phoneme-invariant speaker embedding approach for speaker recognition on extremely short utterances. Intuitively, phonemes are nuisance information for text-independent speaker recognition task since the contents of the speech are usually mismatched between enrolling and testing time. However, many studies have shown that incorporating phoneme information is quite effective to improve the performance of the speaker recognition system. One reasonable explanation for this counter-intuitive result is that the pooling mechanism of segment-based speaker embedding can focus on the specific phonemes which contain rich speaker information, and phoneme information may help this. From this insight, we hypothesize that the pooling mechanism and phoneme-aware training are harmful to extract the speaker embeddings from extremely short utterances. To verify this hypothesis, an adversarial framework is introduced to remove phoneme-variability from the frame-wise speaker embeddings. The experimental results on the Librispeech corpus confirm that our frame-wise, phoneme-adversarial approach outperforms the conventional segment-wise, phoneme-aware approach for short utterances of less than about 1.4 seconds.
Naohiro Tawara, Atsunori Ogawa, Tomoharu Iwata, Marc Delcroix, Tetsuji Ogawa
ICASSP1
2020 Language Model Data Augmentation Based on Text Domain Transfer
Atsunori Ogawa, Naohiro Tawara, Marc Delcroix
INTERSPEECH2
2019 Postfiltering Using an Adversarial Denoising Autoencoder with Noise-aware Training
abstract
An adversarial denoising autoencoder (ADAE) with noise-aware training is proposed and successfully applied to post-filtering for linear noise reduction. The ADAE is effective for attenuating interference sounds, however, it is difficult to learn to handle its various unexpected harmful effects (e.g., various types of noise) using a single network. Legacy speech enhancement was introduced as a pre-processor to make it possible to efficiently train the ADAEs by reducing the unexpected variabilities in the inputs to the ADAEs. Time-frequency masking performed well to suppress the variabilities, however, it induced unpleasant distortion, which is difficult for the ADAE to complement. In this paper, a minimum variance distortionless response (MVDR) beam-former, which can avoid troublesome non-linear distortions, is exploited as a preprocessor, and the MVDR outputs are used as the inputs to the ADAE-based post-filter. In addition, noise-dominant signals derived from the MVDR beamformer can improve the accuracy of the ADAE-based post-filter because the residual noise depends on the original noise signals. Experimental comparisons conducted using multichannel speech enhancement demonstrate that ADAE-based post-filtering yields significant improvements over the MVDR-and ADAE-based speech enhancement systems, and noise-aware training of ADAE works well.
Naohiro Tawara, Hikari Tanabe, Tetsunori Kobayashi, Masaru Fujieda, Kazuhiro Katagiri, Takashi Yazu, Tetsuji Ogawa
ICASSP1
2019 Speaker Adversarial Training of DPGMM-Based Feature Extractor for Zero-Resource Languages
Yosuke Higuchi, Naohiro Tawara, Tetsunori Kobayashi, Tetsuji Ogawa
INTERSPEECH2
2019 Multi-Channel Speech Enhancement Using Time-Domain Convolutional Denoising Autoencoder
Naohiro Tawara, Tetsunori Kobayashi, Tetsuji Ogawa
INTERSPEECH1
2018 Language Model Domain Adaptation Via Recurrent Neural Networks with Domain-Shared and Domain-Specific Representations
abstract
Training recurrent neural network language models (RNNLMs) requires a large amount of data, which is difficult to collect for specific domains such as multiparty conversations. Data augmentation using external resources and model adaptation, which adjusts a model trained on a large amount of data to a target domain, have been proposed for low-resource language modeling. While there are the commonalities and discrepancies between the source and target domains in terms of the statistics of words and their contexts, these methods for domain adaptation make the commonalities and discrepancies jumbled. We propose novel domain adaptation techniques for RNNLM by introducing domain-shared and domain-specific word embedding and contextual features. This explicit modeling of the commonalities and discrepancies would improve the language modeling performance. Experimental comparisons using multiparty conversation data as the target domain augmented by lecture data from the source domain demonstrate that the proposed domain adaptation method exhibits improvements in the perplexity and word error rate over the long short-term memory based language model (LSTMLM) trained using the source and target domain data.
Tsuyoshi Morioka, Naohiro Tawara, Tetsuji Ogawa, Atsunori Ogawa, Tomoharu Iwata, Tetsunori Kobayashi
ICASSP2
2018 Speaker Invariant Feature Extraction for Zero-Resource Languages with Adversarial Learning
abstract
We introduce a novel type of representation learning to obtain a speaker invariant feature for zero-resource languages. Speaker adaptation is an important technique to build a robust acoustic model. For a zero-resource language, however, conventional model-dependent speaker adaptation methods such as constrained maximum likelihood linear regression are insufficient because the acoustic model of the target language is not accessible. Therefore, we introduce a model-independent feature extraction based on a neural network. Specifically, we introduce a multi-task learning to a bottleneck feature-based approach to make bottleneck feature invariant to a change of speakers. The proposed network simultaneously tackles two tasks: phoneme and speaker classifications. This network trains a feature extractor in an adversarial manner to allow it to map input data into a discriminative representation to predict phonemes, whereas it is difficult to predict speakers. We conduct phone discriminant experiments in Zero Resource Speech Challenge 2017. Experimental results showed that our multi-task network yielded more discriminative features eliminating the variety in speakers.
Taira Tsuchiya, Naohiro Tawara, Tetsuji Ogawa, Tetsunori Kobayashi
ICASSP2
2018 Sequential Fish Catch Forecasting Using Bayesian State Space Models
abstract
A new state space model suitable for fixed shore net fishing is proposed and successfully applied to daily fish catch forecasting. Accurate prediction of daily fish catches makes it possible to support fishery workers with decision-making for efficient operations. For that purpose, the predictive model should be intuitive to the fishery workers and provide an estimate with a confidence. In the present paper, a fish catch forecasting method is developed using a state space model that emulates the process of fixed shore net fishing. In this method, the parameter estimation and prediction are sequentially performed using the Hamiltonian Monte Carlo method. The experimental comparisons using actual fish catch data and public meteorological information demonstrated that the proposed forecasting system yielded significant reductions in predictive errors over the systems based on decision-trees and legacy state-space models.
Yuya Kokaki, Naohiro Tawara, Tetsunori Kobayashi, Kazuo Hashimoto, Tetsuji Ogawa
ICPR2
2015 A comparative study of spectral clustering for i-vector-based speaker clustering under noisy conditions
abstract
The present paper dealt with speaker clustering for speech corrupted by noise. In general, the performance of speaker clustering significantly depends on how well the similarities between speech utterances can be measured. The recently proposed i-vector-based cosine similarity has yielded the state-of-the-art performance in speaker clustering systems. However, this similarity often fails to capture the speaker similarity under noisy conditions. Therefore, we attempted to examine the efficiency of spectral clustering on i-vector-based similarity for speech corrupted by noise because spectral clustering can yield robustness against noise by non-linear projection. Experimental comparisons demonstrated that spectral clustering yielded significant improvement from conventional methods, such as agglomerative clustering and k-means clustering, under non-stationary noise conditions.
Naohiro Tawara, Tetsuji Ogawa, Tetsunori Kobayashi
ICASSP1
2012 Fully Bayesian inference of multi-mixture Gaussian model and its evaluation using speaker clustering
abstract
This study aims to verify effective optimization methods for estimating parametric, fully Bayesian models in speech processing. For that purpose, we investigate the impact of the difference in optimization methods for the multi-scale Gaussian mixture model, which is suitable for speaker clustering, on the clustering accuracy. The Markov chain Monte Carlo (MCMC)-based method was compared with the variational Bayesian method in the speaker clustering experiment; with a small amount of data, the MCMC-based method was more effective; with large scale data (more than one million samples), the difference between these methods in terms of the clustering accuracy decreased and the MCMC-based method was computationally efficient.
Naohiro Tawara, Tetsuji Ogawa, Shinji Watanabe 0001, Tetsunori Kobayashi
ICASSP1
2012 Fully Bayesian speaker clustering based on hierarchically structured utterance-oriented Dirichlet process mixture model
Naohiro Tawara, Tetsuji Ogawa, Shinji Watanabe 0001, Atsushi Nakamura, Tetsunori Kobayashi
INTERSPEECH1
2011 Speaker Clustering Based on Utterance-Oriented Dirichlet Process Mixture Model
abstract
This paper provides the analytical solution and algorithm of UO-DPMM based on a non-parametric Bayesian manner, and thus realizes fully Bayesian speaker clustering. We carried out preliminary speaker clustering experiments by using a TIMIT database to compare the proposed method with the conventional Bayesian Information Criterion (BIC) based method, which is an approximate Bayesian approach. The results showed that the proposed method outperformed the conventional one in terms of both computational cost and robustness to changes in tuning parameters.
Naohiro Tawara, Shinji Watanabe 0001, Tetsuji Ogawa, Tetsunori Kobayashi
INTERSPEECH1