EDBT 2026 Demo / reviewers in the wild / expert
Hisashi Kawai
dblp:32/4341
· DBLP profile ↗
141ranked-venue papers
5as first author
28since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 123 · 5 first-author · 24 since 2021Artificial intelligence and machine learning · 91 · 3 first-author · 17 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Layer-wise Analysis for Quality of Multilingual Synthesized SpeechabstractWhile supervised quality predictors for synthesized speech have demonstrated strong correlations with human ratings, their requirement for in-domain labeled training data hinders their generalization ability to new domains. Unsupervised approaches based on pretrained self-supervised learning (SSL) based models and automatic speech recognition (ASR) models are a promising alternative; however, little is known about how these models encode information about speech quality. Towards the goal of better understanding how different aspects of speech quality are encoded in a multilingual setting, we present a layer-wise analysis of multilingual pretrained speech models based on reference modeling. We find that features extracted from early SSL layers show correlations with human ratings of synthesized speech, and later layers of ASR models can predict quality of non-neural systems as well as intelligibility. We also demonstrate the importance of using well-matched reference data. Erica Cooper, Takuma Okamoto, Yamato Ohtani, Tomoki Toda, Hisashi Kawai |
ASRU | 5 |
| 2025 | Voice Factor Control Using FIR-Based Fast Neural Vocoder for Speech Generation ApplicationsabstractWe have proposed a fast neural vocoder based on the source-filter model introducing finite impulse response (FIR) filters called FIRNet. FIRNet is highly compatible with digital signal processing (DSP) and can, therefore, generate waveforms from vocoder parameters and modified voice factors, such as tone, intonation, and timbre, using DSP. Although modern neural waveform generation systems, such as voice conversion and text-to-speech (TTS), have been able to generate human-like synthetic speech and imitate the reference speaker’s timbre, it is challenging for these systems to manually control arbitrary voice factors, unlike traditional TTS systems. By applying FIRNet to modern neural waveform generation systems, they can achieve arbitrary voice factor controllability. We will demonstrate two applications using FIRNet with DSP-based voice factor controls: one is analysis-synthesis, and the other is text-to-speech. Yamato Ohtani, Takuma Okamoto, Tomoki Toda, Hisashi Kawai |
ASRU | 4 |
| 2025 | Mora-Level Prosody Prediction for Text-to-Speech Using Japanese BERT Without Accentual LabelsabstractIn practical text-to-speech (TTS) for pitch accent languages, such as Japanese, high-fidelity synthesis with correct prosody requires not only a phoneme sequence but also accentual information. Although accentual information can be obtained from accent dictionaries, words not included in the dictionaries and accent sandhi are sometimes synthesized with incorrect prosody, and manual registration of huge amounts of accent data is costly. Additionally, previous machine learning-based data-driven accent information estimation approaches for TTS also require huge quantities of handcrafted accentual labels during training. This paper proposes a data-driven prosody prediction method for Japanese TTS that uses Japanese BERT and does not require any accentual labels during training. A Japanese TTS acoustic model with mora-level (katakana sequence) input is first trained and mora-level fundamental frequency values (fo), which directly correspond to the prosody, are extracted for the training data using forced alignment. Then, a pre-trained Japanese BERT is finetuned for the mora-level foprediction task with word sequences including kanji and the corresponding katakana sequences as input and the mora-level foextracted using forced alignment as the prediction target. During TTS inference, the mora-level fosequence predicted by the finetuned Japanese BERT is input to the TTS acoustic model along with the katakana input, and correct prosodic synthesis can be realized thanks to this predicted fosequence. Experimental results demonstrate that the proposed method can realize the same synthesis quality and higher accent correctness compared with conventional neural TTS models with accentual labels. Tadashi Ogura, Takuma Okamoto, Yamato Ohtani, Erica Cooper, Tomoki Toda, Hisashi Kawai |
ICASSP | 6 |
| 2025 | Cross-modal Knowledge Transfer Learning as Graph Matching Based on Optimal Transport for ASR
Xugang Lu, Yu Tsao 0001, Hisashi Kawai |
INTERSPEECH | 4 |
| 2025 | GST-BERT-TTS: Prosody Prediction Without Accentual Labels For Multi-Speaker TTS Using BERT With Global Style Tokens
Tadashi Ogura, Takuma Okamoto, Yamato Ohtani, Erica Cooper, Tomoki Toda, Hisashi Kawai |
INTERSPEECH | 6 |
| 2024 | Hierarchical Cross-Modality Knowledge Transfer with Sinkhorn Attention for CTC-Based ASRabstractDue to the modality discrepancy between textual and acoustic modeling, efficiently transferring linguistic knowledge from a pretrained language model (PLM) to acoustic encoding for automatic speech recognition (ASR) still remains a challenging task. In this study, we propose a cross-modality knowledge transfer (CMKT) learning framework in a temporal connectionist temporal classification (CTC) based ASR system where hierarchical acoustic alignments with the linguistic representation are applied. Additionally, we propose the use of Sinkhorn attention in cross-modality alignment process, where the transformer attention is a special case of this Sinkhorn attention process. The CMKT learning is supposed to compel the acoustic encoder to encode rich linguistic knowledge for ASR. On the AISHELL-1 dataset, with CTC greedy decoding for inference (without using any language model), we achieved state-of-the-art performance with 3.64% and 3.94% character error rates (CERs) for the development and test sets, which corresponding to relative improvements of 34.18% and 34.88% compared to the baseline CTC-ASR system, respectively. Xugang Lu, Yu Tsao 0001, Hisashi Kawai |
ICASSP | 4 |
| 2024 | FIRNet: Fundamental Frequency Controllable Fast Neural Vocoder With Trainable Finite Impulse Response FilterabstractSome neural vocoders with fundamental frequency (f0) control have succeeded in performing real-time inference on a single CPU while preserving the quality of the synthetic speech. However, compared with legacy vocoders based on signal processing, their inference speeds are still low. This paper proposes a neural vocoder based on the source-filter model with trainable time-variant finite impulse response (FIR) filters, to achieve a similar inference speed to legacy vocoders. In the proposed model, FIRNet, multiple FIR coefficients are predicted using the neural networks, and the speech waveform is then generated by convolving a mixed excitation signal with these FIR coefficients. Experimental results show that FIRNet can achieve an inference speed similar to legacy vocoders while maintaining f0controllability and natural speech quality. Yamato Ohtani, Takuma Okamoto, Tomoki Toda, Hisashi Kawai |
ICASSP | 4 |
| 2024 | Convnext-TTS And Convnext-VC: Convnext-Based Fast End-To-End Sequence-To-Sequence Text-To-Speech And Voice ConversionabstractEnd-to-end (E2E) sequence-to-sequence (S2S) neural text-to-speech (TTS) models and E2E-S2S neural voice conversion (VC) models can achieve high-quality speech synthesis with a single neural network. To further improve the synthesis quality of E2E-S2S TTS and VC models and increase their inference speed, we propose a Transformer-free ConvNeXt-based encoder and decoder. Additionally, to further increase the inference speed, we propose ConvNeXt-TTS and ConvNeXt-VC, which include the WaveNeXt neural vocoder. This is also constructed from ConvNeXt blocks and can achieve much faster synthesis than HiFi-GAN. The results of experiments using the Hi-Fi-CAPTAIN corpus for the E2E-S2S-TTS and E2E-S2S-VC conditions demonstrate that the proposed ConvNeXt-based encoder and decoder can perform inference three times faster than a Transformer-based encoder and decoder while improving the synthesis quality. In particular, ConvNeXt-TTS and ConvNeXt-VC can achieve very fast E2E-S2S-TTS and E2E-S2S-VC with a real-time factor of 0.05 using a single-core CPU. Takuma Okamoto, Yamato Ohtani, Tomoki Toda, Hisashi Kawai |
ICASSP | 4 |
| 2024 | Investigating ASR Error Correction with Large Language Model and Multilingual 1-best Hypotheses
Sheng Li 0010, Chen Chen 0075, Kwok Chin Yuen, Chenhui Chu, Chng Eng Siong, Hisashi Kawai |
INTERSPEECH | 6 |
| 2024 | Mobile PresenTra: NICT fast neural text-to-speech system on smartphones with incremental inference of MS-FC-HiFi-GAN for law-latency synthesis
Takuma Okamoto, Yamato Ohtani, Hisashi Kawai |
INTERSPEECH | 3 |
| 2024 | Challenge of Singing Voice Synthesis Using Only Text-To-Speech Corpus With FIRNet Source-Filter Neural Vocoder
Takuma Okamoto, Yamato Ohtani, Sota Shimizu, Tomoki Toda, Hisashi Kawai |
INTERSPEECH | 5 |
| 2024 | Temporal Order Preserved Optimal Transport-Based Cross-Modal Knowledge Transfer Learning for ASRabstractTransferring linguistic knowledge from a pretrained language model (PLM) to an acoustic model has been shown to greatly improve the performance of automatic speech recognition (ASR). However, due to the heterogeneous feature distributions in cross-modalities, designing an effective model for feature alignment and knowledge transfer between linguistic and acoustic sequences remains a challenging task. Optimal transport (OT), which efficiently measures probability distribution discrepancies, holds great potential for aligning and transferring knowledge between acoustic and linguistic modalities. Nonetheless, the original OT treats acoustic and linguistic feature sequences as two unordered sets in alignment and neglects temporal order information during OT coupling estimation. Consequently, a time-consuming pretraining stage is required to learn a good alignment between the acoustic and linguistic representations. In this paper, we propose a Temporal Order Preserved OT (TOT)-based Cross-modal Alignment and Knowledge Transfer (CAKT) (TOT-CAKT) for ASR. In the TOT-CAKT, local neighboring frames of acoustic sequences are smoothly mapped to neighboring regions of linguistic sequences, preserving their temporal order relationship in feature alignment and matching. With the TOT-CAKT model framework, we conduct Mandarin ASR experiments with a pretrained Chinese PLM for linguistic knowledge transfer. Our results demonstrate that the proposed TOT-CAKT significantly improves ASR performance compared to several state-of-the-art models employing linguistic knowledge transfer, and addresses the weaknesses of the original OT-based method in sequential feature alignment for ASR. Xugang Lu, Yu Tsao 0001, Hisashi Kawai |
SLT | 4 |
| 2023 | Cross-Modal Alignment With Optimal Transport For CTC-Based ASRabstractTemporal connectionist temporal classification (CTC)-based automatic speech recognition (ASR) is one of the most successful end to end (E2E) ASR frameworks. However, due to the token independence assumption in decoding, an external language model (LM) is required which destroys its fast parallel decoding property. Several studies have been proposed to transfer linguistic knowledge from a pretrained LM (PLM) to the CTC based ASR. Since the PLM is built from text while the acoustic model is trained with speech, a cross-modal alignment is required in order to transfer the context dependent linguistic knowledge from the PLM to acoustic encoding. In this study, we propose a novel cross-modal alignment algorithm based on optimal transport (OT). In the alignment process, a transport coupling matrix is obtained using OT, which is then utilized to transform a latent acoustic representation for matching the context-dependent linguistic features encoded by the PLM. Based on the alignment, the latent acoustic feature is forced to encode context dependent linguistic information. We integrate this latent acoustic feature to build conformer encoder-based CTC ASR system. On the AISHELL-1 data corpus, our system achieved 3.96 % and 4.27 % character error rate (CER) for dev and test sets, respectively, which corresponds to relative improvements of 28.39 % and 29.42% compared to the baseline conformer CTC ASR system without cross-modal knowledge transfer. Xugang Lu, Yu Tsao 0001, Hisashi Kawai |
ASRU | 4 |
| 2023 | WaveNeXt: ConvNeXt-Based Fast Neural Vocoder Without ISTFT layerabstractA recently proposed neural vocoder, Vocos, can perform inference ten times faster than HiFi-GAN because of its use of ConvNeXt layers that can predict high-resolution short-time Fourier transform (STFT) spectra and an inverse STFT layer. To improve synthesis quality while preserving inference speed, this paper proposes an alternative ConvNeXt-based fast neural vocoder, WaveNeXt, in which the inverse STFT layer in Vocos is replaced with a trainable linear layer that can directly predict speech waveform samples without STFT spectra. Additionally, by integrating the JETS-based end-to-end text-to-speech (E2E TTS) framework, E2E TTS models can also be constructed with Vocos and WaveNeXt. Furthermore, full-band models with a sampling frequency of 48 kHz were investigated. The results of experiments for both the analysis-synthesis and E2E TTS conditions demonstrate that the proposed WaveNeXt can achieve higher quality synthesis than Vocos while preserving its inference speed. Takuma Okamoto, Haruki Yamashita, Yamato Ohtani, Tomoki Toda, Hisashi Kawai |
ASRU | 5 |
| 2023 | Generative Linguistic Representation for Spoken Language IdentificationabstractEffective extraction and application of linguistic features are central to the enhancement of spoken Language IDentification (LID) performance. With the success of recent large models, such as GPT and Whisper, the potential to leverage such pre-trained models for extracting linguistic features for LID tasks has become a promising area of research. In this paper, we explore the utilization of the decoder-based network from the Whisper model to extract linguistic features through its generative mechanism for improving the classification accuracy in LID tasks. We devised two strategies - one based on the language embedding method and the other focusing on direct optimization of LID outputs while simultaneously enhancing the speech recognition tasks. We conducted experiments on the large-scale multilingual datasets MLS, VoxLingua107, and CommonVoice to test our approach. The experimental results demonstrated the effectiveness of the proposed method on both in-domain and out-of-domain datasets for LID tasks. Xuguang Lu, Hisashi Kawai |
ASRU | 3 |
| 2023 | E2E-S2S-VC: End-To-End Sequence-To-Sequence Voice Conversion
Takuma Okamoto, Tomoki Toda, Hisashi Kawai |
INTERSPEECH | 3 |
| 2023 | Homeostatic System Design Based on Understanding the Living Environmental Determinants of FallsabstractFalls among older adults are a serious global challenge. The World Health Organization strongly recommends gait, balance, and functional training, Tai Chi, or home assessment and modification to prevent falls among older people, but these strategies have not changed over the past few decades. The purpose of the present study is to propose an innovative approach to fall prevention naturally embedded in the environment. We first suggest new methods for simplifying and understanding the relationship between falls and everyday life, and describe this relationship as a knowledge graph using emergency transport data on falls among the aged. As a result of a cluster analysis using knowledge graphs, we identified whole-body balance as a key factor and living environmental determinant of falls. We then propose the concept of a homeostatic spatial system and discuss our development of seven homeostatic products as a proof of concept. Finally, we develop an evaluation system to visualize specific locations in a defined area where people maintain whole-body balance using their hands or other body parts. Our findings confirm that the system is useful for evaluating how people change their behaviors to maintain whole-body balance based on homeostatic products and the living environment as a whole. Mikiko Oono, Ayano Nomura, Koji Kitamura, Yoshifumi Nishida, Shunsaburo Nakahara, Hisashi Kawai |
SMC | 6 |
| 2023 | Harmonic-Net: Fundamental Frequency and Speech Rate Controllable Fast Neural VocoderabstractThere is a need to improve the synthesis quality of HiFi-GAN-based real-time neural speech waveform generative models on CPUs while preserving the controllability of fundamental frequency ($f_{\mathrm{o}}$) and speech rate (SR). For this purpose, we propose Harmonic-Net and Harmonic-Net+, which introduce two extended functions into the HiFi-GAN generator. The first extension is a downsampling network, named the excitation signal network, that hierarchically receives multi-channel excitation signals corresponding to$f_{\mathrm{o}}$. The second extension is the layerwise pitch-dependent dilated convolutional network (LW-PDCNN), which can flexibly change its receptive fields depending on the input$f_{\mathrm{o}}$to handle large fluctuations in$f_{\mathrm{o}}$for the upsampling-based HiFi-GAN generator. The proposed explicit input of excitation signals and LW-PDCNNs corresponding to$f_{\mathrm{o}}$are expected to realize high-quality synthesis for the normal and$f_{\mathrm{o}}$-conversion conditions and for the SR-conversion condition. The results of experiments for unseen speaker synthesis, full-band singing voice synthesis, and text-to-speech synthesis show that the proposed method with harmonic waves corresponding to$f_{\mathrm{o}}$can achieve higher synthesis quality than conventional methods in all (i.e., normal,$f_{\mathrm{o}}$-conversion, and SR-conversion) conditions. Keisuke Matsubara, Takuma Okamoto, Ryoichi Takashima, Tetsuya Takiguchi, Tomoki Toda, Hisashi Kawai |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2022 | Transducer-based language embedding for spoken language identificationabstractThe acoustic and linguistic features are important cues for the spoken language identification (LID) task.Recent advanced LID systems mainly use acoustic features that lack the usage of explicit linguistic feature encoding.In this paper, we propose a novel transducer-based language embedding approach for LID tasks by integrating an RNN transducer model into a language embedding framework.Benefiting from the advantages of the RNN transducer's linguistic representation capability, the proposed method can exploit both phonetically-aware acoustic features and explicit linguistic features for LID tasks.Experiments were carried out on the large-scale multilingual LibriSpeech and VoxLingua107 datasets.Experimental results showed the proposed method significantly improves the performance on LID tasks with 12% to 59% and 16% to 24% relative improvement on in-domain and cross-domain datasets, respectively. Xugang Lu, Hisashi Kawai |
INTERSPEECH | 3 |
| 2022 | Pronunciation-Aware Unique Character Encoding for RNN Transducer-Based Mandarin Speech RecognitionabstractFor Mandarin end-to-end (E2E) automatic speech recognition (ASR) tasks, compared to character-based modeling units, pronunciation-based modeling units could improve the sharing of modeling units in model training but meet homophone problems. In this study, we propose to use a novel pronunciation-aware unique character encoding for building E2E RNN-T-based Mandarin ASR systems. The proposed encoding is a combination of pronunciation-base syllable and character index (CI). By introducing the CI, the RNN-T model can overcome the homophone problem while utilizing the pronunciation information for extracting modeling units. With the proposed encoding, the model outputs can be converted into the final recognition result through a one-to-one mapping. We conducted experiments on Aishell and MagicData datasets, and the experimental results showed the effectiveness of the proposed method. Xugang Lu, Hisashi Kawai |
SLT | 3 |
| 2022 | Neural speech-rate conversion with multispeaker WaveNet vocoderabstractSpeech-rate conversion technology, which can expand or compress speech waveforms while preserving the pitch of the sound, is traditionally realized by signal-processing-based approaches. To improve the synthesis quality, this paper proposes a machine-learning-based approach using neural vocoders, to perform neural speech-rate conversion. The proposed approach introduces a multispeaker WaveNet vocoder trained with a multispeaker corpus. Speech-rate conversion for many and unspecified speakers, not included in the training data, is realized by resampling acoustic features or hidden features along the time direction in inference. In experiments, the multispeaker WaveNet vocoder was trained using the JVS corpus and two types of resampling methods were compared. Conventional WSOLA and STRAIGHT were also compared as signal-processing-based baselines. The test sets included Japanese speaker corpora for the monolingual condition, and an English multispeaker corpus (CMU ARCTIC) for the cross-lingual condition. The results of the experiments demonstrate that the proposed approach with resampling of hidden features can achieve higher quality speech-rate conversion than the conventional methods, in both monolingual and cross-lingual conditions, except for speakers with low fundamental frequency in conversion of fast speech. Takuma Okamoto, Keisuke Matsubara, Tomoki Toda, Yoshinori Shiga, Hisashi Kawai |
Speech Commun. | 5 |
| 2021 | Multi-Stream HiFi-GAN with Data-Driven Waveform DecompositionabstractAlthough a HiFi-GAN vocoder can synthesize high-fidelity speech waveforms in real time on CPUs, there is a tradeoff between synthesis quality and inference speed. To increase inference speed while maintaining synthesis quality, a multi-band structure is introduced to HiFi-GAN. However, it cannot be trained well because of the strong constraint imposed by the fixed multi-band structure. As an alternative approach, Multi-stream MelGAN and HiFi-GAN are proposed, in which the fixed synthesis filter in Multi-band MelGAN is replaced by a trainable convolutional layer with the same structure. In contrast to Multi-band MelGAN, the proposed methods use the trainable synthesis filter to decompose speech waveforms in a data-driven manner. To evaluate the proposed Multi-stream HiFi-GAN as an entire real-time neural text-to-speech system on CPUs, a fast acoustic model, based on Parallel Tacotron 2 with forced alignment and accentual label input, was implemented. The results of experiments-using Japanese male, female, and multi-speaker corpora-indicate that Multi-stream HiFi-GAN can increase synthesis speed while improving or maintaining synthesis quality in analysis-synthesis and text-to-speech conditions for single-speaker models and unseen speaker synthesis for multi-speaker models, compared with the original HiFi-GAN. Takuma Okamoto, Tomoki Toda, Hisashi Kawai |
ASRU | 3 |
| 2021 | Unsupervised Neural Adaptation Model Based on Optimal Transport for Spoken Language IdentificationabstractDue to the mismatch of statistical distributions of acoustic speech between training and testing sets, the performance of spoken language identification (SLID) could be drastically degraded. In this paper, we propose an unsupervised neural adaptation model to deal with the distribution mismatch problem for SLID. In our model, we explicitly formulate the adaptation as to reduce the distribution discrepancy on both feature and classifier for training and testing data sets. Moreover, inspired by the strong power of the optimal transport (OT) to measure distribution discrepancy, a Wasserstein distance metric is designed in the adaptation loss. By minimizing the classification loss on the training data set with the adaptation loss on both training and testing data sets, the statistical distribution difference between training and testing domains is reduced. We carried out SLID experiments on the oriental language recognition (OLR) challenge data corpus where the training and testing data sets were collected from different conditions. Our results showed that significant improvements were achieved on the cross domain test tasks. Xugang Lu, Yu Tsao 0001, Hisashi Kawai |
ICASSP | 4 |
| 2021 | High-Intelligibility Speech Synthesis for Dysarthric Speakers with LPCNet-Based TTS and CycleVAE-Based VCabstractThis paper presents a high-intelligibility speech synthesis method for persons with dysarthria caused by athetoid cerebral palsy. The muscular control of such speakers is unstable because of their athetoid symptoms, and their pronunciation is unclear, which makes it difficult for them to communicate. In this paper, we present a method for generating highly intelligible speech that preserves the individuality of dysarthric speakers by combining Transformer-TTS, CycleVAE-VC, and a LPCNet vocoder. Rather than repairing prosody from the dysarthric speech, this method transfers the dysarthric speaker’s individuality to the speech of a healthy person generated by TTS synthesis. This task is both important and challenging. From the results of our evaluation experiments, we confirmed that the proposed method can partially transfer the individuality of the target dysarthric speaker while maintaining the intelligibility of the source speech. Keisuke Matsubara, Takuma Okamoto, Ryoichi Takashima, Tetsuya Takiguchi, Tomoki Toda, Yoshinori Shiga, Hisashi Kawai |
ICASSP | 7 |
| 2021 | Noise Level Limited Sub-Modeling for Diffusion Probabilistic VocodersabstractAlthough diffusion probabilistic vocoders WaveGrad and DiffWave can realize real-time high-fidelity speech synthesis with a simple loss function in training, all noise components with over the full range of noise levels are predicted by one model in all iterations. This paper proposes a simple but effective noise level-limited sub-modeling framework for diffusion probabilistic vocoders Sub-WaveGrad and Sub-DiffWave. In the proposed method, DiffWave conditioned on a continuous noise level like WaveGrad, and spectral enhancement post-filtering are also provided. The proposed Sub-WaveGrad and Sub-DiffWave models are realized using 10 sub-models. These models are separately trained with different noise level limits, and only necessary sub-models are used according to the noise schedule during inference. The results of experiments using a Japanese female speech corpus indicate that both the proposed Sub-WaveGrad and Sub-DiffWave outperform vanilla WaveGrad and DiffWave in terms of the model accuracy and synthesis quality while retaining the inference speed. Takuma Okamoto, Tomoki Toda, Yoshinori Shiga, Hisashi Kawai |
ICASSP | 4 |
| 2021 | Noise Robust Acoustic Modeling for Single-Channel Speech Recognition Based on a Stream-Wise Transformer Architecture
Masakiyo Fujimoto, Hisashi Kawai |
Interspeech | 2 |
| 2021 | Coupling a Generative Model With a Discriminative Learning Framework for Speaker VerificationabstractThe task of speaker verification (SV) is to decide whether an utterance is spoken by a target or an imposter speaker. In most studies of SV, a log-likelihood ratio (LLR) score is estimated based on a generative probability model on speaker features, and compared with a threshold for making a decision. However, the generative model usually focuses on individual feature distributions, does not have the discriminative feature selection ability, and is easy to be distracted by nuisance features. The SV, as a hypothesis test, could be formulated as a binary discrimination task where neural network based discriminative learning could be applied. In discriminative learning, the nuisance features could be removed with the help of label supervision. However, discriminative learning pays more attention to classification boundaries, and is prone to overfitting to a training set which may result in bad generalization on a test set. In this paper, we propose a hybrid learning framework, i.e., coupling a joint Bayesian (JB) generative model structure and parameters with a neural discriminative learning framework for SV. In the hybrid framework, a two-branch Siamese neural network is built with dense layers that are coupled with factorized affine transforms as used in the JB model. The LLR score estimation in the JB model is formulated according to the distance metric in the discriminative learning framework. By initializing the two-branch neural network with the generatively learned model parameters of the JB model, we further train the model parameters with the pairwise samples as a binary discrimination task. Moreover, a direct evaluation metric (DEM) in SV based on minimum empirical Bayes risk (EBR) is designed and integrated as an objective function in the discriminative learning. We carried out SV experiments on Speakers in the wild (SITW) and Voxceleb. Experimental results showed that our proposed model improved the performance with a large margin compared with state of the art models for SV. Xugang Lu, Yu Tsao 0001, Hisashi Kawai |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2021 | Quasi-Periodic Parallel WaveGAN: A Non-Autoregressive Raw Waveform Generative Model With Pitch-Dependent Dilated Convolution Neural NetworkabstractIn this paper, we propose a quasi-periodic parallel WaveGAN (QPPWG) waveform generative model, which applies a quasi-periodic (QP) structure to a parallel WaveGAN (PWG) model using pitch-dependent dilated convolution networks (PDCNNs). PWG is a small-footprint GAN-based raw waveform generative model, whose generation time is much faster than real time because of its compact model and non-autoregressive (non-AR) and non-causal mechanisms. Although PWG achieves high-fidelity speech generation, the generic and simple network architecture lacks pitch controllability for an unseen auxiliary fundamental frequency (F0) feature such as a scaled F0. To improve the pitch controllability and speech modeling capability, we apply a QP structure with PDCNNs to PWG, which introduces pitch information to the network by dynamically changing the network architecture corresponding to the auxiliary F0feature. Both objective and subjective experimental results show that QPPWG outperforms PWG when the auxiliary F0feature is scaled. Moreover, analyses of the intermediate outputs of QPPWG also show better tractability and interpretability of QPPWG, which respectively models spectral and excitation-like signals using the cascaded fixed and adaptive blocks of the QP structure. Yi-Chiao Wu, Tomoki Hayashi, Takuma Okamoto, Hisashi Kawai, Tomoki Toda |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2020 | Transformer-Based Text-to-Speech with Weighted Forced AttentionabstractThis paper investigates state-of-the-art Transformer- and FastSpeech-based high-fidelity neural text-to-speech (TTS) with full-context label input for pitch accent languages. The aim is to realize faster training than conventional Tacotron-based models. Introducing phoneme durations into Tacotron-based TTS models improves both synthesis quality and stability. Therefore, a Transformer-based acoustic model with weighted forced attention obtained from phoneme durations is proposed to improve synthesis accuracy and stability, where both encoder-decoder attention and forced attention are used with a weighting factor. Furthermore, FastSpeech without a duration predictor, in which the phoneme durations are predicted by another conventional model, is also investigated. The results of experiments using a Japanese female corpus and the WaveGlow vocoder indicate that the proposed Transformer using forced attention with a weighting factor of 0.5 outperforms other models, and removing the duration predictor from FastSpeech improves synthesis quality, although the proposed weighted forced attention does not improve synthesis stability. Takuma Okamoto, Tomoki Toda, Yoshinori Shiga, Hisashi Kawai |
ICASSP | 4 |
| 2020 | Investigation of NICT Submission for Short-Duration Speaker Verification Challenge 2020
Xugang Lu, Hisashi Kawai |
INTERSPEECH | 3 |
| 2020 | Quasi-Periodic Parallel WaveGAN Vocoder: A Non-Autoregressive Pitch-Dependent Dilated Convolution Model for Parametric Speech GenerationabstractIn this paper, we propose a parallel WaveGAN (PWG)-like neural vocoder with a quasi-periodic (QP) architecture to improve the pitch controllability of PWG. PWG is a compact non-autoregressive (non-AR) speech generation model, whose generative speed is much faster than real time. While utilizing PWG as a vocoder to generate speech on the basis of acoustic features such as spectral and prosodic features, PWG generates high-fidelity speech. However, when the input acoustic features include unseen pitches, the pitch accuracy of PWG-generated speech degrades because of the fixed and generic network of PWG without prior knowledge of speech periodicity. The proposed QPPWG adopts a pitch-dependent dilated convolution network (PDCNN) module, which introduces the pitch information into PWG via the dynamically changed network architecture, to improve the pitch controllability and speech modeling capability of vanilla PWG. Both objective and subjective evaluation results show the higher pitch accuracy and comparable speech quality of QPPWG-generated speech when the QPPWG model size is only 70 % of that of vanilla PWG. Yi-Chiao Wu, Tomoki Hayashi, Takuma Okamoto, Hisashi Kawai, Tomoki Toda |
INTERSPEECH | 4 |
| 2020 | Knowledge Distillation-Based Representation Learning for Short-Utterance Spoken Language IdentificationabstractWith successful applications of deep feature learning algorithms, spoken language identification (LID) on long utterances obtains satisfactory performance. However, the performance on short utterances is drastically degraded even when the LID system is trained using short utterances. The main reason is due to the large variation of the representation on short utterances which results in high model confusion. To narrow the performance gap between long, and short utterances, we proposed a teacher-student representation learning framework based on a knowledge distillation method to improve LID performance on short utterances. In the proposed framework, in addition to training the student model on short utterances with their true labels, the internal representation from the output of a hidden layer of the student model is supervised with the representation corresponding to their longer utterances. By reducing the distance of internal representations between short, and long utterances, the student model can explore robust discriminative representations for short utterances, which is expected to reduce model confusion. We conducted experiments on our in-house LID dataset, and NIST LRE07 dataset, and showed the effectiveness of the proposed methods for short utterance LID tasks. Xugang Lu, Sheng Li 0010, Hisashi Kawai |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2019 | Tacotron-Based Acoustic Model Using Phoneme Alignment for Practical Neural Text-to-Speech SystemsabstractAlthough sequence-to-sequence (seq2seq) models with attention mechanism in neural text-to-speech (TTS) systems, such as Tacotron 2, can jointly optimize duration and acoustic models, and realize high-fidelity synthesis compared with conventional duration-acoustic pipeline models, these involve a risk that speech samples cannot be sometimes successfully synthesized due to the attention prediction errors. Therefore, these seq2seq models cannot be directly introduced in practical TTS systems. On the other hand, the conventional pipeline models are broadly used in practical TTS systems since there are few crucial prediction errors in the duration model. For realizing high-quality practical TTS systems without attention prediction errors, this paper investigates Tacotron-based acoustic models with phoneme alignment instead of attention. The phoneme durations are first obtained from HMM-based forced alignment and the duration model is a simple bidirectional LSTM-based network. Then, a seq2seq model with forced alignment instead of attention is investigated and an alternative model with Tacotron decoder and phoneme duration is proposed. The results of experiments with full-context label input using WaveGlow vocoder indicate that the proposed model can realize a high-fidelity TTS system for Japanese with a real-time factor of 0.13 using a GPU without attention prediction errors compared with the seq2seq models. Takuma Okamoto, Tomoki Toda, Yoshinori Shiga, Hisashi Kawai |
ASRU | 4 |
| 2019 | HMM-based TTS System FrameworkabstractThe research focuses on the use of Hidden Markov Model (HMM) to build Khmer text-to-speech (TTS) system. Although the system is based on HMM statistic model, language specific functions were newly designed and developed to cope with the orthographical and grammatical nature of Khmer, some of which included word segmentation, grapheme to phoneme conversion, definitions of full context labels and question sets. In total four-thousand phonemically-balanced Khmer sentences were read aloud by an adult male speaker of Khmer, which were in turn served for training a model for Khmer TTS. The system has been incorporated into VoiceTra, a multilingual speech-to-speech translation app that has been developed and maintained by NICT. The app is publicly released for mobile devices and available to download in both App store and Google Play store. Saly Keo, Soky Kak, Yoshinori Shiga, Hiroaki Kato, Hisashi Kawai |
CIFEr | 5 |
| 2019 | Investigations of Real-time Gaussian Fftnet and Parallel Wavenet Neural Vocoders with Simple Acoustic FeaturesabstractThis paper examines four approaches to improving real-time neural vocoders with simple acoustic features (SAF) constructed from fundamental frequency and mel-cepstra rather than mel-spectrograms. The investigations are as follows: 1) the effectiveness of single Gaussian (SG) autoregressive (AR) WaveNet and FFTNet vocoders with SAF, 2) the possibility of SG parallel WaveNet vocoder training and synthesis with SAF, 3) the impact of noise shaping on SG AR neural vocoders, and 4) the efficacy of bandwidth extension to synthesize speech waveforms at a sampling frequency of 24 kHz by SG AR neural vocoders from SAF for that of 16 kHz. The results of experiments indicate that SG AR WaveNet and real-time SG AR FFTNet vocoders with noise shaping using SAF can realize sufficient synthesis quality with bandwidth extension effect. Moreover, a real-time SG parallel WaveNet vocoder can also be trained using SAF. Takuma Okamoto, Tomoki Toda, Yoshinori Shiga, Hisashi Kawai |
ICASSP | 4 |
| 2019 | Interactive Learning of Teacher-student Model for Short Utterance Spoken Language IdentificationabstractShort utterance-based spoken language identification (LID) is a challenging task due to the large variation of its feature representation. Improving feature representation of short utterances using a teacher-student method has been shown its effectiveness for LID tasks. However, conventional teacher-student methods use fixed pre-trained teacher models, that makes it difficult to optimize student models. In this paper, rather than using a fixed pre-trained teacher model, we investigate an interactive teacher-student learning by adjusting the teacher model with reference to the performance of the student model when the student model is stuck in a local minimum. Experiments on a 10-language LID task were carried out to test the algorithm. Our results showed its effectiveness of the proposed algorithm on short utterance LID tasks. Xugang Lu, Sheng Li 0010, Hisashi Kawai |
ICASSP | 4 |
| 2019 | Investigation of Sequence-level Knowledge Distillation Methods for CTC Acoustic ModelsabstractThis paper presents knowledge distillation (KD) methods for training connectionist temporal classification (CTC) acoustic models. In a previous study, we proposed a KD method based on the sequence-level cross-entropy, and showed that the conventional KD method based on the frame-level cross-entropy did not work effectively for CTC acoustic models, whereas the proposed method improved the performance of the models. In this paper, we investigate the implementation of sequence-level KD for CTC models and propose a lattice-based sequence-level KD method. Experiments investigating model compression and the training of a noise-robust model using the Wall Street Journal (WSJ) and CHiME4 datasets demonstrate that the sequence-level KD methods improve the performance of CTC acoustic models on both two tasks, and show that the lattice-based method can compute the sequence-level KD more efficiently than the N-best-based method proposed in our previous work. Ryoichi Takashima, Sheng Li 0010, Hisashi Kawai |
ICASSP | 3 |
| 2019 | End-to-End Articulatory Attribute Modeling for Low-Resource Multilingual Speech Recognition
Sheng Li 0010, Chenchen Ding, Xugang Lu, Tatsuya Kawahara, Hisashi Kawai |
INTERSPEECH | 6 |
| 2019 | One-Pass Single-Channel Noisy Speech Recognition Using a Combination of Noisy and Enhanced Features
Masakiyo Fujimoto, Hisashi Kawai |
INTERSPEECH | 2 |
| 2019 | Investigating Radical-Based End-to-End Speech Recognition Systems for Chinese Dialects and Japanese
Sheng Li 0010, Xugang Lu, Chenchen Ding, Tatsuya Kawahara, Hisashi Kawai |
INTERSPEECH | 6 |
| 2019 | Improving Transformer-Based Speech Recognition Systems with Compressed Structure and Speech Attributes Augmentation
Sheng Li 0010, Raj Dabre, Xugang Lu, Tatsuya Kawahara, Hisashi Kawai |
INTERSPEECH | 6 |
| 2019 | Incorporating Symbolic Sequential Modeling for Speech EnhancementabstractIn a noisy environment, a lossy speech signal can be automatically restored by a listener if he/she knows the language well.That is, with the built-in knowledge of a "language model", a listener may effectively suppress noise interference and retrieve the target speech signals.Accordingly, we argue that familiarity with the underlying linguistic content of spoken utterances benefits speech enhancement (SE) in noisy environments.In this study, in addition to the conventional modeling for learning the acoustic noisy-clean speech mapping, an abstract symbolic sequential modeling is incorporated into the SE framework.This symbolic sequential modeling can be regarded as a "linguistic constraint" in learning the acoustic noisy-clean speech mapping function.In this study, the symbolic sequences for acoustic signals are obtained as discrete representations with a Vector Quantized Variational Autoencoder algorithm.The obtained symbols are able to capture high-level phoneme-like content from speech signals.The experimental results demonstrate that the proposed framework can obtain notable performance improvement in terms of perceptual evaluation of speech quality (PESQ) and short-time objective intelligibility (STOI) on the TIMIT dataset. Chien-Feng Liao, Yu Tsao 0001, Xugang Lu, Hisashi Kawai |
INTERSPEECH | 4 |
| 2019 | Class-Wise Centroid Distance Metric Learning for Acoustic Event Detection
Xugang Lu, Sheng Li 0010, Yu Tsao 0001, Hisashi Kawai |
INTERSPEECH | 5 |
| 2019 | Duration Modeling with Global Phoneme-Duration Vectors
Jinfu Ni, Yoshinori Shiga, Hisashi Kawai |
INTERSPEECH | 3 |
| 2019 | Real-Time Neural Text-to-Speech with Sequence-to-Sequence Acoustic Model and WaveGlow or Single Gaussian WaveRNN Vocoders
Takuma Okamoto, Tomoki Toda, Yoshinori Shiga, Hisashi Kawai |
INTERSPEECH | 4 |
| 2018 | Comparative Evaluations of Various Factored Deep Convolutional Rnn Architectures for Noise Robust Speech RecognitionabstractIn this paper, we present a factored network-based acoustic modeling framework with various deep convolutional recurrent neural network (RNN) architectures for noise-robust automatic speech recognition (ASR). As the factored network-based acoustic model, we have already proposed a deep convolutional neural network (CNN)-based framework. Deep CNNs can emphasize the spatial locality of input speech features, but have no ability to analyze the properties of long-term speech feature sequences. Therefore, we introduce various deep convolutional RNN architectures that achieve both spatial locality and long-term analysis into our proposed factored network-based acoustic modeling framework. Through various comparative evaluations, we reveal that the proposed method successfully improves the accuracy of ASR in noisy environments. Masakiyo Fujimoto, Hisashi Kawai |
ICASSP | 2 |
| 2018 | An Investigation of Subband Wavenet Vocoder Covering Entire Audible Frequency Range with Limited Acoustic FeaturesabstractAlthough a WaveNet vocoder can synthesize more natural-sounding speech waveforms than conventional vocoders with sampling frequencies of 16 and 24 kHz, it is difficult to directly extend the sampling frequency to 48 kHz to cover the entire human audible frequency range for higher-quality synthesis because the model size becomes too large to train with a consumer GPU. For a WaveNet vocoder with a sampling frequency of 48 kHz with a consumer GPU, this paper introduces a subband WaveNet architecture to a speaker-dependent WaveNet vocoder and proposes a subband WaveNet vocoder. In experiments, each conditional subband WaveNet with a sampling frequency of 8 kHz was well trained using a consumer GPU. The results of subjective evaluations with a Japanese male speech corpus indicate that the proposed subband WaveNet vocoder with 36-dimensional simple acoustic features significantly outperformed the conventional source-filter model-based vocoders including STRAIGHT with 86-dimensional features. Takuma Okamoto, Kentaro Tachibana, Tomoki Toda, Yoshinori Shiga, Hisashi Kawai |
ICASSP | 5 |
| 2018 | An Investigation of Noise Shaping with Perceptual Weighting for Wavenet-Based Speech GenerationabstractWe propose a noise shaping method to improve the sound quality of speech signals generated by WaveNet, which is a convolutional neural network (CNN) that predicts a waveform sample sequence as a discrete symbol sequence. Speech signals generated by WaveNet often suffer from noise signals caused by the quantization error generated by representing waveform samples as discrete symbols and the prediction error of the CNN. We analyze these noise signals and show that 1) since the prediction error is much larger than the quantization error, the effect of the quantization error on the noise signals is practically negligible, and 2) noise signals tend to cause large spectral distortion in a high-frequency band. To alleviate the adverse effect of these noise signals on the generated speech signals, the proposed noise shaping method applies a perceptual weighting filter to WaveNet, making it possible to use the frequency masking properties of the human auditory system. We conducted objective and subjective evaluations to investigate the effectiveness of the proposed method and demonstrated that it significantly improved the sound quality of the generated speech signals. Kentaro Tachibana, Tomoki Toda, Yoshinori Shiga, Hisashi Kawai |
ICASSP | 4 |
| 2018 | An Investigation of a Knowledge Distillation Method for CTC Acoustic ModelsabstractEnd-to-end acoustic models, such as connectionist temporal classification (CTC) and the attention model, have been studied, and their speech recognition accuracies come close to those of conventional deep neural network (DNN)-hidden Markov models. However, most high-performance end-to-end models are not suitable for real-time (streaming) speech recognition because they are based on bidirectional recurrent neural networks (RNNs). In this study, to improve the performance of unidirectional RNN-based CTC, which is suitable for real-time processing, we investigate the knowledge distillation (KD)-based model compression method for training a CTC acoustic model. we evaluate a frame-level KD method and a sequence-level KD method for CTC model. The speech recognition experiments on Wall Street Journal tasks demonstrate that, the frame-level KD worsens the WERs ofunidirectional CTC model, whereas sequence-level KD can improve the WERs of the model. Ryoichi Takashima, Sheng Li 0010, Hisashi Kawai |
ICASSP | 3 |
| 2018 | CTC Loss Function with a Unit-Level Ambiguity PenaltyabstractThis paper presents a modified loss function for training connectionist temporal classification (CTC)-based acoustic models. CTC-based acoustic models have been studied as alternatives to conventional hidden Markov models (HMMs), but have often shown worse performance than conventional deep neural network (DNN)-HMM hybrid models. In this paper, we attempt to identify the primary factor preventing CTC-based models from achieving their full potential, and hypothesize this constraint lies in the ambiguity in the identification boundaries among unit-level labels (phonemes or characters). In accordance with this hypothesis, we propose a modified CTC loss function using an ambiguity penalty. This penalty is defined by the conditional entropy and works to increase the separation metrics among unit-level labels. We evaluate the proposed method on the WSJ and CHiME4 tasks, and demonstrate that our modification improves the word error rate compared with that of the conventional CTC-based model when the training dataset is small. Ryoichi Takashima, Sheng Li 0010, Hisashi Kawai |
ICASSP | 3 |
| 2018 | Improving CTC-based Acoustic Model with Very Deep Residual Time-delay Neural Networks
Sheng Li 0010, Xugang Lu, Ryoichi Takashima, Tatsuya Kawahara, Hisashi Kawai |
INTERSPEECH | 6 |
| 2018 | Temporal Attentive Pooling for Acoustic Event Detection
Xugang Lu, Sheng Li 0010, Yu Tsao 0001, Hisashi Kawai |
INTERSPEECH | 5 |
| 2018 | Multilingual Grapheme-to-Phoneme Conversion with Global Character Vectors
Jinfu Ni, Yoshinori Shiga, Hisashi Kawai |
INTERSPEECH | 3 |
| 2018 | Feature Representation of Short Utterances Based on Knowledge Distillation for Spoken Language Identification
Xugang Lu, Sheng Li 0010, Hisashi Kawai |
INTERSPEECH | 4 |
| 2018 | Improving Very Deep Time-Delay Neural Network With Vertical-Attention For Effectively Training CTC-Based ASR SystemsabstractThe very deep neural network has recently been proposed for speech recognition and achieves significant performance. It has excellent potential for integration with end-to-end (E2E) training. Connectionist temporal classification (CTC) has shown great potential in E2E acoustic modeling. In this study, we investigate deep architectures and techniques which are suitable for CTC-based acoustic modeling. We propose a very deep residual time-delay CTC neural network (VResTD-CTC). How to select a suitable deep architecture optimized with the CTC objective function is crucial for obtaining the state of the art performance. Excellent performances can be obtained by selecting deep architecture for non-E2E ASR systems modeling with tied-triphone states. However, these optimized structures do not guarantee to achieve better or comparable performances on E2E (e.g., CTC-based) systems modeling with dynamic acoustic units. For solving this problem and further leveraging the system performance, we introduce the vertical-attention mechanism to reweight the residual blocks at each time step. Speech recognition experiments show our proposed model significantly outperforms the DNN and LSTM-based (both bidirectional and unidirectional) CTC baseline models. Sheng Li 0010, Xugang Lu, Ryoichi Takashima, Tatsuya Kawahara, Hisashi Kawai |
SLT | 6 |
| 2018 | Improving FFTNet Vocoder with Noise Shaping and Subband ApproachesabstractAlthough FFTNet neural vocoders can synthesize speech waveforms in real time, the synthesized speech quality is worse than that of WaveNet vocoders. To improve the synthesized speech quality of FFTNet while ensuring real-time synthesis, residual connections are introduced to enhance the prediction accuracy. Additionally, time-invariant noise shaping and subband approaches, which significantly improve the synthesized speech quality of WaveNet vocoders, are applied. A subband FFTNet vocoder with multiband input is also proposed to directly compensate the phase shift between subbands. The proposed approaches are evaluated through experiments using a Japanese male corpus with a sampling frequency of 16 kHz. The results are compared with those synthesized by the STRAIGHT vocoder without mel-cepstral compression and those from conventional FFTNet and WaveNet vocoders. The proposed approaches are shown to successfully improve the synthesized speech quality of the FFTNet vocoder. In particular, the use of noise shaping enables FFTNet to significantly outperform the STRAIGHT vocoder. Takuma Okamoto, Tomoki Toda, Yoshinori Shiga, Hisashi Kawai |
SLT | 4 |
| 2018 | End-to-End Waveform Utterance Enhancement for Direct Evaluation Metrics Optimization by Fully Convolutional Neural NetworksabstractSpeech enhancement model is used to map a noisy speech to a clean speech. In the training stage, an objective function is often adopted to optimize the model parameters. However, in the existing literature, there is an inconsistency between the model optimization criterion and the evaluation criterion for the enhanced speech. For example, in measuring speech intelligibility, most of the evaluation metric is based on a short-time objective intelligibility (STOI) measure, while the frame based mean square error (MSE) between estimated and clean speech is widely used in optimizing the model. Due to the inconsistency, there is no guarantee that the trained model can provide optimal performance in applications. In this study, we propose an end-to-end utterance-based speech enhancement framework using fully convolutional neural networks (FCN) to reduce the gap between the model optimization and the evaluation criterion. Because of the utterance-based optimization, temporal correlation information of long speech segments, or even at the entire utterance level, can be considered to directly optimize perception-based objective functions. As an example, we implemented the proposed FCN enhancement framework to optimize the STOI measure. Experimental results show that the STOI of a test speech processed by the proposed approach is better than conventional MSE-optimized speech due to the consistency between the training and the evaluation targets. Moreover, by integrating the STOI into model optimization, the intelligibility of human subjects and automatic speech recognition system on the enhanced speech is also substantially improved compared to those generated based on the minimum MSE criterion. Szu-Wei Fu, Taowei Wang, Yu Tsao 0001, Xugang Lu, Hisashi Kawai |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2017 | Incremental training and constructing the very deep convolutional residual network acoustic modelsabstractInspired by the successful applications in image recognition, the very deep convolutional residual network (ResNet) based model has been applied in automatic speech recognition (ASR). However, the computational load is heavy for training the ResNet with a large quantity of data. In this paper, we propose an incremental model training framework to accelerate the training process of the ResNet. The incremental model training framework is based on the unequal importance of each layer and connection in the ResNet. The modules with important layers and connections are regarded as a skeleton model, while those left are regarded as an auxiliary model. The total depth of the skeleton model is quite shallow compared to the very deep full network. In our incremental training, the skeleton model is first trained with the full training data set. Other layers and connections belonging to the auxiliary model are gradually attached to the skeleton model and tuned. Our experiments showed that the proposed incremental training obtained comparable performances and faster training speed compared with the model training as a whole without consideration of the different importance of each layer. Sheng Li 0010, Xugang Lu, Ryoichi Takashima, Tatsuya Kawahara, Hisashi Kawai |
ASRU | 6 |
| 2017 | Subband wavenet with overlapped single-sideband filterbanksabstractCompared with conventional vocoders, deep neural network-based raw audio generative models, such as WaveNet and SampleRNN, can more naturally synthesize speech signals, although the synthesis speed is a problem, especially with high sampling frequency. This paper provides subband WaveNet based on multirate signal processing for high-speed and high-quality synthesis with raw audio generative models. In the training stage, speech waveforms are decomposed and decimated into subband short waveforms with a low sampling rate, and each subband WaveNet network is trained using each subband stream. In the synthesis stage, each generated signal is up-sampled and integrated into a fullband speech signal. The results of objective and subjective experiments for unconditional WaveNet with a sampling frequency of 32 kHz indicate that the proposed subband WaveNet with a square-root Hann window-based overlapped 9-channel single-sideband filterbank can realize about four times the synthesis speed and improve the synthesized speech quality more than the conventional fullband WaveNet. Takuma Okamoto, Kentaro Tachibana, Tomoki Toda, Yoshinori Shiga, Hisashi Kawai |
ASRU | 5 |
| 2017 | Grounded language understanding for manipulation instructions using GAN-based classificationabstractThe target task of this study is grounded language understanding for domestic service robots (DSRs). In particular, we focus on instruction understanding for short sentences where verbs are missing. This task is of critical importance to build communicative DSRs because manipulation is essential for DSRs. Existing instruction understanding methods usually estimate missing information only from non-grounded knowledge; therefore, whether the predicted action is physically executable or not was unclear. In this paper, we present a grounded instruction understanding method to estimate appropriate objects given an instruction and situation. We extend the Generative Adversarial Nets (GAN) and build a GAN-based classifier using latent representations. To quantitatively evaluate the proposed method, we have developed a data set based on the standard data set used for visual question answering (VQA). Experimental results have shown that the proposed method gives the better result than baseline methods. Komei Sugiura, Hisashi Kawai |
ASRU | 2 |
| 2017 | Minimum Bayes risk training of CTC acoustic models in maximum a posteriori based decoding frameworkabstractWhen using connectionist temporal classification (CTC) based acoustic models (AMs) for large vocabulary continuous speech recognition (LVCSR), most previous studies have used a naive interpolation of the CTC-AM score and an additional language model score, although there is no theoretical justification for such an approach. On the other hand, we recently proposed a theoretically more sound decoding framework for CTC-AM called maximum a posteriori (MAP)-based decoding. Although the superiority of the MAP-based decoding framework with CTC-AM has been demonstrated, the effect of additional minimum Bayes risk (MBR) training in the MAP-based decoding framework has not been investigated. In this paper, we report the results of various experiments that examine the effect of MBR training on CTC-AM by comparing two decoding frameworks. Our experiments with English and Japanese LVCSR tasks reveal that the MAP-based decoding framework is superior to the interpolation-based framework, even after the MBR training. In addition, by using about 600 h of training data, we show that the size of the training dataset is a critical factor in achieving good results under CTC-AM. Naoyuki Kanda, Xugang Lu, Hisashi Kawai |
ICASSP | 3 |
| 2017 | Global Syllable Vectors for Building TTS Front-End with Deep Learning
Jinfu Ni, Yoshinori Shiga, Hisashi Kawai |
INTERSPEECH | 3 |
| 2017 | Conditional Generative Adversarial Nets Classifier for Spoken Language Identification
Xugang Lu, Sheng Li 0010, Hisashi Kawai |
INTERSPEECH | 4 |
| 2017 | Regularization of neural network model with distance metric learning for i-vector based spoken language identification
Xugang Lu, Yu Tsao 0001, Hisashi Kawai |
Comput. Speech Lang. | 4 |
| 2017 | Maximum-a-Posteriori-Based Decoding for End-to-End Acoustic ModelsabstractThis paper presents a novel decoding framework for acoustic models (AMs) based on end-to-end neural networks (e.g., connectionist temporal classification). The end-to-end training of AMs has recently demonstrated high accuracy and efficiency in automatic speech recognition (ASR). When using the trained AM in decoding, although a language model (LM) is implicitly involved in such an end-to-end AM, it is still essential to integrate an external LM trained with a large text corpus to achieve the best results. While there is no theoretical justification, most of the studies suggest using a naive interpolation of the end-to-end AM score and the external LM score, empirically. In this paper, we propose a more theoretically sound decoding framework derived from a maximization of the posterior probability of a word sequence given an observation. As a consequence of the theory, the subword LM is newly introduced to seamlessly integrate the external LM score with the end-to-end AM score. Our proposed method can be achieved by a small modification of the conventional weighted finite-state transducer-based implementation, without having to heavily increase the graph size. We tested the proposed decoding framework on ASR experiments with the Corpus of the Wall Street Journal and the Corpus of Spontaneous Japanese. The results showed that the proposed framework achieved significant and consistent improvements over the conventional interpolation-based decoding framework. Naoyuki Kanda, Xugang Lu, Hisashi Kawai |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2016 | Bottleneck linear transformation network adaptation for speaker adaptive training-based hybrid DNN-HMM speech recognizerabstractRecently, a Hybrid DNN-HMM recognizer trained with the Speaker Adaptive Training (SAT) concept was successfully modified to a more effective speaker-adaptation-oriented recognizer whose DNN front-end adopted a Linear Transformation Network (LTN) Speaker Dependent (SD) module. However, the size of SD modules is still large, which incurs high storage costs and the risk of over-training. To alleviate this problem, we analyze the characteristics of an LTN module by focusing on the relation between its size and its feature-representation capability. Moreover, we propose a new SAT-based scheme for reducing the LTN size using SVD-based matrix compression. Evaluation experiments on the TED Talks corpus prove that our LTN size-reduction scheme not only maintains the adaptation performance of the original LTN-embedded, SAT-based DNN-HMM recognizer but also further increases it especially in cases where the speech data available for adaptation training are severely limited. Tsubasa Ochiai, Shigeki Matsuda, Hideyuki Watanabe, Xugang Lu, Hisashi Kawai, Shigeru Katagiri |
ICASSP | 5 |
| 2016 | Local fisher discriminant analysis for spoken language identificationabstractI-vector is a state-of-the-art technique widely used in spoken language identification systems. Since i-vectors include total variability factors, discriminant analysis methods have been introduced to find the most discriminative features while removing the undesired variables for language identification, for example, linear discriminant analysis (LDA) and nonparametric discriminant analysis (NDA). However, these methods either do not consider or use weak local structures of the data. In this study, we introduce a local Fisher discriminant analysis (LFDA) as a post-processing discriminant analysis method to extract the discriminative features from i-vectors. LFDA is a full-rank method which takes the local structure of the data into account for non-Gaussian distribution data, i.e., multimodal. Compared with LDA and NDA, LFDA is a pair-wise local method which enhances the centralization of the distribution of samples in the same class to obtain larger amounts of discriminative features. Experimental results indicate that LFDA is more effective than LDA and NDA for the i-vector-based language identification task. Xugang Lu, Lemao Liu, Hisashi Kawai |
ICASSP | 4 |
| 2016 | Investigation of Semi-Supervised Acoustic Model Training Based on the Committee of Heterogeneous Neural Networks
Naoyuki Kanda, Shoji Harada, Xugang Lu, Hisashi Kawai |
INTERSPEECH | 4 |
| 2016 | Maximum a posteriori Based Decoding for CTC Acoustic Models
Naoyuki Kanda, Xugang Lu, Hisashi Kawai |
INTERSPEECH | 3 |
| 2016 | Pair-Wise Distance Metric Learning of Neural Network Model for Spoken Language Identification
Xugang Lu, Yu Tsao 0001, Hisashi Kawai |
INTERSPEECH | 4 |
| 2016 | Using Zero-Frequency Resonator to Extract Multilingual Intonation Structure
Jinfu Ni, Yoshinori Shiga, Hisashi Kawai |
INTERSPEECH | 3 |
| 2016 | Model Integration for HMM- and DNN-Based Speech Synthesis Using Product-of-Experts Framework
Kentaro Tachibana, Tomoki Toda, Yoshinori Shiga, Hisashi Kawai |
INTERSPEECH | 4 |
| 2016 | F0 Contour Analysis Based on Empirical Mode Decomposition for DNN Acoustic Modeling in Mandarin Speech Recognition
Xiaoyun Wang 0002, Xugang Lu, Hisashi Kawai, Seiichi Yamamoto |
INTERSPEECH | 3 |
| 2016 | Combination of multiple acoustic models with unsupervised adaptation for lecture speech transcription
Xugang Lu, Xinhui Hu, Naoyuki Kanda, Masahiro Saiko, Chiori Hori, Hisashi Kawai |
Speech Commun. | 7 |
| 2015 | Training data pseudo-shuffling and direct decoding framework for recurrent neural network based acoustic modelingabstractWe propose two techniques to enhance the performance of recurrent neural network (RNN)-based acoustic models. The first technique addresses training efficiency. Because RNNs require sequential input, it is difficult to randomly shuffle training samples to accelerate stochastic gradient descent based training. We propose a "pseudo-shuffling" procedure that instead augments training sample unexpectedness by skipping successive samples. The second proposed technique is a novel "direct decoding" framework in which the posterior probability of the RNN is inputted into a decoder without conversion into a hidden Markov model emission probability. In our large vocabulary speech recognition experiments with English lecture recordings, the first technique significantly improved RNN training efficiency, showing a 14.3% relative word error rate (WER) improvement. The second technique further achieved an additional 3.1% relative WER improvement. Our sigmoid-type RNN achieved a 10.7% better WER than same-sized deep neural networks without using long short-term memory cells. Naoyuki Kanda, Mitsuyoshi Tachimori, Xugang Lu, Hisashi Kawai |
ASRU | 4 |
| 2015 | Sparse representation with temporal max-smoothing for acoustic event detection
Xugang Lu, Yu Tsao 0001, Chiori Hori, Hisashi Kawai |
INTERSPEECH | 5 |
| 2015 | HMM based myanmar text to speech system
Ye Kyaw Thu, Win Pa Pa, Jinfu Ni, Yoshinori Shiga, Andrew M. Finch, Chiori Hori, Hisashi Kawai, Eiichiro Sumita |
INTERSPEECH | 7 |
| 2015 | Leveraging social Q&A collections for improving complex question answeringabstractThis paper regards social question-and-answer (Q&A) collections such as Yahoo! Answers as knowledge repositories and investigates techniques to mine knowledge from them to improve sentence-based complex question answering (QA) systems. Specifically, we present a question-type-specific method (QTSM) that extracts question-type-dependent cue expressions from social Q&A pairs in which the question types are the same as the submitted questions. We compare our approach with the question-specific and monolingual translation-based methods presented in previous works. The question-specific method (QSM) extracts question-dependent answer words from social Q&A pairs in which the questions resemble the submitted question. The monolingual translation-based method (MTM) learns word-to-word translation probabilities from all of the social Q&A pairs without considering the question or its type. Experiments on the extension of the NTCIR 2008 Chinese test data set demonstrate that our models that exploit social Q&A collections are significantly more effective than baseline methods such as LexRank. The performance ranking of these methods is QTSM > {QSM, MTM}. The largest F3 improvements in our proposed QTSM over QSM and MTM reach 6.0% and 5.8%, respectively. Youzheng Wu, Chiori Hori, Hideki Kashioka, Hisashi Kawai |
Comput. Speech Lang. | 4 |
| 2014 | Non-monologue HMM-based speech synthesis for service robots: A cloud robotics approachabstractRobot utterances generally sound monotonous, unnatural, and unfriendly because their Text-to-Speech (TTS) systems are not optimized for communication but for text-reading. Here we present a non-monologue speech synthesis for robots. We collected a speech corpus in a non-monologue style in which two professional voice talents read scripted dialogues. Hidden Markov models (HMMs) were then trained with the corpus and used for speech synthesis. We conducted experiments in which the proposed method was evaluated by 24 subjects in three scenarios: text-reading, dialogue, and domestic service robot (DSR) scenarios. In the DSR scenario, we used a physical robot and compared our proposed method with a baseline method using the standard Mean Opinion Score (MOS) criterion. Our experimental results showed that our proposed method's performance was (1) at the same level as the baseline method in the text-reading scenario and (2) exceeded it in the DSR scenario. We deployed our proposed system as a cloud-based speech synthesis service so that it can be used without any cost. Komei Sugiura, Yoshinori Shiga, Hisashi Kawai, Teruhisa Misu, Chiori Hori |
ICRA | 3 |
| 2013 | Multilingual Speech-to-Speech Translation System: VoiceTraabstractThis study presents an overview of VoiceTra, which was developed by NICT and released as the world's first network-based multilingual speech-to-speech translation system for smartphones, and describes in detail its multilingual speech recognition, its multilingual translation, and its multilingual speech synthesis in regards to field experiments. We show the effects of system updates using the data collected from field experiments to improve our acoustic and language models. Shigeki Matsuda, Xinhui Hu, Yoshinori Shiga, Hideki Kashioka, Chiori Hori, Keiji Yasuda, Hideo Okuma, Masao Uchiyama, Eiichiro Sumita, Hisashi Kawai, Satoshi Nakamura 0001 |
MDM (2) | 10 |
| 2012 | An Evaluation of Parameter Generation Methods with Rich Context Models in HMM-Based Speech Synthesis
Shinnosuke Takamichi, Tomoki Toda, Yoshinori Shiga, Hisashi Kawai, Sakriani Sakti, Satoshi Nakamura 0001 |
INTERSPEECH | 4 |
| 2011 | Unsupervised determination of efficient Korean LVCSR units using a Bayesian Dirichlet process modelabstractKorean is an agglutinative language that does not have explicit word boundaries. It is also a highly inflective language that exhibits severe coarticulation effects. These characteristics pose a challenge in developing large-vocabulary continuous speech recognition (LVCSR) systems. Many existing Korean LVCSR systems attempt to overcome these difficulties by defining a set of "word" units using morphological analysis (rule-based) or statistical methods. These approaches usually require a great deal of linguistic knowledge or at least some explicit information about the statistical distribution of the units. However, exceptions or uncommon words (e.g., foreign proper nouns) still exist that cannot be covered by rules alone. In this paper, we investigate the use of an unsupervised, nonparametric Bayesian approach to automatically determining efficient units for a Korean LVCSR system. Specifically, we utilize a Dirichlet process model trained using Bayesian inference through block Gibbs sampling. Our approach provides a principled way of learning units without explicit linguistic knowledge or any static parameters. Experiments were conducted on a travel domain corpus, which includes many foreign words and proper nouns. In our experiments we compared our method to a set of state-of-the-art baseline systems that relied on either morphological analysis or segmentation heuristics. Our system was able to produce a considerably more compact set of "word" units than the best baseline system (the lexical dictionary was approximately half the size), with a recognition accuracy 5.89% higher in terms of the relative word error rate than the best baseline system. Sakriani Sakti, Andrew M. Finch, Ryosuke Isotani, Hisashi Kawai, Satoshi Nakamura 0001 |
ICASSP | 4 |
| 2011 | Increasing discriminative capability on MAP-based mapping function estimation for acoustic model adaptationabstractIn this study, we propose increasing discriminative power on the maximum a posteriori (MAP)-based mapping function estimation for acoustic model adaptation. Based on the effective and stable learning advantages of MAP-based estimation, we incorporate a discriminative term and derive a new objective function. By applying the new function for online mapping function estimation, we developed discriminative maximum a posteriori (DMAP) linear regression (DMAPLR) and DMAP-based ensemble speaker and speaking environment modeling (DMAP-based ESSEM). We evaluate the DMAPLR and DMAP-based ESSEM on the Aurora-2 task in a supervised adaptation mode. The experimental results show that both DMAPLR and DMAP-based ESSEM consistently provide improvements over their ML-based and MAP-based counterparts irrespective of using one, two, or three adaptation utterances. From the improvements, we confirm the strong effect of increasing discriminative capability on the MAP-based mapping function estimation. Moreover, we verify that including multiple knowledge sources in the objective function can efficiently enhance model adaptation performance. When compared with the baseline result DMAP-ESSEM achieves a 15.96% (9.21% to 7.74%) average word error rate (WER) reduction using only one adaptation utterance. Yu Tsao 0001, Ryosuke Isotani, Hisashi Kawai, Satoshi Nakamura 0001 |
ICASSP | 3 |
| 2011 | A sampling-based environment population projection approach for rapid acoustic model adaptationabstractWe propose an environment population projection (EPP) approach for rapid acoustic model adaptation to reduce environment mismatches with limited amounts of adaptation data. This approach consists of two stages: population construction and projection. In the population construction stage, we apply a sampling scheme on the adaptation data to construct an environment population based on acoustic models prepared in the training phase. With this sampling procedure, the environment samples in the population characterize diverse acoustic information embedded in the adaptation data. Next, the projection stage estimates a function to map the environment population into one set of acoustic models that matches the testing condition. With a well constructed environment population, a simple projection function can enable the EPP approach to accurately characterize the testing environment even with a small amount of adaptation data. To examine the rapid adaptation ability of EPP, we used only one adaptation utterance and tested performance in both supervised and unsupervised adaptation modes on Aurora-2 and Aurora-2J tasks. It is found that EPP achieves satisfactory performance under both modes for both tasks. On the Aurora-2J task for example, EPP gives a clear improvement of a 13.87% (8.58% to 7.39%) word error rate (WER) reduction over our baseline in the unsupervised adaptation mode. Yu Tsao 0001, Shigeki Matsuda, Shinsuke Sakai, Ryosuke Isotani, Hisashi Kawai, Satoshi Nakamura 0001 |
ICASSP | 5 |
| 2011 | Improving Related Entity Finding via Incorporating Homepages and Recognizing Fine-grained Entities
Youzheng Wu, Chiori Hori, Hisashi Kawai, Hideki Kashioka |
IJCNLP | 3 |
| 2011 | Answering Complex Questions via Exploiting Social Q&A Collection
Youzheng Wu, Chiori Hori, Hisashi Kawai, Hideki Kashioka |
IJCNLP | 3 |
| 2011 | Speaker-Adaptive Speech Synthesis Based on Eigenvoice Conversion and Language-Dependent Prosodic Conversion in Speech-to-Speech TranslationabstractThis paper describes a novel approach based on voice conversion (VC) to speaker-adaptive speech synthesis for speech-tospeech translation. Voice quality of translated speech in an output language is usually different from that of an input speaker of the translation system since a text-to-speech system is developed with another speaker’s voices in the output language. To render the input speaker’s voice quality in the translated speech, we propose a voice quality control method based on one-tomany eigenvoice conversion (EVC) and language-dependent prosodic conversion. Spectral parameters of the translated speech are effectively converted by one-to-many EVC enabling unsupervised speaker adaptation. Moreover, prosodic parameters are modified considering their global differences between the input and output languages. The effectiveness of the proposed method is confirmed by experimental evaluations on cross-lingual VC among Japanese, English, and Chinese. Index Terms: speech-to-speech translation, speech synthesis, Nobuhiko Hattori, Tomoki Toda, Hisashi Kawai, Hiroshi Saruwatari, Kiyohiro Shikano |
INTERSPEECH | 3 |
| 2011 | Adaptive Regularization Framework for Robust Voice Activity Detection
Xugang Lu, Masashi Unoki, Ryosuke Isotani, Hisashi Kawai, Satoshi Nakamura 0001 |
INTERSPEECH | 4 |
| 2011 | User Study of Spoken Decision Support SystemabstractThis paper presents the results of the user evaluation of spo- ken decision support dialogue systems, which help users select from a set of alternatives. Thus far, we have modeled this deci- sion support dialogue as a partially observable Markov decision process (POMDP), and optimized its dialogue strategy to maxi- mize the value of the user’s decision. In this paper, we present a comparative evaluation of the optimized dialogue strategy with several baseline strategies, and demonstrate that the optimized dialogue strategy that was effective in user simulation experi- ments works well in an evaluation by real users. Teruhisa Misu, Kiyonori Ohtake, Chiori Hori, Hisashi Kawai, Satoshi Nakamura 0001 |
INTERSPEECH | 4 |
| 2011 | Incorporating Regional Information to Enhance MAP-Based Stochastic Feature Compensation for Robust Speech RecognitionabstractIn this study, we propose an environment structuring framework to facilitate suitable prior density preparation for MAP-based stochastic feature matching (SFM) for robust speech recognition. We use a two-stage hierarchical structure to construct the environment structuring framework to characterize the regional information of various speaker and speaking environments. With the regional information, we derive three types of prior densities, namely clustered prior, sequential prior, and hierarchical prior densities. We also designed an integrated prior density to combine the advantages of the above three prior densities. From our experimental results on the Aurora-2 task, we confirmed that with regional information, we can obtain more suitable prior densities and thus enhance the performance of MAP-based SFM. Moreover, we found that by using the integrated prior density, which integrates multiple knowledge sources from the other three, MAP-based SFM gives the best performance. Index Terms: stochastic feature matching, SFM, hierarchical SFM, environment structuring, robust speech recognition. Yu Tsao 0001, Paul R. Dixon, Chiori Hori, Hisashi Kawai |
INTERSPEECH | 4 |
| 2011 | Estimation of Perceptual Spaces for Speaker Identities Based on the Cross-Lingual Discrimination TaskabstractThis paper reconfirms that talker identity can be transmitted across languages. Talker discrimination was examined in the ABX paradigm, where the stimuli A and B were utterances by different talkers in the same language and the stimulus X was an utterance by either of A or B in the different language. The average hit rate of this discrimination task was as high as 0.89. The mutual distance matrices were generated using the discrimination index, ′ d . By applying the multidimensional scaling, three-dimensional perceptual spaces were estimated. The features related with loudness and spectral centroid had high contribution to the perceptual dimensions. Index Terms: talker discrimination, bilingual corpus, MDS, auditory model Minoru Tsuzaki, Keiichi Tokuda, Hisashi Kawai, Jinfu Ni |
INTERSPEECH | 3 |
| 2011 | Toward Construction of Spoken Dialogue System that Evokes Users' Spontaneous Backchannels
Teruhisa Misu, Etsuo Mizukami, Yoshinori Shiga, Shinichi Kawamoto, Hisashi Kawai, Satoshi Nakamura 0001 |
SIGDIAL Conference | 5 |
| 2010 | Brazilian portuguese acoustic model training based on data borrowing from other language
Kazuhiko Abe, Sakriani Sakti, Ryosuke Isotani, Hisashi Kawai, Satoshi Nakamura 0001 |
INTERSPEECH | 4 |
| 2010 | Cluster-based language model for spoken document retrieval using NMF-based document clustering
Xinhui Hu, Ryosuke Isotani, Hisashi Kawai, Satoshi Nakamura 0001 |
INTERSPEECH | 3 |
| 2010 | Construction and evaluations of an annotated Chinese conversational corpus in travel domain for the language model of speech recognition
Xinhui Hu, Ryosuke Isotani, Hisashi Kawai, Satoshi Nakamura 0001 |
INTERSPEECH | 3 |
| 2010 | Voice activity detection in a reguarized reproducing kernel hilbert space
Xugang Lu, Masashi Unoki, Ryosuke Isotani, Hisashi Kawai, Satoshi Nakamura 0001 |
INTERSPEECH | 4 |
| 2010 | An unsupervised approach to creating web audio contents-based HMM voices
Jinfu Ni, Hisashi Kawai |
INTERSPEECH | 2 |
| 2010 | Utilizing a noisy-channel approach for Korean LVCSR
Sakriani Sakti, Ryosuke Isotani, Hisashi Kawai, Satoshi Nakamura 0001 |
INTERSPEECH | 3 |
| 2010 | Improved training of excitation for HMM-based parametric speech synthesisabstractThis paper presents an improved method of training for the unvoiced filter that comprises an excitation model, within the framework of parametric speech synthesis based on hidden Markov models. The conventional approach calculates the unvoiced filter response from the differential signal of the residual and voiced excitation estimate. The differential signal, however, includes the error generated by the voiced excitation estimates. Contaminated by the error, the unvoiced filter tends to be overestimated, which causes the synthetic speech to be noisy. In order for unvoiced filter training to obtain targets that are free from the contamination, the improved approach first separates the non-periodic component of residual signal from the periodic component. The unvoiced filter is then trained from the non-periodic component signals. Experimental results show that unvoiced filter responses trained with the new approach are clearly noiseless, in contrast to the responses trained with the conventional approach. Yoshinori Shiga, Tomoki Toda, Shinsuke Sakai, Hisashi Kawai |
INTERSPEECH | 4 |
| 2010 | Modeling Spoken Decision Making Dialogue and Optimization of its Dialogue Strategy
Teruhisa Misu, Komei Sugiura, Kiyonori Ohtake, Chiori Hori, Hideki Kashioka, Hisashi Kawai, Satoshi Nakamura 0001 |
SIGDIAL Conference | 6 |
| 2010 | Dialogue strategy optimization to assist user's decision for spoken consulting dialogue systemsabstractThis paper addresses a user model and dialogue state definition in spoken consulting dialogue systems that help users in making decision. When selecting from a set of alternatives, users have various decision criteria for making decision. Users often do not have a definite goal or criteria for selection, and thus they may find not only what kind of information the system can provide but their own preference or factors that they should emphasize. In this paper, we model such consulting dialogue as partially observable Markov decision process (POMDP). We then present an optimization of dialogue strategy to help users make better decisions. Teruhisa Misu, Komei Sugiura, Kiyonori Ohtake, Chiori Hori, Hideki Kashioka, Hisashi Kawai, Satoshi Nakamura 0001 |
SLT | 6 |
| 2009 | A close look into the probabilistic concatenation model for corpus-based speech synthesis
Shinsuke Sakai, Ranniery Maia, Hisashi Kawai, Satoshi Nakamura 0001 |
INTERSPEECH | 3 |
| 2008 | Unit database pruning based on the cost degradation criterion for concatenative speech synthesisabstractA novel method of unit database pruning for concatenative speech synthesis is proposed. The proposed method uses sums of the unit preference criterion, which are calculated from cost degradation from the optimal sequence, instead of the appearance frequencies of units, which is used in the conventional method. Therefore, the proposed method is an extension of the conventional method. Since not only the optimal units but also the other candidate units can be taken into account for pruning, unit databases can be pruned with less experimental speech synthesis. The results of a unit selection experiment on 4-hour pruned unit databases built from the original 10.6-hour database indicate that the amount of the experimental speech synthesis can be reduced to 25% of that required for the conventional method without loss of the quality of synthetic speech in terms of average cost. Nobuyuki Nishizawa, Hisashi Kawai |
ICASSP | 2 |
| 2008 | Phone duration modeling using gradient tree boosting
Junichi Yamagishi, Hisashi Kawai, Takao Kobayashi |
Speech Commun. | 2 |
| 2007 | A preselection method based on cost degradation from the optimal sequence for concatenative speech synthesis
Nobuyuki Nishizawa, Hisashi Kawai |
INTERSPEECH | 2 |
| 2006 | Evaluation result of transmission control mechanism for multimedia streams based on the multi-RTCP scheme over multiple IP-based networksabstractA variety of IP-based applications have been developed and provided due to the penetration of the infrastructure of IP communications consisting of hybrid networks combining the wireless IP network and the fixed network. This paper presents an evaluation of the Multi-RTCP Scheme over hybrid networks composed of wireless and fixed networks. We examined a rate- adaptive sending method for RTP packets based on detailed application-level QoS information on the hybrid network. The QoS information which includes each portion of the hybrid network is reported by the Multi-RTCP scheme. We present several advantages of the adaptive sending method using the Multi-RTCP scheme. Our experiments reveal network quality of service and perceived speech quality in practical conditions of the hybrid networks. Norihiro Fukumoto, Hideaki Yamada, Hisashi Kawai |
CCNC | 3 |
| 2006 | Constructing a Phonetic-Rich Speech Corpus While Controlling Time-Dependent Voice Quality Variability for English Speech SynthesisabstractThis paper presents a practical approach to constructing a large-scale speech corpus for corpus-based speech synthesis. This consists of (1) selecting a source text corpus that fits limited target domains; (2) analyzing the source text corpus to obtain the unit statistics; (3) automatically extracting prompt subjects (sentences) from the source text corpus to maximize the intended unit coverage with the given amount of text; and (4) recording prompt subjects while controlling such critical factors that cause undesirable voice variability. This paper describes related computational methods, such as a greedy algorithm for prompt selection, the proximity effects found in a real recording system, and a technique for detecting the timedependent voice variations. While the approach is demonstrated in English, it is also promising for other languages. Jinfu Ni, Toshio Hirai, Hisashi Kawai |
ICASSP (1) | 3 |
| 2006 | A Short-Latency Unit Selection Method with Redundant Search for Concatenative Speech SynthesisabstractA new method for short-latency unit selection is proposed. For prompt response in concatenative speech synthesis systems with large unit databases, waveforms should be output before all speech segment units of an utterance are determined. For that purpose, short-latency unit selection algorithms were introduced in our previous study. However, the short-latency unit selection may cause degradation of quality because units that consist of the optimal unit sequence may be pruned by forcible unit determination on the search. In the proposed method, the degradation of quality is suppressed by redundantly expanded hypotheses based on N-best search. The results of unit selection experiments in a practical configuration indicate that the proposed method is superior to the conventional DP search method when latency in unit selection is set to be short Nobuyuki Nishizawa, Hisashi Kawai |
ICASSP (1) | 2 |
| 2006 | Quick individual fitting methods of simplified hearing compensation for elderly people
Kengo Fujita, Tsuneo Kato, Hisashi Kawai |
INTERSPEECH | 3 |
| 2006 | A text-prompted distributed speaker verification system implemented on a cellular phone and a mobile terminal
Tsuneo Kato, Hisashi Kawai |
INTERSPEECH | 2 |
| 2006 | An evaluation of cost functions sensitively capturing local degradation of naturalness for segment selection in concatenative speech synthesis
Tomoki Toda, Hisashi Kawai, Minoru Tsuzaki, Kiyohiro Shikano |
Speech Commun. | 2 |
| 2006 | The ATR multilingual speech-to-speech translation systemabstractIn this paper, we describe the ATR multilingual speech-to-speech translation (S2ST) system, which is mainly focused on translation between English and Asian languages (Japanese and Chinese). There are three main modules of our S2ST system: large-vocabulary continuous speech recognition, machine text-to-text (T2T) translation, and text-to-speech synthesis. All of them are multilingual and are designed using state-of-the-art technologies developed at ATR. A corpus-based statistical machine learning framework forms the basis of our system design. We use a parallel multilingual database consisting of over 600 000 sentences that cover a broad range of travel-related conversations. Recent evaluation of the overall system showed that speech-to-speech translation quality is high, being at the level of a person having a Test of English for International Communication (TOEIC) score of 750 out of the perfect score of 990. Satoshi Nakamura 0001, Konstantin Markov, Hiromi Nakaiwa, Gen-ichiro Kikui, Hisashi Kawai, Takatoshi Jitsuhiro, Jinsong Zhang 0001, Hirofumi Yamamoto, Eiichiro Sumita, Seiichi Yamamoto |
IEEE Trans. Speech Audio Process. | 5 |
| 2005 | SNR-dependent background noise compensation of PESQ values for cellular phone speech
Kengo Fujita, Tsuneo Kato, Hideaki Yamada, Hisashi Kawai |
INTERSPEECH | 4 |
| 2005 | Analysis of major factors of naturalness degradation in concatenative synthesis
Toshio Hirai, Hisashi Kawai, Minoru Tsuzaki, Nobuyuki Nishizawa |
INTERSPEECH | 2 |
| 2005 | Estimation of intonation variation with constrained tone transformationsabstractThis paper presents a method for quantitatively estimating intonation variation in Mandarin speech. Intonation variation is relative to identical lexical tone structures, and its estimation is performed on two sets of fundamental frequency ( ) contours: one for norms and the other as variants. This is done by transforming target values in pairs from the norms to the variants in which the prosodic contribution to these contours is analyzed as sequences of targets, all of which are confined to the basic elements of the underlying lexical tone structures. The tone transformations are constrained under an assumption of the structural formulation of contours proposed previously. When the norms take the base values of the four lexical tones measured from isolated words in a neutral mood and voice, this method solves acoustic correlations of tone and intonation from the observed contours. The method was implemented on a computer, and its capability of estimating intonation variation was shown through the analysis and synthesis of contours. Jinfu Ni, Hisashi Kawai, Keikichi Hirose |
INTERSPEECH | 2 |
| 2005 | Improvement of rejection performance of keyword spotting using anti-keywords derived from large vocabulary considering acoustical similarity to keywords
Makoto Yamada, Tsuneo Kato, Masaki Naito, Hisashi Kawai |
INTERSPEECH | 4 |
| 2005 | Discriminative training and explicit duration modeling for HMM-based automatic segmentation
Yi-Jian Wu, Hisashi Kawai, Jinfu Ni, Renhua Wang |
Speech Commun. | 2 |
| 2004 | An evaluation of automatic phone segmentation for concatenative speech synthesisabstractThis paper studies the performance of automatic phone segmentation from two viewpoints: (1) temporal precision and (2) effect on the naturalness of synthetic speech. The absolute error of the phone onset time for the best 90 % and worst 10% were 4.6 ms and 25.9 ms, respectively. These values are comparable to discrepancies among human labelers. As the result of perception tests in which naturalness was paircompared between synthetic speeches generated from handsegmented data and from auto-segmented data, it was found that the latter is statistically inferior. 1. Hisashi Kawai, Tomoki Toda |
ICASSP (1) | 1 |
| 2004 | Scaling of waveform segments along the time axis for concatenative speech synthesisabstractWaveform scaling along the time axis is introduced as a pitch and duration conversion method for concatenative speech synthesis. This method will affect F/sub 0,/ duration and spectrum, although no degradation of the naturalness is caused when the scaling ratio is nearly 1. In corpus-based concatenative speech synthesis, when there are many segment candidates with various F/sub 0/ values or durations, excessive scaling may be unnecessary. The result of experiments indicated that the difference in F/sub 0/ and duration between the target and a selected segment became smaller. However, it also showed that the conventional cost function in selection cannot represent the degradation of naturalness by spectral distortion, and that the scaling range without degradation may not be enough for the pitch conversion required in our synthesizer. These problems should be improved by wider range scaling with a new cost function that also considers the degradation. Nobuyuki Nishizawa, Hisashi Kawai |
ICASSP (1) | 2 |
| 2004 | Optimizing sub-cost functions for segment selection based on perceptual evaluations in concatenative speech synthesisabstractIn concatenative speech synthesis, various factors affect the naturalness of synthetic speech. A cost for segment selection is calculated by integrating some sub-costs capturing the degradation of naturalness caused by such factors. In this paper, we optimize each sub-cost function for converting a linguistic feature or an acoustic parameter into a sub-cost based on perceptual evaluations. Two types of perceptual experiments are performed with test sets constructed by controlling the variations of sub-costs to evaluate the independent effect of each sub-cost and the interactions between them. We clarify the effectiveness of perceptually optimizing subcost functions from a result of a preference test comparing synthetic speech before and after the optimization. Tomoki Toda, Hisashi Kawai, Minoru Tsuzaki |
ICASSP (1) | 2 |
| 2004 | Minimum segmentation error based discriminative training for speech synthesis applicationabstractIn the conventional HMM-based segmentation method, the HMM training is based on MLE criteria, which links the segmentation task to the problem of distribution estimation. The HMM are built to identify the phonetic segments, not to detect the boundary. This kind of inconsistency between training and application limited the performance of segmentation. In this paper, we adopt the discriminative training method and introduce a new criterion, named minimum segmentation error (MSGE), for HMM training. In this method, a loss function directly related to the segmentation error is defined. By minimizing the overall empirical loss with the generalized probabilistic descent (GPD) algorithm, the segmentation error is also minimized. From the results on both Chinese and Japanese data, the accuracy of segmentation is improved. Moreover, this method is robust even when we do not have enough knowledge on HMM modeling, e.g. the number of states is not optimized. Yi-Jian Wu, Hisashi Kawai, Jinfu Ni, Renhua Wang |
ICASSP (1) | 2 |
| 2004 | Formulating contextual tonal variations in Mandarin
Jinfu Ni, Hisashi Kawai, Keikichi Hirose |
INTERSPEECH | 2 |
| 2004 | Using a depth-restricted search to reduce delays in unit selection
Nobuyuki Nishizawa, Hisashi Kawai |
INTERSPEECH | 2 |
| 2004 | A study on automatic detection of Japanese vowel devoicing for speech synthesisabstractIn corpus-based speech synthesis, the quality of the synthetic speech critically depends on the speech corpus. Since the high vowel in Japanese might be devoiced in the real speech, we should detect and transcribe them automatically in the corpus construction. In this paper, we apply the HMM-based method, and adopt two kinds of likelihood differences as voicing measures for different focuses. To improve the detection performance, the discriminative training is applied to voiced/ devoiced HMM training. Moreover, some features that can discriminate the voiced/devoiced units, including duration, energy and autocorrelation, are incorporated together with the likelihood differences in several methods. The experiments show different results for each high vowel, i.e. the devoicing is vowel dependent. For the vowel /i/, the discriminative training can improve the detection performance to a certain degree. And by cumulating the voicing features and the likelihood differences with optimized weights, the detection accuracy is improved. But for the vowel /u/, there is very limited improvement, even with the voicing features. Yi-Jian Wu, Hisashi Kawai, Jinfu Ni, Renhua Wang |
INTERSPEECH | 2 |
| 2003 | Tone feature extraction through parametric modeling and analysis-by-synthesis-based pattern matchingabstractA functional fundamental frequency (F/sub 0/) model is applied to extract tone peak and gliding features from Mandarin F/sub 0/ contours aiming at automatic prosodic labeling of a large scale speech corpus. Modeling four lexical tones and representing them in a parametric form based on the F/sub 0/ model, we first cluster baseline tone patterns using the LBG (Linde-Buzo-Gray) algorithm, then perform analysis-by-synthesis-based pattern matching to estimate underlying tone peaks and tone pattern types from observed F/sub 0/ contours and phonetic labels with lexical tones. Tone gliding features are re-estimated after the determination of tone peaks. 94% of the automatically estimated labels were consistent with the manual labels in an open test of 968 utterances from eight native speakers. Also, experimental results indicate that the proposed method is applicable for F/sub 0/ contour smoothing and tone verification. Jinfu Ni, Hisashi Kawai |
ICASSP (1) | 2 |
| 2003 | Segment selection considering local degradation of naturalness in concatenative speech synthesisabstractIn this paper, we investigate the effect of using a novel cost, RMS (root mean square) cost, for segment selection for concatenative text-to-speech synthesis. The RMS cost is affected not only by the total degradation of naturalness but also by the local degradation of naturalness. From the results of experiments comparing this approach with segment selection based on a conventional average cost, it is found that: (1) in the segment selection based on the RMS cost a larger number of concatenations causing slight local degradation are performed in order to avoid concatenations causing greater local degradation; and (2) the effect of the RMS cost has little dependence on the size of the corpus. Moreover, we clarify that the naturalness of synthetic speech can be slightly improved by utilizing the RMS cost. Tomoki Toda, Hisashi Kawai, Minoru Tsuzaki, Kiyohiro Shikano |
ICASSP (1) | 2 |
| 2003 | Tone pattern discrimination combining parametric modeling and maximum likelihood estimation
Jinfu Ni, Hisashi Kawai |
INTERSPEECH | 2 |
| 2003 | Optimizing integrated cost function for segment selection in concatenative speech synthesis based on perceptual evaluationsabstractThis paper describes optimizing a cost function for segment selection in concatenative Text-to-Speech based on perceptual characteristics. We use the norm of a local cost for each segment as an integrated cost function for a segment sequence to consider both the degradation of naturalness over the entire synthetic speech and the local degradation. The cost function is optimized by adjusting not only the power coefficient of the norm but also weights for sub-costs so that the integrated cost corresponds better to perceptual scores determined by perceptual experiments. As a result, it is clarified that the correspondence of the cost can be improved to a greater degree by optimizing both the weights and the power coefficient than by optimizing either the weights or the power coefficient. However, it is also clarified that the correspondence is insufficient after optimizing the integrated cost function. 1. Tomoki Toda, Hisashi Kawai, Minoru Tsuzaki |
INTERSPEECH | 2 |
| 2002 | Unit selection algorithm for Japanese speech synthesis based on both phoneme unit and diphone unitabstractThis paper proposes a novel unit selection algorithm for Japanese Text-To-Speech (TTS) systems. Since Japanese syllables consist of CV (C: Consonant, V: Vowel) or V, except when a vowel is devoiced, CV units are basic to concatenative TTS systems for Japanese. However, speech synthesized with CV units sometimes have discontinuities due to V-V concatenation; In order to alleviate such discontinuities, longer units (CV* or non-uniform units) have been proposed. However, the concatenation between V and V is still unavoidable. To address this problem, we propose a novel unit selection algorithm that incorporates not only phoneme units but also diphone units. The concatenation in the proposed algorithm is performed at the vowel center as well as at the phoneme boundary. Results of evaluation experiments clarify that the proposed algorithm outperforms the conventional algorithm. Tomoki Toda, Hisashi Kawai, Minoru Tsuzaki, Kiyohiro Shikano |
ICASSP | 2 |
| 2002 | Acoustic measures vs. phonetic features as predictors of audible discontinuity in concatenative speech synthesis
Hisashi Kawai, Minoru Tsuzaki |
INTERSPEECH | 1 |
| 2002 | Perceptual evaluation of naturalness due to substitution of Chinese syllable for concatenative speech synthesis
Jinlin Lu, Hisashi Kawai |
INTERSPEECH | 2 |
| 2002 | Design of a Mandarin sentence set for corpus-based speech synthesis by use of a multi-tier algorithm taking account of the varied prosodic and spectral characteristics
Jinfu Ni, Hisashi Kawai |
INTERSPEECH | 2 |
| 2002 | Feature extraction for unit selection in concatenative speech synthesis: comparison between AIM, LPC, and MFCC
Minoru Tsuzaki, Hisashi Kawai |
INTERSPEECH | 2 |
| 2000 | A design method of speech corpus for text-to-speech synthesis taking account of prosody
Hisashi Kawai, Seiichi Yamamoto, Norio Higuchi, Tohru Shimizu |
INTERSPEECH | 1 |
| 1998 | Recognition of connected digit speech in Japanese collected over the telephone network
Hisashi Kawai, Norio Higuchi |
ICSLP | 1 |
| 1994 | Development of a text-to-speech system for Japanese based on waveform splicingabstractA text-to-speech system for Japanese was developed based on waveform splicing. A stored unit is a sequence of phonemes segmented at vowel-consonant boundaries. Four and eight phoneme groups are distinguished for the preceding and succeeding phonemic environment, respectively. An inventory of waveform segments including frequently used 1020 units was constructed based on a statistical analysis of a text database consisting of 20 million phonemes. Each stored unit has, on average, 2.5 waveform segments with different fundamental frequency (F/sub 0/) and phoneme duration. The F/sub 0/ and phoneme duration are modified by a pitch synchronous overlap add (PSOLA) method. A time window which has a flat portion at its center (Tukey window) was adopted in place of an ordinary Hanning window. A preference test indicated that the Tukey window gives better quality when the F/sub 0/ is lowered. The articulation score of an intelligibility test was 89.2%.> Hisashi Kawai, Norio Higuchi, Tohru Shimizu, Seiichi Yamamoto |
ICASSP (1) | 1 |
| 1990 | A system for synthesizing Japanese speech from orthographic textabstractA text-to-speech conversion system for Japanese, developed for the purpose of producing high-quality speech output, is presented. It consists of four processing stages: (1) linguistic processing, (2) phonological processing, (3) control parameter generation, and (4) speech waveform generation. An overview of the whole system is presented. The innovations introduced into the second and the fourth stages, i.e. rules for generating prosodic symbols from the linguistic information in the second stage and the configuration of a new type of terminal analog speech synthesizer designed as the fourth stage, are described. The validity of the approach is confirmed by the improvements both in the prosodic and in the segmental qualities of synthesized speech.> Hiroya Fujisaki, Keikichi Hirose, Hisashi Kawai, Yasuharu Asano |
ICASSP | 3 |
| 1990 | Improvement of the synthetic speech quality of the formant-type speech synthesizer and its subjective evaluation
Norio Higuchi, Hisashi Kawai, Tohru Shimizu, Seiichi Yamamoto |
ICSLP | 2 |
| 1990 | The linguistic processing module for Japanese text-to-speech system
Tohru Shimizu, Norio Higuchi, Hisashi Kawai, Seiichi Yamamoto |
ICSLP | 3 |
| 1988 | Realization of linguistic information in the voice fundamental frequency contour of the spoken JapaneseabstractAlthough it has been well known that prosody plays an important role both in the intelligibility and in the naturalness of speech, the process of generating natural prosody from linguistic information has not been fully understood. The authors first define units of prosody of spoken Japanese on the basis of analysis of fundamental frequency contours. Prosodic words are defined by the presence of an accent component, while prosodic phrases and clauses are defined by the presence/absence of a pause and resetting/addition of a phrase component. It is shown that the accent components, representing the information concerning lexical word accent, are modified systematically by the syntactic and the discourse information. Classifications of prosodic boundaries are also presented, and the relationship between these two kinds of boundary is described.> Hiroya Fujisaki, Hisashi Kawai |
ICASSP | 2 |
| 1986 | Generation of prosodic symbols for rule-synthesis of connected speech of JapaneseabstractIn a text-to-speech conversion system, the prosodic features of speech should be synthesized strictly by rules. In order to construct a set of rules to generate the symbols for the prosodic features, such as the pause duration, the phrase command magnitude, and the accent command amplitude, speech samples of Japanese texts were analyzed quantitatively using the model for fundamental frequency contour generation. The analysis made clear the relationships between the prosodic features and the linguistic information, such as the accent type of a word, the syntactic structure of a sentence, and the discourse structure of a text. Based on these relationships, rules were constructed which generate the prosodic symbols from the linguistic information. Text-to-speech conversion was conducted using these rules. The naturalness of the intonation of the synthesized speech indicated the validity of the rules. Keikichi Hirose, Hiroya Fujisaki, Hisashi Kawai |
ICASSP | 3 |