Kei Hashimoto

dblp:79/8052 · DBLP profile ↗
← Back
43ranked-venue papers
9as first author
7since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 39 · 9 first-author · 6 since 2021Artificial intelligence and machine learning · 19 · 3 first-author · 2 since 2021
YearPublicationVenuePosition
2025 PeriodCodec: A Pitch-Controllable Neural Audio Codec Using Periodic Signals for Singing Voice Synthesis
Masato Takagi, Miku Nishihara, Yukiya Hono, Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda
INTERSPEECH4
2024 PeriodGrad: Towards Pitch-Controllable Neural Vocoder Based on a Diffusion Probabilistic Model
abstract
This paper presents a neural vocoder based on a denoising diffusion probabilistic model (DDPM) incorporating explicit periodic signals as auxiliary conditioning signals. Recently, DDPM-based neural vocoders have gained prominence as non-autoregressive models that can generate high-quality waveforms. The neural vocoders based on DDPM have the advantage of training with a simple time-domain loss. In practical applications, such as singing voice synthesis, there is a demand for neural vocoders to generate high-fidelity speech waveforms with flexible pitch control. However, conventional DDPM-based neural vocoders struggle to generate speech waveforms under such conditions. Our proposed model aims to accurately capture the periodic structure of speech waveforms by incorporating explicit periodic signals. Experimental results show that our model improves sound quality and provides better pitch control than conventional DDPM-based neural vocoders.
Yukiya Hono, Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda
ICASSP2
2023 Singing Voice Synthesis Based on a Musical Note Position-Aware Attention Mechanism
abstract
This paper proposes a novel sequence-to-sequence (seq2seq) model with a musical note position-aware attention mechanism for singing voice synthesis (SVS). A seq2seq modeling approach that can simultaneously perform acoustic and temporal modeling is attractive. However, due to the difficulty of the temporal modeling of singing voices, many recent SVS systems with an encoder-decoder-based model still rely on explicitly on duration information generated by additional modules. Although some studies perform simultaneous modeling using seq2seq models with an attention mechanism, they have insufficient robustness against temporal modeling. The proposed attention mechanism is designed to estimate the attention weights by considering the rhythm given by the musical score. Furthermore, several techniques are also introduced to improve the modeling performance of the singing voice. Experimental results indicated that the proposed model is effective in terms of both naturalness and robustness of timing.
Yukiya Hono, Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda
ICASSP2
2023 Embedding a Differentiable Mel-Cepstral Synthesis Filter to a Neural Speech Synthesis System
abstract
This paper integrates a classic mel-cepstral synthesis filter into a modern neural speech synthesis system towards end-to-end controllable speech synthesis. Since the mel-cepstral synthesis filter is explicitly embedded in neural waveform models in the proposed system, both voice characteristics and the pitch of synthesized speech are highly controlled via a frequency warping parameter and fundamental frequency, respectively. We implement the mel-cepstral synthesis filter as a differentiable and GPU-friendly module to enable the acoustic and waveform models in the proposed system to be simultaneously optimized in an end-to-end manner. Experiments show that the proposed system improves speech quality from a baseline system maintaining controllability. The core PyTorch modules used in the experiments are publicly available on GitHub1.
Takenori Yoshimura, Shinji Takaki, Kazuhiro Nakamura, Keiichiro Oura, Yukiya Hono, Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda
ICASSP6
2022 Autoregressive Variational Autoencoder with a Hidden Semi-Markov Model-Based Structured Attention for Speech Synthesis
abstract
This paper proposes an autoregressive speech synthesis model based on the variational autoencoder incorporating latent sequence representation for acoustic and linguistic features and the structure of a hidden semi-Markov model (HSMM). Although autoregressive models can provide efficient and accurate modeling of acoustic features, they have exposure bias, i.e., the mismatch between training (teacher-forcing) and inference (free-running). To overcome this problem, we introduce an autoregressive latent variable sequence, rather than using autoregressive generation of observations. Latent representation of alignment using HSMM-based structured attention mechanism enables the use of a completely consistent training algorithm for acoustic modeling with explicit duration models. Experimental results indicate that the proposed model outperformed baselines in subjective naturalness.
Takato Fujimoto, Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda
ICASSP2
2021 Periodnet: A Non-Autoregressive Waveform Generation Model with a Structure Separating Periodic and Aperiodic Components
abstract
We propose PeriodNet, a non-autoregressive (non-AR) waveform generation model with a new model structure for modeling periodic and aperiodic components in speech waveforms. The non-AR waveform generation models can generate speech waveforms parallelly and can be used as a speech vocoder by conditioning an acoustic feature. Since a speech waveform contains periodic and aperiodic components, both components should be appropriately modeled to generate a high-quality speech waveform. However, it is difficult to decompose the components from a natural speech waveform in advance. To address this issue, we propose a parallel model and a series model structure separating periodic and aperiodic components. The features of our proposed models are that explicit periodic and aperiodic signals are taken as input, and external periodic/aperiodic decomposition is not needed in training. Experiments using a singing voice corpus show that our proposed structure improves the naturalness of the generated waveform. We also show that the speech waveforms with a pitch outside of the training data range can be generated with more naturalness.
Yukiya Hono, Shinji Takaki, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda
ICASSP3
2021 Sinsy: A Deep Neural Network-Based Singing Voice Synthesis System
abstract
This paper presents Sinsy, a deep neural network (DNN)-based singing voice synthesis (SVS) system. In recent years, DNNs have been utilized in statistical parametric SVS systems, and DNN-based SVS systems have demonstrated better performance than conventional hidden Markov model-based ones. SVS systems are required to synthesize a singing voice with pitch and timing that strictly follow a given musical score. Additionally, singing expressions that are not described on the musical score, such as vibrato and timing fluctuations, should be reproduced. The proposed system is composed of four modules: a time-lag model, a duration model, an acoustic model, and a vocoder, and singing voices can be synthesized taking these characteristics of singing voices into account. To better model a singing voice, the proposed system incorporates improved approaches to modeling pitch and vibrato and better training criteria into the acoustic model. In addition, we incorporated PeriodNet, a non-autoregressive neural vocoder with robustness for the pitch, into our systems to generate a high-fidelity singing voice waveform. Moreover, we propose automatic pitch correction techniques for DNN-based SVS to synthesize singing voices with correct pitch even if the training data has out-of-tune phrases. Experimental results show our system can synthesize a singing voice with better timing, more natural vibrato, and correct pitch, and it can achieve better mean opinion scores in subjective evaluation tests.
Yukiya Hono, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda
IEEE ACM Trans. Audio Speech Lang. Process.2
2020 Semi-Supervised Learning Based on Hierarchical Generative Models for End-to-End Speech Synthesis
abstract
This paper proposes a general framework of semi-supervised learning based on hierarchical generative models and adapts it to a Japanese end-to-end text-to-speech (TTS) system. In English TTS, several end-to-end systems have recently achieved sound quality close to that of natural human speech. However, in non-alphabetic languages such as Japanese, it is difficult to realize true text-input end-to-end TTS due to character diversity and pitch accents. To address this problem, we propose end-to-end TTS based on semi-supervised learning that makes the most of existing data consisting of any combination of text, phoneme, and waveform as training data. To demonstrate the effectiveness of the proposed system, listening tests were conducted for pronunciation and naturalness. Our results show that the proposed system improves both pronunciation and naturalness.
Takato Fujimoto, Shinji Takaki, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda
ICASSP3
2020 Fast and High-Quality Singing Voice Synthesis System Based on Convolutional Neural Networks
abstract
The present paper describes singing voice synthesis based on convolutional neural networks (CNNs). Singing voice synthesis systems based on deep neural networks (DNNs) are currently being proposed and are improving the naturalness of synthesized singing voices. As singing voices represent a rich form of expression, a powerful technique to model them accurately is required. In the proposed technique, long-term dependencies of singing voices are modeled by CNNs. An acoustic feature sequence is generated for each segment that consists of long-term frames, and a natural trajectory is obtained without the parameter generation algorithm. Furthermore, a computational complexity reduction technique, which drives the DNNs in different time units depending on type of musical score features, is proposed. Experimental results show that the proposed method can synthesize natural sounding singing voices much faster than the conventional method.
Kazuhiro Nakamura, Shinji Takaki, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda
ICASSP3
2020 Hierarchical Multi-Grained Generative Model for Expressive Speech Synthesis
abstract
This paper proposes a hierarchical generative model with a multi-grained latent variable to synthesize expressive speech. In recent years, fine-grained latent variables are introduced into the text-to-speech synthesis that enable the fine control of the prosody and speaking styles of synthesized speech. However, the naturalness of speech degrades when these latent variables are obtained by sampling from the standard Gaussian prior. To solve this problem, we propose a novel framework for modeling the fine-grained latent variables, considering the dependence on an input text, a hierarchical linguistic structure, and a temporal structure of latent variables. This framework consists of a multi-grained variational autoencoder, a conditional prior, and a multi-level auto-regressive latent converter to obtain the different time-resolution latent variables and sample the finer-level latent variables from the coarser-level ones by taking into account the input text. Experimental results indicate an appropriate method of sampling fine-grained latent variables without the reference signal at the synthesis stage. Our proposed framework also provides the controllability of speaking style in an entire utterance.
Yukiya Hono, Kazuna Tsuboi, Kei Sawada, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda
INTERSPEECH4
2019 Singing Voice Synthesis Based on Generative Adversarial Networks
abstract
This paper proposes a generative adversarial training method for deep neural network (DNN)-based singing voice synthesis. The DNN-based approach has been used in statistical parametric singing voice synthesis and improved the naturalness of the synthesized singing voice [1]. Recently, generative adversarial networks (GANs) [2] have attracted significant attention in various machine learning research areas including speech synthesis [3]. GANs have achieved great success in modeling the distributions of complex data, and they have the potential to alleviate over-smoothing problem on the generated speech parameters in speech synthesis. In this paper, we propose a DNN-based singing voice synthesis system incorporating the GAN. Experimental results show that the proposed method outperforms the conventional method in the naturalness of the synthesized singing voice.
Yukiya Hono, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda
ICASSP2
2019 Speaker-dependent Wavenet-based Delay-free Adpcm Speech Coding
abstract
This paper proposes a WaveNet-based delay-free adaptive differential pulse code modulation (ADPCM) speech coding system. The WaveNet generative model, which is a state-of-the-art model for neural-network-based speech waveform synthesis, is used as the adaptive predictor in ADPCM. To further improve speech quality, mel-cepstrum-based noise shaping and postfiltering were integrated with the proposed ADPCM system. Both objective and subjective evaluation results indicate that the proposed ADPCM system outperformed not only the conventional ADPCM system based on ITU-T Recommendation G.726 but also the ADPCM system based on adaptive mel-cepstral analysis.
Takenori Yoshimura, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda
ICASSP2
2018 Image Recognition Based on Separable Lattice Hmms Using a Deep Neural Network for Output Probability Distributions
abstract
This paper proposes an image recognition method based on separable lattice hidden Markov models (SLHMMs) using a deep neural network (DNN) for output probability distributions. The geometric variations of the object to be recognized, e.g., size and location, are essential in image recognition. SLHMMs, which have been proposed to reduce the effect of geometric variations, can perform elastic matching both horizontally and vertically. Gaussian distributions are typical for modeling the output distribution of SLHMMs. However, these distributions may not be sufficient to represent patterns of image regions. Our method integrates SLHMMs and a DNN and can be used to model an image effectively by explicit modeling of the generative process based on SLHMMs and advanced feature classification based on a DNN. image recognition experiments showed that the proposed method improves recognition performance.
Eiji Ichikawa, Kei Sawada, Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda
ICASSP3
2018 Statistical Voice Conversion Based on Wavenet
abstract
This paper proposes a voice conversion technique based on WaveNet to directly generate target audio waveforms from acoustic features of a source speaker. In voice conversion based on statistical models, the relation between acoustic features, such as spectral parameters, extracted from source and target audio waveforms is generally modeled using statistical models, such as Gaussian mixture models and neural networks. Although modeling the relation between acoustic features is reasonable and efficient, these models are not optimized for predicting target audio waveforms because the vocoder parameters are used as intermediate representations. To overcome this problem, we developed a voice conversion method to model the relation between target audio waveforms and acoustic features extracted from source audio waveforms using WaveNet, which is a generative model for audio waveforms. The proposed model can directly generate converted audio waveforms without vocoders. Experimental results indicate that the proposed method can generate a more naturally sounding converted speech than that using a conventional DNN method.
Jumpei Niwa, Takenori Yoshimura, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda
ICASSP3
2018 WaveNet-Based Zero-Delay Lossless Speech Coding
abstract
This paper presents a WaveNet-based zero-delay lossless speech coding technique for high-quality communications. The WaveNet generative model, which is a state-of-the-art model for neural-network-based speech waveform synthesis, is used in both the encoder and decoder. In the encoder, discrete speech signals are losslessly compressed using sample-by-sample entropy coding. The decoder fully reconstructs the original speech signals from the compressed signals without algorithmic delay. Experimental results show that the proposed coding technique can transmit speech audio waveforms with 50% their original bit rate and the WaveNet-based speech coder remains effective for unknown speakers.
Takenori Yoshimura, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda
SLT2
2018 Mel-Cepstrum-Based Quantization Noise Shaping Applied to Neural-Network-Based Speech Waveform Synthesis
abstract
This paper presents a mel-cepstrum-based quantization noise shaping method for improving the quality of synthetic speech generated by neural-network-based speech waveform synthesis systems. Since mel-cepstral coefficients closely match the characteristics of human auditory perception, the proposed method effectively masks the white noise introduced by the quantization typically used in neural-network-based speech waveform synthesis systems. The paper also describes a computationally efficient implementation of the proposed method using the structure of the mel-log spectrum approximation filter. Experiments using the WaveNet generative model, which is a state-of-the-art model for neural-network-based speech waveform synthesis, showed that speech quality is significantly improved by the proposed method.
Takenori Yoshimura, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda
IEEE ACM Trans. Audio Speech Lang. Process.2
2017 Image recognition based on discriminative models using features generated from separable lattice HMMS
abstract
This paper presents an image recognition technique based on discriminative models using features generated from separable lattice hidden Markov models (SL-HMMs). A major problem in image recognition is that the recognition performance is degraded by geometric variations such as that in position and size of the object to be recognized. SL-HMMs have been proposed to solve this problem. SL-HMMs are an extension of HMMs with size and locational invariances based on state transitions. An SL-HMM is a generative model and can represent generation processes of observations well. However, there is a possibility that the recognition performance of generative models is inferior to that of discriminative models because discriminative models are specialized to identification. In this paper, we propose image recognition based on log linear models (LLMs) using features extracted from SL-HMMs. The proposed method can extract features invariant to geometric variations by using SL-HMMs and built an accurate classifier based on discriminative models with the extracted features. Face recognition experiments showed that the proposed method obtained higher recognition rates than SL-HMMs and convolutional neural networks based methods.
Yoshinari Tsuzuki, Kei Sawada, Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda
ICASSP3
2017 Articulatory Text-to-Speech Synthesis Using the Digital Waveguide Mesh Driven by a Deep Neural Network
abstract
Following recent advances in direct modeling of the speech waveform using a deep neural network, we propose a novel method that directly estimates a physical model of the vocal tract from the speech waveform, rather than magnetic resonance imaging data. This provides a clear relationship between the model and the size and shape of the vocal tract, offering considerable flexibility in terms of speech characteristics such as age and gender. Initial tests indicate that despite a highly simplified physical model, intelligible synthesized speech is obtained. This illustrates the potential of the combined technique for the control of physical models in general, and hence the generation of more natural-sounding synthetic speech.
Amelia Jane Gully, Takenori Yoshimura, Damian T. Murphy, Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda
INTERSPEECH4
2017 Simultaneous Optimization of Multiple Tree-Based Factor Analyzed HMM for Speech Synthesis
abstract
This paper proposes a novel method to build multiple decision trees as a structure of factor analyzed hidden Markov model for speech synthesis. In the proposed method, the multiple decision trees grow simultaneously rather than sequentially to take into account the relationship between the trees. However, the simultaneous growing is computationally infeasible due to an exponential increase in the number of tree structures to be evaluated. To solve the problem, we further propose two computational complexity reduction algorithms that achieve a significant reduction in the computational time. Experimental results show that the proposed method outperforms the conventional one based on a single decision tree.
Takenori Yoshimura, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda
IEEE ACM Trans. Audio Speech Lang. Process.2
2016 Trajectory training considering global variance for speech synthesis based on neural networks
abstract
This paper proposes a new training method of deep neural networks (DNNs) for statistical parametric speech synthesis. DNNs are recently used as acoustic models that represent mapping functions from linguistic features to acoustic features in statistical parametric speech synthesis. There are problems to be solved in conventional DNN-based speech synthesis: 1) the inconsistency between the training and synthesis criteria; and 2) the over-smoothing of the generated parameter trajectories. In this paper, we introduce the parameter trajectory generation process considering the global variance (GV) into the training of DNNs. A unified framework which consistently uses the same criterion in both training and synthesis can be obtained and the model parameters are optimized for parameter generation considering the GV in the proposed method. Experimental results show that the proposed method outperforms the conventional method in the naturalness of synthesized speech.
Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda
ICASSP1
2016 Privacy-preserving sound to degrade automatic speaker verification performance
abstract
In this paper, a privacy protection method to prevent speaker identification from recorded speech is proposed and evaluated. Although many techniques for preserving various private information included in speech have been proposed, their impacts on human speech communication in physical space are not taken into account. To overcome this problem, this paper proposes privacy-preserving sound as a privacy protection method. The privacy-preserving sound can degrade speaker verification performance without interfering with human speech communication in physical space. To make a first step toward solving this problem, suitable sound characteristics for preserving privacy are evaluated in terms of the speaker verification performance and speech intelligibility. The experimental results show that appropriate sound can efficiently degrade the speaker verification performance without degrading speech intelligibility.
Kei Hashimoto, Junichi Yamagishi, Isao Echizen
ICASSP1
2016 Redefining the Linguistic Context Feature Set for HMM and DNN TTS Through Position and Parsing
Rasmus Dall, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda
INTERSPEECH2
2016 Voice Conversion Based on Trajectory Model Training of Neural Networks Considering Global Variance
Naoki Hosaka, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda
INTERSPEECH2
2016 Singing Voice Synthesis Based on Deep Neural Networks
Masanari Nishimura, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda
INTERSPEECH2
2015 The effect of neural networks in statistical parametric speech synthesis
abstract
This paper investigates how to use neural networks in statistical parametric speech synthesis. Recently, deep neural networks (DNNs) have been used for statistical parametric speech synthesis. However, the specific way how DNNs should be used in statistical parametric speech synthesis has not been studied thoroughly. A generation process of statistical parametric speech synthesis based on generative models can be divided into several components, and those components can be represented by DNNs. In this paper, the effect of DNNs for each component is investigated by comparing DNNs with generative models. Experimental results show that the use of a DNN as acoustic models is effective and the parameter generation combined with a DNN improves the naturalness of synthesized speech.
Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda
ICASSP1
2015 Simultaneous optimization of multiple tree structures for factor analyzed HMM-based speech synthesis
Takenori Yoshimura, Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda
INTERSPEECH2
2014 Integration of speaker and pitch adaptive training for HMM-based singing voice synthesis
abstract
A statistical parametric approach to singing voice synthesis based on hidden Markov models (HMMs) has been growing in popularity over the last few years. The spectrum, excitation, vibrato, and duration of the singing voice in this approach are simultaneously modeled with context-dependent HMMs and waveforms are generated from the HMMs themselves. Since HMM-based singing voice synthesis systems are “corpus-based,” the HMMs corresponding to contextual factors that rarely appear in the training data cannot be well-trained. However, it may be difficult to prepare a large enough quantity of singing voice data sung by one singer. Furthermore, the pitch included in each song is imbalanced, and there is the vocal range of the singer. In this paper, we propose “singer adaptive training” which can solve the data sparse-ness problem. Experimental results demonstrated that the proposed technique improved the quality of the synthesized singing voices.
Kanako Shirota, Kazuhiro Nakamura, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda
ICASSP3
2014 A mel-cepstral analysis technique restoring high frequency components from low-sampling-rate speech
Kazuhiro Nakamura, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda
INTERSPEECH2
2013 Separable lattice 2-D HMMS introducing state duration control for recognition of images with various variations
abstract
In this paper, an extension of separable lattice HMMs (SL-HMM) is described that introduces state duration control for dealing with images with various variations. SL-HMM are generative models that have size and location invariances based on state transition of HMMs. An extended model that has the structure of hidden semi-Markov models (HSMMs) in which the state duration probability is explicitly modeled by parametric distributions is also proposed. However, in this model, each state duration in a Markov chain is independent. It is supposed that each state duration should have a correlation. Therefore, in this paper, we propose a novel model that solves this problem by introducing variables representing the correlation among the state durations. Face recognition experiments show that the proposed model improved the recognition performance for images with size, locational, and rotational variations.
Takaya Makino, Shinji Takaki, Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda
ICASSP3
2013 Integration of acoustic modeling and mel-cepstral analysis for HMM-based speech synthesis
abstract
In this paper, a novel approach for integrating acoustic modeling and mel-cepstral analysis is proposed. The aim of HMM-based speech synthesis is to model speech waveforms with a statistical model. However, the conventional techniques divide the modeling process into two steps: the frame by frame feature extraction step and the acoustic modeling step. Although it is reasonably effective, the deterioration of speech quality is caused by the divide of the objective function. In this paper, we propose an approach to modeling them as an integrative model and show the possibility of improving synthesized speech.
Kazuhiro Nakamura, Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda
ICASSP2
2012 Face recognition based on separable lattice 2-D HMMS using variational bayesian method
abstract
This paper proposes an image recognition technique based on separable lattice 2-D HMMs (SL2D-HMMs) using the variational Bayesian method. SL2D-HMMs have been proposed to reduce the effect of geometric variations, e.g., size and location. The maximum likelihood criterion had previously been used in training SL2D-HMMs. However, in many image recognition tasks, it is difficult to use sufficient training data, and it suffers from the over-fitting problem. A higher generalization ability based on model marginalization is expected by applying the Bayesian criterion and useful prior information on model parameters can be utilized as prior distributions. Experiments on face recognition indicated that the proposed method improved image recognition.
Kei Sawada, Akira Tamamori, Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda
ICASSP3
2012 A model structure integration based on a Bayesian framework for speech recognition
abstract
This paper proposes an acoustic modeling technique based on Bayesian framework using multiple model structures for speech recognition. The Bayesian approach is a statistical technique for estimating reliable predictive distributions by marginalizing model parameters, and its effectiveness in HMM-based speech recognition has been reported. Although the basic idea underlying the Bayesian approach is to treat all parameters as random variables, only one model structure is still selected in the conventional method. Multiple model structures are treated as latent variables in the proposed method and integrated based on the Bayesian framework. Furthermore, we applied deterministic annealing to the training algorithm to estimate appropriate acoustic models. The proposed method effectively utilizes multiple model structures, especially in the early stage of training and this leads to better predictive distributions and improvement of recognition performance.
Sayaka Shiota, Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda
ICASSP2
2012 A Bayesian Approach to Speaker Recognition Based on GMMs Using Multiple Model Structures
Takafumi Hattori, Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda
INTERSPEECH2
2012 Impacts of machine translation and speech synthesis on speech-to-speech translation
Kei Hashimoto, Junichi Yamagishi, William J. Byrne, Simon King 0001, Keiichi Tokuda
Speech Commun.1
2011 An analysis of machine translation and speech synthesis in speech-to-speech translation system
abstract
This paper provides an analysis of the impacts of machine translation and speech synthesis on speech-to-speech translation systems. The speech-to-speech translation system consists of three components: speech recognition, machine translation and speech synthesis. Many techniques for integration of speech recognition and machine translation have been proposed. However, speech synthesis has not yet been considered. Therefore, in this paper, we focus on machine translation and speech synthesis, and report a subjective evaluation to analyze the impact of each component. The results of these analyses show that the naturalness and intelligibility of synthesized speech are strongly affected by the fluency of the translated sentences.
Kei Hashimoto, Junichi Yamagishi, William J. Byrne, Simon King 0001, Keiichi Tokuda
ICASSP1
2011 Multi-Speaker Modeling with Shared Prior Distributions and Model Structures for Bayesian Speech Synthesis
abstract
This paper investigates a multi-speaker modeling technique with shared prior distributions and model structures for Bayesian speech synthesis. The quality of synthesized speech is improved by selecting appropriate model structures in HMMbased speech synthesis. Bayesian approach is known to work for such model selection. However, the result is strongly affected by prior distributions of model parameters. Therefore, determination of prior distributions and selection of model structures should be performed simultaneously. This paper investigates prior distributions and model structures in the situation where training data of multiple speakers are available. The prior distributions and model structures which represent acoustic features common to every speakers can be obtained by sharing them between multiple speaker-dependent models. Index Terms: speech synthesis, Bayesian approach, prior distribution, context clustering, multi-speaker modeling A statistical parametric speech synthesis system based on hidden Markov models (HMMs) was recently developed. In HMM-based speech synthesis, the spectrum, excitation, and duration of speech are simultaneously modeled with HMMs, and speech parameter sequences are generated from the HMMs themselves [1]. The maximum likelihood (ML) criterion has typically been used for training HMMs and generating speech parameters. The ML criterion guarantees that the ML estimates approach the true values of the parameters. However, since the ML criterion produces a point estimate of the model parameters, its estimation accuracy may degrade when the amount of training data is insufficient. In the Bayesian approach, all variables introduced when the models are parameterized, such as model parameters and latent variables, are treated as random variables, and their posterior distributions are obtained by the Bayes theorem. The Bayesian approach can generally construct a more robust model than the ML approach by estimating posterior distributions. Recently, Bayesian speech synthesis has been proposed as a Bayesian framework for statistical parametric speech synthesis (e.g., HMM-based speech synthesis), and it shows good performance [2]. In Bayesian speech synthesis, all processes for constructing the system can be derived from a single predictive distribution that directly represents the problem of speech synthesis. The quality of synthesized speech is improved by selecting appropriate model structures in HMM-based speech synthesis. Although the Bayesian approach is known to work for such model selection, the results are strongly affected by prior distributions of the model parameters. Therefore, in Bayesian speech synthesis, determination of prior distributions and selection of model structures should be performed simultaneously. To overcome this problem, we have proposed Bayesian context clustering using cross validation [3]. In this method, prior distributions are determined by using a part of training data, and model structures are evaluated by using the determined prior distribution based on cross validation. In this paper, we investigates prior distributions and model
Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda
INTERSPEECH1
2010 A Deterministic Annealing-Based Training Algorithm For Statistical Machine Translation Models
Pascual Martínez-Gómez, Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda, Germán Sanchis-Trilles
EAMT2
2009 A Bayesian approach to HMM-based speech synthesis
abstract
This paper proposes a new framework of speech synthesis based on the Bayesian approach. The Bayesian method is a statistical technique for estimating reliable predictive distributions by marginalizing model parameters. In the proposed framework, all processes for constructing the system can be derived from one single predictive distribution which represents the basic problem of speech synthesis directly. Using HMM as the likelihood function and assuming some approximations, it can be regarded as an application of the variational Bayesian method to the HMM-based speech synthesis. Experimental results show that the proposed method outperforms the conventional one in a subjective test.
Kei Hashimoto, Heiga Zen, Yoshihiko Nankaku, Takashi Masuko, Keiichi Tokuda
ICASSP1
2009 A Bayesian approach to Hidden Semi-Markov Model based speech synthesis
abstract
This paper proposes a Bayesian approach to hidden semi-Markov model (HSMM) based speech synthesis. Recently, hid-den Markov model (HMM) based speech synthesis based on the Bayesian approach was proposed. The Bayesian approach is a statistical technique for estimating reliable predictive distribu-tions by treating model parameters as random variables. In the Bayesian approach, all processes for constructing the system are derived from one single predictive distribution which exactly represents the problem of speech synthesis. However, there is an inconsistency between training and synthesis: although the speech is synthesized from HMMs with explicit state duration probability distributions, HMMs are trained without them. In this paper, we introduce an HSMM, which is an HMM with explicit state duration probability distributions, into the HMM-based Bayesian speech synthesis system. Experimental results show that the use of HSMM improves the naturalness of the synthesized speech. Index Terms: speech synthesis, HSMM, Bayesian approach 1.
Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda
INTERSPEECH1
2009 Deterministic annealing based training algorithm for Bayesian speech recognition
abstract
This paper proposes a deterministic annealing based training algorithm for Bayesian speech recognition. The Bayesian method is a statistical technique for estimating reliable predictive distributions by marginalizing model parameters. However, the local maxima problem in the Bayesian method is more serious than in the ML-based approach, because the Bayesian method treats not only state sequences but also model parameters as latent variables. The deterministic annealing EM (DAEM) algorithm has been proposed to improve the local maxima problem in the EM algorithm, and its effectiveness has been reported in HMMbased speech recognition using ML criterion. In this paper, the DAEM algorithm is applied to Bayesian speech recognition to relax the local maxima problem. Speech recognition experiments show that the proposed method achieved a higher performance than the conventional methods.
Sayaka Shiota, Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda
INTERSPEECH2
2008 Bayesian context clustering using cross valid prior distribution for HMM-based speech recognition
abstract
This paper proposes a prior distribution determination tech-nique using cross validation for speech recognition based on the Bayesian approach. The Bayesian method is a statisti-cal technique for estimating reliable predictive distributions by marginalizing model parameters and its approximate version, the variational Bayesian method has been applied to HMM-based speech recognition. Since prior distributions represent-ing prior information about model parameters affect the pos-terior distributions and model selection, the determination of prior distributions is an important problem. However, it has not been thoroughly investigate in speech recognition. The pro-posed method can determine reliable prior distributions with-out tuning parameters and select an appropriate model struc-ture dependently on the amount of training data. Continu-ous phoneme recognition experiments show that the proposed method achieved a higher performance than the conventional methods. Index Terms: variational Bayes, cross validation, context clus-tering, continuous phoneme recognition
Kei Hashimoto, Heiga Zen, Yoshihiko Nankaku, Akinobu Lee, Keiichi Tokuda
INTERSPEECH1
2008 Speaker recognition based on variational Bayesian method
Tatsuya Ito, Kei Hashimoto, Yoshihiko Nankaku, Akinobu Lee, Keiichi Tokuda
INTERSPEECH2
2008 Acoustic modeling based on model structure annealing for speech recognition
abstract
This paper proposes an HMM training technique using multiple phonetic decision trees and evaluates it in speech recognition. In the use of context dependent models, the decision tree based context clustering is applied to find a parameter tying structure. However, the clustering is usually performed based on statistics of HMM state sequences which are obtained by unreliable models without context clustering. To avoid this problem, we optimize the decision trees and HMM state sequences simultaneously. In the proposed method, this is performed by maximum likelihood (ML) estimation of a newly defined statistical model which includes multiple decision trees as hidden variables. Applying the deterministic annealing expectation maximization (DAEM) algorithm and using multiple decision trees in early stage of model training, state sequences are reliably estimated. In continuous phoneme recognition experiments, the proposed method can improve the recognition performance.
Sayaka Shiota, Kei Hashimoto, Heiga Zen, Yoshihiko Nankaku, Akinobu Lee, Keiichi Tokuda
INTERSPEECH2