EDBT 2026 Demo / reviewers in the wild / expert
Keiichi Tokuda
dblp:25/369
· DBLP profile ↗
185ranked-venue papers
15as first author
9since 2021 · last 2025
0000-0001-6143-0133ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 169 · 13 first-author · 8 since 2021Artificial intelligence and machine learning · 93 · 6 first-author · 3 since 2021Systems, architecture and hardware · 1Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | PeriodCodec: A Pitch-Controllable Neural Audio Codec Using Periodic Signals for Singing Voice Synthesis
Masato Takagi, Miku Nishihara, Yukiya Hono, Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda |
INTERSPEECH | 6 |
| 2024 | PeriodGrad: Towards Pitch-Controllable Neural Vocoder Based on a Diffusion Probabilistic ModelabstractThis paper presents a neural vocoder based on a denoising diffusion probabilistic model (DDPM) incorporating explicit periodic signals as auxiliary conditioning signals. Recently, DDPM-based neural vocoders have gained prominence as non-autoregressive models that can generate high-quality waveforms. The neural vocoders based on DDPM have the advantage of training with a simple time-domain loss. In practical applications, such as singing voice synthesis, there is a demand for neural vocoders to generate high-fidelity speech waveforms with flexible pitch control. However, conventional DDPM-based neural vocoders struggle to generate speech waveforms under such conditions. Our proposed model aims to accurately capture the periodic structure of speech waveforms by incorporating explicit periodic signals. Experimental results show that our model improves sound quality and provides better pitch control than conventional DDPM-based neural vocoders. Yukiya Hono, Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda |
ICASSP | 4 |
| 2023 | Singing Voice Synthesis Based on a Musical Note Position-Aware Attention MechanismabstractThis paper proposes a novel sequence-to-sequence (seq2seq) model with a musical note position-aware attention mechanism for singing voice synthesis (SVS). A seq2seq modeling approach that can simultaneously perform acoustic and temporal modeling is attractive. However, due to the difficulty of the temporal modeling of singing voices, many recent SVS systems with an encoder-decoder-based model still rely on explicitly on duration information generated by additional modules. Although some studies perform simultaneous modeling using seq2seq models with an attention mechanism, they have insufficient robustness against temporal modeling. The proposed attention mechanism is designed to estimate the attention weights by considering the rhythm given by the musical score. Furthermore, several techniques are also introduced to improve the modeling performance of the singing voice. Experimental results indicated that the proposed model is effective in terms of both naturalness and robustness of timing. Yukiya Hono, Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda |
ICASSP | 4 |
| 2023 | Lightweight, Multi-Speaker, Multi-Lingual Indic Text-to-SpeechabstractThe Lightweight, Multi-speaker, Multi-lingual Indic Text-to-Speech (LIMMITS’23) challenge is organized as part of the ICASSP 2023 signal processing grand challenge. LIMMITS’23 aims at the development of a lightweight, multi-speaker, multi-lingual Text to Speech (TTS) model using datasets in Marathi, Hindi, and Telugu. The challenge encourages the advancement of TTS in Indian Languages as well as the development of techniques involved in TTS data selection and model compression. The 3 tracks of LIMMITS’23 have provided an opportunity for various researchers and practitioners around the world to explore the state of the art in TTS research. Abhayjeet Singh, Amala Nagireddi, Deekshitha G, Jesuraja Bandekar, Roopa R., Sandhya Badiger, Sathvik Udupa, Prasanta Kumar Ghosh, Hema A. Murthy, Heiga Zen, Pranaw Kumar, Kamal Kant, Amol Bole, Bira Chandra Singh, Keiichi Tokuda, Mark Hasegawa-Johnson, Philipp Olbrich |
ICASSP | 15 |
| 2023 | Embedding a Differentiable Mel-Cepstral Synthesis Filter to a Neural Speech Synthesis SystemabstractThis paper integrates a classic mel-cepstral synthesis filter into a modern neural speech synthesis system towards end-to-end controllable speech synthesis. Since the mel-cepstral synthesis filter is explicitly embedded in neural waveform models in the proposed system, both voice characteristics and the pitch of synthesized speech are highly controlled via a frequency warping parameter and fundamental frequency, respectively. We implement the mel-cepstral synthesis filter as a differentiable and GPU-friendly module to enable the acoustic and waveform models in the proposed system to be simultaneously optimized in an end-to-end manner. Experiments show that the proposed system improves speech quality from a baseline system maintaining controllability. The core PyTorch modules used in the experiments are publicly available on GitHub1. Takenori Yoshimura, Shinji Takaki, Kazuhiro Nakamura, Keiichiro Oura, Yukiya Hono, Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda |
ICASSP | 8 |
| 2022 | Autoregressive Variational Autoencoder with a Hidden Semi-Markov Model-Based Structured Attention for Speech SynthesisabstractThis paper proposes an autoregressive speech synthesis model based on the variational autoencoder incorporating latent sequence representation for acoustic and linguistic features and the structure of a hidden semi-Markov model (HSMM). Although autoregressive models can provide efficient and accurate modeling of acoustic features, they have exposure bias, i.e., the mismatch between training (teacher-forcing) and inference (free-running). To overcome this problem, we introduce an autoregressive latent variable sequence, rather than using autoregressive generation of observations. Latent representation of alignment using HSMM-based structured attention mechanism enables the use of a completely consistent training algorithm for acoustic modeling with explicit duration models. Experimental results indicate that the proposed model outperformed baselines in subjective naturalness. Takato Fujimoto, Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda |
ICASSP | 4 |
| 2022 | End-to-End Text-to-Speech Based on Latent Representation of Speaking Styles Using Spontaneous DialogueabstractThe recent text-to-speech (TTS) has achieved quality comparable to that of humans; however, its application in spoken dialogue has not been widely studied. This study aims to realize a TTS that closely resembles human dialogue. First, we record and transcribe actual spontaneous dialogues. Then, the proposed dialogue TTS is trained in two stages: first stage, variational autoencoder (VAE)-VITS or Gaussian mixture variational autoencoder (GMVAE)-VITS is trained, which introduces an utterance-level latent variable into variational inference with adversarial learning for end-to-end text-to-speech (VITS), a recently proposed end-to-end TTS model. A style encoder that extracts a latent speaking style representation from speech is trained jointly with TTS. In the second stage, a style predictor is trained to predict the speaking style to be synthesized from dialogue history. During inference, by passing the speaking style representation predicted by the style predictor to VAE/GMVAE-VITS, speech can be synthesized in a style appropriate to the context of the dialogue. Subjective evaluation results demonstrate that the proposed method outperforms the original VITS in terms of dialogue-level naturalness. Kentaro Mitsui, Tianyu Zhao 0001, Kei Sawada, Yukiya Hono, Yoshihiko Nankaku, Keiichi Tokuda |
INTERSPEECH | 6 |
| 2021 | Periodnet: A Non-Autoregressive Waveform Generation Model with a Structure Separating Periodic and Aperiodic ComponentsabstractWe propose PeriodNet, a non-autoregressive (non-AR) waveform generation model with a new model structure for modeling periodic and aperiodic components in speech waveforms. The non-AR waveform generation models can generate speech waveforms parallelly and can be used as a speech vocoder by conditioning an acoustic feature. Since a speech waveform contains periodic and aperiodic components, both components should be appropriately modeled to generate a high-quality speech waveform. However, it is difficult to decompose the components from a natural speech waveform in advance. To address this issue, we propose a parallel model and a series model structure separating periodic and aperiodic components. The features of our proposed models are that explicit periodic and aperiodic signals are taken as input, and external periodic/aperiodic decomposition is not needed in training. Experiments using a singing voice corpus show that our proposed structure improves the naturalness of the generated waveform. We also show that the speech waveforms with a pitch outside of the training data range can be generated with more naturalness. Yukiya Hono, Shinji Takaki, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda |
ICASSP | 6 |
| 2021 | Sinsy: A Deep Neural Network-Based Singing Voice Synthesis SystemabstractThis paper presents Sinsy, a deep neural network (DNN)-based singing voice synthesis (SVS) system. In recent years, DNNs have been utilized in statistical parametric SVS systems, and DNN-based SVS systems have demonstrated better performance than conventional hidden Markov model-based ones. SVS systems are required to synthesize a singing voice with pitch and timing that strictly follow a given musical score. Additionally, singing expressions that are not described on the musical score, such as vibrato and timing fluctuations, should be reproduced. The proposed system is composed of four modules: a time-lag model, a duration model, an acoustic model, and a vocoder, and singing voices can be synthesized taking these characteristics of singing voices into account. To better model a singing voice, the proposed system incorporates improved approaches to modeling pitch and vibrato and better training criteria into the acoustic model. In addition, we incorporated PeriodNet, a non-autoregressive neural vocoder with robustness for the pitch, into our systems to generate a high-fidelity singing voice waveform. Moreover, we propose automatic pitch correction techniques for DNN-based SVS to synthesize singing voices with correct pitch even if the training data has out-of-tune phrases. Experimental results show our system can synthesize a singing voice with better timing, more natural vibrato, and correct pitch, and it can achieve better mean opinion scores in subjective evaluation tests. Yukiya Hono, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2020 | Semi-Supervised Learning Based on Hierarchical Generative Models for End-to-End Speech SynthesisabstractThis paper proposes a general framework of semi-supervised learning based on hierarchical generative models and adapts it to a Japanese end-to-end text-to-speech (TTS) system. In English TTS, several end-to-end systems have recently achieved sound quality close to that of natural human speech. However, in non-alphabetic languages such as Japanese, it is difficult to realize true text-input end-to-end TTS due to character diversity and pitch accents. To address this problem, we propose end-to-end TTS based on semi-supervised learning that makes the most of existing data consisting of any combination of text, phoneme, and waveform as training data. To demonstrate the effectiveness of the proposed system, listening tests were conducted for pronunciation and naturalness. Our results show that the proposed system improves both pronunciation and naturalness. Takato Fujimoto, Shinji Takaki, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda |
ICASSP | 6 |
| 2020 | Fast and High-Quality Singing Voice Synthesis System Based on Convolutional Neural NetworksabstractThe present paper describes singing voice synthesis based on convolutional neural networks (CNNs). Singing voice synthesis systems based on deep neural networks (DNNs) are currently being proposed and are improving the naturalness of synthesized singing voices. As singing voices represent a rich form of expression, a powerful technique to model them accurately is required. In the proposed technique, long-term dependencies of singing voices are modeled by CNNs. An acoustic feature sequence is generated for each segment that consists of long-term frames, and a natural trajectory is obtained without the parameter generation algorithm. Furthermore, a computational complexity reduction technique, which drives the DNNs in different time units depending on type of musical score features, is proposed. Experimental results show that the proposed method can synthesize natural sounding singing voices much faster than the conventional method. Kazuhiro Nakamura, Shinji Takaki, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda |
ICASSP | 6 |
| 2020 | Hierarchical Multi-Grained Generative Model for Expressive Speech SynthesisabstractThis paper proposes a hierarchical generative model with a multi-grained latent variable to synthesize expressive speech. In recent years, fine-grained latent variables are introduced into the text-to-speech synthesis that enable the fine control of the prosody and speaking styles of synthesized speech. However, the naturalness of speech degrades when these latent variables are obtained by sampling from the standard Gaussian prior. To solve this problem, we propose a novel framework for modeling the fine-grained latent variables, considering the dependence on an input text, a hierarchical linguistic structure, and a temporal structure of latent variables. This framework consists of a multi-grained variational autoencoder, a conditional prior, and a multi-level auto-regressive latent converter to obtain the different time-resolution latent variables and sample the finer-level latent variables from the coarser-level ones by taking into account the input text. Experimental results indicate an appropriate method of sampling fine-grained latent variables without the reference signal at the synthesis stage. Our proposed framework also provides the controllability of speaking style in an entire utterance. Yukiya Hono, Kazuna Tsuboi, Kei Sawada, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda |
INTERSPEECH | 7 |
| 2020 | A Vector Quantized Variational Autoencoder (VQ-VAE) Autoregressive Neural F0 Model for Statistical Parametric Speech SynthesisabstractRecurrent neural networks (RNNs) can predict fundamental frequency (F0) for statistical parametric speech synthesis systems, given linguistic features as input. However, these models assume conditional independence between consecutive F0values, given the RNN state. In a previous study, we proposed autoregressive (AR) neural F0models to capture the causal dependency of successive F0values. In subjective evaluations, a deep AR model (DAR) outperformed an RNN. Here, we propose a Vector Quantized Variational Autoencoder (VQ-VAE) neural F0model that is both more efficient and more interpretable than the DAR. This model has two stages: one uses the VQ-VAE framework to learn a latent code for the F0contour of each linguistic unit, and other learns to map from linguistic features to latent codes. In contrast to the DAR and RNN, which process the input linguistic features frame-by-frame, the new model converts one linguistic feature vector into one latent code for each linguistic unit. The new model achieves better objective scores than the DAR, has a smaller memory footprint and is computationally faster. Visualization of the latent codes for phones and moras reveals that each latent code represents an F0shape for a linguistic unit. Xin Wang 0037, Shinji Takaki, Junichi Yamagishi, Simon King 0001, Keiichi Tokuda |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2019 | Singing Voice Synthesis Based on Generative Adversarial NetworksabstractThis paper proposes a generative adversarial training method for deep neural network (DNN)-based singing voice synthesis. The DNN-based approach has been used in statistical parametric singing voice synthesis and improved the naturalness of the synthesized singing voice [1]. Recently, generative adversarial networks (GANs) [2] have attracted significant attention in various machine learning research areas including speech synthesis [3]. GANs have achieved great success in modeling the distributions of complex data, and they have the potential to alleviate over-smoothing problem on the generated speech parameters in speech synthesis. In this paper, we propose a DNN-based singing voice synthesis system incorporating the GAN. Experimental results show that the proposed method outperforms the conventional method in the naturalness of the synthesized singing voice. Yukiya Hono, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda |
ICASSP | 5 |
| 2019 | Speaker-dependent Wavenet-based Delay-free Adpcm Speech CodingabstractThis paper proposes a WaveNet-based delay-free adaptive differential pulse code modulation (ADPCM) speech coding system. The WaveNet generative model, which is a state-of-the-art model for neural-network-based speech waveform synthesis, is used as the adaptive predictor in ADPCM. To further improve speech quality, mel-cepstrum-based noise shaping and postfiltering were integrated with the proposed ADPCM system. Both objective and subjective evaluation results indicate that the proposed ADPCM system outperformed not only the conventional ADPCM system based on ITU-T Recommendation G.726 but also the ADPCM system based on adaptive mel-cepstral analysis. Takenori Yoshimura, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda |
ICASSP | 5 |
| 2019 | Statistical Approach to Speech Synthesis: Past, Present and Future
Keiichi Tokuda |
INTERSPEECH | 1 |
| 2018 | Image Recognition Based on Separable Lattice Hmms Using a Deep Neural Network for Output Probability DistributionsabstractThis paper proposes an image recognition method based on separable lattice hidden Markov models (SLHMMs) using a deep neural network (DNN) for output probability distributions. The geometric variations of the object to be recognized, e.g., size and location, are essential in image recognition. SLHMMs, which have been proposed to reduce the effect of geometric variations, can perform elastic matching both horizontally and vertically. Gaussian distributions are typical for modeling the output distribution of SLHMMs. However, these distributions may not be sufficient to represent patterns of image regions. Our method integrates SLHMMs and a DNN and can be used to model an image effectively by explicit modeling of the generative process based on SLHMMs and advanced feature classification based on a DNN. image recognition experiments showed that the proposed method improves recognition performance. Eiji Ichikawa, Kei Sawada, Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda |
ICASSP | 5 |
| 2018 | Statistical Voice Conversion Based on WavenetabstractThis paper proposes a voice conversion technique based on WaveNet to directly generate target audio waveforms from acoustic features of a source speaker. In voice conversion based on statistical models, the relation between acoustic features, such as spectral parameters, extracted from source and target audio waveforms is generally modeled using statistical models, such as Gaussian mixture models and neural networks. Although modeling the relation between acoustic features is reasonable and efficient, these models are not optimized for predicting target audio waveforms because the vocoder parameters are used as intermediate representations. To overcome this problem, we developed a voice conversion method to model the relation between target audio waveforms and acoustic features extracted from source audio waveforms using WaveNet, which is a generative model for audio waveforms. The proposed model can directly generate converted audio waveforms without vocoders. Experimental results indicate that the proposed method can generate a more naturally sounding converted speech than that using a conventional DNN method. Jumpei Niwa, Takenori Yoshimura, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda |
ICASSP | 6 |
| 2018 | WaveNet-Based Zero-Delay Lossless Speech CodingabstractThis paper presents a WaveNet-based zero-delay lossless speech coding technique for high-quality communications. The WaveNet generative model, which is a state-of-the-art model for neural-network-based speech waveform synthesis, is used in both the encoder and decoder. In the encoder, discrete speech signals are losslessly compressed using sample-by-sample entropy coding. The decoder fully reconstructs the original speech signals from the compressed signals without algorithmic delay. Experimental results show that the proposed coding technique can transmit speech audio waveforms with 50% their original bit rate and the WaveNet-based speech coder remains effective for unknown speakers. Takenori Yoshimura, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda |
SLT | 5 |
| 2018 | Mel-Cepstrum-Based Quantization Noise Shaping Applied to Neural-Network-Based Speech Waveform SynthesisabstractThis paper presents a mel-cepstrum-based quantization noise shaping method for improving the quality of synthetic speech generated by neural-network-based speech waveform synthesis systems. Since mel-cepstral coefficients closely match the characteristics of human auditory perception, the proposed method effectively masks the white noise introduced by the quantization typically used in neural-network-based speech waveform synthesis systems. The paper also describes a computationally efficient implementation of the proposed method using the structure of the mel-log spectrum approximation filter. Experiments using the WaveNet generative model, which is a state-of-the-art model for neural-network-based speech waveform synthesis, showed that speech quality is significantly improved by the proposed method. Takenori Yoshimura, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2017 | The blizzard machine learning challenge 2017abstractThis paper describes the Blizzard Machine Learning Challenge (BMLC) 2017, which is a spin-off of the Blizzard Challenge. The annual Blizzard Challenges 2005-2017 were held to better understand and compare research techniques in building corpus-based text-to-speech (TTS) systems on the same data. The series of Blizzard Challenges has helped us measure progress in TTS technology. However, to get competitive performance, a lot time has to be spent on skilled tasks. This may make the Blizzard Challenge unattractive to machine learning researchers from other fields. Therefore, we recommend that the BMLC not involve these speech-specific tasks and that it allow participants to concentrate on the acoustic modeling task, framed as a straightforward machine learning problem, with a fixed dataset. In the BMLC 2017, two types of datasets consisting of four hours of speech data suitable for machine learning problems were distributed. This paper summarizes the purpose, design, and whole process of the challenge and its results. Kei Sawada, Keiichi Tokuda, Simon King 0001, Alan W. Black |
ASRU | 2 |
| 2017 | Image recognition based on discriminative models using features generated from separable lattice HMMSabstractThis paper presents an image recognition technique based on discriminative models using features generated from separable lattice hidden Markov models (SL-HMMs). A major problem in image recognition is that the recognition performance is degraded by geometric variations such as that in position and size of the object to be recognized. SL-HMMs have been proposed to solve this problem. SL-HMMs are an extension of HMMs with size and locational invariances based on state transitions. An SL-HMM is a generative model and can represent generation processes of observations well. However, there is a possibility that the recognition performance of generative models is inferior to that of discriminative models because discriminative models are specialized to identification. In this paper, we propose image recognition based on log linear models (LLMs) using features extracted from SL-HMMs. The proposed method can extract features invariant to geometric variations by using SL-HMMs and built an accurate classifier based on discriminative models with the extracted features. Face recognition experiments showed that the proposed method obtained higher recognition rates than SL-HMMs and convolutional neural networks based methods. Yoshinari Tsuzuki, Kei Sawada, Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda |
ICASSP | 5 |
| 2017 | Articulatory Text-to-Speech Synthesis Using the Digital Waveguide Mesh Driven by a Deep Neural NetworkabstractFollowing recent advances in direct modeling of the speech waveform using a deep neural network, we propose a novel method that directly estimates a physical model of the vocal tract from the speech waveform, rather than magnetic resonance imaging data. This provides a clear relationship between the model and the size and shape of the vocal tract, offering considerable flexibility in terms of speech characteristics such as age and gender. Initial tests indicate that despite a highly simplified physical model, intelligible synthesized speech is obtained. This illustrates the potential of the combined technique for the control of physical models in general, and hence the generation of more natural-sounding synthetic speech. Amelia Jane Gully, Takenori Yoshimura, Damian T. Murphy, Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda |
INTERSPEECH | 6 |
| 2017 | Simultaneous Optimization of Multiple Tree-Based Factor Analyzed HMM for Speech SynthesisabstractThis paper proposes a novel method to build multiple decision trees as a structure of factor analyzed hidden Markov model for speech synthesis. In the proposed method, the multiple decision trees grow simultaneously rather than sequentially to take into account the relationship between the trees. However, the simultaneous growing is computationally infeasible due to an exponential increase in the number of tree structures to be evaluated. To solve the problem, we further propose two computational complexity reduction algorithms that achieve a significant reduction in the computational time. Experimental results show that the proposed method outperforms the conventional one based on a single decision tree. Takenori Yoshimura, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2016 | Trajectory training considering global variance for speech synthesis based on neural networksabstractThis paper proposes a new training method of deep neural networks (DNNs) for statistical parametric speech synthesis. DNNs are recently used as acoustic models that represent mapping functions from linguistic features to acoustic features in statistical parametric speech synthesis. There are problems to be solved in conventional DNN-based speech synthesis: 1) the inconsistency between the training and synthesis criteria; and 2) the over-smoothing of the generated parameter trajectories. In this paper, we introduce the parameter trajectory generation process considering the global variance (GV) into the training of DNNs. A unified framework which consistently uses the same criterion in both training and synthesis can be obtained and the model parameters are optimized for parameter generation considering the GV in the proposed method. Experimental results show that the proposed method outperforms the conventional method in the naturalness of synthesized speech. Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda |
ICASSP | 4 |
| 2016 | Directly modeling voiced and unvoiced components in speech waveforms by neural networksabstractThis paper proposes a novel acoustic model based on neural networks for statistical parametric speech synthesis. The neural network outputs parameters of a non-zero mean Gaussian process, which defines a probability density function of a speech waveform given linguistic features. The mean and covariance functions of the Gaussian process represent deterministic (voiced) and stochastic (unvoiced) components of a speech waveform, whereas the previous approach considered the unvoiced component only. Experimental results show that the proposed approach can generate speech waveforms approximating natural speech waveforms. Keiichi Tokuda, Heiga Zen |
ICASSP | 1 |
| 2016 | Redefining the Linguistic Context Feature Set for HMM and DNN TTS Through Position and Parsing
Rasmus Dall, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda |
INTERSPEECH | 5 |
| 2016 | Voice Conversion Based on Trajectory Model Training of Neural Networks Considering Global Variance
Naoki Hosaka, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda |
INTERSPEECH | 5 |
| 2016 | Singing Voice Synthesis Based on Deep Neural Networks
Masanari Nishimura, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda |
INTERSPEECH | 5 |
| 2016 | A Hierarchical Predictor of Synthetic Speech Naturalness Using Neural NetworksabstractA problem when developing and tuning speech synthesis systems is that there is no well-established method of automatically rating the quality of the synthetic speech. This research attempts to obtain a new automated measure which is trained on the result of large-scale subjective evaluations employing many human listeners, i.e., the Blizzard Challenge. To exploit the data, we experiment with linear regression, feed-forward and convolutional neural network models, and combinations of them to regress from synthetic speech to the perceptual scores obtained from listeners. The biggest improvements were seen when combining stimulus- and system-level predictions. Takenori Yoshimura, Gustav Eje Henter, Oliver Watts, Mirjam Wester, Junichi Yamagishi, Keiichi Tokuda |
INTERSPEECH | 6 |
| 2015 | The effect of neural networks in statistical parametric speech synthesisabstractThis paper investigates how to use neural networks in statistical parametric speech synthesis. Recently, deep neural networks (DNNs) have been used for statistical parametric speech synthesis. However, the specific way how DNNs should be used in statistical parametric speech synthesis has not been studied thoroughly. A generation process of statistical parametric speech synthesis based on generative models can be divided into several components, and those components can be represented by DNNs. In this paper, the effect of DNNs for each component is investigated by comparing DNNs with generative models. Experimental results show that the use of a DNN as acoustic models is effective and the parameter generation combined with a DNN improves the naturalness of synthesized speech. Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda |
ICASSP | 4 |
| 2015 | Directly modeling speech waveforms by neural networks for statistical parametric speech synthesisabstractThis paper proposes a novel approach for directly-modeling speech at the waveform level using a neural network. This approach uses the neural network-based statistical parametric speech synthesis framework with a specially designed output layer. As acoustic feature extraction is integrated to acoustic model training, it can overcome the limitations of conventional approaches, such as two-step (feature extraction and acoustic modeling) optimization, use of spectra rather than waveforms as targets, use of overlapping and shifting frames as unit, and fixed decision tree structure. Experimental results show that the proposed approach can directly maximize the likelihood defined at the waveform domain. Keiichi Tokuda, Heiga Zen |
ICASSP | 1 |
| 2015 | Simultaneous optimization of multiple tree structures for factor analyzed HMM-based speech synthesis
Takenori Yoshimura, Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda |
INTERSPEECH | 4 |
| 2015 | Using speaker adaptive training to realize Mandarin-Tibetan cross-lingual speech synthesis
Hongwu Yang, Keiichiro Oura, Zhenye Gan, Keiichi Tokuda |
Multim. Tools Appl. | 5 |
| 2014 | Voice interaction system with 3D-CG virtual agent for stand-alone smartphonesabstractIn this paper, we propose a voice interaction system using 3D-CG virtual agents for stand-alone smartphones. Because the proposed system can handle speech recognition and speech synthesis on a stand-alone smartphone differently from the existing mobile voice interaction systems, this system enables us to talk naturally without encountering delays caused by network communications. Moreover, proposed system can be fully customized by dialogue scripts, Java-based plugins, and Android APIs. Therefore, developers can make original voice interaction systems for smartphones easily based on proposed system. We have made a subset of the proposed system available as open-source software. We expect that this system will contribute to studies of human-agent interaction using smartphones. Daisuke Yamamoto, Keiichiro Oura, Ryota Nishimura, Takahiro Uchiya, Akinobu Lee, Ichi Takumi, Keiichi Tokuda |
HAI | 7 |
| 2014 | HMM-Based singing voice synthesis and its application to Japanese and EnglishabstractThe present paper describes Japanese and English singing voice synthesis systems based on hidden Markov models (HMMs). In this approach, the spectrum, excitation, and vibrato of the singing voice are simultaneously modeled by context-dependent HMMs, and waveforms are generated by the HMMs themselves. Japanese singing voice synthesis systems have already been developed and used to create variable musical contents. To extend this system to English, language independent contexts are designed. Furthermore, methods for matching musical notes and pronunciation of English lyrics are presented and evaluated in subjective experiments. Then, Japanese and English singing voice synthesis systems are compared. Kazuhiro Nakamura, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda |
ICASSP | 4 |
| 2014 | Integration of speaker and pitch adaptive training for HMM-based singing voice synthesisabstractA statistical parametric approach to singing voice synthesis based on hidden Markov models (HMMs) has been growing in popularity over the last few years. The spectrum, excitation, vibrato, and duration of the singing voice in this approach are simultaneously modeled with context-dependent HMMs and waveforms are generated from the HMMs themselves. Since HMM-based singing voice synthesis systems are “corpus-based,” the HMMs corresponding to contextual factors that rarely appear in the training data cannot be well-trained. However, it may be difficult to prepare a large enough quantity of singing voice data sung by one singer. Furthermore, the pitch included in each song is imbalanced, and there is the vocal range of the singer. In this paper, we propose “singer adaptive training” which can solve the data sparse-ness problem. Experimental results demonstrated that the proposed technique improved the quality of the synthesized singing voices. Kanako Shirota, Kazuhiro Nakamura, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda |
ICASSP | 6 |
| 2014 | A mel-cepstral analysis technique restoring high frequency components from low-sampling-rate speech
Kazuhiro Nakamura, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda |
INTERSPEECH | 5 |
| 2013 | Mmdagent - A fully open-source toolkit for voice interaction systemsabstractThis paper describes development of an open-source toolkit which makes it possible to explore a vast variety of aspects in speech interactions at spoken dialog systems and speech interfaces. The toolkit tightly incorporates recent speech recognition and synthesis technologies with a 3-D CG rendering module that can manipulates expressive embodied agent characters. The software design and its interfaces are carefully designed to be fully open toolkit. Ongoing demonstration experiments to public indicates that it is promoting related researches and developments of voice interaction systems in various scenes. Akinobu Lee, Keiichiro Oura, Keiichi Tokuda |
ICASSP | 3 |
| 2013 | Separable lattice 2-D HMMS introducing state duration control for recognition of images with various variationsabstractIn this paper, an extension of separable lattice HMMs (SL-HMM) is described that introduces state duration control for dealing with images with various variations. SL-HMM are generative models that have size and location invariances based on state transition of HMMs. An extended model that has the structure of hidden semi-Markov models (HSMMs) in which the state duration probability is explicitly modeled by parametric distributions is also proposed. However, in this model, each state duration in a Markov chain is independent. It is supposed that each state duration should have a correlation. Therefore, in this paper, we propose a novel model that solves this problem by introducing variables representing the correlation among the state durations. Face recognition experiments show that the proposed model improved the recognition performance for images with size, locational, and rotational variations. Takaya Makino, Shinji Takaki, Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda |
ICASSP | 5 |
| 2013 | Integration of acoustic modeling and mel-cepstral analysis for HMM-based speech synthesisabstractIn this paper, a novel approach for integrating acoustic modeling and mel-cepstral analysis is proposed. The aim of HMM-based speech synthesis is to model speech waveforms with a statistical model. However, the conventional techniques divide the modeling process into two steps: the frame by frame feature extraction step and the acoustic modeling step. Although it is reasonably effective, the deterioration of speech quality is caused by the divide of the objective function. In this paper, we propose an approach to modeling them as an integrative model and show the possibility of improving synthesized speech. Kazuhiro Nakamura, Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda |
ICASSP | 4 |
| 2013 | Contextual partial additive structure for HMM-based speech synthesisabstractThis paper proposes a spectral modeling technique based on a contextual partial additive structure for HMM-based speech synthesis. To represent complicated context dependencies, contextual additive structure models assume multiple independent components which have different context dependencies to form acoustic features. In additive structure models, there is a constraint that a fixed number of additive components are used for generating acoustic features. However, it is natural to assume that the number of components depends on contexts. In the proposed technique, partial additive components affecting arbitrary contextual sub-spaces are created on demand to increase the likelihood. Then, the number of components for each context can be automatically determined with the training data. Experimental results show that the proposed technique outperformed the standard technique in a subjective test. Shinji Takaki, Yoshihiko Nankaku, Keiichi Tokuda |
ICASSP | 3 |
| 2013 | Image recognition based on separable lattice trajectory 2-D HMMSabstractIn this paper, a novel statistical model for image recognition based on separable lattice 2-D HMMs (SL2D-HMMs) is proposed. Although SL2D-HMMs can model invariance to size and location deformation, its modeling accuracy is still insufficient because of the following two assumptions: i) the statistics of each state are constant and ii) the state output probabilities are conditionally independent. In this paper, SL2D-HMMs are reformulated as a trajectory model that can capture dependencies between adjacent observations. The effectiveness of the proposed model was demonstrated in face recognition and image alignment experiments. Akira Tamamori, Yoshihiko Nankaku, Keiichi Tokuda |
ICASSP | 3 |
| 2013 | Personalising speech-to-speech translation: Unsupervised cross-lingual speaker adaptation for HMM-based speech synthesis
John Dines, Lakshmi Babu Saheer, Matthew Gibson, William J. Byrne, Keiichiro Oura, Keiichi Tokuda, Junichi Yamagishi, Simon King 0001, Mirjam Wester, Teemu Hirsimäki, Reima Karhila, Mikko Kurimo |
Comput. Speech Lang. | 7 |
| 2013 | Speech Synthesis Based on Hidden Markov ModelsabstractThis paper gives a general overview of hidden Markov model (HMM)-based speech synthesis, which has recently been demonstrated to be very effective in synthesizing speech. The main advantage of this approach is its flexibility in changing speaker identities, emotions, and speaking styles. This paper also discusses the relation between the HMM-based approach and the more conventional unit-selection approach that has dominated over the last decades. Finally, advanced techniques for future developments are described. Keiichi Tokuda, Yoshihiko Nankaku, Tomoki Toda, Heiga Zen, Junichi Yamagishi, Keiichiro Oura |
Proc. IEEE | 1 |
| 2012 | Face recognition based on extended separable lattice 2-D HMMSabstractThis paper proposes an extension of separable lattice 2-D hidden Markov models (SL-HMMs) for dealing with image rotation and local deformation. It is important to reduce the effect of geometrical variations in image recognition, e.g., location, size, and rotation. SLHMMs are one of the most efficient structures to accomplish invariance to size and location variations. However, since SL-HMMs only have one state sequence in each direction, they cannot deal with rotation or local deformation. The proposed models have state sequences corresponding to all rows and columns of an input image, and the complicated state alignments can represent rotation and local deformation. The effectiveness of the proposed models was demonstrated in face recognition experiments. Keisuke Kumaki, Yoshihiko Nankaku, Keiichi Tokuda |
ICASSP | 3 |
| 2012 | Pitch adaptive training for hmm-based singing voice synthesisabstractA statistical parametric approach to singing voice synthesis based on hidden Markov Models (HMMs) has been growing in popularity over the last few years. The spectrum, excitation, vibrato, and duration of singing voices in this approach are simultaneously modeled with context-dependent HMMs and waveforms are generated from the HMMs themselves. HMM-based singing voice synthesis systems are heavily based on the training data in performance because these systems are “corpus-based.” Therefore, HMMs corresponding to contextual factors that hardly ever appear in the training data cannot be well-trained. Pitch should especially be correctly covered since generated F0trajectories have a great impact on the subjective quality of synthesized singing voices. We applied the method of “speaker adaptive training” (SAT) to “pitch adaptive training,” which is discussed in this paper. This technique made it possible to normalize pitch based on musical notes in the training process. The experimental results demonstrated that the proposed technique could alleviate the data sparseness problem. Keiichiro Oura, Ayami Mase, Yoshihiko Nankaku, Keiichi Tokuda |
ICASSP | 4 |
| 2012 | Face recognition based on separable lattice 2-D HMMS using variational bayesian methodabstractThis paper proposes an image recognition technique based on separable lattice 2-D HMMs (SL2D-HMMs) using the variational Bayesian method. SL2D-HMMs have been proposed to reduce the effect of geometric variations, e.g., size and location. The maximum likelihood criterion had previously been used in training SL2D-HMMs. However, in many image recognition tasks, it is difficult to use sufficient training data, and it suffers from the over-fitting problem. A higher generalization ability based on model marginalization is expected by applying the Bayesian criterion and useful prior information on model parameters can be utilized as prior distributions. Experiments on face recognition indicated that the proposed method improved image recognition. Kei Sawada, Akira Tamamori, Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda |
ICASSP | 5 |
| 2012 | A model structure integration based on a Bayesian framework for speech recognitionabstractThis paper proposes an acoustic modeling technique based on Bayesian framework using multiple model structures for speech recognition. The Bayesian approach is a statistical technique for estimating reliable predictive distributions by marginalizing model parameters, and its effectiveness in HMM-based speech recognition has been reported. Although the basic idea underlying the Bayesian approach is to treat all parameters as random variables, only one model structure is still selected in the conventional method. Multiple model structures are treated as latent variables in the proposed method and integrated based on the Bayesian framework. Furthermore, we applied deterministic annealing to the training algorithm to estimate appropriate acoustic models. The proposed method effectively utilizes multiple model structures, especially in the early stage of training and this leads to better predictive distributions and improvement of recognition performance. Sayaka Shiota, Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda |
ICASSP | 4 |
| 2012 | A Bayesian Approach to Speaker Recognition Based on GMMs Using Multiple Model Structures
Takafumi Hattori, Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda |
INTERSPEECH | 4 |
| 2012 | Cross-lingual Speaker Adaptation for HMM-based Speech Synthesis based on Perceptual Characteristics and Speaker Interpolation
Viviane de Franca Oliveira, Sayaka Shiota, Yoshihiko Nankaku, Keiichi Tokuda |
INTERSPEECH | 4 |
| 2012 | Impacts of machine translation and speech synthesis on speech-to-speech translation
Kei Hashimoto, Junichi Yamagishi, William J. Byrne, Simon King 0001, Keiichi Tokuda |
Speech Commun. | 5 |
| 2012 | Analysis of unsupervised cross-lingual speaker adaptation for HMM-based speech synthesis using KLD-based transform mapping
Keiichiro Oura, Junichi Yamagishi, Mirjam Wester, Simon King 0001, Keiichi Tokuda |
Speech Commun. | 5 |
| 2012 | Product of Experts for Statistical Parametric Speech SynthesisabstractMultiple acoustic models are often combined in statistical parametric speech synthesis. Both linear and non-linear functions of an observation sequence are used as features to be modeled. This paper shows that this combination of multiple acoustic models can be expressed as a product of experts (PoE); the likelihoods from the models are scaled, multiplied together, and then normalized. Normally these models are individually trained and only combined at the synthesis stage. This paper discusses a more consistent PoE framework where the models are jointly trained. A training algorithm for PoEs based on linear feature functions and Gaussian experts is derived by generalizing the training algorithm for trajectory HMMs. However for non-linear feature functions or non-Gaussian experts this is not possible, so a scheme based on contrastive divergence learning is described. Experimental results show that the PoE framework provides both a mathematically elegant way to train multiple acoustic models jointly and significant improvements in the quality of the synthesized speech. Heiga Zen, Mark J. F. Gales, Yoshihiko Nankaku, Keiichi Tokuda |
IEEE Trans. Speech Audio Process. | 4 |
| 2011 | An analysis of machine translation and speech synthesis in speech-to-speech translation systemabstractThis paper provides an analysis of the impacts of machine translation and speech synthesis on speech-to-speech translation systems. The speech-to-speech translation system consists of three components: speech recognition, machine translation and speech synthesis. Many techniques for integration of speech recognition and machine translation have been proposed. However, speech synthesis has not yet been considered. Therefore, in this paper, we focus on machine translation and speech synthesis, and report a subjective evaluation to analyze the impact of each component. The results of these analyses show that the naturalness and intelligibility of synthesized speech are strongly affected by the fluency of the translated sentences. Kei Hashimoto, Junichi Yamagishi, William J. Byrne, Simon King 0001, Keiichi Tokuda |
ICASSP | 5 |
| 2011 | Global variance modeling on frequency domain delta LSP for HMM-based speech synthesisabstractThe speech parameter generation algorithm considering global variance (GV) for HMM-based speech synthesis proved to be effective against the over-smoothing problem. However, the correlation between dimensions of parameter vector is not sufficiently considered in the current GV model. For some parameters, e.g., Line Spectral Pairs (LSP), the difference of adjacent LSPs has strong influence on the spectral envelope. Considering this important feature, the paper proposes a GV modeling on the difference of adjacent LSPs, i.e., GV on frequency domain delta LSP. By improving the GV likelihood on frequency domain delta LSP, the over-smoothing effect of generated parameter trajectory is better alleviated than conventional one. The result of a perceptual evaluation shows the proposed method outperforms the conventional one, and the naturalness of synthetic speech is improved. Shifeng Pan, Yoshihiko Nankaku, Keiichi Tokuda, Jianhua Tao 0001 |
ICASSP | 3 |
| 2011 | An optimization algorithm of independent mean and variance parameter tying structures for HMM-based speech synthesisabstractThis paper proposes a technique for constructing independent parameter tying structures of mean and variance in HMM based speech synthesis. Conventionally, mean and variance parameters are assumed to have the same tying structure. However, it has been reported that a clustering technique of mean vectors while tying all variance matrices improves the quality of synthesized speech. This indicates that mean and variance parameters should have different optimal tying structures. In the proposed technique, the decision trees for mean and variance parameters are simultaneously grown by taking into account the dependency on mean and variance parameters. Experimental results show that the proposed technique outperforms the conventional one. Shinji Takaki, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda |
ICASSP | 4 |
| 2011 | Estimation of Window Coefficients for Dynamic Feature Extraction for HMM-Based Speech Synthesis
Linghui Chen, Yoshihiko Nankaku, Heiga Zen, Keiichi Tokuda, Zhen-Hua Ling, Li-Rong Dai 0001 |
INTERSPEECH | 4 |
| 2011 | Multi-Speaker Modeling with Shared Prior Distributions and Model Structures for Bayesian Speech SynthesisabstractThis paper investigates a multi-speaker modeling technique with shared prior distributions and model structures for Bayesian speech synthesis. The quality of synthesized speech is improved by selecting appropriate model structures in HMMbased speech synthesis. Bayesian approach is known to work for such model selection. However, the result is strongly affected by prior distributions of model parameters. Therefore, determination of prior distributions and selection of model structures should be performed simultaneously. This paper investigates prior distributions and model structures in the situation where training data of multiple speakers are available. The prior distributions and model structures which represent acoustic features common to every speakers can be obtained by sharing them between multiple speaker-dependent models. Index Terms: speech synthesis, Bayesian approach, prior distribution, context clustering, multi-speaker modeling A statistical parametric speech synthesis system based on hidden Markov models (HMMs) was recently developed. In HMM-based speech synthesis, the spectrum, excitation, and duration of speech are simultaneously modeled with HMMs, and speech parameter sequences are generated from the HMMs themselves [1]. The maximum likelihood (ML) criterion has typically been used for training HMMs and generating speech parameters. The ML criterion guarantees that the ML estimates approach the true values of the parameters. However, since the ML criterion produces a point estimate of the model parameters, its estimation accuracy may degrade when the amount of training data is insufficient. In the Bayesian approach, all variables introduced when the models are parameterized, such as model parameters and latent variables, are treated as random variables, and their posterior distributions are obtained by the Bayes theorem. The Bayesian approach can generally construct a more robust model than the ML approach by estimating posterior distributions. Recently, Bayesian speech synthesis has been proposed as a Bayesian framework for statistical parametric speech synthesis (e.g., HMM-based speech synthesis), and it shows good performance [2]. In Bayesian speech synthesis, all processes for constructing the system can be derived from a single predictive distribution that directly represents the problem of speech synthesis. The quality of synthesized speech is improved by selecting appropriate model structures in HMM-based speech synthesis. Although the Bayesian approach is known to work for such model selection, the results are strongly affected by prior distributions of the model parameters. Therefore, in Bayesian speech synthesis, determination of prior distributions and selection of model structures should be performed simultaneously. To overcome this problem, we have proposed Bayesian context clustering using cross validation [3]. In this method, prior distributions are determined by using a part of training data, and model structures are evaluated by using the determined prior distribution based on cross validation. In this paper, we investigates prior distributions and model Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda |
INTERSPEECH | 3 |
| 2011 | Large-Scale Subjective Evaluations of Speech Rate Control Methods for HMM-Based Speech SynthesizersabstractThree speech rate control methods for HMM-based speech synthesis were compared by large-scale subjective evaluations. The methods are 1) synthesizing speech sounds based on HMMs trained from corpora at a target speech rate, 2) stretching or shrinking utterance durations proportionally in waveform generation, and 3) determining state durations based on ML criterion under a restriction of utterance duration. The results indicated that the proportional shrinking had significant advantages for fast rate, whereas HMMs trained from slow speech sounds had a slight advantage for slow rate. We also found an advantage of proportionally shrunk speech from a synthesizer trained from slow speech corpora. Tsuneo Kato, Makoto Yamada, Nobuyuki Nishizawa, Keiichiro Oura, Keiichi Tokuda |
INTERSPEECH | 5 |
| 2011 | A Bayesian Approach to Voice Conversion Based on GMMs Using Multiple Model Structures
Yoshihiko Nankaku, Keiichi Tokuda |
INTERSPEECH | 3 |
| 2011 | GMM-Based Missing-Feature Reconstruction on Multi-Frame Windows
Ulpu Remes, Yoshihiko Nankaku, Keiichi Tokuda |
INTERSPEECH | 3 |
| 2011 | Estimation of Perceptual Spaces for Speaker Identities Based on the Cross-Lingual Discrimination TaskabstractThis paper reconfirms that talker identity can be transmitted across languages. Talker discrimination was examined in the ABX paradigm, where the stimuli A and B were utterances by different talkers in the same language and the stimulus X was an utterance by either of A or B in the different language. The average hit rate of this discrimination task was as high as 0.89. The mutual distance matrices were generated using the discrimination index, ′ d . By applying the multidimensional scaling, three-dimensional perceptual spaces were estimated. The features related with loudness and spectral centroid had high contribution to the perceptual dimensions. Index Terms: talker discrimination, bilingual corpus, MDS, auditory model Minoru Tsuzaki, Keiichi Tokuda, Hisashi Kawai, Jinfu Ni |
INTERSPEECH | 2 |
| 2011 | Continuous Stochastic Feature Mapping Based on Trajectory HMMsabstractThis paper proposes a technique of continuous stochastic feature mapping based on trajectory hidden Markov models (HMMs), which have been derived from HMMs by imposing explicit relationships between static and dynamic features. Although Gaussian mixture model (GMM)- or HMM-based feature-mapping techniques work effectively, their accuracy occasionally degrades due to inappropriate dynamic characteristics caused by frame-by-frame mapping. While the use of dynamic-feature constraints at the mapping stage can alleviate this problem, it also introduces inconsistencies between training and mapping. The technique we propose can eliminate these inconsistencies while retaining the benefits of using dynamic-feature constraints, and it offers entire sequence-level transformation rather than frame-by-frame mapping. The results obtained from speaker-conversion, acoustic-to-articulatory inversion-mapping, and noise-compensation experiments demonstrated that our new approach outperformed the conventional one. Heiga Zen, Yoshihiko Nankaku, Keiichi Tokuda |
IEEE Trans. Speech Audio Process. | 3 |
| 2010 | A Deterministic Annealing-Based Training Algorithm For Statistical Machine Translation Models
Pascual Martínez-Gómez, Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda, Germán Sanchis-Trilles |
EAMT | 4 |
| 2010 | Factor analyzed voice models for HMM-based speech synthesisabstractThis paper describes factor analyzed voice models for realizing various voice characteristics in the HMM-based speech synthesis. The eigenvoice method can synthesize speech with arbitrary voice characteristics by interpolating representative HMM sets. However, the objective of PCA is to accurately reconstruct each speaker-dependent HMM set, and this is not equivalent to estimating models which represent training data accurately. To overcome this problem, we propose a general speech model which generates speech utterances with various voice characteristics directly. In the proposed method, the HMM states, factors representing voice characteristics and contextual decision trees are simultaneously optimized within a unified framework. Kyosuke Kazumi, Yoshihiko Nankaku, Keiichi Tokuda |
ICASSP | 3 |
| 2010 | Unsupervised cross-lingual speaker adaptation for HMM-based speech synthesisabstractIn the EMIME project, we are developing a mobile device that performs personalized speech-to-speech translation such that a user's spoken input in one language is used to produce spoken output in another language, while continuing to sound like the user's voice. We integrate two techniques, unsupervised adaptation for HMM-based TTS using a word-based large-vocabulary continuous speech recognizer and cross-lingual speaker adaptation for HMM-based TTS, into a single architecture. Thus, an unsupervised cross-lingual speaker adaptation system can be developed. Listening tests show very promising results, demonstrating that adapted voices sound similar to the target speaker and that differences between supervised and unsupervised cross-lingual speaker adaptation are small. Keiichiro Oura, Keiichi Tokuda, Junichi Yamagishi, Simon King 0001, Mirjam Wester |
ICASSP | 2 |
| 2010 | Face recognition based on separable lattice 2-D HMM with state duration modelingabstractThis paper describes an extension of separable lattice 2-D HMMs (SL-HMMs) using state duration models for image recognition. SL-HMMs are generative models which have size and location invariances based on state transition of HMMs. However, the state duration probability of HMMs exponentially decreases with increasing duration, therefore it may not be appropriate for modeling image variations accuratelty. To overcome this problem, we employ the structure of hidden semi Markov models (HSMMs) in which the state duration probability is explicitly modeled by parametric distributions. Face recognition experiments show that the proposed model improved the performance for images with size and location variations. Yoshiaki Takahashi, Akira Tamamori, Yoshihiko Nankaku, Keiichi Tokuda |
ICASSP | 4 |
| 2010 | An extension of Separable Lattice 2-D HMMS for rotational data variationsabstractThis paper proposes a new generative model which can deal with rotational data variations by extending Separable Lattice 2-D HMMs (SL2D-HMMs). In image recognition, geometrical variations such as size, location and rotation degrade the performance, therefore normalization is required. SL2D-HMMs can perform an elastic matching in both horizontal and vertical directions; this makes it possible to model invariances to size and location. To deal with rotational variations, we introduce additional HMM states which represent the shifts of the state alignments of the observation lines in a particular direction. Face recognition experiments show that the proposed method improves the performance significantly for rotational variation data. Akira Tamamori, Yoshihiko Nankaku, Keiichi Tokuda |
ICASSP | 3 |
| 2010 | Statistical parametric speech synthesis based on product of expertsabstractMultiple-level acoustic models (AMs) are often combined in statistical parametric speech synthesis. Both linear and non-linear functions of the observation sequence are used as features in these AMs. This combination of multiple-level AMs can be expressed as a product of experts (PoE); the likelihoods from the AMs are scaled, multiplied together and then normalized. Currently these multiple-level AMs are individually trained and only combined at the synthesis stage. This paper discusses a more consistent PoE framework where the AMs are jointly trained. A generalization of trajectory HMM training can be used for multiple-level Gaussian AMs based on linear functions. However for the non-linear case this is not possible, so a scheme based on contrastive divergence learning is described. Experimental results show that the proposed technique provides both a mathematically elegant way to train multiple-level AMs and statistically significant improvements in the quality of synthesized speech. Heiga Zen, Mark J. F. Gales, Yoshihiko Nankaku, Keiichi Tokuda |
ICASSP | 4 |
| 2010 | Speaker adaptation based on nonlinear spectral transform for speech recognition
Toyohiro Hayashi, Yoshihiko Nankaku, Akinobu Lee, Keiichi Tokuda |
INTERSPEECH | 4 |
| 2010 | HMM-based singing voice synthesis system using pitch-shifted pseudo training data
Ayami Mase, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda |
INTERSPEECH | 4 |
| 2010 | Voice activity detection based on conditional random fields using multiple featuresabstractThis paper proposes a Voice Activity Detection (VAD) algorithm based on Conditional Random Fields (CRF) using multiple features.VAD is a technique used to distinguish between speech and non-speech in noisy environments and is an important component in many real-world speech applications.The posterior probability of output labels in the proposed method is directly modeled by the weighted sum of the feature functions.Effective features are automatically selected by estimating appropriate weight parameters to improve the accuracy of VAD.Experimental results on the CENSREC-1-C database revealed that the proposed approach can decrease error rates by using CRF. Akira Saito, Yoshihiko Nankaku, Akinobu Lee, Keiichi Tokuda |
INTERSPEECH | 4 |
| 2010 | Thousands of Voices for HMM-Based Speech Synthesis-Analysis and Application of TTS Systems Built on Various ASR CorporaabstractIn conventional speech synthesis, large amounts of phonetically balanced speech data recorded in highly controlled recording studio environments are typically required to build a voice. Although using such data is a straightforward solution for high quality synthesis, the number of voices available will always be limited, because recording costs are high. On the other hand, our recent experiments with HMM-based speech synthesis systems have demonstrated that speaker-adaptive HMM-based speech synthesis (which uses an “average voice model” plus model adaptation) is robust to non-ideal speech data that are recorded under various conditions and with varying microphones, that are not perfectly clean, and/or that lack phonetic balance. This enables us to consider building high-quality voices on “non-TTS” corpora such as ASR corpora. Since ASR corpora generally include a large number of speakers, this leads to the possibility of producing an enormous number of voices automatically. In this paper, we demonstrate the thousands of voices for HMM-based speech synthesis that we have made from several popular ASR corpora such as the Wall Street Journal (WSJ0, WSJ1, and WSJCAM0), Resource Management, Globalphone, and SPEECON databases. We also present the results of associated analysis based on perceptual evaluation, and discuss remaining issues. Junichi Yamagishi, Bela Usabaev, Simon King 0001, Oliver Watts, John Dines, Jilei Tian, Rile Hu, Keiichiro Oura, Yi-Jian Wu, Keiichi Tokuda, Reima Karhila, Mikko Kurimo |
IEEE Trans. Speech Audio Process. | 11 |
| 2009 | A Bayesian approach to HMM-based speech synthesisabstractThis paper proposes a new framework of speech synthesis based on the Bayesian approach. The Bayesian method is a statistical technique for estimating reliable predictive distributions by marginalizing model parameters. In the proposed framework, all processes for constructing the system can be derived from one single predictive distribution which represents the basic problem of speech synthesis directly. Using HMM as the likelihood function and assuming some approximations, it can be regarded as an application of the variational Bayesian method to the HMM-based speech synthesis. Experimental results show that the proposed method outperforms the conventional one in a subjective test. Kei Hashimoto, Heiga Zen, Yoshihiko Nankaku, Takashi Masuko, Keiichi Tokuda |
ICASSP | 5 |
| 2009 | Full covariance state duration modeling for HMM-based speech synthesisabstractThis paper proposes a state duration modeling method using full covariance matrix for HMM-based speech synthesis. In this method, a full covariance matrix instead of the conventional diagonal covariance matrix is adopted in the multi-dimensional Gaussian distribution to model the state duration of each context-dependent phoneme. At synthesis stage, the state durations are predicted using the clustered context-dependent distributions with full covariance matrices. Experimental results show that the synthesized speech using full-covariance state duration models is more natural than the conventional method when we change the speaking rate of synthesized speech. Heng Lu 0002, Yi-Jian Wu, Keiichi Tokuda, Li-Rong Dai 0001, Renhua Wang |
ICASSP | 3 |
| 2009 | Minimum generation error training by using original spectrum as reference for log spectral distortion measureabstractThis paper improves a minimum generation error (MGE) based HMM training technique for HMM-based speech synthesis by directly using the original spectrum instead of line spectral pairs (LSPs) as reference spectrum for log spectral distortion (LSD) measure. Two types of original reference spectra for LSD calculation are investigated, including the spectrum extracted from speech waveform by STRAIGHT, and the short-time FFT spectrum calculated from speech waveforms. Since only the harmonics of the FFT spectrum are coincident with the underlying spectral envelope, the LSD between generated LSPs and original FFT spectrum is calculated by sampling at the harmonic frequencies, and a weighting function is designed to simulate the sampling strategy on LSPs. From the experimental results, the MGE-LSD training using the FFT spectrum as reference spectrum achieved the best performance. Yi-Jian Wu, Keiichi Tokuda |
ICASSP | 2 |
| 2009 | Voice conversion based on simultaneous modelling of spectrum and F0abstractThis paper proposes a simultaneous modeling of spectrum and F0 for voice conversion based on MSD (multi-space probability distribution) models. As a conventional technique, a spectral conversion based on GMM (Gaussian mixture model) has been proposed. Although this technique converts spectral feature sequences nonlinearly based on GMM, F0 sequences are usually converted by a simple linear function. This is because F0 is undefined in unvoiced segments. To overcome this problem, we apply MSD models. The MSD-GMM allows to model continuous F0 values in voiced frames and a discrete symbol representing unvoiced frames within an unified framework. Furthermore, the MSD-HMM is adopted to model long term correlations in F0 sequences. Kaori Yutani, Yosuke Uto, Yoshihiko Nankaku, Akinobu Lee, Keiichi Tokuda |
ICASSP | 5 |
| 2009 | Stereo-based stochastic noise compensation based on trajectory GMMSabstractThis paper proposes a novel stereo-based stochastic noise compensation technique based on trajectory GMMs. Although the GMM-based noise compensation techniques such as SPLICE work effective, their performance sometimes degrades due to the inappropriate dynamic characteristics caused by the frame-by-frame mapping. While the use of dynamic feature constraints on the mapping stage can alleviate this problem, it also introduces an inconsistency between training and mapping. The recently proposed trajectory GMM-based feature mapping technique can solve this inconsistency while keeping the benefits of the use of dynamic features, and offers an entire sequence-level transformation rather than the frame-by-frame mapping. Results from a noise compensation experiment on the AURORA-2 task show that the proposed trajectory GMM-based noise compensation technique outperforms the conventional ones. Heiga Zen, Yoshihiko Nankaku, Keiichi Tokuda |
ICASSP | 3 |
| 2009 | A Bayesian approach to Hidden Semi-Markov Model based speech synthesisabstractThis paper proposes a Bayesian approach to hidden semi-Markov model (HSMM) based speech synthesis. Recently, hid-den Markov model (HMM) based speech synthesis based on the Bayesian approach was proposed. The Bayesian approach is a statistical technique for estimating reliable predictive distribu-tions by treating model parameters as random variables. In the Bayesian approach, all processes for constructing the system are derived from one single predictive distribution which exactly represents the problem of speech synthesis. However, there is an inconsistency between training and synthesis: although the speech is synthesized from HMMs with explicit state duration probability distributions, HMMs are trained without them. In this paper, we introduce an HSMM, which is an HMM with explicit state duration probability distributions, into the HMM-based Bayesian speech synthesis system. Experimental results show that the use of HSMM improves the naturalness of the synthesized speech. Index Terms: speech synthesis, HSMM, Bayesian approach 1. Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda |
INTERSPEECH | 3 |
| 2009 | A decision tree-based clustering approach to state definition in an excitation modeling framework for HMM-based speech synthesisabstractThis paper presents a decision tree-based algorithm to cluster residual segments assuming an excitation model based on statedependent filtering of pulse train and white noise. The decision tree construction principle is the same as the one applied to speech recognition. Here parent nodes are split using the residual maximum likelihood criterion. Once these excitation decision trees are constructed for residual signals segmented by full context models, using questions related to the full context of the training sentences, they can be utilized for excitation modeling in speech synthesis based on hidden Markov models (HMM). Experimental results have shown that the algorithm in question is very effective in terms of clustering residual signals given segmentation, pitch marks and full context questions, resulting in filters with good residual modeling properties. Ranniery Maia, Tomoki Toda, Keiichi Tokuda, Shinsuke Sakai, Satoshi Nakamura 0001 |
INTERSPEECH | 3 |
| 2009 | Tying covariance matrices to reduce the footprint of HMM-based speech synthesis systemsabstractThis paper proposes a technique of reducing footprint of HMMbased speech synthesis systems by tying all covariance matrices. HMM-based speech synthesis systems usually consume smaller footprint than unit-selection synthesis systems because statistics rather than speech waveforms are stored. However, further reduction is essential to put them on embedded devices which have very small memory. According to the empirical knowledge that covariance matrices have smaller impact for the quality of synthesized speech than mean vectors, here we propose a clustering technique of mean vectors while tying all covariance matrices. Subjective listening test results show that the proposed technique can shrink the footprint of an HMM-based speech synthesis system while retaining the quality of synthesized speech. Index Terms: HMM, speech synthesis, decision tree, contextclustering, MDL criterion, embedded device Keiichiro Oura, Heiga Zen, Yoshihiko Nankaku, Akinobu Lee, Keiichi Tokuda |
INTERSPEECH | 5 |
| 2009 | Deterministic annealing based training algorithm for Bayesian speech recognitionabstractThis paper proposes a deterministic annealing based training algorithm for Bayesian speech recognition. The Bayesian method is a statistical technique for estimating reliable predictive distributions by marginalizing model parameters. However, the local maxima problem in the Bayesian method is more serious than in the ML-based approach, because the Bayesian method treats not only state sequences but also model parameters as latent variables. The deterministic annealing EM (DAEM) algorithm has been proposed to improve the local maxima problem in the EM algorithm, and its effectiveness has been reported in HMMbased speech recognition using ML criterion. In this paper, the DAEM algorithm is applied to Bayesian speech recognition to relax the local maxima problem. Speech recognition experiments show that the proposed method achieved a higher performance than the conventional methods. Sayaka Shiota, Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda |
INTERSPEECH | 4 |
| 2009 | State mapping based method for cross-lingual speaker adaptation in HMM-based speech synthesisabstractA phone mapping-based method had been introduced for cross-lingual speaker adaptation in HMM-based speech syn-thesis. In this paper, we continue to propose a state mapping based method for cross-lingual speaker adaptation, where the state mapping between voice models in source and target lan-guages is established under minimum Kullback-Leibler diver-gence (KLD) criterion. We introduce two approaches to use the established mapping information for cross-lingual speaker adaptation, including data mapping and transform mapping ap-proaches. From the experimental results, the state mapping based method outperformed the phone mapping based method. In addition, the data mapping approach achieved better speaker similarity, and the transform mapping approach achieved better speech quality after cross-lingual speaker adaptation. Index Terms: Speech synthesis, HMM, speaker adaptation, minimum generation error, linear regression Yi-Jian Wu, Yoshihiko Nankaku, Keiichi Tokuda |
INTERSPEECH | 3 |
| 2009 | An improved minimum generation error based model adaptation for HMM-based speech synthesisabstractA minimum generation error (MGE) criterion had been proposed for model training in HMM-based speech synthesis. In this paper, we apply the MGE criterion to model adaptation for HMM-based speech synthesis, and introduce an MGE linear regression (MGELR) based model adaptation algorithm, where the regression matrices used to transform source models are optimized so as to minimize the generation errors of adaptation data. In addition, we incorporate the recent improvements of MGE criterion into MGELR-based model adaptation, including state alignment under MGE criterion and using a log spectral distortion (LSD) instead of Euclidean distance for spectral distortion measure. From the experimental results, the adaptation performance was improved after incorporating these two techniques, and the formal listening tests showed that the quality and speaker similarity of synthesized speech after MGELRbased adaptation were significantly improved over the original MLLR-based adaptation. Index Terms: Speech synthesis, HMM, speaker adaptation, minimum generation error, linear regression Yi-Jian Wu, Keiichi Tokuda |
INTERSPEECH | 3 |
| 2009 | Thousands of voices for HMM-based speech synthesisabstractOur recent experiments with HMM-based speech synthesis systems have demonstrated that speaker-adaptive HMM-based speech synthesis (which uses an 'average voice model' plus model adaptation) is robust to non-ideal speech data that are recorded under various conditions and with varying microphones, that are not perfectly clean, and/or that lack of phonetic balance. This enables us consider building high-quality voices on 'non-TTS' corpora such as ASR corpora. Since ASR corpora generally include a large number of speakers, this leads to the possibility of producing an enormous number of voices automatically. In this paper we show thousands of voices for HMM-based speech synthesis that we have made from several popular ASR corpora such as the Wall Street Journal databases (WSJ0/WSJ1/WSJCAM0), Resource Management, Globalphone and Speecon. We report some perceptual evaluation results and outline the outstanding issues. Junichi Yamagishi, Bela Usabaev, Simon King 0001, Oliver Watts, John Dines, Jilei Tian, Rile Hu, Keiichiro Oura, Keiichi Tokuda, Reima Karhila, Mikko Kurimo |
INTERSPEECH | 10 |
| 2009 | Statistical parametric speech synthesis
Heiga Zen, Keiichi Tokuda, Alan W. Black |
Speech Commun. | 2 |
| 2009 | Robust Speaker-Adaptive HMM-Based Text-to-Speech SynthesisabstractThis paper describes a speaker-adaptive HMM-based speech synthesis system. The new system, called ldquoHTS-2007,rdquo employs speaker adaptation (CSMAPLR+MAP), feature-space adaptive training, mixed-gender modeling, and full-covariance modeling using CSMAPLR transforms, in addition to several other techniques that have proved effective in our previous systems. Subjective evaluation results show that the new system generates significantly better quality synthetic speech than speaker-dependent approaches with realistic amounts of speech data, and that it bears comparison with speaker-dependent approaches even when large amounts of speech data are available. In addition, a comparison study with several speech synthesis techniques shows the new system is very robust: It is able to build voices from less-than-ideal speech data and synthesize good-quality speech even for out-of-domain sentences. Junichi Yamagishi, Takashi Nose, Heiga Zen, Zhen-Hua Ling, Tomoki Toda, Keiichi Tokuda, Simon King 0001, Steve Renals |
IEEE Trans. Speech Audio Process. | 6 |
| 2008 | On the state definition for a trainable excitation model in HMM-based speech synthesisabstractOne of the issues of speech synthesizers based on hidden Markov models concerns the vocoded quality of the synthesized speech. From the principle of analysis-by-synthesis speech coders a trainable excitation model has been proposed to improve naturalness, where the method consists in the design of a set of state-dependent filters in a way to minimize the distortion between residual and synthetic excitation. Although this approach seems successful, state definition still represents an open issue. This paper describes a method for state definition wherein bottom-up clustering is performed on full context decision trees, using the likelihood of the residual database as merging criterion. Experiments have shown that improvement on residual modeling through better filter design can be achieved. Ranniery Maia, Tomoki Toda, Keiichi Tokuda, Shinichi Sakai, Shun Nakamura |
ICASSP | 3 |
| 2008 | Acoustic modeling with contextual additive structure for HMM-based speech recognitionabstractThis paper proposes an acoustic modeling technique based on an additive structure of context dependencies for HMM-based speech recognition. Typical context dependent models, e.g., triphone HMMs, have direct dependencies of phonetic contexts, i.e., if a phonetic context is given, the Gaussian distribution is specified immediately. This paper assumes a more complex structure, an additive structure of acoustic feature components which have different context dependencies. Since the output probability distribution is composed of additive component distributions, a number of different distributions can be efficiently represented by a combination of fewer distributions. To automatically extract additive components, this paper presents a context clustering algorithm for the additive structure model in which multiple decision trees are constructed simultaneously. Experimental results show that the proposed technique improves phoneme recognition accuracy with fewer number of distributions than the conventional triphone HMMs. Yoshihiko Nankaku, Kazuhiro Nakamura, Heiga Zen, Keiichi Tokuda |
ICASSP | 4 |
| 2008 | Statistical approach to vocal tract transfer function estimation based on factor analyzed trajectory HMMabstractIn this paper, we describe a novel statistical approach to the vocal tract transfer function (VTTF) estimation of a speech signal based on a factor analyzed trajectory hidden Markov model (HMM). Because speech is a quasi-periodic signal, there are many missing frequency components between adjacent F0harmonics. The proposed method determines a time-varying VTTF sequence based on the maximum a posteriori (MAP) estimation considering not only harmonic components observed at each analyzed frame but also those at other frames for stochastically interpolating the missing frequency parts. Tomoki Toda, Keiichi Tokuda |
ICASSP | 2 |
| 2008 | Performance evaluation of the speaker-independent HMM-based speech synthesis system "HTS 2007" for the Blizzard Challenge 2007abstractThis paper describes a speaker-independent/adaptive HMM-based speech synthesis system developed for the Blizzard Challenge 2007. The new system, named “HTS-2007”, employs speaker adaptation (CSMAPLR+MAP), feature-space adaptive training, mixed-gender modeling, and full-covariance modeling using CSMAPLR transforms, in addition to several other techniques that have proved effective in our previous systems. Subjective evaluation results show that the new system generates significantly better quality synthetic speech than that of speaker-dependent approaches with realistic amounts of speech data, and that it bears comparison with speaker-dependent approaches even when large amounts of speech data are available. Junichi Yamagishi, Takashi Nose, Heiga Zen, Tomoki Toda, Keiichi Tokuda |
ICASSP | 5 |
| 2008 | Bayesian context clustering using cross valid prior distribution for HMM-based speech recognitionabstractThis paper proposes a prior distribution determination tech-nique using cross validation for speech recognition based on the Bayesian approach. The Bayesian method is a statisti-cal technique for estimating reliable predictive distributions by marginalizing model parameters and its approximate version, the variational Bayesian method has been applied to HMM-based speech recognition. Since prior distributions represent-ing prior information about model parameters affect the pos-terior distributions and model selection, the determination of prior distributions is an important problem. However, it has not been thoroughly investigate in speech recognition. The pro-posed method can determine reliable prior distributions with-out tuning parameters and select an appropriate model struc-ture dependently on the amount of training data. Continu-ous phoneme recognition experiments show that the proposed method achieved a higher performance than the conventional methods. Index Terms: variational Bayes, cross validation, context clus-tering, continuous phoneme recognition Kei Hashimoto, Heiga Zen, Yoshihiko Nankaku, Akinobu Lee, Keiichi Tokuda |
INTERSPEECH | 5 |
| 2008 | Speaker recognition based on variational Bayesian method
Tatsuya Ito, Kei Hashimoto, Yoshihiko Nankaku, Akinobu Lee, Keiichi Tokuda |
INTERSPEECH | 5 |
| 2008 | Unsupervised adaptation for HMM-based speech synthesisabstractIt is now possible to synthesise speech using HMMs with a comparable quality to unit-selection techniques. Generating speech from a model has many potential advantages over concatenating waveforms. The most exciting is model adaptation. It has been shown that supervised speaker adaptation can yield highquality synthetic voices with an order of magnitude less data than required to train a speaker-dependent model or to build a basic unit-selection system. Such supervised methods require labelled adaptation data for the target speaker. In this paper, we introduce a method capable of unsupervised adaptation, using only speech from the target speaker without any labelling. Index Terms: speech synthesis, HMM-based speech synthesis, HTS, trajectory HMMs, speaker adaptation, MLLR Simon King 0001, Keiichi Tokuda, Heiga Zen, Junichi Yamagishi |
INTERSPEECH | 2 |
| 2008 | Acoustic modeling based on model structure annealing for speech recognitionabstractThis paper proposes an HMM training technique using multiple phonetic decision trees and evaluates it in speech recognition. In the use of context dependent models, the decision tree based context clustering is applied to find a parameter tying structure. However, the clustering is usually performed based on statistics of HMM state sequences which are obtained by unreliable models without context clustering. To avoid this problem, we optimize the decision trees and HMM state sequences simultaneously. In the proposed method, this is performed by maximum likelihood (ML) estimation of a newly defined statistical model which includes multiple decision trees as hidden variables. Applying the deterministic annealing expectation maximization (DAEM) algorithm and using multiple decision trees in early stage of model training, state sequences are reliably estimated. In continuous phoneme recognition experiments, the proposed method can improve the recognition performance. Sayaka Shiota, Kei Hashimoto, Heiga Zen, Yoshihiko Nankaku, Akinobu Lee, Keiichi Tokuda |
INTERSPEECH | 6 |
| 2008 | Minimum generation error training with direct log spectral distortion on LSPs for HMM-based speech synthesisabstractA minimum generation error (MGE) criterion had been proposed to solve the issues related to maximum likelihood (ML) based HMM training in HMM-based speech synthesis. In this paper, we improve the MGE criterion by imposing a log spectral distortion (LSD) instead of the Euclidean distance to define the generation error between the original and generated line spectral pair (LSP) coefficients. Moreover, we investigate the effect of different sampling strategies to calculate the integration of the LSD function. From the experimental results, using the LSDs calculated by sampling at LSPs achieved the best performance, and the quality of synthesized speech after the MGE-LSD training was improved over the original MGE training. Yi-Jian Wu, Keiichi Tokuda |
INTERSPEECH | 2 |
| 2008 | Probabilistic answer selection based on conditional random fields for spoken dialog system
Yoshitaka Yoshimi, Ryota Kakitsuba, Yoshihiko Nankaku, Akinobu Lee, Keiichi Tokuda |
INTERSPEECH | 5 |
| 2008 | Simultaneous conversion of duration and spectrum based on statistical models including time-sequence matchingabstractThis paper describes a simultaneous conversion technique of duration and spectrum based on a statistical model including time-sequence matching. Conventional GMM-based approaches cannot perform spectral conversion taking account of speaking rate because it assumes one to one frame matching between source and target features. However, speaker characteristics may appear in speaking rates. In order to perform duration conversion, we attach duration models to statistical models including time-sequence matching (DPGMM). Since DPGMM can represent two different length sequences directly, the conversion of spectrum and duration can be performed within an integrated framework. In the proposed technique, each mixture component of DPGMM has different duration transformation functions, therefore durations are converted nonlinearly and dependently on spectral information. In the subjective DMOS test, the proposed method is superior to the conventional method. Index Terms: voice conversion, GMM, duration conversion 1. Kaori Yutani, Yosuke Uto, Yoshihiko Nankaku, Tomoki Toda, Keiichi Tokuda |
INTERSPEECH | 5 |
| 2008 | Probabilistic feature mapping based on trajectory HMMs
Heiga Zen, Yoshihiko Nankaku, Keiichi Tokuda |
INTERSPEECH | 3 |
| 2008 | Statistical mapping between articulatory movements and acoustic spectrum using a Gaussian mixture model
Tomoki Toda, Alan W. Black, Keiichi Tokuda |
Speech Commun. | 3 |
| 2007 | Statistical Parametric Speech SynthesisabstractThis paper gives a general overview of techniques in statistical parametric speech synthesis. One of the instances of these techniques, called HMM-based generation synthesis (or simply HMM-based synthesis), has recently been shown to be very effective in generating acceptable speech synthesis. This paper also contrasts these techniques with the more conventional unit selection technology that has dominated speech synthesis over the last ten years. Advantages and disadvantages of statistical parametric synthesis are highlighted as well as identifying where we expect the key developments to appear in the immediate future. Alan W. Black, Heiga Zen, Keiichi Tokuda |
ICASSP (4) | 3 |
| 2007 | Face Recognition using Hidden Markov Eigenface ModelsabstractThis paper proposes hidden Markov eigenface models (HMEMs) in which the eigenfaces are integrated into separable lattice hidden Markov models (SL-HMMs). SL-HMMs have been proposed for modeling multi-dimensional data, e.g., images, image sequences, 3-D objects. In its application to face recognition, SL-HMMs can perform an elastic image matching in both horizontal and vertical directions. However, SL-HMMs still have a limitation that the observations are assumed to be generated independently from corresponding states; it is insufficient to represent variations in face images, e.g., lighting conditions, facial expressions, etc. To overcome this problem, the structure of probabilistic principal component analysis (PPCA) and factor analysis (FA) is used as a probabilistic representation of eigenfaces. The proposed model has good properties of both PPCA/FA and SL-HMMs: a linear feature extraction and invariances to size and location of images. In face recognition experiments on the XM2VTS database, the proposed model improved the performance significantly. Yoshihiko Nankaku, Keiichi Tokuda |
ICASSP (2) | 2 |
| 2007 | A trainable excitation model for HMM-based speech synthesisabstractThis paper introduces a novel excitation approach for speech synthesizers in which the final waveform is generated through parameters directly obtained from Hidden Markov Models (HMMs). Despite the attractiveness of the HMM-based speech synthesis technique, namely utilization of small corpora and flexibility concerning the achievement of different voice styles, synthesized speech presents a characteristic buzziness caused by the simple excitation model which is employed during the speech production. This paper presents an innovative scheme where mixed excitation is modeled through closed-loop training of a set of state-dependent filters and pulse trains, with minimization of the error between excitation and residual sequences. The proposed method shows effectiveness, yielding synthesized speech with quality far superior to the simple excitation baseline and comparable to the best excitation schemes thus far reported for HMM-based speech synthesis. Ranniery Maia, Tomoki Toda, Heiga Zen, Yoshihiko Nankaku, Keiichi Tokuda |
INTERSPEECH | 5 |
| 2007 | Model-space MLLR for trajectory HMMsabstractThis paper proposes model-space Maximum Likelihood Linear Regression (mMLLR) based speaker adaptation technique for trajectory HMMs, which have been derived from HMMs by imposing explicit relationships between static and dynamic features. This model can alleviate two limitations of the HMM: constant statistics within a state and conditional independence assumption of state out-put probabilities without increasing the number of model parameters. Results in a continuous speech recognition experiments show that the proposed algorithm can adapt trajectory HMMs to a specific speaker and improve the performance of a trajectory HMM-based speech recogni-tion system. Index Terms: trajectory HMM, speaker adaptation, model-space MLLR Heiga Zen, Yoshihiko Nankaku, Keiichi Tokuda |
INTERSPEECH | 3 |
| 2007 | Reformulating the HMM as a trajectory model by imposing explicit relationships between static and dynamic feature vector sequences
Heiga Zen, Keiichi Tokuda, Tadashi Kitamura |
Comput. Speech Lang. | 2 |
| 2007 | Voice Conversion Based on Maximum-Likelihood Estimation of Spectral Parameter TrajectoryabstractIn this paper, we describe a novel spectral conversion method for voice conversion (VC). A Gaussian mixture model (GMM) of the joint probability density of source and target features is employed for performing spectral conversion between speakers. The conventional method converts spectral parameters frame by frame based on the minimum mean square error. Although it is reasonably effective, the deterioration of speech quality is caused by some problems: 1) appropriate spectral movements are not always caused by the frame-based conversion process, and 2) the converted spectra are excessively smoothed by statistical modeling. In order to address those problems, we propose a conversion method based on the maximum-likelihood estimation of a spectral parameter trajectory. Not only static but also dynamic feature statistics are used for realizing the appropriate converted spectrum sequence. Moreover, the oversmoothing effect is alleviated by considering a global variance feature of the converted spectra. Experimental results indicate that the performance of VC can be dramatically improved by the proposed method in view of both speech quality and conversion accuracy for speaker individuality. Tomoki Toda, Alan W. Black, Keiichi Tokuda |
IEEE Trans. Speech Audio Process. | 3 |
| 2006 | Face Recognition Based on Separable Lattice HMMSabstractIn this paper, we propose separable lattice hidden Markov models, in which multiple hidden state sequences interact to model the observation on a lattice. The proposed model can be efficiently applied for modeling images, image sequences, 3-D object models and higher dimensional applications, due to the composite structure of Markov chains which reduces the complexity while retaining good properties for multi-dimensional data. In case of 2-D lattices, the proposed model performs an elastic matching in both horizontal and vertical directions; this makes it possible to model not only invariances to the size and location of an object but also nonlinear warping in each dimension. We present a training algorithm for separable lattice HMMs based on a variational approximation. Moreover, the deterministic annealing EM (DAEM) algorithm was applied to the variational algorithm for separable lattice HMMs. Face recognition experiments on the XM2VTS database show that the proposed model has good properties for face image modeling. Daisuke Kurata, Yoshihiko Nankaku, Keiichi Tokuda, Tadashi Kitamura, Zoubin Ghahramani |
ICASSP (5) | 3 |
| 2006 | On the Use of Phonetic Information for Mapping from Articulatory Movements to Vocal Tract SpectrumabstractThis paper describes a method for determining the vocal tract spectrum from articulatory movements using an hidden Markov models (HMMs). In the proposed system, articulatory parameters are generated from a TTS system and converted to acoustic features to be synthesized. Comparing with conventional GMM-based systems, the proposed system has two additional properties: 1) phonetic information given input texts is available for the conversion, 2) the use of HMMs allows us to utilize the temporal structure of speech. In this paper, we investigate the optimal structure of HMMs for the conversion. Experimental results show that using phonetic and temporal information can improve the mapping accuracy in a spectral distortion measure Kenichi Nakamura, Tomoki Toda, Yoshihiko Nankaku, Keiichi Tokuda |
ICASSP (1) | 4 |
| 2006 | Hidden Semi-Markov Model Based Speech Recognition System using Weighted Finite-State TransducerabstractIn hidden Markov models (HMMs), state duration probabilities decrease exponentially with time. It would be inappropriate representation of temporal structure of speech. One of the solutions for this problem is integrating state duration probability distributions explicitly into the HMM. This form is known as a hidden semi-Markov model (HSMM) [1]. Although a number of attempts to use explicit duration models in speech recognition systems have been proposed, they are not consistent because various approximations were used in both training and decoding. In the present paper, a fully consistent speech recognition system based on the HSMM framework is proposed. In a speaker-dependent continuous speech recognition experiment, HSMM-based speech recognition system achieved about 5.9% relative error reduction over the corresponding HMM-based one. Keiichiro Oura, Heiga Zen, Yoshihiko Nankaku, Akinobu Lee, Keiichi Tokuda |
ICASSP (1) | 5 |
| 2006 | Estimating Trajectory Hmm Parameters Using Monte Carlo Em With Gibbs SamplerabstractIn the present paper, the Monte Carlo EM (MCEM) algorithm with a Gibbs sampler is applied for estimating parameters of a trajectory HMM, which has been derived from an HMM by imposing explicit relationships between static and dynamic features. The trajectory HMM can alleviate two limitations of the HMM, which are i) constant statistics within a state, and ii) conditional independence of state output probabilities, without increasing the number of model parameters. In a speaker-dependent continuous speech recognition experiment, trajectory HMMs estimated by the MCEM algorithm achieved significant improvements over the corresponding HMMs trained by the EM (Baum-Welch) algorithm. Heiga Zen, Yoshihiko Nankaku, Keiichi Tokuda, Tadashi Kitamura |
ICASSP (1) | 3 |
| 2006 | Reducing computation on parallel decoding using frame-wise confidence scores
Tomohiro Hakamata, Akinobu Lee, Yoshihiko Nankaku, Keiichi Tokuda |
INTERSPEECH | 4 |
| 2006 | An HMM-based singing voice synthesis systemabstractAbstract The present paper describes a corpus-based singing voice syn-thesis system based on hidden Markov models (HMMs). Thissystem employs the HMM-based speech synthesis to synthesizesingingvoice. Musical information such aslyrics, tones, durationsis modeled simultaneously in a unified framework of the context-dependent HMM. It can mimic the voice quality and singing styleof the original singer. Results of a singing voice synthesis exper-iment show that the proposed system can synthesize smooth andnatural-sounding singing voice. Index Terms : singing voice synthesis, HMM, time-lag model. 1. Introduction In recent years, various applications of speech synthesis systemshave been proposed and investigated. Singing voice synthesis isone of the hot topics in this area [1–5]. However, only a fewcorpus-based singing voice synthesis systems which can be con-structed automatically have been proposed.Currently, there are two main paradigms in the corpus-basedspeech synthesis area: sample-based approach and statistical ap-proach. The sample-based approach such as unit selection [6]can synthesize high-quality speech. However, it requires a hugeamountoftrainingdatatorealizevariousvoicecharacteristics. Onthe other hand, the quality of statistical approach such as HMM-basedspeechsynthesis[7]isbuzzybecauseitisbasedonavocod-ingtechnique. However,itissmoothandstable,anditsvoicechar-acteristics can easily be modified by transforming HMM parame-ters appropriately. For singing voice synthesis, applying the unitselection seems to be difficult because a huge amount of singingspeech which covers vast combinations of contextual factors thataffect singing voice has to be recorded. On the other hand, theHMM-based system can be constructed using a relatively smallamount of training data. From this point of view, the HMM-basedapproach seems to be more suitable for the singing voice synthe-sizer. In the present paper, we apply the HMM-based synthesisapproach to singing voice synthesis.Although the singing voice synthesis system proposed in thepresent paper is quite similar to the HMM-based text-to-speechsynthesissystem[7],therearetwomaindifferencesbetweenthem.In the HMM-based text-to-speech synthesis system, contextualfactors which may affect reading speech (e.g. phonemes, sylla-bles, words, phrases, etc.) are taken into account. However, con-textual factors which may affect singing voice should be different Keijiro Saino, Heiga Zen, Yoshihiko Nankaku, Akinobu Lee, Keiichi Tokuda |
INTERSPEECH | 5 |
| 2006 | Voice conversion based on mixtures of factor analyzersabstractThis paper describes the voice conversion based on the Mixtures of Factor Analyzers (MFA) which can provide an efficient modeling with a limited amount of training data. As a typical spectral conversion method, a mapping algorithm based on the Gaussian Mixture Model (GMM) has been proposed. In this method two kinds of covariance matrix structures are often used : the diagonal and full covariance matrices. GMM with diagonal covariance matrices requires a large number of mixture components for accurately estimating spectral features. On the other hand, GMM with full covariance matrices needs sufficient training data to estimate model parameters. In order to cope with these problems, we apply MFA to voice conversion. MFA can be regarded as intermediate model between GMM with diagonal covariance and with full covariance. Experimental results show that MFA can improve the conversion accuracy compared with the conventional GMM. Yosuke Uto, Yoshihiko Nankaku, Tomoki Toda, Akinobu Lee, Keiichi Tokuda |
INTERSPEECH | 5 |
| 2006 | Speaker adaptation of trajectory HMMs using feature-space MLLRabstractAbstract Recently, a trajectory model, derived from the hiddenMarkov model (HMM) by imposing explicit relationshipsbetween static and dynamic features, has been proposed.The derived model, named trajectory HMM , can alleviatetwo limitations of the HMM: constant statistics within astate and conditional independence assumption of state out-put probabilities. In the present paper, a speaker adapta-tion algorithm for the trajectory HMM based on feature-space Maximum Likelihood Linear Regression (fMLLR)is derived and evaluated. Results of a simple continu-ous speech recognition experiment shows that adapting tra-jectory HMMs using the derived adaptation algorithm im-proves the speech recognition performance. Index Terms : trajectory HMM, adaptation, fMLLR. 1. Introduction Speech recognition technologies have achieved significantprogress with the introduction of hidden Markov models(HMMs). Their tractability and efficient implementationsare achieved by a number of assumptions, such as constantstatistics within an HMM state, conditional independenceof state output probabilities. Although these assumptionsmake the HMM practically useful, they are not realistic formodeling sequences of speech spectra, especially in spon-taneous speech. To overcome these shortcomings of theHMM, a variety of alternative models have been proposed,e.g., [1–3]. Although these models can improve the speechrecognition performance, they generally require an increaseinthenumberofmodelparametersandcomputationalcom-plexity. Alternatively, the use of dynamic features (deltaand delta-delta features) [4] also improves the performanceof HMM-based speech recognizers. It can be viewed as asimple mechanism to capture time dependencies. However,it has been thought of as an ad hoc rather than an essentialsolution. Generally, dynamic features are calculated as re-gression coefficients from their neighboring static features.Therefore, relationshipsbetweenstaticanddynamicfeaturevector sequences are deterministic. However, usually theserelationships are ignored and the static and dynamic fea-tures are modeled as independent random variables. Ignor-ing these dependencies allows inconsistency between thestatic and dynamic features when the HMM is used as agenerative model in the obvious way.Recently, a trajectory model, derived from the HMM byimposing the explicit relationships between static and dy-namic features, has been proposed [5]. The derived model,named Heiga Zen, Yoshihiko Nankaku, Keiichi Tokuda, Tadashi Kitamura |
INTERSPEECH | 3 |
| 2005 | Minimum Classification Error Interactive Training for Speaker IdentificationabstractThis paper describes an online discriminative training algorithm aiming at achieving speaker identification on interactive robots. A robot incrementally acquires speakers' voice characteristics during the interaction with the speakers. We simulate the situation that the speakers never give their IDs and the robot can only know whether the identification decision was correct or not from the speaker's positive or negative behavioral reaction. The speaker models are adjusted based on this limited information using minimum classification error (MCE) training consisting of positive and negative adaptation. In cases of correct identification, the conventional MCE training algorithm can be used. We compare three kinds of negative adaptation algorithms for the cases of incorrect identification. Experimental results show that the combination of the positive and negative adaptation achieves faster convergence, and negative adaptation which adjusts only a misclassified speaker model reaches an identification rate of 80% four times faster than the positive adaptation alone. Yusuke Kida, Hiroyoshi Yamamoto, Chiyomi Miyajima, Keiichi Tokuda, Tadashi Kitamura |
ICASSP (1) | 4 |
| 2005 | Sparse KPCA for Feature Extraction in Speech RecognitionabstractThis paper presents an analysis of the applicability of sparse kernel principal component analysis (SKPCA) for feature extraction in speech recognition, as well as a proposed approach to make the SKPCA technique realizable for a large amount of training data, which is a usual context in speech recognition systems. Although the KPCA (kernel principal component analysis) has proved to be an efficient technique for being applied to speech recognition, it has the disadvantage of requiring training data reduction, when its amount is excessively large. The standard approach to perform this data reduction is to randomly choose frames from the original data set, which does not necessarily provide a good statistical representation of the original data set. In order to solve this problem a likelihood related re-estimation procedure was applied to the KPCA framework, thus creating the SKPCA. The experimental results show the efficiency of SKPCA technique with the proposed approach over the KPCA with the standard sparse solution using randomly chosen frames and the standard feature extraction techniques. Amaro A. de Lima, Heiga Zen, Yoshihiko Nankaku, Keiichi Tokuda, Tadashi Kitamura, Fernando Gil Vianna Resende Jr. |
ICASSP (1) | 4 |
| 2005 | Spectral Conversion Based on Maximum Likelihood Estimation Considering Global Variance of Converted ParameterabstractThe paper describes a novel spectral conversion method for voice transformation. We perform spectral conversion between speakers using a Gaussian mixture model (GMM) on the joint probability density of source and target features. A smooth spectral sequence can be estimated by applying maximum likelihood (ML) estimation to the GMM-based mapping using dynamic features. However, there is still degradation of the converted speech quality due to an over-smoothing of the converted spectra, which is inevitable in conventional ML-based parameter estimation. In order to alleviate the over-smoothing, we propose an ML-based conversion taking account of the global variance of the converted parameter in each utterance. Experimental results show that the performance of the voice conversion can be improved by using the global variance information. Moreover, it is demonstrated that the proposed algorithm is more effective than spectral enhancement by postfiltering. Tomoki Toda, Alan W. Black, Keiichi Tokuda |
ICASSP (1) | 3 |
| 2005 | HMM-based european Portuguese TTS system
Maria João Barros, Ranniery Maia, Keiichi Tokuda, Fernando Gil Vianna Resende Jr., Diamantino Freitas |
INTERSPEECH | 3 |
| 2005 | The blizzard challenge - 2005: evaluating corpus-based speech synthesis on common datasetsabstractIn order to better understand different speech synthesis techniques on a common dataset, we devised a challenge that will help us better compare research techniques in building corpusbased speech synthesizers. In 2004, we released the first two 1200-utterance single-speaker databases from the CMU ARC-TIC speech databases, and challenged current groups working in speech synthesis around the world to build their best voices from these databases. In January of 2005, we released two further databases and a set of 50 utterance texts from each of five genres and asked the participants to synthesize these utterances. Their resulting synthesized utterances were then presented to three groups of listeners: speech experts, volunteers, and US English-speaking undergraduates. This paper summarizes the purpose, design, and whole process of the challenge. Alan W. Black, Keiichi Tokuda |
INTERSPEECH | 2 |
| 2005 | Speech parameter generation algorithm considering global variance for HMM-based speech synthesisabstractThis paper describes a novel parameter generation algorithm for the HMM-based speech synthesis. The conventional algorithm generates a trajectory of static features that maximizes an output probability of a parameter sequence consisting of the static and dynamic features from HMMs under an actual constraint between the two features. The generated trajectory is often excessively smoothed due to the statistical processing. Using the over-smoothed trajectory causes the muffled sound. In order to alleviate the over-smoothing effect, we propose the generation algorithm considering not only the output probability used for the conventional method but also that of a global variance (GV) of the generated trajectory. The latter probability works as a penalty for a reduction of the variance of the generated trajectory. A result of a perceptual evaluation demonstrates that the proposed method causes large improvements of the naturalness of synthetic speech. 1. Tomoki Toda, Keiichi Tokuda |
INTERSPEECH | 2 |
| 2004 | Parameter sharing and minimum classification error training of mixtures of factor analyzers for speaker identificationabstractThis paper investigates the parameter tying strategies of mixtures of factor analyzers (MFA) and discriminative training of MFA for speaker identification. The parameters of factor loading matrices or diagonal matrices are shared in different mixtures of MFA. The minimum classification error (MCE) training is applied to the MFA parameters to enhance the discrimination abilities. The results of text-independent speaker identification experiments show that MFA outperforms the conventional Gaussian mixture models (GMM) with diagonal or full covariance matrices and achieves the best performance when sharing the diagonal matrices, resulting in a relative gain of 26% over the GMM with diagonal covariance matrices. The recognition performance is further improved by the MCE training with an additional 3% error reduction. Hiroyoshi Yamamoto, Yoshihiko Nankaku, Chiyomi Miyajima, Keiichi Tokuda, Tadashi Kitamura |
ICASSP (1) | 4 |
| 2004 | A Viterbi algorithm for a trajectory model derived from HMM with explicit relationship between static and dynamic featuresabstractThis paper introduces a Viterbi algorithm to obtain a sub-optimal state sequence for trajectory-HMM, which is derived from HMM with explicit relationship between static and dynamic features. The trajectory-HMM can alleviate some limitations of HMM, which are (i) constant statistics within HMM state and (ii) conditional independence of observations given the state sequence, without increasing the number of model parameters. The proposed algorithm was applied to state-boundary optimization for Viterbi training and N-best rescoring. In a speaker-dependent continuous speech recognition experiment, trajectory-HMM with the proposed algorithm achieved about 14% error reduction over the standard HMM with the conventional Viterbi algorithm. Heiga Zen, Keiichi Tokuda, Tadashi Kitamura |
ICASSP (1) | 2 |
| 2004 | Deterministic annealing EM algorithm in parameter estimation for acoustic modelabstractABSTRACT This paper investigates the effectiveness of the DAEM (Determin-istic Annealing EM) algorithm in acoustic modeling for speakerand speech recognition. Although the EM algorithm has beenwidely used to approximate the ML estimates, it has the problemof initialization dependence. To relax this problem, the DAEMalgorithm has been proposed and confirmed the effectiveness insmall tasks. In this paper, we applied the DAEM algorithm tospeakerrecognitionbasedonGMMsandcontinuousspeechrecog-nitionbasedonHMMs. ExperimentalresultsshowthattheDAEMalgorithm can improve the recognition performance as comparedtotheordinaryEMalgorithmwithconventionalinitializationmeth-ods,especiallyintheflatstarttrainingforcontinuousspeechrecog-nition. 1. INTRODUCTION The EM (Expectation-Maximization) algorithm [1] is widelyused for parameter estimation of statistical models with hiddenvariables. This algorithm provides a simple iterative proceduretoobtainapproximateML(maximumlikelihood)estimates. How-ever, since the EM algorithm is a hill-climbing approach, it suffersfrom the local maxima problem.On the other hand, GMMs (Gaussian mixture models) [2] andHMMs (hidden Markov models) [3] have been commonly usedin acoustic modeling for speaker and speech recognition, respec-tively. In conventional approaches, the LBG algorithm for GMMsand the segmental k-means algorithm for HMMs have been em-ployed to obtain initial model parameters before applying the EMalgorithm. However these initial values are not guaranteed to benear the true maximum likelihood point, and the posterior den-sity becomes unreliable at an early stage of training. Especiallyin continuous speech recognition, it is difficult to obtain accuratephoneme boundaries for all training data. Hence, the embeddedtraininghasbeenusedinwhichphonemeboundariesarealsodealtas hidden variables, and estimated based on the EM algorithm.Furthermore, in the worse case that the boundary information isnot available, a method called the flat start training is often ap-plied. In this method, initial parameters of HMMs are given bymaking all states of all models equal, and then carry out the em-bedded training. In these situations, we do not have enough priorknowledge to obtain a good initial values for the EM algorithm,and it would converge to one of the local maxima or saddle points Yohei Itaya, Heiga Zen, Yoshihiko Nankaku, Chiyomi Miyajima, Keiichi Tokuda, Tadashi Kitamura |
INTERSPEECH | 5 |
| 2004 | Decision-tree backing-off in HMM-based speech synthesisabstractThis paper proposes a decision-tree backing-off technique for an HMM-based speech synthesis system. In the system, a decision-tree based context clustering technique is used for constructing parameter tying structures. In the context clustering, the MDL criterion has been used as a stopping criterion. In this paper, however, huge decision-trees are constructed without any stopping criterion. In the synthesis phase, decision-trees obtained in this way are used in the proposed backing-off scheme. This enables us to adjust the cluster size dynamically at runtime according to the text to be synthesized. Results of subjective listening tests show that the proposed technique improves the synthesized speech quality. Shunsuke Kataoka, Nobuaki Mizutani, Keiichi Tokuda, Tadashi Kitamura |
INTERSPEECH | 3 |
| 2004 | Acoustic-to-articulatory inversion mapping with Gaussian mixture modelabstractThis paper describes the acoustic-to-articulatory inversion mapping using a Gaussian Mixture Model (GMM).Correspondence of an acoustic parameter and an articulatory parameter is modeled by the GMM trained using the parallel acousticarticulatory data.We measure the performance of the GMMbased mapping and investigate the effectiveness of using multiple acoustic frames as an input feature and using multiple mixtures.As a result, it is shown that although increasing the number of mixtures is useful for reducing the estimation error, it causes many discontinuities in the estimated articulatory trajectories.In order to address this problem, we apply maximum likelihood estimation (MLE) considering articulatory dynamic features to the GMM-based mapping.Experimental results demonstrate that the MLE using dynamic features can estimate more appropriate articulatory movements compared with the GMM-based mapping applied smoothing by lowpass filter. Tomoki Toda, Alan W. Black, Keiichi Tokuda |
INTERSPEECH | 3 |
| 2004 | Constructing emotional speech synthesizers with limited speech databaseabstractThis paper describes an emotional speech synthesis system based on HMMs and related modeling techniques. For concatenative speech synthesis, we require all of the concatenation units that will be used to be recorded beforehand and made available at synthesis time. To adopt this approach for synthesizing the wide variety of human emotions possible in speech, implies that this process should be repeated for every targeted emotion making this task challenging and time consuming. In this paper, we propose an emotional speech synthesis technique based on HMMs, especially for the case where only limited amount of training data is available, directly incorporating subjective evaluation results performed on the training data. Listening results performed on the synthesized speech suggest that the proposed technique helps to improve the emotional content of synthesized speech. Heiga Zen, Tadashi Kitamura, Murtaza Bulut, Shri Narayanan, Ryosuke Tsuzuki, Keiichi Tokuda |
INTERSPEECH | 6 |
| 2004 | Hidden semi-Markov model based speech synthesisabstractIn the present paper, a hidden-semi Markov model (HSMM) based speech synthesis system is proposed. In a hidden Markov model (HMM) based speech synthesis system which we have proposed, rhythm and tempo are controlled by state duration probability distributions modeled by single Gaussian distributions. To synthesis speech, it constructs a sentence HMM corresponding to an arbitralily given text and determine state durations maximizing their probabilities, then a speech parameter vector sequence is generated for the given state sequence. However, there is an inconsistency: although the speech is synthesized from HMMs with explicit state duration probability distributions, HMMs are trained without them. In the present paper, we introduce an HSMM, which is an HMM with explicit state duration probability distributions, into the HMM-based speech synthesis system. Experimental results show that the use of HSMM training improves the naturalness of the synthesized speech. Heiga Zen, Keiichi Tokuda, Takashi Masuko, Takao Kobayashi, Tadashi Kitamura |
INTERSPEECH | 2 |
| 2003 | Improving the performance of HMM-based very low bit rate speech codingabstractIn this paper, we define an F0 quantization scheme for a very low bit rate speech coder based on HMM (hidden Markov model). In the coding system, the encoder carries out phoneme recognition, and transmits phoneme indices, state durations and F0 information to the decoder. In the decoder, phoneme HMM are concatenated according to the phoneme indices, and a sequence of mel-cepstral coefficient vectors is generated from the concatenated HMM. Finally we obtain synthetic speech by using the MLSA (mel log spectrum approximation) filter according to the mel-cepstral coefficients and F0 information. In addition to the F0 quantization, we investigate encoding methods for other parameters to reduce the bit rate, yet keeping the subjective speech quality. A subjective listening test shows that the performance of the proposed coder at about 100/spl sim/150 bit/s is superior to a VQ-based vocoder at 600 bit/s (mel-cepstrum: 6 bit/frame/spl times/50 frame/s, F0: 6 bit/frame/spl times/50 frame/s). Takahiro Hoshiya, Shinji Sako, Heiga Zen, Keiichi Tokuda, Takashi Masuko, Takao Kobayashi, Tadashi Kitamura |
ICASSP (1) | 4 |
| 2003 | Speech recognition using voice-characteristic-dependent acoustic modelsabstractThis paper proposes a speech recognition technique based on acoustic models considering voice characteristic variations. Context-dependent acoustic models, which are typically triphone HMM, are often used in continuous speech recognition systems. This work hypothesizes that the speaker voice characteristics that humans can perceive by listening are also factors in acoustic variation for construction of acoustic models, and a tree-based clustering technique is also applied to speaker voice characteristics to construct voice-characteristic-dependent acoustic models. In speech recognition using triphone models, the neighboring phonetic context is given from the linguistic-phonetic knowledge in advance; in contrast, the voice characteristics of input speech are unknown in recognition using voice-characteristic-dependent acoustic models. This paper proposes a method of recognizing speech even under conditions where the voice characteristics of the input speech are unknown. The result of a gender-dependent speech recognition experiment shows that the proposed method achieves higher recognition performance in comparison to conventional methods. Hiroyuki Suzuki, Heiga Zen, Yoshihiko Nankaku, Chiyomi Miyajima, Keiichi Tokuda, Tadashi Kitamura |
ICASSP (1) | 5 |
| 2003 | A training method for average voice model based on shared decision tree context clustering and speaker adaptive trainingabstractThis paper describes a new training method of average voice model for speech synthesis in which an arbitrary speaker's voice is generated based on speaker adaptation. When the amount of training data is limited, the distributions of average voice model often have bias depending on speaker and/or gender and this will degrade the quality of synthetic speech. In the proposed method, to reduce the influence of speaker dependence, we incorporate a context clustering technique called shared decision tree context clustering and speaker adaptive training into the training procedure of the average voice model. From the results of subjective tests, we show that the average voice model trained using the proposed method generates more natural sounding speech than the conventional average voice model. Moreover, it is shown that voice characteristics of synthetic speech generated from the adapted model using the proposed method are closer to the target speaker than the conventional method. Junichi Yamagishi, Takashi Masuko, Keiichi Tokuda, Takao Kobayashi |
ICASSP (1) | 3 |
| 2003 | On the use of kernel PCA for feature extraction in speech recognitionabstractThis paper describes an approachfor feature extraction in speech recognition systems using kernel principal componentanalysis (KPCA). This approachconsists in representing speech features as the projection of the extracted speech features mapped into a feature space via a nonlinear mapping onto the principal components. The nonlinear mapping is implicitly performed using the kerneltrick, which is an useful way of not mapping the input space into a featurespace explicitly,makingthis mapping computationally feasible. Better results were obtained by using this approach when compared to the standard technique. Amaro A. de Lima, Heiga Zen, Yoshihiko Nankaku, Chiyomi Miyajima, Keiichi Tokuda, Tadashi Kitamura |
INTERSPEECH | 5 |
| 2003 | Towards the development of a brazilian portuguese text-to-speech system based on HMMabstractThis paper describes the development of a Brazilian Portuguese text-to-speech system which applies a technique wherein speech is directly synthesized from hidden Markov models. In order to build the synthesizer a speech database was recorded and phonetically segmented. Furthermore, contextual informations about syllables, words, phrases, and utterances were determined, as well as questions for decision tree-based context clustering algorithms. The resulting system presents a fair reproduction of the prosody even when a small database is used for training. Ranniery Maia, Heiga Zen, Keiichi Tokuda, Tadashi Kitamura, Fernando Gil Vianna Resende Jr. |
INTERSPEECH | 3 |
| 2003 | Trajectory modeling based on HMMs with the explicit relationship between static and dynamic featuresabstractThis paper shows that the HMM whose state output vector includes static and dynamic feature parameters can be reformulated as a trajectory model by imposing the explicit relationship between the static and dynamic features. The derived model, named trajectory HMM, can alleviate the limitations of HMMs: i) constant statistics within an HMM state and ii) independence assumption of state output probabilities. We also derive a Viterbi-type training algorithm for the trajectory HMM. A preliminary speech recognition experiment based on N-best rescoring demonstrates that the training algorithm can improve the recognition performance significantly even though the trajectory HMM has the same parameterization as the standard HMM. Keiichi Tokuda, Heiga Zen, Tadashi Kitamura |
INTERSPEECH | 1 |
| 2003 | Decision tree-based simultaneous clustering of phonetic contexts, dimensions, and state positions for acoustic modelingabstractIn this paper, a new decision tree-based clustering technique called Phonetic, Dimensional and State Positional Decision Tree (PDS-DT) is proposed. In PDS-DT, phonetic contexts, dimensions and state positions are grouped simultaneously during decision tree construction. PDS-DT provides a complicate distribution sharing structure without any external control parameters. In speaker-independent continuous speech recognition experiments, PDS-DT achieved about 13%--15% error reduction over the phonetic decision tree-based state-tying technique. Heiga Zen, Keiichi Tokuda, Tadashi Kitamura |
INTERSPEECH | 2 |
| 2002 | Eigenvoices for HMM-based speech synthesis
Kengo Shichiri, Atsushi Sawabe, Takayoshi Yoshimura, Keiichi Tokuda, Takashi Masuko, Takao Kobayashi, Tadashi Kitamura |
INTERSPEECH | 4 |
| 2002 | A context clustering technique for average voice model in HMM-based speech synthesis
Junichi Yamagishi, Masatsune Tamura, Takashi Masuko, Keiichi Tokuda, Takao Kobayashi |
INTERSPEECH | 4 |
| 2002 | Decision tree distribution tying based on a dimensional split techniqueabstractSplit Phonetic Decision Tree (DS-PDT) is proposed. In DSPDT, state distributions are split dimensionally when applying phonetic question. This technique is an extension of the decision tree based acoustic modeling. It gives a proper context-dependent sharing structure of each dimension automatically while maintaining the correlations among the dimensions. In speaker-independent continuous speech recognition experiments, DS-PDT achieved about 8% error reduction over the phonetic decision tree clustering. Heiga Zen, Keiichi Tokuda, Tadashi Kitamura |
INTERSPEECH | 2 |
| 2001 | Speaker identification using Gaussian mixture models based on multi-space probability distributionabstractPresents an approach to modeling speech spectra and pitch for text-independent speaker identification using Gaussian mixture models based on multi-space probability distribution (MSD-GMM). The MSD-GMM allows us to model continuous pitch values for voiced frames and discrete symbols representing unvoiced frames in a unified framework. Spectral and pitch features are jointly modeled by a two-stream MSD-GMM. We derive maximum likelihood estimation formulae for the MSD-GMM parameters, and the MSD-GMM speaker models are evaluated for text-independent speaker identification tasks. Experimental results, show that the MSD-GMM can efficiently model spectral and pitch features of each speaker and outperforms conventional speaker models. Chiyomi Miyajima, Yosuke Hattori, Keiichi Tokuda, Takashi Masuko, Takao Kobayashi, Tadashi Kitamura |
ICASSP | 3 |
| 2001 | Adaptation of pitch and spectrum for HMM-based speech synthesis using MLLRabstractDescribes a technique for synthesizing speech with arbitrary speaker characteristics using speaker independent speech units, which we call "average voice" units. The technique is based on an HMM-based text-to-speech (TTS) system and maximum likelihood linear regression (MLLR) adaptation algorithm. In the HMM-based TTS system, speech synthesis units are modeled by multi-space probability distribution (MSD) HMMs which can model spectrum and pitch simultaneously in a unified framework. We derive an extension of the MLLR algorithm to apply it to MSD-HMMs. We demonstrate that a few sentences uttered by a target speaker are sufficient to adapt not only voice characteristics but also prosodic features. Synthetic speech generated from adapted models using only four sentences is very close to that from speaker dependent models trained using 450 sentences. Masatsune Tamura, Takashi Masuko, Keiichi Tokuda, Takao Kobayashi |
ICASSP | 3 |
| 2001 | Minimum classification error training for speaker identification using Gaussian mixture models based on multi-space probability distributionabstractIn our previous work, we have proposed a speaker modeling technique using spectral and pitch features for text-independent speaker identification based on Multi-Space Probability Distribution Gaussian Mixture Models (MSD-GMMs). We have presented a maximum likelihood (ML) estimation procedure for the MSD-GMM parameters and demonstrated its high recognition performance. In this paper, we describe an minimum classification error (MCE) training procedure for the MSDGMM speaker models. MCE training is also applied to automatically estimate mixture-dependent stream weights for spectral and pitch streams. The MCE-based MSD-GMM speaker models are evaluated for a text-independent speaker identification task. Experimental results show that MCE training of the MSD-GMM parameters significantly reduces identification errors and system performance is further improved by appropriately weighting spectral and pitch streams using MCE training. Chiyomi Miyajima, Keiichi Tokuda, Tadashi Kitamura |
INTERSPEECH | 2 |
| 2001 | A robust speaker verification system against imposture using an HMM-based speech synthesis system
Takayuki Satoh, Takashi Masuko, Takao Kobayashi, Keiichi Tokuda |
INTERSPEECH | 4 |
| 2001 | Text-to-speech synthesis with arbitrary speaker's voice from average voiceabstractThis paper describes a technique for synthesizing speech with any desired voice. The technique is based on an HMM-based text-to-speech (TTS) system and MLLR adaptation algorithm. To generate speech of an arbitrarily given target speaker, speaker-independent speech units, i.e., average voice models, is adapted to the target speaker using MLLR framework. In addition to spectrum and pitch adaptation, we derive an algorithm for adaptation of state duration. We demonstrate that a few sentences uttered by a target speaker are sufficient to adapt not only voice characteristics but also prosodic features. Synthetic speech generated from adapted models using only four sentences is very close to that from speaker dependent models trained using a large amount of speech data. Masatsune Tamura, Takashi Masuko, Keiichi Tokuda, Takao Kobayashi |
INTERSPEECH | 3 |
| 2001 | Mixed excitation for HMM-based speech synthesisabstractThis paper describes improvements on the excitation model of an HMM-based text-to-speech system. In our previous work, natural sounding speech can be synthesized from trained HMMs. However, it has a typical quality of “vocoded speech” since the system uses a traditional excitation model with either a periodic impulse train or white noise. In this paper, in order to reduce the synthetic quality, a mixed excitation model used in MELP is incorporated into the system. Excitation parameters used in mixed excitation are modeled by HMMs, and generated from HMMs by a parameter generation algorithm in the synthesis phase. The result of a listening test shows that the mixed excitation model significantly improves quality of synthesized speech as compared with the traditional excitation model. Takayoshi Yoshimura, Keiichi Tokuda, Takashi Masuko, Takao Kobayashi, Tadashi Kitamura |
INTERSPEECH | 2 |
| 2001 | A new approach to designing a feature extractor in speaker identification based on discriminative feature extraction
Chiyomi Miyajima, Hideyuki Watanabe, Keiichi Tokuda, Tadashi Kitamura, Shigeru Katagiri |
Speech Commun. | 3 |
| 2000 | Speech parameter generation algorithms for HMM-based speech synthesisabstractThis paper derives a speech parameter generation algorithm for HMM-based speech synthesis, in which the speech parameter sequence is generated from HMMs whose observation vector consists of a spectral parameter vector and its dynamic feature vectors. In the algorithm, we assume that the state sequence (state and mixture sequence for the multi-mixture case) or a part of the state sequence is unobservable (i.e., hidden or latent). As a result, the algorithm iterates the forward-backward algorithm and the parameter generation algorithm for the case where the state sequence is given. Experimental results show that by using the algorithm, we can reproduce clear formant structure from multi-mixture HMMs as compared with that produced from single-mixture HMMs. Keiichi Tokuda, Takayoshi Yoshimura, Takashi Masuko, Takao Kobayashi, Tadashi Kitamura |
ICASSP | 1 |
| 2000 | Normalized Training for HMM-Based Visual Speech RecognitionabstractThis paper presents an approach to estimating the parameters of continuous density HMMs for visual speech recognition. One of the key issues of image-based visual speech recognition is normalization of lip location and lighting conditions prior to estimating the parameters of HMMs. We presented a normalized training method in which the normalization process is integrated in the model training. This paper extends it for contrast normalization in addition to average-intensity and location normalization. The proposed method provides a theoretically-well-defined algorithm based on a maximum likelihood formulation, hence the likelihood for the training data is guaranteed to increase at each iteration of the normalized training. Experiments on the M2VTS database show that the recognition performance can be significantly improved by the normalized training. Yoshihiko Nankaku, Keiichi Tokuda, Tadashi Kitamura, Takao Kobayashi |
ICIP | 2 |
| 2000 | Imposture using synthetic speech against speaker verification based on spectrum and pitchabstractThis paper describes security of speaker verification systems against imposture using synthetic speech. We propose a text-prompted speaker verification technique which utilizes pitch information in addition to spectral information, and investigate whether synthetic speech is rejected. Experimental results show that pitch information is not necessarily useful for rejection of synthetic speech, and it is required to develop techniques to discriminate synthetic speech from natural speech. Takashi Masuko, Keiichi Tokuda, Takao Kobayashi |
INTERSPEECH | 2 |
| 2000 | Audio-visual speech recognition using MCE-based hmms and model-dependent stream weights
Chiyomi Miyajima, Keiichi Tokuda, Tadashi Kitamura |
INTERSPEECH | 2 |
| 2000 | HMM-based text-to-audio-visual speech synthesis
Shinji Sako, Keiichi Tokuda, Takashi Masuko, Takao Kobayashi, Tadashi Kitamura |
INTERSPEECH | 2 |
| 2000 | RLS-type two-dimensional adaptive filter with a t-distribution assumption
Junibakti Sanubari, Keiichi Tokuda |
Signal Process. | 2 |
| 1999 | Hidden Markov models based on multi-space probability distribution for pitch pattern modelingabstractThis paper discusses a hidden Markov model (HMM) based on multi-space probability distribution (MSD). The HMMs are widely-used statistical models to characterize the sequence of speech spectra and have successfully been applied to speech recognition systems. From these facts, it is considered that the HMM is useful for modeling pitch patterns of speech. However, we cannot apply the conventional discrete or continuous HMMs to pitch pattern modeling since the observation sequence of the pitch pattern is composed of one-dimensional continuous values and a discrete symbol which represents "unvoiced". MSD-HMM includes discrete HMMs and continuous mixture HMMs as special cases, and further can model the sequence of observation vectors with variable dimension including zero-dimensional observations, i.e., discrete symbols. As a result, MSD-HMMs can model pitch patterns without heuristic assumption. We derive a reestimation algorithm for the extended HMM and show that it can find a critical point of the likelihood function. Keiichi Tokuda, Takashi Masuko, Noboru Miyazaki, Takao Kobayashi |
ICASSP | 1 |
| 1999 | Image Modeling Using Two Dimensional Exponential SystemsabstractIn this paper a new image modeling is proposed. We propose the usage of the exponential (EXP) model as an alternative for the existing AR models which have stability problem. Since the EXP model is always stable, the obtained model can always be utilized to synthesize the original image. Since the EXP systems have an infinite impulse response, it is suitable to model images which contain not only poles, but also images with poles and zeroes. Therefore, it is expected that the EXP system is more suitable to model more variations of images than that of AR system. The simulation results show the the power gain (PG) of proposed EXP and the conventional AR model are comparable. Junibakti Sanubari, Keiichi Tokuda |
ICIP (4) | 2 |
| 1999 | Location Normalization of HMM-Based Lip Reading: Experiments for the M2VTS DatabaseabstractThis paper describes an HMM-based lip location normalization process, in order to improve the recognition performance in automatic lip-reading. This paper uses the image-based method in order to represent the lip visual information. One of the most critical factors which affect the recognition results in image-based method is the position of lips in frames. This paper describes a method to normalize the lip location which is similar to SAT (speaker adaptive training), and presents several experiments which were carried out in order to measure the effectiveness of the proposed method. Experiments of isolated words with and without the original movement from speakers were carried out on the M2VTS database. Oscar Vanegas, Keiichi Tokuda, Tadashi Kitamura |
ICIP (2) | 2 |
| 1999 | On the security of HMM-based speaker verification systems against imposture using synthetic speech
Takashi Masuko, Takafumi Hitotsumatsu, Keiichi Tokuda, Takao Kobayashi |
EUROSPEECH | 3 |
| 1999 | Intensity- and location-normalized training for HMM-based visual speech recognition
Yoshihiko Nankaku, Keiichi Tokuda, Tadashi Kitamura |
EUROSPEECH | 2 |
| 1999 | Simultaneous modeling of spectrum, pitch and duration in HMM-based speech synthesis
Takayoshi Yoshimura, Keiichi Tokuda, Takashi Masuko, Takao Kobayashi, Tadashi Kitamura |
EUROSPEECH | 2 |
| 1998 | A wideband CELP speech coder at 16 kbit/s based on mel-generalized cepstral analysisabstractThis paper proposes a wideband CELP coder using frequency warping. Instead of linear prediction, the proposed coder adopts the mel-generalized cepstral analysis, and encodes the fullband of the speech signal through a warped frequency scale. It is shown that the subjective quality of the proposed coder at 16 kbit/s is better than that of the ITU-T G.722 at 64 kbit/s. Furthermore, the proposed coder gives a much smaller difference in performance for male and female speakers than the conventional CELP coder. These results indicate that the frequency warping makes a large contribution to the improvement of the subjective quality for wideband speech coding. Kazuhito Koishida, Gou Hirabayashi, Keiichi Tokuda, Takao Kobayashi |
ICASSP | 3 |
| 1998 | Text-to-visual speech synthesis based on parameter generation from HMMabstractThis paper presents a new technique for synthesizing visual speech from arbitrarily given text. The technique is based on an algorithm for parameter generation from HMM with dynamic features, which has been successfully applied to text-to-speech synthesis. In the training phase, syllable HMMs are trained with visual speech parameter sequences that represent lip movements. In the synthesis phase, a sentence HMM is constructed by concatenating syllable HMMs corresponding to the phonetic transcription for the input text. Then an optimum visual speech parameter sequence is generated from the sentence HMM in an ML sense. The proposed technique can generate synchronized lip movements with speech in a unified framework. Furthermore, coarticulation is implicitly incorporated into the generated mouth shapes. As a result, synthetic lip motion becomes smooth and realistic. Takashi Masuko, Takao Kobayashi, Masatsune Tamura, Jun Masubuchi, Keiichi Tokuda |
ICASSP | 5 |
| 1998 | A very low bit rate speech coder using HMM-based speech recognition/synthesis techniquesabstractThis paper presents a very low bit rate speech coder based on HMM (hidden Markov model). The encoder carries out phoneme recognition, and transmits phoneme indexes, state durations and pitch information to the decoder. In the decoder, phoneme HMMs are concatenated according to the phoneme indexes, and a sequence of mel-cepstral coefficient vectors is generated from the concatenated HMM by using an ML-based speech parameter generation technique. Finally we obtain synthetic speech by exciting the MLSA (mel log spectrum approximation) filter, whose coefficients are given by mel-cepstral coefficients, according to the pitch information. A subjective listening test shows that the performance of the proposed coder at about 150 bit/s (for the test data including 26% silence region) is comparable to a VQ-based vocoder at 400 bit/s (=8 bit/frame/spl times/50 frame/s) without pitch quantization for both coders. Keiichi Tokuda, Takashi Masuko, Jun Hiroi, Takao Kobayashi, Tadashi Kitamura |
ICASSP | 1 |
| 1998 | A 16 kbit/s wideband CELP coder using MEL-generalized cepstral analysis and its subjective evaluationabstractWe have proposed a wideband CELP coder, called MGC-CELP, which provides high quality speech by utilizing mel-generalized cepstral (MGC) analysis instead of linear prediction (LP). In this paper, we investigate the performance of the wideband MGCCELP coder at 16 kbit/s in terms of short-term predictor order, i.e., order of MGC analysis. Subjective tests show that the MGCCELP coder with a predictor of order 20 gives better performance than ITU-T G.722 at 64 kbit/s. It is also found that the MGCCELP coder with 12th order achieves comparable quality to the 64 kbit/s G.722, and outperforms the 16 kbit/s conventional CELP coder using 20th-order LP analysis under the same conditions. 1. INTRODUCTION Recently several schemes for high-quality wideband speech coding at low bit rates have been developed. Most of the work in this field uses either transform/subband coding or CELP (Code Excited Linear Prediction) coding. At the bit rates around 16 kbit/s, CELP coding has received much attention s... Kazuhito Koishida, Gou Hirabayashi, Keiichi Tokuda, Takao Kobayashi |
ICSLP | 3 |
| 1998 | A very low bit rate speech coder using HMM with speaker adaptationabstractThis paper describes a speaker adaptation technique for a phonetic vocoder based on HMM. In the vocoder, the encoder performs phoneme recognition and transmits phoneme indexes and state durations to the decoder, and the decoder synthesizes speech using HMM-based speech synthesis technique. One of the main problems of this vocoder is that the voice characteristics of synthetic speech depend on HMMs used in the decoder, and are therefore fixed regardless of a variety of input speakers. To overcome this problem, we adapt HMMs to input speech by transmitting transfer vectors, information on mismatch between the input speech and HMMs. The results of the subjective tests show that the performance of the proposed vocoder without quantization of transfer vectors is comparable to that of a speaker dependent vocoder. 1. INTRODUCTION To code speech at rates on the order of 100 bit/s, phonetic or segment vocoders are the most popular techniques [1]-[6]. These coders decompose speech into a seque... Takashi Masuko, Keiichi Tokuda, Takao Kobayashi |
ICSLP | 2 |
| 1998 | HMM-based visual speech recognition using intensity and location normalizationabstractThis paper describes intensity and location normalization techniques for improving the performance of visual speech recognizers used in audio-visual speech recognition. For auditory speech recognition, there exist many methods for dealing with channel characteristics and speaker individualities, e.g., CMN (cepstral mean normalization), SAT (speaker adaptive training). We present two techniques similar to CMN and SAT, respectively, for intensity and location normalization in visual speech recognition. Word recognition experiments based on HMM show that a significant improvement in recogniton performance is achieved by combining the two techniques. Oscar Vanegas, Akiji Tanaka, Keiichi Tokuda, Tadashi Kitamura |
ICSLP | 3 |
| 1998 | Duration modeling for HMM-based speech synthesis
Takayoshi Yoshimura, Keiichi Tokuda, Takashi Masuko, Takao Kobayashi, Tadashi Kitamura |
ICSLP | 2 |
| 1997 | Efficient encoding of mel-generalized cepstrum for CELP codersabstractThe performance of several algorithms for the quantization of the mel-generalized cepstral coefficients is studied. First, the objective and subjective performance of two-stage vector quantization (VQ) is measured. It is shown that the subjective quality for the mel-generalized cepstral coefficients is higher than that for LSP. Secondly, interframe prediction is introduced in the encoding of mel-generalized cepstral coefficients. By utilizing interframe moving average (MA) prediction, the mel-generalized cepstral coefficients can be encoded more efficiently than LSP in terms of cepstral distortion. Finally, we implement a CELP coder based on mel-generalized cepstral analysis in which mel-generalized cepstral coefficients are quantized using MA prediction. This coder has a higher objective quality than conventional CELP. Kazuhito Koishida, Keiichi Tokuda, Takao Kobayashi, Satoshi Imai |
ICASSP | 2 |
| 1997 | Voice characteristics conversion for HMM-based speech synthesis systemabstractWe describe an approach to voice characteristics conversion for an HMM-based text-to-speech synthesis system. Since this speech synthesis system uses phoneme HMMs as speech units, voice characteristics conversion is achieved by changing the HMM parameters appropriately. To transform the voice characteristics of synthesized speech to the target speaker, we applied the maximum a posteriori estimation and vector field smoothing (MAP/VFS) algorithm to the phoneme HMMs. Using 5 or 8 sentences as adaptation data, speech samples synthesized from a set of adapted tied triphone HMMs, which have approximately 2,000 distributions, are judged to be closer to the target speaker by 79.7% or 90.6%, respectively, in an ABX listening test. Takashi Masuko, Keiichi Tokuda, Takao Kobayashi, Satoshi Imai |
ICASSP | 2 |
| 1997 | HMM compensation for noisy speech recognition based on cepstral parameter generation
Takao Kobayashi, Takashi Masuko, Keiichi Tokuda |
EUROSPEECH | 3 |
| 1997 | Speaker interpolation in HMM-based speech synthesis system
Takayoshi Yoshimura, Takashi Masuko, Keiichi Tokuda, Takao Kobayashi, Tadashi Kitamura |
EUROSPEECH | 3 |
| 1996 | Speech synthesis using HMMs with dynamic featuresabstractThis paper presents a new text-to-speech synthesis system based on HMM which includes dynamic features, i.e., delta and delta-delta parameters of speech. The system uses triphone HMMs as the synthesis units. The triphone HMMs share less than 2,000 clustered states, each of which is modelled by a single Gaussian distribution. For a given text to be synthesized, a sentence HMM is constructed by concatenating the triphone HMMs. Speech parameters are generated from the sentence HMM in such a way that the output probability is maximized. The speech signal is synthesized directly from the obtained parameters using the mel log spectral approximation (MLSA) filter. Without dynamic features, the discontinuity of the generated speech spectra causes glitches in the synthesized speech. On the other hand, with dynamic features, the synthesized speech becomes quite smooth and natural even if the number of clustered states is small. Takashi Masuko, Keiichi Tokuda, Takao Kobayashi, Satoshi Imai |
ICASSP | 2 |
| 1996 | Robust two dimensional spectral estimation based on AR model excited by a t-distribution processabstractA new robust two dimensional (2-D) spectral estimation method based on an AR model is proposed. The optimal coefficient is selected by assuming that the excitation signal is a t-distribution t(/spl alpha/) with /spl alpha/ degrees of freedom. When /spl alpha/=/spl infin/, we get the conventional least square (L/sub 2/) method. Thus, the proposed method can be regarded as a generalization of the L/sub 2/ method. Simulation results show that the obtained estimates using the proposed method with small /spl alpha/ are more efficient, the standard deviation (SD) of the estimation results are smaller, and more accurate than that with large /spl alpha/. The proposed estimator with small /spl alpha/ is more efficient and more accurate than the recursive method based on Huber's (1981) M-estimate. Junibakti Sanubari, Keiichi Tokuda, Mahoki Onoda |
ICASSP | 2 |
| 1996 | CELP coding system based on mel-generalized cepstral analysis
Kazuhito Koishida, Keiichi Tokuda, Takao Kobayashi, Satoshi Imai |
ICSLP | 2 |
| 1995 | CELP coding based on mel-cepstral analysisabstractWe propose a CELP coder based on mel-cepstral analysis. In the coder, since the transfer functions of perceptual weighting and postfiltering are defined through mel-cepstral coefficients, the effects of perceptual weighting and postfiltering should fit with the characteristics of the human auditory sensation. We use a basic CELP structure without adaptive codebook, and the subjective speech quality of the proposed coder in terms of the opinion equivalent Q is measured and compared with that of the conventional CELP coder. It is shown that the improvement of more than 1.8 dB is achieved by the proposed coder over the conventional CELP coder. Kazuhito Koishida, Keiichi Tokuda, Takao Kobayashi, Satoshi Imai |
ICASSP | 2 |
| 1995 | Speech parameter generation from HMM using dynamic featuresabstractThis paper proposes an algorithm for speech parameter generation from HMMs which include the dynamic features. The performance of speech recognition based on HMMs has been improved by introducing the dynamic features of speech. Thus we surmise that, if there is a method for speech parameter generation from HMMs which include the dynamic features, it will be useful for speech synthesis by rule. It is shown that the parameter generation from HMMs using the dynamic features results in searching for the optimum state sequence and solving a set of linear equations for each possible state sequence. We derive a fast algorithm for the solution by the analogy of the RLS algorithm for adaptive filtering. We also show the effect of incorporating the dynamic features by an example of speech parameter generation. Keiichi Tokuda, Takao Kobayashi, Satoshi Imai |
ICASSP | 1 |
| 1995 | An algorithm for speech parameter generation from continuous mixture HMMs with dynamic featuresabstractThis paper proposes an algorithm for speech parameter generation from continuous mixture HMMs which include dynamic features, i.e., delta and delta-delta parameters of speech. We showthatthe parameter generation from HMMs using the dynamic features results in searching for the optimal state sequence and solving a set of linear equations for each possible state sequence. Tosolve the problem, we derivea fast algorithm on the analogy of the RLS algorithm for adaptive #ltering. We show that the generated speech parameter vectors re#ect not only the means of static and dynamic feature vectors but also the covariances of those. An example presenting e#ectiveness of the proposed algorithm in speech synthesis is given. 1. INTRODUCTION The hidden Markov models #HMMs# can model sequences of speech spectra with well-de#ned algorithms, and have successfully been applied to speech recognition systems. From these facts, we surmise that HMMs are also useful for speech synthesis. Actually, some at... Keiichi Tokuda, Takashi Masuko, Tetsuya Yamada, Takao Kobayashi, Satoshi Imai |
EUROSPEECH | 1 |
| 1995 | Adaptive cepstral analysis of speechabstractThis paper proposes an algorithm for adaptive cepstral analysis based on the UELS (unbiased estimation of log spectrum). In the UELS, the model spectrum is represented by cepstral coefficients and the mean square of the inverse filter output is minimized with respect to the cepstral coefficients. By introducing an instantaneous gradient estimate of the criterion in a similar manner of the LMS algorithm, we develop an adaptive cepstral analysis algorithm. In the analysis system, an IIR adaptive filter whose coefficients are given by cepstral coefficients is realized using the log magnitude approximation (LMA) filter. The filter approximates an exponential transfer function and its stability is guaranteed for approximation of speech spectra. To implement the M th order cepstral analysis, the algorithm requires O(M) operations per sample. It is shown that the algorithm has fast convergence properties in comparison with the LMS algorithm. Several examples of the adaptive cepstral analysis for synthetic signal and natural speech are shown to demonstrate the effectiveness of the algorithm. Keiichi Tokuda, Takao Kobayashi, Satoshi Imai |
IEEE Trans. Speech Audio Process. | 1 |
| 1994 | Robust recursive spectral estimation based on AR model excited by a t-distribution processabstractIn this paper a new robust spectral estimation method based on an AR model is proposed. The optimal coefficient is selected by assuming that the excitation signal is t-distribution t(/spl alpha/) with /spl alpha/ degrees of freedom. The calculation is done by using a recursive algorithm. When /spl alpha/=/spl infin/, we get the RLS method. Simulation results show that the obtained estimates using the proposed method with small /spl alpha/ are more efficient, the standard deviation (SD) of the estimation results are smaller, and more accurate than that with large /spl alpha/. The proposed estimator with small /spl alpha/ is more efficient and more accurate then the recursive method based on Huber's M-estimate.> Junibakti Sanubari, Keiichi Tokuda, Mahoki Onoda |
ICASSP (3) | 2 |
| 1994 | Speech coding based on adaptive mel-cepstral analysisabstractWe propose an ADPCM coder which uses a backward adaptive predictor based on the adaptive mel-cepstral analysis. The spectrum represented by the mel-cepstral coefficients has frequency resolution similar to that of the human ear which has high resolution at low frequencies. In the coder, since the transfer functions of noise shaping and postfiltering are also defined through the mel-cepstral coefficients, the effects of nose shaping and postfiltering should fit with characteristics of the human auditory sensation. We incorporate a pitch predictor into the ADPCM coder, and evaluate the speech quality based on objective and subjective performance tests. It is shown that the coder at 16 kb/s can produce a high quality speech comparable with that of the CCITT G.721 ADPCM coder at 32 kb/s with no algorithmic delay.> Keiichi Tokuda, Hidetoshi Matsumura, Takao Kobayashi, Satoshi Imai |
ICASSP (1) | 1 |
| 1994 | Speech coding based on adaptive MEL-cepstral analysis for noisy channels
Kazuhito Koishida, Keiichi Tokuda, Takao Kobayashi, Satoshi Imai |
ICSLP | 2 |
| 1994 | Mel-generalized cepstral analysis - a unified approach to speech spectral estimation
Keiichi Tokuda, Takao Kobayashi, Takashi Masuko, Satoshi Imai |
ICSLP | 1 |
| 1994 | AR Spectrum Estimation Based on Wavelet RepresentationabstractA new adaptive AR spectrum estimation method is proposed. The cost function is defined by using the discrete-time wavelet transform coefficients of the linear prediction error. Instead of a single window throughout the whole frequency spectrum, a wavelet-like windowing method is used to increase the frequency resolution of the low-frequency components and to improve the time resolution of the high-frequency components. Special properties of the covariance matrix are used to derive an RLS algorithm which requires O(M/sup 2/) operations. Simulation results show that the wavelet based spectrum estimation method gives fine frequency resolution at low frequencies and good time resolution at high frequencies, while with conventional methods it is possible to have only one of these characteristics.> Fernando Gil Vianna Resende Jr., Keiichi Tokuda, Mineo Kaneko |
ISCAS | 2 |
| 1992 | An adaptive algorithm for mel-cepstral analysis of speechabstractThe authors describe a mel-cepstral analysis method and its adaptive algorithm. In the proposed method, the authors apply the criterion used in the unbiased estimation of log spectrum to the spectral model represented by the mel-cepstral coefficients. To solve the nonlinear minimization problem involved in the method, they give an iterative algorithm whose convergence is guaranteed. Furthermore, they derive an adaptive algorithm for the mel-cepstral analysis by introducing an instantaneous estimate for gradient of the criterion. The adaptive mel-cepstral analysis system is implemented with an IIR adaptive filter which has an exponential transfer function, and whose stability is guaranteed. The authors also present examples of speech analysis and results of an isolated word recognition experiment.> Toshiaki Fukada, Keiichi Tokuda, Takao Kobayashi, Satoshi Imai |
ICASSP | 2 |
| 1992 | Design of stable two-dimensional IIR digital filters with arbitrary magnitude functionabstractA technique for designing two-dimensional (2-D) digital filters which can approximate arbitrary magnitude functions is proposed. A 2-D spectral factorization technique is used to obtain the recursively computable and stable system with nonsymmetric half-plane support from a given 2-D magnitude function. A new class of realizable 2-D digital filters referred to as 2-D log magnitude approximation (2-D LMA) filters are used to approximate the system obtained by the 2-D factorization. The design procedure is straightforward and computationally efficient. A simple stability condition which guarantees the stability of the designed 2-D LMA filter is given. An efficient network structure of 2-D LMA filters for parallel implementation is also discussed.> Takao Kobayashi, Kazuyoshi Fukushi, Keiichi Tokuda, Satoshi Imai |
ICASSP | 3 |
| 1992 | Spectral estimation based on AR-model excited by t-distribution processabstractA new spectral estimation method is proposed. Since in the least square L/sub 2/ method the obtained estimates are very much affected by the large signal portions, in the proposed method a loss function which assigns large weighting factor for the small residual portions and vice versa is used. The loss function is based on an assumption that the residual signal has an identical and independent t-distribution t( alpha ) with alpha degrees of freedom to achieve accurate and efficient (low standard deviation) estimates. When alpha = infinity , the conventional L/sub 2/ method is obtained. In the calculation, the loss function is modified in a way similar to the autocorrelation method, so that the proposed method can be seen as a generalization of the autocorrelation method. The optimal solution is selected by the Newton-Raphson method. The simulation results show that only a few iterations are needed to reach a stationary point, the stationary point is always a local minimum, and the obtained predictor is stable.> Junibakti Sanubari, Keiichi Tokuda, Mahoki Onoda |
ICASSP | 2 |
| 1990 | Adaptive filtering based on cepstral representation-adaptive cepstral analysis of speechabstractAn adaptive cepstral analysis method based on an unbiased estimation of the log spectrum is proposed. In the method, an infinite impulse response adaptive filter whose coefficients are given by cepstral coefficients is realized using the log magnitude approximation (LMA) filter. To implement the Mth-order cepstral analysis, the algorithm requires O(M) operations per sample. It is shown that the algorithm has fast convergence properties in comparison with the least-mean-square algorithm. A real-time analysis system is implemented with a general-purpose digital signal processor, and an example of natural speech analysis is shown to demonstrate the convergence.> Keiichi Tokuda, Takao Kobayashi, Shoji Shiomoto, Satoshi Imai |
ICASSP | 1 |
| 1990 | Generalized cepstral analysis of speech - unified approach to LPC and cepstral method
Keiichi Tokuda, Takao Kobayashi, Satoshi Imai |
ICSLP | 1 |