VLDB 2026 Research / reviewers in the wild / expert
Yoshihiko Nankaku
dblp:90/302
· DBLP profile ↗
91ranked-venue papers
4as first author
9since 2021 · last 2025
0009-0000-3978-5130ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 83 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 46 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | PeriodCodec: A Pitch-Controllable Neural Audio Codec Using Periodic Signals for Singing Voice Synthesis
Masato Takagi, Miku Nishihara, Yukiya Hono, Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda |
INTERSPEECH | 5 |
| 2025 | Enhancing Social Presence in Dyadic Text-Chatting with a Robot Avatar Expressing Users' ActionsabstractText-based Computer-Mediated Communication (CMC) diminishes the social presence of the conversational partner owing to the lack of nonverbal cues exchanged in face-to-face (FtF) communication. This study aims to enhance social presence in a dyadic text chat by employing a robot avatar that supplements nonverbal cues. In particular, we propose a chat system with a robot avatar that can express gestures in response to the user’s actions (typing, sending messages) and read the user’s messages aloud. In the evaluation experiment, nine pairs of participants (for a total of 18 participants) used two types of chat system: the proposed system and the baseline system in which a robot avatar only reads messages aloud. As a result, the proposed robot avatar significantly increased social presence compared to the baseline system. Furthermore, the proposed gesture expressions significantly improved the ease of chatting. Specifically, our results suggest that robot gestures in response to typing might have an effect similar to nonverbal cues in FtF communication. Yasutaka Nakamura, Seiichi Harata, Takuto Sakuma, Yoshihiro Tanaka, Yoshihiko Nankaku, Shohei Kato |
RO-MAN | 5 |
| 2024 | PeriodGrad: Towards Pitch-Controllable Neural Vocoder Based on a Diffusion Probabilistic ModelabstractThis paper presents a neural vocoder based on a denoising diffusion probabilistic model (DDPM) incorporating explicit periodic signals as auxiliary conditioning signals. Recently, DDPM-based neural vocoders have gained prominence as non-autoregressive models that can generate high-quality waveforms. The neural vocoders based on DDPM have the advantage of training with a simple time-domain loss. In practical applications, such as singing voice synthesis, there is a demand for neural vocoders to generate high-fidelity speech waveforms with flexible pitch control. However, conventional DDPM-based neural vocoders struggle to generate speech waveforms under such conditions. Our proposed model aims to accurately capture the periodic structure of speech waveforms by incorporating explicit periodic signals. Experimental results show that our model improves sound quality and provides better pitch control than conventional DDPM-based neural vocoders. Yukiya Hono, Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda |
ICASSP | 3 |
| 2023 | Singing Voice Synthesis Based on a Musical Note Position-Aware Attention MechanismabstractThis paper proposes a novel sequence-to-sequence (seq2seq) model with a musical note position-aware attention mechanism for singing voice synthesis (SVS). A seq2seq modeling approach that can simultaneously perform acoustic and temporal modeling is attractive. However, due to the difficulty of the temporal modeling of singing voices, many recent SVS systems with an encoder-decoder-based model still rely on explicitly on duration information generated by additional modules. Although some studies perform simultaneous modeling using seq2seq models with an attention mechanism, they have insufficient robustness against temporal modeling. The proposed attention mechanism is designed to estimate the attention weights by considering the rhythm given by the musical score. Furthermore, several techniques are also introduced to improve the modeling performance of the singing voice. Experimental results indicated that the proposed model is effective in terms of both naturalness and robustness of timing. Yukiya Hono, Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda |
ICASSP | 3 |
| 2023 | Embedding a Differentiable Mel-Cepstral Synthesis Filter to a Neural Speech Synthesis SystemabstractThis paper integrates a classic mel-cepstral synthesis filter into a modern neural speech synthesis system towards end-to-end controllable speech synthesis. Since the mel-cepstral synthesis filter is explicitly embedded in neural waveform models in the proposed system, both voice characteristics and the pitch of synthesized speech are highly controlled via a frequency warping parameter and fundamental frequency, respectively. We implement the mel-cepstral synthesis filter as a differentiable and GPU-friendly module to enable the acoustic and waveform models in the proposed system to be simultaneously optimized in an end-to-end manner. Experiments show that the proposed system improves speech quality from a baseline system maintaining controllability. The core PyTorch modules used in the experiments are publicly available on GitHub1. Takenori Yoshimura, Shinji Takaki, Kazuhiro Nakamura, Keiichiro Oura, Yukiya Hono, Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda |
ICASSP | 7 |
| 2022 | Autoregressive Variational Autoencoder with a Hidden Semi-Markov Model-Based Structured Attention for Speech SynthesisabstractThis paper proposes an autoregressive speech synthesis model based on the variational autoencoder incorporating latent sequence representation for acoustic and linguistic features and the structure of a hidden semi-Markov model (HSMM). Although autoregressive models can provide efficient and accurate modeling of acoustic features, they have exposure bias, i.e., the mismatch between training (teacher-forcing) and inference (free-running). To overcome this problem, we introduce an autoregressive latent variable sequence, rather than using autoregressive generation of observations. Latent representation of alignment using HSMM-based structured attention mechanism enables the use of a completely consistent training algorithm for acoustic modeling with explicit duration models. Experimental results indicate that the proposed model outperformed baselines in subjective naturalness. Takato Fujimoto, Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda |
ICASSP | 3 |
| 2022 | End-to-End Text-to-Speech Based on Latent Representation of Speaking Styles Using Spontaneous DialogueabstractThe recent text-to-speech (TTS) has achieved quality comparable to that of humans; however, its application in spoken dialogue has not been widely studied. This study aims to realize a TTS that closely resembles human dialogue. First, we record and transcribe actual spontaneous dialogues. Then, the proposed dialogue TTS is trained in two stages: first stage, variational autoencoder (VAE)-VITS or Gaussian mixture variational autoencoder (GMVAE)-VITS is trained, which introduces an utterance-level latent variable into variational inference with adversarial learning for end-to-end text-to-speech (VITS), a recently proposed end-to-end TTS model. A style encoder that extracts a latent speaking style representation from speech is trained jointly with TTS. In the second stage, a style predictor is trained to predict the speaking style to be synthesized from dialogue history. During inference, by passing the speaking style representation predicted by the style predictor to VAE/GMVAE-VITS, speech can be synthesized in a style appropriate to the context of the dialogue. Subjective evaluation results demonstrate that the proposed method outperforms the original VITS in terms of dialogue-level naturalness. Kentaro Mitsui, Tianyu Zhao 0001, Kei Sawada, Yukiya Hono, Yoshihiko Nankaku, Keiichi Tokuda |
INTERSPEECH | 5 |
| 2021 | Periodnet: A Non-Autoregressive Waveform Generation Model with a Structure Separating Periodic and Aperiodic ComponentsabstractWe propose PeriodNet, a non-autoregressive (non-AR) waveform generation model with a new model structure for modeling periodic and aperiodic components in speech waveforms. The non-AR waveform generation models can generate speech waveforms parallelly and can be used as a speech vocoder by conditioning an acoustic feature. Since a speech waveform contains periodic and aperiodic components, both components should be appropriately modeled to generate a high-quality speech waveform. However, it is difficult to decompose the components from a natural speech waveform in advance. To address this issue, we propose a parallel model and a series model structure separating periodic and aperiodic components. The features of our proposed models are that explicit periodic and aperiodic signals are taken as input, and external periodic/aperiodic decomposition is not needed in training. Experiments using a singing voice corpus show that our proposed structure improves the naturalness of the generated waveform. We also show that the speech waveforms with a pitch outside of the training data range can be generated with more naturalness. Yukiya Hono, Shinji Takaki, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda |
ICASSP | 5 |
| 2021 | Sinsy: A Deep Neural Network-Based Singing Voice Synthesis SystemabstractThis paper presents Sinsy, a deep neural network (DNN)-based singing voice synthesis (SVS) system. In recent years, DNNs have been utilized in statistical parametric SVS systems, and DNN-based SVS systems have demonstrated better performance than conventional hidden Markov model-based ones. SVS systems are required to synthesize a singing voice with pitch and timing that strictly follow a given musical score. Additionally, singing expressions that are not described on the musical score, such as vibrato and timing fluctuations, should be reproduced. The proposed system is composed of four modules: a time-lag model, a duration model, an acoustic model, and a vocoder, and singing voices can be synthesized taking these characteristics of singing voices into account. To better model a singing voice, the proposed system incorporates improved approaches to modeling pitch and vibrato and better training criteria into the acoustic model. In addition, we incorporated PeriodNet, a non-autoregressive neural vocoder with robustness for the pitch, into our systems to generate a high-fidelity singing voice waveform. Moreover, we propose automatic pitch correction techniques for DNN-based SVS to synthesize singing voices with correct pitch even if the training data has out-of-tune phrases. Experimental results show our system can synthesize a singing voice with better timing, more natural vibrato, and correct pitch, and it can achieve better mean opinion scores in subjective evaluation tests. Yukiya Hono, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2020 | Semi-Supervised Learning Based on Hierarchical Generative Models for End-to-End Speech SynthesisabstractThis paper proposes a general framework of semi-supervised learning based on hierarchical generative models and adapts it to a Japanese end-to-end text-to-speech (TTS) system. In English TTS, several end-to-end systems have recently achieved sound quality close to that of natural human speech. However, in non-alphabetic languages such as Japanese, it is difficult to realize true text-input end-to-end TTS due to character diversity and pitch accents. To address this problem, we propose end-to-end TTS based on semi-supervised learning that makes the most of existing data consisting of any combination of text, phoneme, and waveform as training data. To demonstrate the effectiveness of the proposed system, listening tests were conducted for pronunciation and naturalness. Our results show that the proposed system improves both pronunciation and naturalness. Takato Fujimoto, Shinji Takaki, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda |
ICASSP | 5 |
| 2020 | Fast and High-Quality Singing Voice Synthesis System Based on Convolutional Neural NetworksabstractThe present paper describes singing voice synthesis based on convolutional neural networks (CNNs). Singing voice synthesis systems based on deep neural networks (DNNs) are currently being proposed and are improving the naturalness of synthesized singing voices. As singing voices represent a rich form of expression, a powerful technique to model them accurately is required. In the proposed technique, long-term dependencies of singing voices are modeled by CNNs. An acoustic feature sequence is generated for each segment that consists of long-term frames, and a natural trajectory is obtained without the parameter generation algorithm. Furthermore, a computational complexity reduction technique, which drives the DNNs in different time units depending on type of musical score features, is proposed. Experimental results show that the proposed method can synthesize natural sounding singing voices much faster than the conventional method. Kazuhiro Nakamura, Shinji Takaki, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda |
ICASSP | 5 |
| 2020 | Hierarchical Multi-Grained Generative Model for Expressive Speech SynthesisabstractThis paper proposes a hierarchical generative model with a multi-grained latent variable to synthesize expressive speech. In recent years, fine-grained latent variables are introduced into the text-to-speech synthesis that enable the fine control of the prosody and speaking styles of synthesized speech. However, the naturalness of speech degrades when these latent variables are obtained by sampling from the standard Gaussian prior. To solve this problem, we propose a novel framework for modeling the fine-grained latent variables, considering the dependence on an input text, a hierarchical linguistic structure, and a temporal structure of latent variables. This framework consists of a multi-grained variational autoencoder, a conditional prior, and a multi-level auto-regressive latent converter to obtain the different time-resolution latent variables and sample the finer-level latent variables from the coarser-level ones by taking into account the input text. Experimental results indicate an appropriate method of sampling fine-grained latent variables without the reference signal at the synthesis stage. Our proposed framework also provides the controllability of speaking style in an entire utterance. Yukiya Hono, Kazuna Tsuboi, Kei Sawada, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda |
INTERSPEECH | 6 |
| 2019 | Singing Voice Synthesis Based on Generative Adversarial NetworksabstractThis paper proposes a generative adversarial training method for deep neural network (DNN)-based singing voice synthesis. The DNN-based approach has been used in statistical parametric singing voice synthesis and improved the naturalness of the synthesized singing voice [1]. Recently, generative adversarial networks (GANs) [2] have attracted significant attention in various machine learning research areas including speech synthesis [3]. GANs have achieved great success in modeling the distributions of complex data, and they have the potential to alleviate over-smoothing problem on the generated speech parameters in speech synthesis. In this paper, we propose a DNN-based singing voice synthesis system incorporating the GAN. Experimental results show that the proposed method outperforms the conventional method in the naturalness of the synthesized singing voice. Yukiya Hono, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda |
ICASSP | 4 |
| 2019 | Speaker-dependent Wavenet-based Delay-free Adpcm Speech CodingabstractThis paper proposes a WaveNet-based delay-free adaptive differential pulse code modulation (ADPCM) speech coding system. The WaveNet generative model, which is a state-of-the-art model for neural-network-based speech waveform synthesis, is used as the adaptive predictor in ADPCM. To further improve speech quality, mel-cepstrum-based noise shaping and postfiltering were integrated with the proposed ADPCM system. Both objective and subjective evaluation results indicate that the proposed ADPCM system outperformed not only the conventional ADPCM system based on ITU-T Recommendation G.726 but also the ADPCM system based on adaptive mel-cepstral analysis. Takenori Yoshimura, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda |
ICASSP | 4 |
| 2018 | Image Recognition Based on Separable Lattice Hmms Using a Deep Neural Network for Output Probability DistributionsabstractThis paper proposes an image recognition method based on separable lattice hidden Markov models (SLHMMs) using a deep neural network (DNN) for output probability distributions. The geometric variations of the object to be recognized, e.g., size and location, are essential in image recognition. SLHMMs, which have been proposed to reduce the effect of geometric variations, can perform elastic matching both horizontally and vertically. Gaussian distributions are typical for modeling the output distribution of SLHMMs. However, these distributions may not be sufficient to represent patterns of image regions. Our method integrates SLHMMs and a DNN and can be used to model an image effectively by explicit modeling of the generative process based on SLHMMs and advanced feature classification based on a DNN. image recognition experiments showed that the proposed method improves recognition performance. Eiji Ichikawa, Kei Sawada, Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda |
ICASSP | 4 |
| 2018 | Statistical Voice Conversion Based on WavenetabstractThis paper proposes a voice conversion technique based on WaveNet to directly generate target audio waveforms from acoustic features of a source speaker. In voice conversion based on statistical models, the relation between acoustic features, such as spectral parameters, extracted from source and target audio waveforms is generally modeled using statistical models, such as Gaussian mixture models and neural networks. Although modeling the relation between acoustic features is reasonable and efficient, these models are not optimized for predicting target audio waveforms because the vocoder parameters are used as intermediate representations. To overcome this problem, we developed a voice conversion method to model the relation between target audio waveforms and acoustic features extracted from source audio waveforms using WaveNet, which is a generative model for audio waveforms. The proposed model can directly generate converted audio waveforms without vocoders. Experimental results indicate that the proposed method can generate a more naturally sounding converted speech than that using a conventional DNN method. Jumpei Niwa, Takenori Yoshimura, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda |
ICASSP | 5 |
| 2018 | WaveNet-Based Zero-Delay Lossless Speech CodingabstractThis paper presents a WaveNet-based zero-delay lossless speech coding technique for high-quality communications. The WaveNet generative model, which is a state-of-the-art model for neural-network-based speech waveform synthesis, is used in both the encoder and decoder. In the encoder, discrete speech signals are losslessly compressed using sample-by-sample entropy coding. The decoder fully reconstructs the original speech signals from the compressed signals without algorithmic delay. Experimental results show that the proposed coding technique can transmit speech audio waveforms with 50% their original bit rate and the WaveNet-based speech coder remains effective for unknown speakers. Takenori Yoshimura, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda |
SLT | 4 |
| 2018 | Mel-Cepstrum-Based Quantization Noise Shaping Applied to Neural-Network-Based Speech Waveform SynthesisabstractThis paper presents a mel-cepstrum-based quantization noise shaping method for improving the quality of synthetic speech generated by neural-network-based speech waveform synthesis systems. Since mel-cepstral coefficients closely match the characteristics of human auditory perception, the proposed method effectively masks the white noise introduced by the quantization typically used in neural-network-based speech waveform synthesis systems. The paper also describes a computationally efficient implementation of the proposed method using the structure of the mel-log spectrum approximation filter. Experiments using the WaveNet generative model, which is a state-of-the-art model for neural-network-based speech waveform synthesis, showed that speech quality is significantly improved by the proposed method. Takenori Yoshimura, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2017 | Image recognition based on discriminative models using features generated from separable lattice HMMSabstractThis paper presents an image recognition technique based on discriminative models using features generated from separable lattice hidden Markov models (SL-HMMs). A major problem in image recognition is that the recognition performance is degraded by geometric variations such as that in position and size of the object to be recognized. SL-HMMs have been proposed to solve this problem. SL-HMMs are an extension of HMMs with size and locational invariances based on state transitions. An SL-HMM is a generative model and can represent generation processes of observations well. However, there is a possibility that the recognition performance of generative models is inferior to that of discriminative models because discriminative models are specialized to identification. In this paper, we propose image recognition based on log linear models (LLMs) using features extracted from SL-HMMs. The proposed method can extract features invariant to geometric variations by using SL-HMMs and built an accurate classifier based on discriminative models with the extracted features. Face recognition experiments showed that the proposed method obtained higher recognition rates than SL-HMMs and convolutional neural networks based methods. Yoshinari Tsuzuki, Kei Sawada, Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda |
ICASSP | 4 |
| 2017 | Articulatory Text-to-Speech Synthesis Using the Digital Waveguide Mesh Driven by a Deep Neural NetworkabstractFollowing recent advances in direct modeling of the speech waveform using a deep neural network, we propose a novel method that directly estimates a physical model of the vocal tract from the speech waveform, rather than magnetic resonance imaging data. This provides a clear relationship between the model and the size and shape of the vocal tract, offering considerable flexibility in terms of speech characteristics such as age and gender. Initial tests indicate that despite a highly simplified physical model, intelligible synthesized speech is obtained. This illustrates the potential of the combined technique for the control of physical models in general, and hence the generation of more natural-sounding synthetic speech. Amelia Jane Gully, Takenori Yoshimura, Damian T. Murphy, Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda |
INTERSPEECH | 5 |
| 2017 | Simultaneous Optimization of Multiple Tree-Based Factor Analyzed HMM for Speech SynthesisabstractThis paper proposes a novel method to build multiple decision trees as a structure of factor analyzed hidden Markov model for speech synthesis. In the proposed method, the multiple decision trees grow simultaneously rather than sequentially to take into account the relationship between the trees. However, the simultaneous growing is computationally infeasible due to an exponential increase in the number of tree structures to be evaluated. To solve the problem, we further propose two computational complexity reduction algorithms that achieve a significant reduction in the computational time. Experimental results show that the proposed method outperforms the conventional one based on a single decision tree. Takenori Yoshimura, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2016 | Trajectory training considering global variance for speech synthesis based on neural networksabstractThis paper proposes a new training method of deep neural networks (DNNs) for statistical parametric speech synthesis. DNNs are recently used as acoustic models that represent mapping functions from linguistic features to acoustic features in statistical parametric speech synthesis. There are problems to be solved in conventional DNN-based speech synthesis: 1) the inconsistency between the training and synthesis criteria; and 2) the over-smoothing of the generated parameter trajectories. In this paper, we introduce the parameter trajectory generation process considering the global variance (GV) into the training of DNNs. A unified framework which consistently uses the same criterion in both training and synthesis can be obtained and the model parameters are optimized for parameter generation considering the GV in the proposed method. Experimental results show that the proposed method outperforms the conventional method in the naturalness of synthesized speech. Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda |
ICASSP | 3 |
| 2016 | Redefining the Linguistic Context Feature Set for HMM and DNN TTS Through Position and Parsing
Rasmus Dall, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda |
INTERSPEECH | 4 |
| 2016 | Voice Conversion Based on Trajectory Model Training of Neural Networks Considering Global Variance
Naoki Hosaka, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda |
INTERSPEECH | 4 |
| 2016 | Singing Voice Synthesis Based on Deep Neural Networks
Masanari Nishimura, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda |
INTERSPEECH | 4 |
| 2015 | The effect of neural networks in statistical parametric speech synthesisabstractThis paper investigates how to use neural networks in statistical parametric speech synthesis. Recently, deep neural networks (DNNs) have been used for statistical parametric speech synthesis. However, the specific way how DNNs should be used in statistical parametric speech synthesis has not been studied thoroughly. A generation process of statistical parametric speech synthesis based on generative models can be divided into several components, and those components can be represented by DNNs. In this paper, the effect of DNNs for each component is investigated by comparing DNNs with generative models. Experimental results show that the use of a DNN as acoustic models is effective and the parameter generation combined with a DNN improves the naturalness of synthesized speech. Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda |
ICASSP | 3 |
| 2015 | Prosodically-enhanced recurrent neural network language modelsabstractRecurrent neural network language models have been shown to consistently reduce the word error rates (WERs) of large vocabulary speech recognition tasks. In this work we propose to enhance the RNNLMs with prosodic features computed using the context of the current word. Since it is plausible to compute the prosody features at the word and syllable level we have trained the models on prosody features computed at both these levels. To investigate the effectiveness of proposed models we report perplexity and WER for two speech recognition tasks, Switchboard and TED. We observed substantial improvements in perplexity and small improvements in WER. Index Terms: RNNLMs, 3-gram, prosody features, pause duration, duration of the word, syllable duration, syllable F0, GMMHMM, DNN-HMM, Switchboard conversations and TED lectures Siva Reddy Gangireddy, Steve Renals, Yoshihiko Nankaku, Akinobu Lee |
INTERSPEECH | 3 |
| 2015 | Simultaneous optimization of multiple tree structures for factor analyzed HMM-based speech synthesis
Takenori Yoshimura, Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda |
INTERSPEECH | 3 |
| 2014 | HMM-Based singing voice synthesis and its application to Japanese and EnglishabstractThe present paper describes Japanese and English singing voice synthesis systems based on hidden Markov models (HMMs). In this approach, the spectrum, excitation, and vibrato of the singing voice are simultaneously modeled by context-dependent HMMs, and waveforms are generated by the HMMs themselves. Japanese singing voice synthesis systems have already been developed and used to create variable musical contents. To extend this system to English, language independent contexts are designed. Furthermore, methods for matching musical notes and pronunciation of English lyrics are presented and evaluated in subjective experiments. Then, Japanese and English singing voice synthesis systems are compared. Kazuhiro Nakamura, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda |
ICASSP | 3 |
| 2014 | Integration of speaker and pitch adaptive training for HMM-based singing voice synthesisabstractA statistical parametric approach to singing voice synthesis based on hidden Markov models (HMMs) has been growing in popularity over the last few years. The spectrum, excitation, vibrato, and duration of the singing voice in this approach are simultaneously modeled with context-dependent HMMs and waveforms are generated from the HMMs themselves. Since HMM-based singing voice synthesis systems are “corpus-based,” the HMMs corresponding to contextual factors that rarely appear in the training data cannot be well-trained. However, it may be difficult to prepare a large enough quantity of singing voice data sung by one singer. Furthermore, the pitch included in each song is imbalanced, and there is the vocal range of the singer. In this paper, we propose “singer adaptive training” which can solve the data sparse-ness problem. Experimental results demonstrated that the proposed technique improved the quality of the synthesized singing voices. Kanako Shirota, Kazuhiro Nakamura, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda |
ICASSP | 5 |
| 2014 | A mel-cepstral analysis technique restoring high frequency components from low-sampling-rate speech
Kazuhiro Nakamura, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda |
INTERSPEECH | 4 |
| 2013 | Separable lattice 2-D HMMS introducing state duration control for recognition of images with various variationsabstractIn this paper, an extension of separable lattice HMMs (SL-HMM) is described that introduces state duration control for dealing with images with various variations. SL-HMM are generative models that have size and location invariances based on state transition of HMMs. An extended model that has the structure of hidden semi-Markov models (HSMMs) in which the state duration probability is explicitly modeled by parametric distributions is also proposed. However, in this model, each state duration in a Markov chain is independent. It is supposed that each state duration should have a correlation. Therefore, in this paper, we propose a novel model that solves this problem by introducing variables representing the correlation among the state durations. Face recognition experiments show that the proposed model improved the recognition performance for images with size, locational, and rotational variations. Takaya Makino, Shinji Takaki, Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda |
ICASSP | 4 |
| 2013 | Integration of acoustic modeling and mel-cepstral analysis for HMM-based speech synthesisabstractIn this paper, a novel approach for integrating acoustic modeling and mel-cepstral analysis is proposed. The aim of HMM-based speech synthesis is to model speech waveforms with a statistical model. However, the conventional techniques divide the modeling process into two steps: the frame by frame feature extraction step and the acoustic modeling step. Although it is reasonably effective, the deterioration of speech quality is caused by the divide of the objective function. In this paper, we propose an approach to modeling them as an integrative model and show the possibility of improving synthesized speech. Kazuhiro Nakamura, Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda |
ICASSP | 3 |
| 2013 | Contextual partial additive structure for HMM-based speech synthesisabstractThis paper proposes a spectral modeling technique based on a contextual partial additive structure for HMM-based speech synthesis. To represent complicated context dependencies, contextual additive structure models assume multiple independent components which have different context dependencies to form acoustic features. In additive structure models, there is a constraint that a fixed number of additive components are used for generating acoustic features. However, it is natural to assume that the number of components depends on contexts. In the proposed technique, partial additive components affecting arbitrary contextual sub-spaces are created on demand to increase the likelihood. Then, the number of components for each context can be automatically determined with the training data. Experimental results show that the proposed technique outperformed the standard technique in a subjective test. Shinji Takaki, Yoshihiko Nankaku, Keiichi Tokuda |
ICASSP | 2 |
| 2013 | Image recognition based on separable lattice trajectory 2-D HMMSabstractIn this paper, a novel statistical model for image recognition based on separable lattice 2-D HMMs (SL2D-HMMs) is proposed. Although SL2D-HMMs can model invariance to size and location deformation, its modeling accuracy is still insufficient because of the following two assumptions: i) the statistics of each state are constant and ii) the state output probabilities are conditionally independent. In this paper, SL2D-HMMs are reformulated as a trajectory model that can capture dependencies between adjacent observations. The effectiveness of the proposed model was demonstrated in face recognition and image alignment experiments. Akira Tamamori, Yoshihiko Nankaku, Keiichi Tokuda |
ICASSP | 2 |
| 2013 | Speech Synthesis Based on Hidden Markov ModelsabstractThis paper gives a general overview of hidden Markov model (HMM)-based speech synthesis, which has recently been demonstrated to be very effective in synthesizing speech. The main advantage of this approach is its flexibility in changing speaker identities, emotions, and speaking styles. This paper also discusses the relation between the HMM-based approach and the more conventional unit-selection approach that has dominated over the last decades. Finally, advanced techniques for future developments are described. Keiichi Tokuda, Yoshihiko Nankaku, Tomoki Toda, Heiga Zen, Junichi Yamagishi, Keiichiro Oura |
Proc. IEEE | 2 |
| 2012 | Face recognition based on extended separable lattice 2-D HMMSabstractThis paper proposes an extension of separable lattice 2-D hidden Markov models (SL-HMMs) for dealing with image rotation and local deformation. It is important to reduce the effect of geometrical variations in image recognition, e.g., location, size, and rotation. SLHMMs are one of the most efficient structures to accomplish invariance to size and location variations. However, since SL-HMMs only have one state sequence in each direction, they cannot deal with rotation or local deformation. The proposed models have state sequences corresponding to all rows and columns of an input image, and the complicated state alignments can represent rotation and local deformation. The effectiveness of the proposed models was demonstrated in face recognition experiments. Keisuke Kumaki, Yoshihiko Nankaku, Keiichi Tokuda |
ICASSP | 2 |
| 2012 | Pitch adaptive training for hmm-based singing voice synthesisabstractA statistical parametric approach to singing voice synthesis based on hidden Markov Models (HMMs) has been growing in popularity over the last few years. The spectrum, excitation, vibrato, and duration of singing voices in this approach are simultaneously modeled with context-dependent HMMs and waveforms are generated from the HMMs themselves. HMM-based singing voice synthesis systems are heavily based on the training data in performance because these systems are “corpus-based.” Therefore, HMMs corresponding to contextual factors that hardly ever appear in the training data cannot be well-trained. Pitch should especially be correctly covered since generated F0trajectories have a great impact on the subjective quality of synthesized singing voices. We applied the method of “speaker adaptive training” (SAT) to “pitch adaptive training,” which is discussed in this paper. This technique made it possible to normalize pitch based on musical notes in the training process. The experimental results demonstrated that the proposed technique could alleviate the data sparseness problem. Keiichiro Oura, Ayami Mase, Yoshihiko Nankaku, Keiichi Tokuda |
ICASSP | 3 |
| 2012 | Face recognition based on separable lattice 2-D HMMS using variational bayesian methodabstractThis paper proposes an image recognition technique based on separable lattice 2-D HMMs (SL2D-HMMs) using the variational Bayesian method. SL2D-HMMs have been proposed to reduce the effect of geometric variations, e.g., size and location. The maximum likelihood criterion had previously been used in training SL2D-HMMs. However, in many image recognition tasks, it is difficult to use sufficient training data, and it suffers from the over-fitting problem. A higher generalization ability based on model marginalization is expected by applying the Bayesian criterion and useful prior information on model parameters can be utilized as prior distributions. Experiments on face recognition indicated that the proposed method improved image recognition. Kei Sawada, Akira Tamamori, Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda |
ICASSP | 4 |
| 2012 | A model structure integration based on a Bayesian framework for speech recognitionabstractThis paper proposes an acoustic modeling technique based on Bayesian framework using multiple model structures for speech recognition. The Bayesian approach is a statistical technique for estimating reliable predictive distributions by marginalizing model parameters, and its effectiveness in HMM-based speech recognition has been reported. Although the basic idea underlying the Bayesian approach is to treat all parameters as random variables, only one model structure is still selected in the conventional method. Multiple model structures are treated as latent variables in the proposed method and integrated based on the Bayesian framework. Furthermore, we applied deterministic annealing to the training algorithm to estimate appropriate acoustic models. The proposed method effectively utilizes multiple model structures, especially in the early stage of training and this leads to better predictive distributions and improvement of recognition performance. Sayaka Shiota, Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda |
ICASSP | 3 |
| 2012 | A Bayesian Approach to Speaker Recognition Based on GMMs Using Multiple Model Structures
Takafumi Hattori, Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda |
INTERSPEECH | 3 |
| 2012 | Cross-lingual Speaker Adaptation for HMM-based Speech Synthesis based on Perceptual Characteristics and Speaker Interpolation
Viviane de Franca Oliveira, Sayaka Shiota, Yoshihiko Nankaku, Keiichi Tokuda |
INTERSPEECH | 3 |
| 2012 | Product of Experts for Statistical Parametric Speech SynthesisabstractMultiple acoustic models are often combined in statistical parametric speech synthesis. Both linear and non-linear functions of an observation sequence are used as features to be modeled. This paper shows that this combination of multiple acoustic models can be expressed as a product of experts (PoE); the likelihoods from the models are scaled, multiplied together, and then normalized. Normally these models are individually trained and only combined at the synthesis stage. This paper discusses a more consistent PoE framework where the models are jointly trained. A training algorithm for PoEs based on linear feature functions and Gaussian experts is derived by generalizing the training algorithm for trajectory HMMs. However for non-linear feature functions or non-Gaussian experts this is not possible, so a scheme based on contrastive divergence learning is described. Experimental results show that the PoE framework provides both a mathematically elegant way to train multiple acoustic models jointly and significant improvements in the quality of the synthesized speech. Heiga Zen, Mark J. F. Gales, Yoshihiko Nankaku, Keiichi Tokuda |
IEEE Trans. Speech Audio Process. | 3 |
| 2011 | Global variance modeling on frequency domain delta LSP for HMM-based speech synthesisabstractThe speech parameter generation algorithm considering global variance (GV) for HMM-based speech synthesis proved to be effective against the over-smoothing problem. However, the correlation between dimensions of parameter vector is not sufficiently considered in the current GV model. For some parameters, e.g., Line Spectral Pairs (LSP), the difference of adjacent LSPs has strong influence on the spectral envelope. Considering this important feature, the paper proposes a GV modeling on the difference of adjacent LSPs, i.e., GV on frequency domain delta LSP. By improving the GV likelihood on frequency domain delta LSP, the over-smoothing effect of generated parameter trajectory is better alleviated than conventional one. The result of a perceptual evaluation shows the proposed method outperforms the conventional one, and the naturalness of synthetic speech is improved. Shifeng Pan, Yoshihiko Nankaku, Keiichi Tokuda, Jianhua Tao 0001 |
ICASSP | 2 |
| 2011 | An optimization algorithm of independent mean and variance parameter tying structures for HMM-based speech synthesisabstractThis paper proposes a technique for constructing independent parameter tying structures of mean and variance in HMM based speech synthesis. Conventionally, mean and variance parameters are assumed to have the same tying structure. However, it has been reported that a clustering technique of mean vectors while tying all variance matrices improves the quality of synthesized speech. This indicates that mean and variance parameters should have different optimal tying structures. In the proposed technique, the decision trees for mean and variance parameters are simultaneously grown by taking into account the dependency on mean and variance parameters. Experimental results show that the proposed technique outperforms the conventional one. Shinji Takaki, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda |
ICASSP | 3 |
| 2011 | Estimation of Window Coefficients for Dynamic Feature Extraction for HMM-Based Speech Synthesis
Linghui Chen, Yoshihiko Nankaku, Heiga Zen, Keiichi Tokuda, Zhen-Hua Ling, Li-Rong Dai 0001 |
INTERSPEECH | 2 |
| 2011 | Multi-Speaker Modeling with Shared Prior Distributions and Model Structures for Bayesian Speech SynthesisabstractThis paper investigates a multi-speaker modeling technique with shared prior distributions and model structures for Bayesian speech synthesis. The quality of synthesized speech is improved by selecting appropriate model structures in HMMbased speech synthesis. Bayesian approach is known to work for such model selection. However, the result is strongly affected by prior distributions of model parameters. Therefore, determination of prior distributions and selection of model structures should be performed simultaneously. This paper investigates prior distributions and model structures in the situation where training data of multiple speakers are available. The prior distributions and model structures which represent acoustic features common to every speakers can be obtained by sharing them between multiple speaker-dependent models. Index Terms: speech synthesis, Bayesian approach, prior distribution, context clustering, multi-speaker modeling A statistical parametric speech synthesis system based on hidden Markov models (HMMs) was recently developed. In HMM-based speech synthesis, the spectrum, excitation, and duration of speech are simultaneously modeled with HMMs, and speech parameter sequences are generated from the HMMs themselves [1]. The maximum likelihood (ML) criterion has typically been used for training HMMs and generating speech parameters. The ML criterion guarantees that the ML estimates approach the true values of the parameters. However, since the ML criterion produces a point estimate of the model parameters, its estimation accuracy may degrade when the amount of training data is insufficient. In the Bayesian approach, all variables introduced when the models are parameterized, such as model parameters and latent variables, are treated as random variables, and their posterior distributions are obtained by the Bayes theorem. The Bayesian approach can generally construct a more robust model than the ML approach by estimating posterior distributions. Recently, Bayesian speech synthesis has been proposed as a Bayesian framework for statistical parametric speech synthesis (e.g., HMM-based speech synthesis), and it shows good performance [2]. In Bayesian speech synthesis, all processes for constructing the system can be derived from a single predictive distribution that directly represents the problem of speech synthesis. The quality of synthesized speech is improved by selecting appropriate model structures in HMM-based speech synthesis. Although the Bayesian approach is known to work for such model selection, the results are strongly affected by prior distributions of the model parameters. Therefore, in Bayesian speech synthesis, determination of prior distributions and selection of model structures should be performed simultaneously. To overcome this problem, we have proposed Bayesian context clustering using cross validation [3]. In this method, prior distributions are determined by using a part of training data, and model structures are evaluated by using the determined prior distribution based on cross validation. In this paper, we investigates prior distributions and model Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda |
INTERSPEECH | 2 |
| 2011 | Evaluation of Tree-Trellis Based Decoding in Over-Million LVCSR
Naoaki Ito, Yoshihiko Nankaku, Akinobu Lee |
INTERSPEECH | 2 |
| 2011 | A Bayesian Approach to Voice Conversion Based on GMMs Using Multiple Model Structures
Yoshihiko Nankaku, Keiichi Tokuda |
INTERSPEECH | 2 |
| 2011 | GMM-Based Missing-Feature Reconstruction on Multi-Frame Windows
Ulpu Remes, Yoshihiko Nankaku, Keiichi Tokuda |
INTERSPEECH | 2 |
| 2011 | Continuous Stochastic Feature Mapping Based on Trajectory HMMsabstractThis paper proposes a technique of continuous stochastic feature mapping based on trajectory hidden Markov models (HMMs), which have been derived from HMMs by imposing explicit relationships between static and dynamic features. Although Gaussian mixture model (GMM)- or HMM-based feature-mapping techniques work effectively, their accuracy occasionally degrades due to inappropriate dynamic characteristics caused by frame-by-frame mapping. While the use of dynamic-feature constraints at the mapping stage can alleviate this problem, it also introduces inconsistencies between training and mapping. The technique we propose can eliminate these inconsistencies while retaining the benefits of using dynamic-feature constraints, and it offers entire sequence-level transformation rather than frame-by-frame mapping. The results obtained from speaker-conversion, acoustic-to-articulatory inversion-mapping, and noise-compensation experiments demonstrated that our new approach outperformed the conventional one. Heiga Zen, Yoshihiko Nankaku, Keiichi Tokuda |
IEEE Trans. Speech Audio Process. | 2 |
| 2010 | A Deterministic Annealing-Based Training Algorithm For Statistical Machine Translation Models
Pascual Martínez-Gómez, Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda, Germán Sanchis-Trilles |
EAMT | 3 |
| 2010 | Factor analyzed voice models for HMM-based speech synthesisabstractThis paper describes factor analyzed voice models for realizing various voice characteristics in the HMM-based speech synthesis. The eigenvoice method can synthesize speech with arbitrary voice characteristics by interpolating representative HMM sets. However, the objective of PCA is to accurately reconstruct each speaker-dependent HMM set, and this is not equivalent to estimating models which represent training data accurately. To overcome this problem, we propose a general speech model which generates speech utterances with various voice characteristics directly. In the proposed method, the HMM states, factors representing voice characteristics and contextual decision trees are simultaneously optimized within a unified framework. Kyosuke Kazumi, Yoshihiko Nankaku, Keiichi Tokuda |
ICASSP | 2 |
| 2010 | Face recognition based on separable lattice 2-D HMM with state duration modelingabstractThis paper describes an extension of separable lattice 2-D HMMs (SL-HMMs) using state duration models for image recognition. SL-HMMs are generative models which have size and location invariances based on state transition of HMMs. However, the state duration probability of HMMs exponentially decreases with increasing duration, therefore it may not be appropriate for modeling image variations accuratelty. To overcome this problem, we employ the structure of hidden semi Markov models (HSMMs) in which the state duration probability is explicitly modeled by parametric distributions. Face recognition experiments show that the proposed model improved the performance for images with size and location variations. Yoshiaki Takahashi, Akira Tamamori, Yoshihiko Nankaku, Keiichi Tokuda |
ICASSP | 3 |
| 2010 | An extension of Separable Lattice 2-D HMMS for rotational data variationsabstractThis paper proposes a new generative model which can deal with rotational data variations by extending Separable Lattice 2-D HMMs (SL2D-HMMs). In image recognition, geometrical variations such as size, location and rotation degrade the performance, therefore normalization is required. SL2D-HMMs can perform an elastic matching in both horizontal and vertical directions; this makes it possible to model invariances to size and location. To deal with rotational variations, we introduce additional HMM states which represent the shifts of the state alignments of the observation lines in a particular direction. Face recognition experiments show that the proposed method improves the performance significantly for rotational variation data. Akira Tamamori, Yoshihiko Nankaku, Keiichi Tokuda |
ICASSP | 2 |
| 2010 | Statistical parametric speech synthesis based on product of expertsabstractMultiple-level acoustic models (AMs) are often combined in statistical parametric speech synthesis. Both linear and non-linear functions of the observation sequence are used as features in these AMs. This combination of multiple-level AMs can be expressed as a product of experts (PoE); the likelihoods from the AMs are scaled, multiplied together and then normalized. Currently these multiple-level AMs are individually trained and only combined at the synthesis stage. This paper discusses a more consistent PoE framework where the AMs are jointly trained. A generalization of trajectory HMM training can be used for multiple-level Gaussian AMs based on linear functions. However for the non-linear case this is not possible, so a scheme based on contrastive divergence learning is described. Experimental results show that the proposed technique provides both a mathematically elegant way to train multiple-level AMs and statistically significant improvements in the quality of synthesized speech. Heiga Zen, Mark J. F. Gales, Yoshihiko Nankaku, Keiichi Tokuda |
ICASSP | 3 |
| 2010 | Speaker adaptation based on nonlinear spectral transform for speech recognition
Toyohiro Hayashi, Yoshihiko Nankaku, Akinobu Lee, Keiichi Tokuda |
INTERSPEECH | 2 |
| 2010 | HMM-based singing voice synthesis system using pitch-shifted pseudo training data
Ayami Mase, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda |
INTERSPEECH | 3 |
| 2010 | Voice activity detection based on conditional random fields using multiple featuresabstractThis paper proposes a Voice Activity Detection (VAD) algorithm based on Conditional Random Fields (CRF) using multiple features.VAD is a technique used to distinguish between speech and non-speech in noisy environments and is an important component in many real-world speech applications.The posterior probability of output labels in the proposed method is directly modeled by the weighted sum of the feature functions.Effective features are automatically selected by estimating appropriate weight parameters to improve the accuracy of VAD.Experimental results on the CENSREC-1-C database revealed that the proposed approach can decrease error rates by using CRF. Akira Saito, Yoshihiko Nankaku, Akinobu Lee, Keiichi Tokuda |
INTERSPEECH | 2 |
| 2009 | A Bayesian approach to HMM-based speech synthesisabstractThis paper proposes a new framework of speech synthesis based on the Bayesian approach. The Bayesian method is a statistical technique for estimating reliable predictive distributions by marginalizing model parameters. In the proposed framework, all processes for constructing the system can be derived from one single predictive distribution which represents the basic problem of speech synthesis directly. Using HMM as the likelihood function and assuming some approximations, it can be regarded as an application of the variational Bayesian method to the HMM-based speech synthesis. Experimental results show that the proposed method outperforms the conventional one in a subjective test. Kei Hashimoto, Heiga Zen, Yoshihiko Nankaku, Takashi Masuko, Keiichi Tokuda |
ICASSP | 3 |
| 2009 | Voice conversion based on simultaneous modelling of spectrum and F0abstractThis paper proposes a simultaneous modeling of spectrum and F0 for voice conversion based on MSD (multi-space probability distribution) models. As a conventional technique, a spectral conversion based on GMM (Gaussian mixture model) has been proposed. Although this technique converts spectral feature sequences nonlinearly based on GMM, F0 sequences are usually converted by a simple linear function. This is because F0 is undefined in unvoiced segments. To overcome this problem, we apply MSD models. The MSD-GMM allows to model continuous F0 values in voiced frames and a discrete symbol representing unvoiced frames within an unified framework. Furthermore, the MSD-HMM is adopted to model long term correlations in F0 sequences. Kaori Yutani, Yosuke Uto, Yoshihiko Nankaku, Akinobu Lee, Keiichi Tokuda |
ICASSP | 3 |
| 2009 | Stereo-based stochastic noise compensation based on trajectory GMMSabstractThis paper proposes a novel stereo-based stochastic noise compensation technique based on trajectory GMMs. Although the GMM-based noise compensation techniques such as SPLICE work effective, their performance sometimes degrades due to the inappropriate dynamic characteristics caused by the frame-by-frame mapping. While the use of dynamic feature constraints on the mapping stage can alleviate this problem, it also introduces an inconsistency between training and mapping. The recently proposed trajectory GMM-based feature mapping technique can solve this inconsistency while keeping the benefits of the use of dynamic features, and offers an entire sequence-level transformation rather than the frame-by-frame mapping. Results from a noise compensation experiment on the AURORA-2 task show that the proposed trajectory GMM-based noise compensation technique outperforms the conventional ones. Heiga Zen, Yoshihiko Nankaku, Keiichi Tokuda |
ICASSP | 2 |
| 2009 | A Bayesian approach to Hidden Semi-Markov Model based speech synthesisabstractThis paper proposes a Bayesian approach to hidden semi-Markov model (HSMM) based speech synthesis. Recently, hid-den Markov model (HMM) based speech synthesis based on the Bayesian approach was proposed. The Bayesian approach is a statistical technique for estimating reliable predictive distribu-tions by treating model parameters as random variables. In the Bayesian approach, all processes for constructing the system are derived from one single predictive distribution which exactly represents the problem of speech synthesis. However, there is an inconsistency between training and synthesis: although the speech is synthesized from HMMs with explicit state duration probability distributions, HMMs are trained without them. In this paper, we introduce an HSMM, which is an HMM with explicit state duration probability distributions, into the HMM-based Bayesian speech synthesis system. Experimental results show that the use of HSMM improves the naturalness of the synthesized speech. Index Terms: speech synthesis, HSMM, Bayesian approach 1. Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda |
INTERSPEECH | 2 |
| 2009 | Tying covariance matrices to reduce the footprint of HMM-based speech synthesis systemsabstractThis paper proposes a technique of reducing footprint of HMMbased speech synthesis systems by tying all covariance matrices. HMM-based speech synthesis systems usually consume smaller footprint than unit-selection synthesis systems because statistics rather than speech waveforms are stored. However, further reduction is essential to put them on embedded devices which have very small memory. According to the empirical knowledge that covariance matrices have smaller impact for the quality of synthesized speech than mean vectors, here we propose a clustering technique of mean vectors while tying all covariance matrices. Subjective listening test results show that the proposed technique can shrink the footprint of an HMM-based speech synthesis system while retaining the quality of synthesized speech. Index Terms: HMM, speech synthesis, decision tree, contextclustering, MDL criterion, embedded device Keiichiro Oura, Heiga Zen, Yoshihiko Nankaku, Akinobu Lee, Keiichi Tokuda |
INTERSPEECH | 3 |
| 2009 | Deterministic annealing based training algorithm for Bayesian speech recognitionabstractThis paper proposes a deterministic annealing based training algorithm for Bayesian speech recognition. The Bayesian method is a statistical technique for estimating reliable predictive distributions by marginalizing model parameters. However, the local maxima problem in the Bayesian method is more serious than in the ML-based approach, because the Bayesian method treats not only state sequences but also model parameters as latent variables. The deterministic annealing EM (DAEM) algorithm has been proposed to improve the local maxima problem in the EM algorithm, and its effectiveness has been reported in HMMbased speech recognition using ML criterion. In this paper, the DAEM algorithm is applied to Bayesian speech recognition to relax the local maxima problem. Speech recognition experiments show that the proposed method achieved a higher performance than the conventional methods. Sayaka Shiota, Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda |
INTERSPEECH | 3 |
| 2009 | State mapping based method for cross-lingual speaker adaptation in HMM-based speech synthesisabstractA phone mapping-based method had been introduced for cross-lingual speaker adaptation in HMM-based speech syn-thesis. In this paper, we continue to propose a state mapping based method for cross-lingual speaker adaptation, where the state mapping between voice models in source and target lan-guages is established under minimum Kullback-Leibler diver-gence (KLD) criterion. We introduce two approaches to use the established mapping information for cross-lingual speaker adaptation, including data mapping and transform mapping ap-proaches. From the experimental results, the state mapping based method outperformed the phone mapping based method. In addition, the data mapping approach achieved better speaker similarity, and the transform mapping approach achieved better speech quality after cross-lingual speaker adaptation. Index Terms: Speech synthesis, HMM, speaker adaptation, minimum generation error, linear regression Yi-Jian Wu, Yoshihiko Nankaku, Keiichi Tokuda |
INTERSPEECH | 2 |
| 2008 | Acoustic modeling with contextual additive structure for HMM-based speech recognitionabstractThis paper proposes an acoustic modeling technique based on an additive structure of context dependencies for HMM-based speech recognition. Typical context dependent models, e.g., triphone HMMs, have direct dependencies of phonetic contexts, i.e., if a phonetic context is given, the Gaussian distribution is specified immediately. This paper assumes a more complex structure, an additive structure of acoustic feature components which have different context dependencies. Since the output probability distribution is composed of additive component distributions, a number of different distributions can be efficiently represented by a combination of fewer distributions. To automatically extract additive components, this paper presents a context clustering algorithm for the additive structure model in which multiple decision trees are constructed simultaneously. Experimental results show that the proposed technique improves phoneme recognition accuracy with fewer number of distributions than the conventional triphone HMMs. Yoshihiko Nankaku, Kazuhiro Nakamura, Heiga Zen, Keiichi Tokuda |
ICASSP | 1 |
| 2008 | Bayesian context clustering using cross valid prior distribution for HMM-based speech recognitionabstractThis paper proposes a prior distribution determination tech-nique using cross validation for speech recognition based on the Bayesian approach. The Bayesian method is a statisti-cal technique for estimating reliable predictive distributions by marginalizing model parameters and its approximate version, the variational Bayesian method has been applied to HMM-based speech recognition. Since prior distributions represent-ing prior information about model parameters affect the pos-terior distributions and model selection, the determination of prior distributions is an important problem. However, it has not been thoroughly investigate in speech recognition. The pro-posed method can determine reliable prior distributions with-out tuning parameters and select an appropriate model struc-ture dependently on the amount of training data. Continu-ous phoneme recognition experiments show that the proposed method achieved a higher performance than the conventional methods. Index Terms: variational Bayes, cross validation, context clus-tering, continuous phoneme recognition Kei Hashimoto, Heiga Zen, Yoshihiko Nankaku, Akinobu Lee, Keiichi Tokuda |
INTERSPEECH | 3 |
| 2008 | Speaker recognition based on variational Bayesian method
Tatsuya Ito, Kei Hashimoto, Yoshihiko Nankaku, Akinobu Lee, Keiichi Tokuda |
INTERSPEECH | 3 |
| 2008 | Acoustic modeling based on model structure annealing for speech recognitionabstractThis paper proposes an HMM training technique using multiple phonetic decision trees and evaluates it in speech recognition. In the use of context dependent models, the decision tree based context clustering is applied to find a parameter tying structure. However, the clustering is usually performed based on statistics of HMM state sequences which are obtained by unreliable models without context clustering. To avoid this problem, we optimize the decision trees and HMM state sequences simultaneously. In the proposed method, this is performed by maximum likelihood (ML) estimation of a newly defined statistical model which includes multiple decision trees as hidden variables. Applying the deterministic annealing expectation maximization (DAEM) algorithm and using multiple decision trees in early stage of model training, state sequences are reliably estimated. In continuous phoneme recognition experiments, the proposed method can improve the recognition performance. Sayaka Shiota, Kei Hashimoto, Heiga Zen, Yoshihiko Nankaku, Akinobu Lee, Keiichi Tokuda |
INTERSPEECH | 4 |
| 2008 | Probabilistic answer selection based on conditional random fields for spoken dialog system
Yoshitaka Yoshimi, Ryota Kakitsuba, Yoshihiko Nankaku, Akinobu Lee, Keiichi Tokuda |
INTERSPEECH | 3 |
| 2008 | Simultaneous conversion of duration and spectrum based on statistical models including time-sequence matchingabstractThis paper describes a simultaneous conversion technique of duration and spectrum based on a statistical model including time-sequence matching. Conventional GMM-based approaches cannot perform spectral conversion taking account of speaking rate because it assumes one to one frame matching between source and target features. However, speaker characteristics may appear in speaking rates. In order to perform duration conversion, we attach duration models to statistical models including time-sequence matching (DPGMM). Since DPGMM can represent two different length sequences directly, the conversion of spectrum and duration can be performed within an integrated framework. In the proposed technique, each mixture component of DPGMM has different duration transformation functions, therefore durations are converted nonlinearly and dependently on spectral information. In the subjective DMOS test, the proposed method is superior to the conventional method. Index Terms: voice conversion, GMM, duration conversion 1. Kaori Yutani, Yosuke Uto, Yoshihiko Nankaku, Tomoki Toda, Keiichi Tokuda |
INTERSPEECH | 3 |
| 2008 | Probabilistic feature mapping based on trajectory HMMs
Heiga Zen, Yoshihiko Nankaku, Keiichi Tokuda |
INTERSPEECH | 2 |
| 2007 | Face Recognition using Hidden Markov Eigenface ModelsabstractThis paper proposes hidden Markov eigenface models (HMEMs) in which the eigenfaces are integrated into separable lattice hidden Markov models (SL-HMMs). SL-HMMs have been proposed for modeling multi-dimensional data, e.g., images, image sequences, 3-D objects. In its application to face recognition, SL-HMMs can perform an elastic image matching in both horizontal and vertical directions. However, SL-HMMs still have a limitation that the observations are assumed to be generated independently from corresponding states; it is insufficient to represent variations in face images, e.g., lighting conditions, facial expressions, etc. To overcome this problem, the structure of probabilistic principal component analysis (PPCA) and factor analysis (FA) is used as a probabilistic representation of eigenfaces. The proposed model has good properties of both PPCA/FA and SL-HMMs: a linear feature extraction and invariances to size and location of images. In face recognition experiments on the XM2VTS database, the proposed model improved the performance significantly. Yoshihiko Nankaku, Keiichi Tokuda |
ICASSP (2) | 1 |
| 2007 | A trainable excitation model for HMM-based speech synthesisabstractThis paper introduces a novel excitation approach for speech synthesizers in which the final waveform is generated through parameters directly obtained from Hidden Markov Models (HMMs). Despite the attractiveness of the HMM-based speech synthesis technique, namely utilization of small corpora and flexibility concerning the achievement of different voice styles, synthesized speech presents a characteristic buzziness caused by the simple excitation model which is employed during the speech production. This paper presents an innovative scheme where mixed excitation is modeled through closed-loop training of a set of state-dependent filters and pulse trains, with minimization of the error between excitation and residual sequences. The proposed method shows effectiveness, yielding synthesized speech with quality far superior to the simple excitation baseline and comparable to the best excitation schemes thus far reported for HMM-based speech synthesis. Ranniery Maia, Tomoki Toda, Heiga Zen, Yoshihiko Nankaku, Keiichi Tokuda |
INTERSPEECH | 4 |
| 2007 | Model-space MLLR for trajectory HMMsabstractThis paper proposes model-space Maximum Likelihood Linear Regression (mMLLR) based speaker adaptation technique for trajectory HMMs, which have been derived from HMMs by imposing explicit relationships between static and dynamic features. This model can alleviate two limitations of the HMM: constant statistics within a state and conditional independence assumption of state out-put probabilities without increasing the number of model parameters. Results in a continuous speech recognition experiments show that the proposed algorithm can adapt trajectory HMMs to a specific speaker and improve the performance of a trajectory HMM-based speech recogni-tion system. Index Terms: trajectory HMM, speaker adaptation, model-space MLLR Heiga Zen, Yoshihiko Nankaku, Keiichi Tokuda |
INTERSPEECH | 2 |
| 2006 | Face Recognition Based on Separable Lattice HMMSabstractIn this paper, we propose separable lattice hidden Markov models, in which multiple hidden state sequences interact to model the observation on a lattice. The proposed model can be efficiently applied for modeling images, image sequences, 3-D object models and higher dimensional applications, due to the composite structure of Markov chains which reduces the complexity while retaining good properties for multi-dimensional data. In case of 2-D lattices, the proposed model performs an elastic matching in both horizontal and vertical directions; this makes it possible to model not only invariances to the size and location of an object but also nonlinear warping in each dimension. We present a training algorithm for separable lattice HMMs based on a variational approximation. Moreover, the deterministic annealing EM (DAEM) algorithm was applied to the variational algorithm for separable lattice HMMs. Face recognition experiments on the XM2VTS database show that the proposed model has good properties for face image modeling. Daisuke Kurata, Yoshihiko Nankaku, Keiichi Tokuda, Tadashi Kitamura, Zoubin Ghahramani |
ICASSP (5) | 2 |
| 2006 | On the Use of Phonetic Information for Mapping from Articulatory Movements to Vocal Tract SpectrumabstractThis paper describes a method for determining the vocal tract spectrum from articulatory movements using an hidden Markov models (HMMs). In the proposed system, articulatory parameters are generated from a TTS system and converted to acoustic features to be synthesized. Comparing with conventional GMM-based systems, the proposed system has two additional properties: 1) phonetic information given input texts is available for the conversion, 2) the use of HMMs allows us to utilize the temporal structure of speech. In this paper, we investigate the optimal structure of HMMs for the conversion. Experimental results show that using phonetic and temporal information can improve the mapping accuracy in a spectral distortion measure Kenichi Nakamura, Tomoki Toda, Yoshihiko Nankaku, Keiichi Tokuda |
ICASSP (1) | 3 |
| 2006 | Hidden Semi-Markov Model Based Speech Recognition System using Weighted Finite-State TransducerabstractIn hidden Markov models (HMMs), state duration probabilities decrease exponentially with time. It would be inappropriate representation of temporal structure of speech. One of the solutions for this problem is integrating state duration probability distributions explicitly into the HMM. This form is known as a hidden semi-Markov model (HSMM) [1]. Although a number of attempts to use explicit duration models in speech recognition systems have been proposed, they are not consistent because various approximations were used in both training and decoding. In the present paper, a fully consistent speech recognition system based on the HSMM framework is proposed. In a speaker-dependent continuous speech recognition experiment, HSMM-based speech recognition system achieved about 5.9% relative error reduction over the corresponding HMM-based one. Keiichiro Oura, Heiga Zen, Yoshihiko Nankaku, Akinobu Lee, Keiichi Tokuda |
ICASSP (1) | 3 |
| 2006 | Estimating Trajectory Hmm Parameters Using Monte Carlo Em With Gibbs SamplerabstractIn the present paper, the Monte Carlo EM (MCEM) algorithm with a Gibbs sampler is applied for estimating parameters of a trajectory HMM, which has been derived from an HMM by imposing explicit relationships between static and dynamic features. The trajectory HMM can alleviate two limitations of the HMM, which are i) constant statistics within a state, and ii) conditional independence of state output probabilities, without increasing the number of model parameters. In a speaker-dependent continuous speech recognition experiment, trajectory HMMs estimated by the MCEM algorithm achieved significant improvements over the corresponding HMMs trained by the EM (Baum-Welch) algorithm. Heiga Zen, Yoshihiko Nankaku, Keiichi Tokuda, Tadashi Kitamura |
ICASSP (1) | 2 |
| 2006 | Reducing computation on parallel decoding using frame-wise confidence scores
Tomohiro Hakamata, Akinobu Lee, Yoshihiko Nankaku, Keiichi Tokuda |
INTERSPEECH | 3 |
| 2006 | An HMM-based singing voice synthesis systemabstractAbstract The present paper describes a corpus-based singing voice syn-thesis system based on hidden Markov models (HMMs). Thissystem employs the HMM-based speech synthesis to synthesizesingingvoice. Musical information such aslyrics, tones, durationsis modeled simultaneously in a unified framework of the context-dependent HMM. It can mimic the voice quality and singing styleof the original singer. Results of a singing voice synthesis exper-iment show that the proposed system can synthesize smooth andnatural-sounding singing voice. Index Terms : singing voice synthesis, HMM, time-lag model. 1. Introduction In recent years, various applications of speech synthesis systemshave been proposed and investigated. Singing voice synthesis isone of the hot topics in this area [1–5]. However, only a fewcorpus-based singing voice synthesis systems which can be con-structed automatically have been proposed.Currently, there are two main paradigms in the corpus-basedspeech synthesis area: sample-based approach and statistical ap-proach. The sample-based approach such as unit selection [6]can synthesize high-quality speech. However, it requires a hugeamountoftrainingdatatorealizevariousvoicecharacteristics. Onthe other hand, the quality of statistical approach such as HMM-basedspeechsynthesis[7]isbuzzybecauseitisbasedonavocod-ingtechnique. However,itissmoothandstable,anditsvoicechar-acteristics can easily be modified by transforming HMM parame-ters appropriately. For singing voice synthesis, applying the unitselection seems to be difficult because a huge amount of singingspeech which covers vast combinations of contextual factors thataffect singing voice has to be recorded. On the other hand, theHMM-based system can be constructed using a relatively smallamount of training data. From this point of view, the HMM-basedapproach seems to be more suitable for the singing voice synthe-sizer. In the present paper, we apply the HMM-based synthesisapproach to singing voice synthesis.Although the singing voice synthesis system proposed in thepresent paper is quite similar to the HMM-based text-to-speechsynthesissystem[7],therearetwomaindifferencesbetweenthem.In the HMM-based text-to-speech synthesis system, contextualfactors which may affect reading speech (e.g. phonemes, sylla-bles, words, phrases, etc.) are taken into account. However, con-textual factors which may affect singing voice should be different Keijiro Saino, Heiga Zen, Yoshihiko Nankaku, Akinobu Lee, Keiichi Tokuda |
INTERSPEECH | 3 |
| 2006 | Voice conversion based on mixtures of factor analyzersabstractThis paper describes the voice conversion based on the Mixtures of Factor Analyzers (MFA) which can provide an efficient modeling with a limited amount of training data. As a typical spectral conversion method, a mapping algorithm based on the Gaussian Mixture Model (GMM) has been proposed. In this method two kinds of covariance matrix structures are often used : the diagonal and full covariance matrices. GMM with diagonal covariance matrices requires a large number of mixture components for accurately estimating spectral features. On the other hand, GMM with full covariance matrices needs sufficient training data to estimate model parameters. In order to cope with these problems, we apply MFA to voice conversion. MFA can be regarded as intermediate model between GMM with diagonal covariance and with full covariance. Experimental results show that MFA can improve the conversion accuracy compared with the conventional GMM. Yosuke Uto, Yoshihiko Nankaku, Tomoki Toda, Akinobu Lee, Keiichi Tokuda |
INTERSPEECH | 2 |
| 2006 | Speaker adaptation of trajectory HMMs using feature-space MLLRabstractAbstract Recently, a trajectory model, derived from the hiddenMarkov model (HMM) by imposing explicit relationshipsbetween static and dynamic features, has been proposed.The derived model, named trajectory HMM , can alleviatetwo limitations of the HMM: constant statistics within astate and conditional independence assumption of state out-put probabilities. In the present paper, a speaker adapta-tion algorithm for the trajectory HMM based on feature-space Maximum Likelihood Linear Regression (fMLLR)is derived and evaluated. Results of a simple continu-ous speech recognition experiment shows that adapting tra-jectory HMMs using the derived adaptation algorithm im-proves the speech recognition performance. Index Terms : trajectory HMM, adaptation, fMLLR. 1. Introduction Speech recognition technologies have achieved significantprogress with the introduction of hidden Markov models(HMMs). Their tractability and efficient implementationsare achieved by a number of assumptions, such as constantstatistics within an HMM state, conditional independenceof state output probabilities. Although these assumptionsmake the HMM practically useful, they are not realistic formodeling sequences of speech spectra, especially in spon-taneous speech. To overcome these shortcomings of theHMM, a variety of alternative models have been proposed,e.g., [1–3]. Although these models can improve the speechrecognition performance, they generally require an increaseinthenumberofmodelparametersandcomputationalcom-plexity. Alternatively, the use of dynamic features (deltaand delta-delta features) [4] also improves the performanceof HMM-based speech recognizers. It can be viewed as asimple mechanism to capture time dependencies. However,it has been thought of as an ad hoc rather than an essentialsolution. Generally, dynamic features are calculated as re-gression coefficients from their neighboring static features.Therefore, relationshipsbetweenstaticanddynamicfeaturevector sequences are deterministic. However, usually theserelationships are ignored and the static and dynamic fea-tures are modeled as independent random variables. Ignor-ing these dependencies allows inconsistency between thestatic and dynamic features when the HMM is used as agenerative model in the obvious way.Recently, a trajectory model, derived from the HMM byimposing the explicit relationships between static and dy-namic features, has been proposed [5]. The derived model,named Heiga Zen, Yoshihiko Nankaku, Keiichi Tokuda, Tadashi Kitamura |
INTERSPEECH | 2 |
| 2005 | Sparse KPCA for Feature Extraction in Speech RecognitionabstractThis paper presents an analysis of the applicability of sparse kernel principal component analysis (SKPCA) for feature extraction in speech recognition, as well as a proposed approach to make the SKPCA technique realizable for a large amount of training data, which is a usual context in speech recognition systems. Although the KPCA (kernel principal component analysis) has proved to be an efficient technique for being applied to speech recognition, it has the disadvantage of requiring training data reduction, when its amount is excessively large. The standard approach to perform this data reduction is to randomly choose frames from the original data set, which does not necessarily provide a good statistical representation of the original data set. In order to solve this problem a likelihood related re-estimation procedure was applied to the KPCA framework, thus creating the SKPCA. The experimental results show the efficiency of SKPCA technique with the proposed approach over the KPCA with the standard sparse solution using randomly chosen frames and the standard feature extraction techniques. Amaro A. de Lima, Heiga Zen, Yoshihiko Nankaku, Keiichi Tokuda, Tadashi Kitamura, Fernando Gil Vianna Resende Jr. |
ICASSP (1) | 3 |
| 2004 | Parameter sharing and minimum classification error training of mixtures of factor analyzers for speaker identificationabstractThis paper investigates the parameter tying strategies of mixtures of factor analyzers (MFA) and discriminative training of MFA for speaker identification. The parameters of factor loading matrices or diagonal matrices are shared in different mixtures of MFA. The minimum classification error (MCE) training is applied to the MFA parameters to enhance the discrimination abilities. The results of text-independent speaker identification experiments show that MFA outperforms the conventional Gaussian mixture models (GMM) with diagonal or full covariance matrices and achieves the best performance when sharing the diagonal matrices, resulting in a relative gain of 26% over the GMM with diagonal covariance matrices. The recognition performance is further improved by the MCE training with an additional 3% error reduction. Hiroyoshi Yamamoto, Yoshihiko Nankaku, Chiyomi Miyajima, Keiichi Tokuda, Tadashi Kitamura |
ICASSP (1) | 2 |
| 2004 | Deterministic annealing EM algorithm in parameter estimation for acoustic modelabstractABSTRACT This paper investigates the effectiveness of the DAEM (Determin-istic Annealing EM) algorithm in acoustic modeling for speakerand speech recognition. Although the EM algorithm has beenwidely used to approximate the ML estimates, it has the problemof initialization dependence. To relax this problem, the DAEMalgorithm has been proposed and confirmed the effectiveness insmall tasks. In this paper, we applied the DAEM algorithm tospeakerrecognitionbasedonGMMsandcontinuousspeechrecog-nitionbasedonHMMs. ExperimentalresultsshowthattheDAEMalgorithm can improve the recognition performance as comparedtotheordinaryEMalgorithmwithconventionalinitializationmeth-ods,especiallyintheflatstarttrainingforcontinuousspeechrecog-nition. 1. INTRODUCTION The EM (Expectation-Maximization) algorithm [1] is widelyused for parameter estimation of statistical models with hiddenvariables. This algorithm provides a simple iterative proceduretoobtainapproximateML(maximumlikelihood)estimates. How-ever, since the EM algorithm is a hill-climbing approach, it suffersfrom the local maxima problem.On the other hand, GMMs (Gaussian mixture models) [2] andHMMs (hidden Markov models) [3] have been commonly usedin acoustic modeling for speaker and speech recognition, respec-tively. In conventional approaches, the LBG algorithm for GMMsand the segmental k-means algorithm for HMMs have been em-ployed to obtain initial model parameters before applying the EMalgorithm. However these initial values are not guaranteed to benear the true maximum likelihood point, and the posterior den-sity becomes unreliable at an early stage of training. Especiallyin continuous speech recognition, it is difficult to obtain accuratephoneme boundaries for all training data. Hence, the embeddedtraininghasbeenusedinwhichphonemeboundariesarealsodealtas hidden variables, and estimated based on the EM algorithm.Furthermore, in the worse case that the boundary information isnot available, a method called the flat start training is often ap-plied. In this method, initial parameters of HMMs are given bymaking all states of all models equal, and then carry out the em-bedded training. In these situations, we do not have enough priorknowledge to obtain a good initial values for the EM algorithm,and it would converge to one of the local maxima or saddle points Yohei Itaya, Heiga Zen, Yoshihiko Nankaku, Chiyomi Miyajima, Keiichi Tokuda, Tadashi Kitamura |
INTERSPEECH | 3 |
| 2003 | Speech recognition using voice-characteristic-dependent acoustic modelsabstractThis paper proposes a speech recognition technique based on acoustic models considering voice characteristic variations. Context-dependent acoustic models, which are typically triphone HMM, are often used in continuous speech recognition systems. This work hypothesizes that the speaker voice characteristics that humans can perceive by listening are also factors in acoustic variation for construction of acoustic models, and a tree-based clustering technique is also applied to speaker voice characteristics to construct voice-characteristic-dependent acoustic models. In speech recognition using triphone models, the neighboring phonetic context is given from the linguistic-phonetic knowledge in advance; in contrast, the voice characteristics of input speech are unknown in recognition using voice-characteristic-dependent acoustic models. This paper proposes a method of recognizing speech even under conditions where the voice characteristics of the input speech are unknown. The result of a gender-dependent speech recognition experiment shows that the proposed method achieves higher recognition performance in comparison to conventional methods. Hiroyuki Suzuki, Heiga Zen, Yoshihiko Nankaku, Chiyomi Miyajima, Keiichi Tokuda, Tadashi Kitamura |
ICASSP (1) | 3 |
| 2003 | On the use of kernel PCA for feature extraction in speech recognitionabstractThis paper describes an approachfor feature extraction in speech recognition systems using kernel principal componentanalysis (KPCA). This approachconsists in representing speech features as the projection of the extracted speech features mapped into a feature space via a nonlinear mapping onto the principal components. The nonlinear mapping is implicitly performed using the kerneltrick, which is an useful way of not mapping the input space into a featurespace explicitly,makingthis mapping computationally feasible. Better results were obtained by using this approach when compared to the standard technique. Amaro A. de Lima, Heiga Zen, Yoshihiko Nankaku, Chiyomi Miyajima, Keiichi Tokuda, Tadashi Kitamura |
INTERSPEECH | 3 |
| 2000 | Normalized Training for HMM-Based Visual Speech RecognitionabstractThis paper presents an approach to estimating the parameters of continuous density HMMs for visual speech recognition. One of the key issues of image-based visual speech recognition is normalization of lip location and lighting conditions prior to estimating the parameters of HMMs. We presented a normalized training method in which the normalization process is integrated in the model training. This paper extends it for contrast normalization in addition to average-intensity and location normalization. The proposed method provides a theoretically-well-defined algorithm based on a maximum likelihood formulation, hence the likelihood for the training data is guaranteed to increase at each iteration of the normalized training. Experiments on the M2VTS database show that the recognition performance can be significantly improved by the normalized training. Yoshihiko Nankaku, Keiichi Tokuda, Tadashi Kitamura, Takao Kobayashi |
ICIP | 1 |
| 1999 | Intensity- and location-normalized training for HMM-based visual speech recognition
Yoshihiko Nankaku, Keiichi Tokuda, Tadashi Kitamura |
EUROSPEECH | 1 |