Heiga Zen

dblp:42/7014 · DBLP profile ↗
← Back
88ranked-venue papers
25as first author
21since 2021 · last 2025
0000-0002-8959-5471ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 73 · 21 first-author · 16 since 2021Artificial intelligence and machine learning · 56 · 16 first-author · 13 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 SimulTron: On-Device Simultaneous Speech to Speech Translation
abstract
Simultaneous speech-to-speech translation (S2ST) holds the promise of breaking down communication barriers and enabling fluid conversations across languages. However, achieving accurate, real-time translation through mobile devices remains a major challenge. We introduce SimulTron, a novel S2ST architecture designed to tackle this task. SimulTron is a lightweight direct S2ST model that uses the strengths of the Translatotron framework while incorporating key modifications for streaming operation, and an adjustable fixed delay. Our experiments show that SimulTron surpasses Translatotron 2 in offline evaluations. Furthermore, real-time evaluations reveal that SimulTron improves upon the performance achieved by Translatotron 1. Additionally, SimulTron achieves superior BLEU scores and latency compared to previous real-time S2ST method on the MuST-C dataset. Significantly, we have successfully deployed SimulTron on a Pixel 7 Pro device, show its potential for simultaneous S2ST on-device.
Alex Agranovich, Eliya Nachmani, Oleg Rybakov, Yifan Ding 0004, Ye Jia, Nadav Bar, Heiga Zen, Michelle Tadmor Ramanovich
ICASSP7
2024 Translatotron 3: Speech to Speech Translation with Monolingual Data
abstract
This paper presents Translatotron 3, a novel approach to unsupervised direct speech-to-speech translation from monolingual speech-text datasets by combining masked autoencoder, unsupervised embedding mapping, and back-translation. Experimental results in speech-to-speech translation tasks between Spanish and English show that Translatotron 3 outperforms a baseline cascade system, reporting 18.14 BLEU points improvement on the synthesized Unpaired-Conversational dataset. In contrast to supervised approaches that necessitate real paired data, or specialized modeling to replicate para-/non-linguistic information such as pauses, speaking rates, and speaker identity, Translatotron 3 showcases its capability to retain it.
Eliya Nachmani, Alon Levkovitch, Yifan Ding 0004, Chulayuth Asawaroengchai, Heiga Zen, Michelle Tadmor Ramanovich
ICASSP5
2024 Extending Multilingual Speech Synthesis to 100+ Languages without Transcribed Data
abstract
Collecting high-quality studio recordings of audio is challenging, which limits the language coverage of text-to-speech (TTS) systems. This paper proposes a framework for scaling a multilingual TTS model to 100+ languages using found data without supervision. The proposed framework combines speech-text encoder pretraining with unsupervised training using untranscribed speech and unspoken text data sources, thereby leveraging massively multilingual joint speech and text representation learning. Without any transcribed speech in a new language, this TTS model can generate intelligible speech in >30 unseen languages (CER difference of <10% to ground truth). With just 15 minutes of transcribed, found data, we can reduce the intelligibility difference to 1% or less from the ground-truth, and achieve naturalness scores that match the ground-truth in several languages.
Takaaki Saeki, Gary Wang, Nobuyuki Morioka, Isaac Elias, Kyle Kastner, Andrew Rosenberg, Bhuvana Ramabhadran, Heiga Zen, Françoise Beaufays, Hadar Shemtov
ICASSP8
2024 FLEURS-R: A Restored Multilingual Speech Corpus for Generation Tasks
Yuma Koizumi, Shigeki Karita, Heiga Zen, Jason Riesa, Haruko Ishikawa, Michiel Bacchiani
INTERSPEECH4
2024 Geometric-Averaged Preference Optimization for Soft Preference Labels
abstract
Many algorithms for aligning LLMs with human preferences assume that human preferences are binary and deterministic. However, human preferences can vary across individuals, and therefore should be represented distributionally. In this work, we introduce the distributional soft preference labels and improve Direct Preference Optimization (DPO) with a weighted geometric average of the LLM output likelihood in the loss function. This approach adjusts the scale of learning loss based on the soft labels such that the loss would approach zero when the responses are closer to equally preferred. This simple modification can be easily applied to any DPO-based methods and mitigate over-optimization and objective mismatch, which prior works suffer from. Our experiments simulate the soft preference labels with AI feedback from LLMs and demonstrate that geometric averaging consistently improves performance on standard benchmarks for alignment research. In particular, we observe more preferable responses than binary labels and significant improvements where modestly-confident labels are in the majority.
Hiroki Furuta, Kuang-Huei Lee, Shixiang Gu, Yutaka Matsuo, Aleksandra Faust, Heiga Zen, Izzeddin Gur
NeurIPS6
2023 Virtuoso: Massive Multilingual Speech-Text Joint Semi-Supervised Learning for Text-to-Speech
abstract
This paper proposes Virtuoso, a massively multilingual speech–text joint semi-supervised learning framework for text-to-speech synthesis (TTS) models. Existing multilingual TTS typically supports tens of languages, which are a small fraction of the thousands of languages in the world. One difficulty to scale multilingual TTS to hundreds of languages is collecting high-quality speech–text paired data in low-resource languages. This study extends Maestro, a speech–text joint pretraining framework for automatic speech recognition (ASR), to speech generation tasks. To train a TTS model from various types of speech and text data, different training schemes are designed to handle supervised (paired TTS and ASR data) and unsupervised (untranscribed speech and unspoken text) datasets. Experimental evaluation shows that 1) multilingual TTS models trained on Virtuoso can achieve significantly better naturalness and intelligibility than baseline ones in seen languages, and 2) they can synthesize reasonably intelligible and naturally sounding speech for unseen languages where no high-quality paired TTS data is available.
Takaaki Saeki, Heiga Zen, Zhehuai Chen, Nobuyuki Morioka, Gary Wang, Yu Zhang 0033, Ankur Bapna, Andrew Rosenberg, Bhuvana Ramabhadran
ICASSP2
2023 Lightweight, Multi-Speaker, Multi-Lingual Indic Text-to-Speech
abstract
The Lightweight, Multi-speaker, Multi-lingual Indic Text-to-Speech (LIMMITS’23) challenge is organized as part of the ICASSP 2023 signal processing grand challenge. LIMMITS’23 aims at the development of a lightweight, multi-speaker, multi-lingual Text to Speech (TTS) model using datasets in Marathi, Hindi, and Telugu. The challenge encourages the advancement of TTS in Indian Languages as well as the development of techniques involved in TTS data selection and model compression. The 3 tracks of LIMMITS’23 have provided an opportunity for various researchers and practitioners around the world to explore the state of the art in TTS research.
Abhayjeet Singh, Amala Nagireddi, Deekshitha G, Jesuraja Bandekar, Roopa R., Sandhya Badiger, Sathvik Udupa, Prasanta Kumar Ghosh, Hema A. Murthy, Heiga Zen, Pranaw Kumar, Kamal Kant, Amol Bole, Bira Chandra Singh, Keiichi Tokuda, Mark Hasegawa-Johnson, Philipp Olbrich
ICASSP10
2023 LibriTTS-R: A Restored Multi-Speaker Text-to-Speech Corpus
Yuma Koizumi, Heiga Zen, Shigeki Karita, Yifan Ding 0004, Kohei Yatabe, Nobuyuki Morioka, Michiel Bacchiani, Yu Zhang 0033, Wei Han 0002, Ankur Bapna
INTERSPEECH2
2023 Extracting representative subset from extensive text data for training pre-trained language models
abstract
This paper investigates the existence of a representative subset obtained from a large original dataset that can achieve the same performance level obtained using the entire dataset in the context of training neural language models. We employ the likelihood-based scoring method based on two distinct types of pre-trained language models to select a representative subset. We conduct our experiments on widely used 17 natural language processing datasets with 24 evaluation metrics. The experimental results showed that the representative subset obtained using the likelihood difference score can achieve the 90% performance level even when the size of the dataset is reduced to approximately two to three orders of magnitude smaller than the original dataset. We also compare the performance with the models trained with the same amount of subset selected randomly to show the effectiveness of the representative subset.
Jun Suzuki 0001, Heiga Zen, Hideto Kazawa
Inf. Process. Manag.2
2023 Guest Editorial: Special Issue on Affective Speech and Language Synthesis, Generation, and Conversion
abstract
The papers in this special section focus on affective speech and language synthesis, generation, and conversion. As an inseparable and crucial part of spoken language, emotions play a substantial role in human-human and human-technology conversation. They convey information about a person’s needs, how one feels about the objectives of a conversation, the trustworthiness of one’s verbal communication, and more. Accordingly, substantial efforts have been made to generate affective text and speech for conversational AI, artificial storytelling, and machine translation. Similarly, there is a push for converting the affect in text and speech, ideally, in real-time and fully preserving intelligibility, e. g., to hide one’s emotion, for creative applications and in entertainment, or even to augment training data for affect analyzing AI.
Shahin Amiriparian, Björn W. Schuller, Nabiha Asghar, Heiga Zen, Felix Burkhardt
IEEE Trans. Affect. Comput.4
2022 MAESTRO: Matched Speech Text Representations through Modality Matching
abstract
We present Maestro, a self-supervised training method to unify representations learnt from speech and text modalities.Self-supervised learning from speech signals aims to learn the latent structure inherent in the signal, while self-supervised learning from text attempts to capture lexical information.Learning aligned representations from unpaired speech and text sequences is a challenging task.Previous work either implicitly enforced the representations learnt from these two modalities to be aligned in the latent space through multitasking and parameter sharing or explicitly through conversion of modalities via speech synthesis.While the former suffers from interference between the two modalities, the latter introduces additional complexity.In this paper, we propose Maestro, a novel algorithm to learn unified representations from both these modalities simultaneously that can transfer to diverse downstream tasks such as Automated Speech Recognition (ASR) and Speech Translation (ST).Maestro learns unified representations through sequence alignment, duration prediction and matching embeddings in the learned space through an aligned masked-language model loss.We establish a new state-of-the-art (SOTA) on VoxPopuli multilingual ASR with a 8% relative reduction in Word Error Rate (WER), multidomain SpeechStew ASR (3.7% relative) and 21 languages to English multilingual ST on CoVoST 2 with an improvement of 2.8 BLEU averaged over 21 languages.
Zhehuai Chen, Yu Zhang 0033, Andrew Rosenberg, Bhuvana Ramabhadran, Pedro J. Moreno 0001, Ankur Bapna, Heiga Zen
INTERSPEECH7
2022 Training Text-To-Speech Systems From Synthetic Data: A Practical Approach For Accent Transfer Tasks
abstract
Transfer tasks in text-to-speech (TTS) synthesis -where one or more aspects of the speech of one set of speakers is transferred to another set of speakers that do not feature these aspects originally -remains a challenging task.One of the challenges is that models that have high-quality transfer capabilities can have issues in stability, making them impractical for user-facing critical tasks.This paper demonstrates that transfer can be obtained by training a robust TTS system on data generated by a less robust TTS system designed for a high-quality transfer task; in particular, a CHiVE-BERT monolingual TTS system is trained on the output of a Tacotron model designed for accent transfer.While some quality loss is inevitable with this approach, experimental results show that the models trained on synthetic data this way can produce high quality audio displaying accent transfer, while preserving speaker characteristics such as speaking style.
Lev Finkelstein, Heiga Zen, Norman Casagrande, Chun-an Chan, Ye Jia, Tom Kenter, Alexey Petelin, Jonathan Shen, Vincent Wan, Yu Zhang 0033, Rob Clark
INTERSPEECH2
2022 SpecGrad: Diffusion Probabilistic Model based Neural Vocoder with Adaptive Noise Spectral Shaping
abstract
Neural vocoder using denoising diffusion probabilistic model (DDPM) has been improved by adaptation of the diffusion noise distribution to given acoustic features.In this study, we propose SpecGrad that adapts the diffusion noise so that its timevarying spectral envelope becomes close to the conditioning log-mel spectrogram.This adaptation by time-varying filtering improves the sound quality especially in the high-frequency bands.It is processed in the time-frequency domain to keep the computational cost almost the same as the conventional DDPMbased neural vocoders.Experimental results showed that Spec-Grad generates higher-fidelity speech waveform than conventional DDPM-based neural vocoders in both analysis-synthesis and speech enhancement scenarios.Audio demos are available at wavegrad.github.io/specgrad/.
Yuma Koizumi, Heiga Zen, Kohei Yatabe, Nanxin Chen, Michiel Bacchiani
INTERSPEECH2
2022 CVSS Corpus and Massively Multilingual Speech-to-Speech Translation
abstract
We introduce CVSS, a massively multilingual-to-English speech-to-speech translation (S2ST) corpus, covering sentence-level parallel S2ST pairs from 21 languages into English. CVSS is derived from the Common Voice speech corpus and the CoVoST 2 speech-to-text translation (ST) corpus, by synthesizing the translation text from CoVoST 2 into speech using state-of-the-art TTS systems. Two versions of translation speech in English are provided: 1) CVSS-C: All the translation speech is in a single high-quality canonical voice; 2) CVSS-T: The translation speech is in voices transferred from the corresponding source speech. In addition, CVSS provides normalized translation text which matches the pronunciation in the translation speech. On each version of CVSS, we built baseline multilingual direct S2ST models and cascade S2ST models, verifying the effectiveness of the corpus. To build strong cascade S2ST baselines, we trained an ST model on CoVoST 2, which outperforms the previous state-of-the-art trained on the corpus without extra data by 5.8 BLEU. Nevertheless, the performance of the direct S2ST models approaches the strong cascade baselines when trained from scratch, and with only 0.1 or 0.7 BLEU difference on ASR transcribed translation when initialized from matching ST models.
Ye Jia, Michelle Tadmor Ramanovich, Heiga Zen
LREC4
2022 Wavefit: an Iterative and Non-Autoregressive Neural Vocoder Based on Fixed-Point Iteration
abstract
Denoising diffusion probabilistic models (DDPMs) and generative adversarial networks (GANs) are popular generative models for neural vocoders. The DDPMs and GANs can be characterized by the iterative denoising framework and adversarial training, respectively. This study proposes a fast and high-quality neural vocoder called WaveFit, which integrates the essence of GANs into a DDPM-like iterative framework based on fixed-point iteration. WaveFit iteratively denoises an input signal, and trains a deep neural network (DNN) for minimizing an adversarial loss calculated from intermediate outputs at all iterations. Subjective (side-by-side) listening tests showed no statistically significant differences in naturalness between human natural speech and those synthesized by WaveFit with five iterations. Furthermore, the inference speed of WaveFit was more than 240 times faster than WaveRNN. Audio demos are available at google.github.io/df-conformer/wavefit/.
Yuma Koizumi, Kohei Yatabe, Heiga Zen, Michiel Bacchiani
SLT3
2021 Parallel Tacotron: Non-Autoregressive and Controllable TTS
abstract
Although neural end-to-end text-to-speech models can synthesize highly natural speech, there is still room for improvements to its efficiency and naturalness. This paper proposes a non-autoregressive neural text-to-speech model augmented with a variational autoencoder-based residual encoder. This model, called Parallel Tacotron, is highly parallelizable during both training and inference, allowing efficient synthesis on modern parallel hardware. The use of the variational autoencoder relaxes the one-to-many mapping nature of the text-to-speech problem and improves naturalness. To further improve the naturalness, we use lightweight convolutions, which can efficiently capture local contexts, and introduce an iterative spectrogram loss inspired by iterative refinement. Experimental results show that Parallel Tacotron matches a strong autoregressive baseline in subjective evaluations with significantly decreased inference time.
Isaac Elias, Heiga Zen, Jonathan Shen, Yu Zhang 0033, Ye Jia, Ron J. Weiss
ICASSP2
2021 WaveGrad: Estimating Gradients for Waveform Generation
Nanxin Chen, Yu Zhang 0033, Heiga Zen, Ron J. Weiss, Mohammad Norouzi 0002
ICLR3
2021 Semi-Supervision in ASR: Sequential MixMatch and Factorized TTS-Based Augmentation
Zhehuai Chen, Andrew Rosenberg, Yu Zhang 0033, Heiga Zen, Mohammadreza Ghodsi, Jesse Emond, Gary Wang, Bhuvana Ramabhadran, Pedro J. Moreno 0001
Interspeech4
2021 WaveGrad 2: Iterative Refinement for Text-to-Speech Synthesis
abstract
This paper introduces WaveGrad 2, a non-autoregressive generative model for text-to-speech synthesis. WaveGrad 2 is trained to estimate the gradient of the log conditional density of the waveform given a phoneme sequence. The model takes an input phoneme sequence, and through an iterative refinement process, generates an audio waveform. This contrasts to the original WaveGrad vocoder which conditions on mel-spectrogram features, generated by a separate model. The iterative refinement process starts from Gaussian noise, and through a series of refinement steps (e.g., 50 steps), progressively recovers the audio sequence. WaveGrad 2 offers a natural way to trade-off between inference speed and sample quality, through adjusting the number of refinement steps. Experiments show that the model can generate high fidelity audio, approaching the performance of a state-of-the-art neural TTS system. We also report various ablation studies over different model configurations. Audio samples are available at this https URL.
Nanxin Chen, Yu Zhang 0033, Heiga Zen, Ron J. Weiss, Mohammad Norouzi 0002, Najim Dehak
Interspeech3
2021 Parallel Tacotron 2: A Non-Autoregressive Neural TTS Model with Differentiable Duration Modeling
abstract
This paper introduces Parallel Tacotron 2, a non-autoregressive neural text-to-speech model with a fully differentiable duration model which does not require supervised duration signals.The duration model is based on a novel attention mechanism and an iterative reconstruction loss based on Soft Dynamic Time Warping, this model can learn token-frame alignments as well as token durations automatically.Experimental results show that Parallel Tacotron 2 outperforms baselines in subjective naturalness in several diverse multi speaker evaluations.
Isaac Elias, Heiga Zen, Jonathan Shen, Yu Zhang 0033, Ye Jia, R. J. Skerry-Ryan
Interspeech2
2021 PnG BERT: Augmented BERT on Phonemes and Graphemes for Neural TTS
abstract
This paper introduces PnG BERT, a new encoder model for neural TTS.This model is augmented from the original BERT model, by taking both phoneme and grapheme representations of text as input, as well as the word-level alignment between them.It can be pre-trained on a large text corpus in a selfsupervised manner, and fine-tuned in a TTS task.Experimental results show that a neural TTS model using a pre-trained PnG BERT as its encoder yields more natural prosody and more accurate pronunciation than a baseline model using only phoneme input with no pre-training.Subjective side-by-side preference evaluations show that raters have no statistically significant preference between the speech synthesized using a PnG BERT and ground truth recordings from professional speakers.
Ye Jia, Heiga Zen, Jonathan Shen, Yu Zhang 0033
Interspeech2
2020 Generating Diverse and Natural Text-to-Speech Samples Using a Quantized Fine-Grained VAE and Autoregressive Prosody Prior
abstract
Recent neural text-to-speech (TTS) models with fine-grained latent features enable precise control of the prosody of synthesized speech. Such models typically incorporate a fine-grained variational autoencoder (VAE) structure, extracting latent features at each input token (e.g., phonemes). However, generating samples with the standard VAE prior often results in unnatural and discontinuous speech, with dramatic prosodic variation between tokens. This paper proposes a sequential prior in a discrete latent space which can generate more naturally sounding samples. This is accomplished by discretizing the latent features using vector quantization (VQ), and separately training an autoregressive (AR) prior model over the result. We evaluate the approach using listening tests, objective metrics of automatic speech recognition (ASR) performance, and measurements of prosody attributes. Experimental results show that the proposed model significantly improves the naturalness in random sample generation. Furthermore, initial experiments demonstrate that randomly sampling from the proposed model can be used as data augmentation to improve the ASR performance.
Guangzhi Sun, Yu Zhang 0033, Ron J. Weiss, Yuan Cao 0007, Heiga Zen, Andrew Rosenberg, Bhuvana Ramabhadran
ICASSP5
2020 Fully-Hierarchical Fine-Grained Prosody Modeling For Interpretable Speech Synthesis
abstract
This paper proposes a hierarchical, fine-grained and interpretable latent variable model for prosody based on the Tacotron 2 text-to-speech model. It achieves multi-resolution modeling of prosody by conditioning finer level representations on coarser level ones. Additionally, it imposes hierarchical conditioning across all latent dimensions using a conditional variational auto-encoder (VAE) with an auto-regressive structure. Evaluation of reconstruction performance illustrates that the new structure does not degrade the model while allowing better interpretability. Interpretations of prosody attributes are provided together with the comparison between word-level and phone-level prosody representations. Moreover, both qualitative and quantitative evaluations are used to demonstrate the improvement in the disentanglement of the latent dimensions.
Guangzhi Sun, Yu Zhang 0033, Ron J. Weiss, Yuan Cao 0007, Heiga Zen
ICASSP5
2019 Sample Efficient Adaptive Text-to-Speech
Yutian Chen 0001, Yannis M. Assael, Brendan Shillingford, David Budden, Scott E. Reed, Heiga Zen, Luis C. Cobo, Andrew Trask, Ben Laurie, Caglar Gulcehre, Aäron van den Oord, Oriol Vinyals, Nando de Freitas
ICLR (Poster)6
2019 Hierarchical Generative Modeling for Controllable Speech Synthesis
Wei-Ning Hsu, Yu Zhang 0033, Ron J. Weiss, Heiga Zen, Yuxuan Wang 0002, Yuan Cao 0007, Ye Jia, Jonathan Shen, Patrick Nguyen, Ruoming Pang
ICLR (Poster)4
2019 LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech
abstract
This paper introduces a new speech corpus called "LibriTTS" designed for text-to-speech use.It is derived from the original audio and text materials of the LibriSpeech corpus, which has been used for training and evaluating automatic speech recognition systems.The new corpus inherits desired properties of the LibriSpeech corpus while addressing a number of issues which make LibriSpeech less than ideal for text-to-speech work.The released corpus consists of 585 hours of speech data at 24kHz sampling rate from 2,456 speakers and the corresponding texts.Experimental results show that neural end-to-end TTS models trained from the LibriTTS corpus achieved above 4.0 in mean opinion scores in naturalness in five out of six evaluation speakers.
Heiga Zen, Viet Dang, Rob Clark, Yu Zhang 0033, Ron J. Weiss, Ye Jia
INTERSPEECH1
2019 Learning to Speak Fluently in a Foreign Language: Multilingual Speech Synthesis and Cross-Language Voice Cloning
abstract
We present a multispeaker, multilingual text-to-speech (TTS) synthesis model based on Tacotron that is able to produce high quality speech in multiple languages.Moreover, the model is able to transfer voices across languages, e.g.synthesize fluent Spanish speech using an English speaker's voice, without training on any bilingual or parallel examples.Such transfer works across distantly related languages, e.g.English and Mandarin.Critical to achieving this result are: 1. using a phonemic input representation to encourage sharing of model capacity across languages, and 2. incorporating an adversarial loss term to encourage the model to disentangle its representation of speaker identity (which is perfectly correlated with language in the training data) from the speech content.Further scaling up the model by training on multiple speakers of each language, and incorporating an autoencoding input to help stabilize attention during training, results in a model which can be used to consistently synthesize intelligible speech for training speakers in all languages seen during training, and in native or foreign accents.
Yu Zhang 0033, Ron J. Weiss, Heiga Zen, R. J. Skerry-Ryan, Ye Jia, Andrew Rosenberg, Bhuvana Ramabhadran
INTERSPEECH3
2018 Parallel WaveNet: Fast High-Fidelity Speech Synthesis
abstract
The recently-developed WaveNet architecture is the current state of the art in realistic speech synthesis, consistently rated as more natural sounding for many different languages than any previous system. However, because WaveNet relies on sequential generation of one audio sample at a time, it is poorly suited to today’s massively parallel computers, and therefore hard to deploy in a real-time production setting. This paper introduces Probability Density Distillation, a new method for training a parallel feed-forward network from a trained WaveNet with no significant difference in quality. The resulting system is capable of generating high-fidelity speech samples at more than 20 times faster than real-time, a 1000x speed up relative to the original WaveNet, and capable of serving multiple English and Japanese voices in a production setting.
Aäron van den Oord, Yazhe Li, Igor Babuschkin, Karen Simonyan, Oriol Vinyals, Koray Kavukcuoglu, George van den Driessche 0002, Edward Lockhart, Luis C. Cobo, Florian Stimberg, Norman Casagrande, Dominik Grewe, Seb Noury, Sander Dieleman, Erich Elsen, Nal Kalchbrenner, Heiga Zen, Alex Graves, Helen King, Tom Walters, Daniel Belov, Demis Hassabis
ICML17
2018 Sequence-to-sequence Neural Network Model with 2D Attention for Learning Japanese Pitch Accents
Antoine Bruguier, Heiga Zen, Arkady Arkhangorodsky
INTERSPEECH2
2016 Directly modeling voiced and unvoiced components in speech waveforms by neural networks
abstract
This paper proposes a novel acoustic model based on neural networks for statistical parametric speech synthesis. The neural network outputs parameters of a non-zero mean Gaussian process, which defines a probability density function of a speech waveform given linguistic features. The mean and covariance functions of the Gaussian process represent deterministic (voiced) and stochastic (unvoiced) components of a speech waveform, whereas the previous approach considered the unvoiced component only. Experimental results show that the proposed approach can generate speech waveforms approximating natural speech waveforms.
Keiichi Tokuda, Heiga Zen
ICASSP2
2016 Multi-Language Multi-Speaker Acoustic Modeling for LSTM-RNN Based Statistical Parametric Speech Synthesis
Bo Li 0028, Heiga Zen
INTERSPEECH2
2016 Fast, Compact, and High Quality LSTM-RNN Based Statistical Parametric Speech Synthesizers for Mobile Devices
abstract
Acoustic models based on long short-term memory recurrent neural networks (LSTM-RNNs) were applied to statistical parametric speech synthesis (SPSS) and showed significant improvements in naturalness and latency over those based on hidden Markov models (HMMs). This paper describes further optimizations of LSTM-RNN-based SPSS for deployment on mobile devices; weight quantization, multi-frame inference, and robust inference using an ε-contaminated Gaussian loss function. Experimental results in subjective listening tests show that these optimizations can make LSTM-RNN-based SPSS comparable to HMM-based SPSS in runtime speed while maintaining naturalness. Evaluations between LSTM-RNN- based SPSS and HMM-driven unit selection speech synthesis are also presented.
Heiga Zen, Yannis Agiomyrgiannakis, Niels Egberts, Fergus Henderson, Przemyslaw Szczepaniak
INTERSPEECH1
2015 Directly modeling speech waveforms by neural networks for statistical parametric speech synthesis
abstract
This paper proposes a novel approach for directly-modeling speech at the waveform level using a neural network. This approach uses the neural network-based statistical parametric speech synthesis framework with a specially designed output layer. As acoustic feature extraction is integrated to acoustic model training, it can overcome the limitations of conventional approaches, such as two-step (feature extraction and acoustic modeling) optimization, use of spectra rather than waveforms as targets, use of overlapping and shifting frames as unit, and fixed decision tree structure. Experimental results show that the proposed approach can directly maximize the likelihood defined at the waveform domain.
Keiichi Tokuda, Heiga Zen
ICASSP2
2015 Unidirectional long short-term memory recurrent neural network with recurrent output layer for low-latency speech synthesis
abstract
Long short-term memory recurrent neural networks (LSTM-RNNs) have been applied to various speech applications including acoustic modeling for statistical parametric speech synthesis. One of the concerns for applying them to text-to-speech applications is its effect on latency. To address this concern, this paper proposes a low-latency, streaming speech synthesis architecture using unidirectional LSTM-RNNs with a recurrent output layer. The use of unidirectional RNN architecture allows frame-synchronous streaming inference of output acoustic features given input linguistic features. The recurrent output layer further encourages smooth transition between acoustic features at consecutive frames. Experimental results in subjective listening tests show that the proposed architecture can synthesize natural sounding speech without requiring utterance-level batch processing.
Heiga Zen, Hasim Sak
ICASSP1
2014 Deep mixture density networks for acoustic modeling in statistical parametric speech synthesis
abstract
Statistical parametric speech synthesis (SPSS) using deep neural networks (DNNs) has shown its potential to produce naturally-sounding synthesized speech. However, there are limitations in the current implementation of DNN-based acoustic modeling for speech synthesis, such as the unimodal nature of its objective function and its lack of ability to predict variances. To address these limitations, this paper investigates the use of a mixture density output layer. It can estimate full probability density functions over real-valued output features conditioned on the corresponding input features. Experimental results in objective and subjective evaluations show that the use of the mixture density output layer improves the prediction accuracy of acoustic features and the naturalness of the synthesized speech.
Heiga Zen, Andrew W. Senior
ICASSP1
2013 Statistical parametric speech synthesis using deep neural networks
abstract
Conventional approaches to statistical parametric speech synthesis typically use decision tree-clustered context-dependent hidden Markov models (HMMs) to represent probability densities of speech parameters given texts. Speech parameters are generated from the probability densities to maximize their output probabilities, then a speech waveform is reconstructed from the generated parameters. This approach is reasonably effective but has a couple of limitations, e.g. decision trees are inefficient to model complex context dependencies. This paper examines an alternative scheme that is based on a deep neural network (DNN). The relationship between input texts and their acoustic realizations is modeled by a DNN. The use of the DNN can address some limitations of the conventional approach. Experimental results show that the DNN-based systems outperformed the HMM-based systems with similar numbers of parameters.
Heiga Zen, Andrew W. Senior, Mike Schuster
ICASSP1
2013 Speech Synthesis Based on Hidden Markov Models
abstract
This paper gives a general overview of hidden Markov model (HMM)-based speech synthesis, which has recently been demonstrated to be very effective in synthesizing speech. The main advantage of this approach is its flexibility in changing speaker identities, emotions, and speaking styles. This paper also discusses the relation between the HMM-based approach and the more conventional unit-selection approach that has dominated over the last decades. Finally, advanced techniques for future developments are described.
Keiichi Tokuda, Yoshihiko Nankaku, Tomoki Toda, Heiga Zen, Junichi Yamagishi, Keiichiro Oura
Proc. IEEE4
2013 Autoregressive Models for Statistical Parametric Speech Synthesis
abstract
We propose using the autoregressive hidden Markov model (HMM) for speech synthesis. The autoregressive HMM uses the same model for parameter estimation and synthesis in a consistent way, in contrast to the standard approach to statistical parametric speech synthesis. It supports easy and efficient parameter estimation using expectation maximization, in contrast to the trajectory HMM. At the same time its similarities to the standard approach allow use of established high quality synthesis algorithms such as speech parameter generation considering global variance. The autoregressive HMM also supports a speech parameter generation algorithm not available for the standard approach or the trajectory HMM and which has particular advantages in the domain of real-time, low latency synthesis. We show how to do efficient parameter estimation and synthesis with the autoregressive HMM and look at some of the similarities and differences between the standard approach, the trajectory HMM and the autoregressive HMM. We compare the three approaches in subjective and objective evaluations. We also systematically investigate which choices of parameters such as autoregressive order and number of states are optimal for the autoregressive HMM.
Matt Shannon, Heiga Zen, William J. Byrne
IEEE Trans. Speech Audio Process.2
2012 Cepstral analysis based on the glimpse proportion measure for improving the intelligibility of HMM-based synthetic speech in noise
abstract
In this paper we introduce a new cepstral coefficient extraction method based on an intelligibility measure for speech in noise, the Glimpse Proportion measure. This new method aims to increase the intelligibility of speech in noise by modifying the clean speech, and has applications in scenarios such as public announcement and car navigation systems. We first explain how the Glimpse Proportion measure operates and further show how we approximated it to integrate it into an existing spectral envelope parameter extraction method commonly used in the HMM-based speech synthesis framework. We then demonstrate how this new method changes the modelled spectrum according to the characteristics of the noise and show results for a listening test with vocoded and HMM-based synthetic speech. The test indicates that the proposed method can significantly improve intelligibility of synthetic speech in speech shaped noise.
Cassia Valentini-Botinhao, Ranniery Maia, Junichi Yamagishi, Simon King 0001, Heiga Zen
ICASSP5
2012 Combining multiple high quality corpora for improving HMM-TTS
abstract
The most reliable way to build synthetic voices for end-products is to start with high quality recordings from professional voice talents. This paper describes the application of average voice models (AVMs) and a novel application of cluster adaptive training (CAT) to combine a small number of these high quality corpora to make best use of them and improve overall voice quality in hidden Markov model based text-to-speech (HMMTTS) systems. It is shown that integrated training by both CAT and AVM approaches, yields better sounding voices than speaker dependent modelling. It is also shown that CAT has an advantage over AVMs when adapting to a new speaker. Given a limited amount of adaptation data CAT maintains a much higher voice quality even when adapted to tiny amounts of speech.
Vincent Wan, Javier Latorre, K. K. Chin, Langzhou Chen, Mark J. F. Gales, Heiga Zen, Kate M. Knill, Masami Akamine
INTERSPEECH6
2012 Statistical Parametric Speech Synthesis Based on Speaker and Language Factorization
abstract
An increasingly common scenario in building speech synthesis and recognition systems is training on inhomogeneous data. This paper proposes a new framework for estimating hidden Markov models on data containing both multiple speakers and multiple languages. The proposed framework, speaker and language factorization, attempts to factorize speaker-/language-specific characteristics in the data and then model them using separate transforms. Language-specific factors in the data are represented by transforms based on cluster mean interpolation with cluster-dependent decision trees. Acoustic variations caused by speaker characteristics are handled by transforms based on constrained maximum-likelihood linear regression. Experimental results on statistical parametric speech synthesis show that the proposed framework enables data from multiple speakers in different languages to be used to: train a synthesis system; synthesize speech in a language using speaker characteristics estimated in a different language; and adapt to a new language.
Heiga Zen, Norbert Braunschweiler, Sabine Buchholz, Mark J. F. Gales, Kate M. Knill, Sacha Krstulovic, Javier Latorre
IEEE Trans. Speech Audio Process.1
2012 Product of Experts for Statistical Parametric Speech Synthesis
abstract
Multiple acoustic models are often combined in statistical parametric speech synthesis. Both linear and non-linear functions of an observation sequence are used as features to be modeled. This paper shows that this combination of multiple acoustic models can be expressed as a product of experts (PoE); the likelihoods from the models are scaled, multiplied together, and then normalized. Normally these models are individually trained and only combined at the synthesis stage. This paper discusses a more consistent PoE framework where the models are jointly trained. A training algorithm for PoEs based on linear feature functions and Gaussian experts is derived by generalizing the training algorithm for trajectory HMMs. However for non-linear feature functions or non-Gaussian experts this is not possible, so a scheme based on contrastive divergence learning is described. Experimental results show that the PoE framework provides both a mathematically elegant way to train multiple acoustic models jointly and significant improvements in the quality of the synthesized speech.
Heiga Zen, Mark J. F. Gales, Yoshihiko Nankaku, Keiichi Tokuda
IEEE Trans. Speech Audio Process.1
2011 Decision tree-based context clustering based on cross validation and hierarchical priors
abstract
The standard, ad-hoc stopping criteria used in decision tree-based context clustering are known to be sub-optimal and require parameters to be tuned. This paper proposes a new approach for decision tree-based context clustering based on cross validation and hierarchical priors. Combination of cross validation and hierarchical priors within decision tree-based context clustering offers better model selection and more robust parameter estimation than conventional approaches, with no tuning parameters. Experimental results on HMM-based speech synthesis show that the proposed approach achieved significant improvements in naturalness of synthesized speech over the conventional approaches.
Heiga Zen, Mark J. F. Gales
ICASSP1
2011 Estimation of Window Coefficients for Dynamic Feature Extraction for HMM-Based Speech Synthesis
Linghui Chen, Yoshihiko Nankaku, Heiga Zen, Keiichi Tokuda, Zhen-Hua Ling, Li-Rong Dai 0001
INTERSPEECH3
2011 Multipulse Sequences for Residual Signal Modeling
abstract
In source-filter models of speech production, the residual signal - what remains after passing the speech signal through the inverse filter - contains important information for the generation of naturally sounding re-synthesized speech. Typically, the voiced regions of residual signals are regarded as a mixture of glottal pulse and noise. This paper introduces a novel approach to represent the noise component of voiced regions of residual signals through autoregressive filtering of multipulse sequences. The positions and amplitudes of the non-zero samples of these multipulse signals are optimized through a closed-loop procedure. The method in question is applied to excitation modeling in statistical parametric synthesis. Experimental results indicate that the use of multipulse-based noise component construction eliminates the necessity of run-time ad hoc procedures such as high-pass filtering and time modulation, common on excitation models for statistical parametric synthesizers, with no loss of synthesized speech quality. Copyright © 2011 ISCA.
Ranniery Maia, Heiga Zen, Kate M. Knill, Mark J. F. Gales, Sabine Buchholz
INTERSPEECH2
2011 Gaussian Process Experts for Voice Conversion
abstract
Conventional approaches to voice conversion typically use a GMM to represent the joint probability density of source and target features. This model is then used to perform spectral conversion between speakers. This approach is reasonably effective but can be prone to overfitting and oversmoothing of the target spectra. This paper proposes an alternative scheme that uses a collection of Gaussian process experts to perform the spectral conversion. Gaussian processes are robust to overfitting and oversmoothing and can predict the target spectra more accurately. Experimental results indicate that the objective performance of voice conversion can be improved using the proposed approach. Copyright © 2011 ISCA.
Nicholas Pilkington, Heiga Zen, Mark J. F. Gales
INTERSPEECH2
2011 The Effect of Using Normalized Models in Statistical Speech Synthesis
abstract
The standard approach to HMM-based speech synthesis is inconsistent in the enforcement of the deterministic constraints between static and dynamic features. The trajectory HMM and autoregressive HMM have been proposed as normalized models which rectify this inconsistency. This paper investigates the practical effects of using these normalized models, and examines the strengths and weaknesses of the different models as probabilistic models of speech. The most striking difference observed is that the standard approach greatly underestimates predictive variance. We argue that the normalized models have better predictive distributions than the standard approach, but that all the models we consider are still far from satisfactory probabilistic models of speech. We also present evidence that better intra-frame correlation modelling goes some way towards improving existing normalized models. Index terms: HMM-based speech synthesis, acoustic modelling, autoregressive HMM, trajectory HMM, normalization 1.
Matt Shannon, Heiga Zen, William J. Byrne
INTERSPEECH2
2011 Context adaptive training with factorized decision trees for HMM-based statistical parametric speech synthesis
Kai Yu 0004, Heiga Zen, François Mairesse, Steve J. Young
Speech Commun.2
2011 Continuous Stochastic Feature Mapping Based on Trajectory HMMs
abstract
This paper proposes a technique of continuous stochastic feature mapping based on trajectory hidden Markov models (HMMs), which have been derived from HMMs by imposing explicit relationships between static and dynamic features. Although Gaussian mixture model (GMM)- or HMM-based feature-mapping techniques work effectively, their accuracy occasionally degrades due to inappropriate dynamic characteristics caused by frame-by-frame mapping. While the use of dynamic-feature constraints at the mapping stage can alleviate this problem, it also introduces inconsistencies between training and mapping. The technique we propose can eliminate these inconsistencies while retaining the benefits of using dynamic-feature constraints, and it offers entire sequence-level transformation rather than frame-by-frame mapping. The results obtained from speaker-conversion, acoustic-to-articulatory inversion-mapping, and noise-compensation experiments demonstrated that our new approach outperformed the conventional one.
Heiga Zen, Yoshihiko Nankaku, Keiichi Tokuda
IEEE Trans. Speech Audio Process.1
2010 Statistical parametric speech synthesis based on product of experts
abstract
Multiple-level acoustic models (AMs) are often combined in statistical parametric speech synthesis. Both linear and non-linear functions of the observation sequence are used as features in these AMs. This combination of multiple-level AMs can be expressed as a product of experts (PoE); the likelihoods from the AMs are scaled, multiplied together and then normalized. Currently these multiple-level AMs are individually trained and only combined at the synthesis stage. This paper discusses a more consistent PoE framework where the AMs are jointly trained. A generalization of trajectory HMM training can be used for multiple-level Gaussian AMs based on linear functions. However for the non-linear case this is not possible, so a scheme based on contrastive divergence learning is described. Experimental results show that the proposed technique provides both a mathematically elegant way to train multiple-level AMs and statistically significant improvements in the quality of synthesized speech.
Heiga Zen, Mark J. F. Gales, Yoshihiko Nankaku, Keiichi Tokuda
ICASSP1
2010 Training a parametric-based logF0 model with the minimum generation error criterion
abstract
This paper describes an approach for improving a statistical parametric-based logF0 model using minimum-generationerror (MGE) training. Compared with the previous scheme based on decision tree clustering, MGE allows the minimisation of the error in the generated logF0 to take into account not only each cluster by itself, but also the way in which the clusters interact with each other in the generation of the F0 over the whole sentence. Moreover, the “weights” of each component of the model, which previously were adjusted manually, are optimized automatically by the MGE training during the re-estimation of the model covariances. Objective evaluation indicated that, although the logF0 contours generated by the models trained with MGE have approximately the same root mean square error and correlation factor as those generated with the baseline models, they present a higher dynamic range. The subjective evaluation shows a small but significant preference for the system trained with MGE.
Javier Latorre, Mark J. F. Gales, Heiga Zen
INTERSPEECH3
2010 An implementation of decision tree-based context clustering on graphics processing units
Nicholas Pilkington, Heiga Zen
INTERSPEECH2
2010 Context adaptive training with factorized decision trees for HMM-based speech synthesis
abstract
To achieve natural high quality synthesised speech in HMMbased speech synthesis, the effective modelling of complex acoustic and linguistic contexts is critical. Traditional approaches use context-dependent HMMs with decision tree based parameter clustering to model the full combination of contexts. However, weak contexts, such as word-level emphasis in neutral speech, are difficult to capture using this approach. To effectively model weak contexts and reduce the data sparsity problem, weak and normal contexts should be treated independently. Context adaptive training provides a structured framework for this whereby standard HMMs represent normal contexts and linear transforms represent additional effects of weak contexts. In contrast to speaker adaptive training, separate decision trees have to be built for the weak and normal context factors. This paper describes the general framework of context adaptive training and investigates three concrete forms: MLLR, CMLLR and CAT based systems. Experiments on a word-level emphasis synthesis task show that all context adaptive training approaches can outperform the standard full-context-dependent HMM approach. However, the MLLR based system achieved the best performance. Index Terms: HMM-based speech synthesis, context adaptive training, factorized decision tree
Kai Yu 0004, Heiga Zen, François Mairesse, Steve J. Young
INTERSPEECH2
2010 Speaker and language adaptive training for HMM-based polyglot speech synthesis
Heiga Zen
INTERSPEECH1
2009 A Bayesian approach to HMM-based speech synthesis
abstract
This paper proposes a new framework of speech synthesis based on the Bayesian approach. The Bayesian method is a statistical technique for estimating reliable predictive distributions by marginalizing model parameters. In the proposed framework, all processes for constructing the system can be derived from one single predictive distribution which represents the basic problem of speech synthesis directly. Using HMM as the likelihood function and assuming some approximations, it can be regarded as an application of the variational Bayesian method to the HMM-based speech synthesis. Experimental results show that the proposed method outperforms the conventional one in a subjective test.
Kei Hashimoto, Heiga Zen, Yoshihiko Nankaku, Takashi Masuko, Keiichi Tokuda
ICASSP2
2009 Stereo-based stochastic noise compensation based on trajectory GMMS
abstract
This paper proposes a novel stereo-based stochastic noise compensation technique based on trajectory GMMs. Although the GMM-based noise compensation techniques such as SPLICE work effective, their performance sometimes degrades due to the inappropriate dynamic characteristics caused by the frame-by-frame mapping. While the use of dynamic feature constraints on the mapping stage can alleviate this problem, it also introduces an inconsistency between training and mapping. The recently proposed trajectory GMM-based feature mapping technique can solve this inconsistency while keeping the benefits of the use of dynamic features, and offers an entire sequence-level transformation rather than the frame-by-frame mapping. Results from a noise compensation experiment on the AURORA-2 task show that the proposed trajectory GMM-based noise compensation technique outperforms the conventional ones.
Heiga Zen, Yoshihiko Nankaku, Keiichi Tokuda
ICASSP1
2009 Tying covariance matrices to reduce the footprint of HMM-based speech synthesis systems
abstract
This paper proposes a technique of reducing footprint of HMMbased speech synthesis systems by tying all covariance matrices. HMM-based speech synthesis systems usually consume smaller footprint than unit-selection synthesis systems because statistics rather than speech waveforms are stored. However, further reduction is essential to put them on embedded devices which have very small memory. According to the empirical knowledge that covariance matrices have smaller impact for the quality of synthesized speech than mean vectors, here we propose a clustering technique of mean vectors while tying all covariance matrices. Subjective listening test results show that the proposed technique can shrink the footprint of an HMM-based speech synthesis system while retaining the quality of synthesized speech. Index Terms: HMM, speech synthesis, decision tree, contextclustering, MDL criterion, embedded device
Keiichiro Oura, Heiga Zen, Yoshihiko Nankaku, Akinobu Lee, Keiichi Tokuda
INTERSPEECH2
2009 Context-dependent additive log f_0 model for HMM-based speech synthesis
Heiga Zen, Norbert Braunschweiler
INTERSPEECH1
2009 Statistical parametric speech synthesis
Heiga Zen, Keiichi Tokuda, Alan W. Black
Speech Commun.1
2009 Robust Speaker-Adaptive HMM-Based Text-to-Speech Synthesis
abstract
This paper describes a speaker-adaptive HMM-based speech synthesis system. The new system, called ldquoHTS-2007,rdquo employs speaker adaptation (CSMAPLR+MAP), feature-space adaptive training, mixed-gender modeling, and full-covariance modeling using CSMAPLR transforms, in addition to several other techniques that have proved effective in our previous systems. Subjective evaluation results show that the new system generates significantly better quality synthetic speech than speaker-dependent approaches with realistic amounts of speech data, and that it bears comparison with speaker-dependent approaches even when large amounts of speech data are available. In addition, a comparison study with several speech synthesis techniques shows the new system is very robust: It is able to build voices from less-than-ideal speech data and synthesize good-quality speech even for out-of-domain sentences.
Junichi Yamagishi, Takashi Nose, Heiga Zen, Zhen-Hua Ling, Tomoki Toda, Keiichi Tokuda, Simon King 0001, Steve Renals
IEEE Trans. Speech Audio Process.3
2008 Acoustic modeling with contextual additive structure for HMM-based speech recognition
abstract
This paper proposes an acoustic modeling technique based on an additive structure of context dependencies for HMM-based speech recognition. Typical context dependent models, e.g., triphone HMMs, have direct dependencies of phonetic contexts, i.e., if a phonetic context is given, the Gaussian distribution is specified immediately. This paper assumes a more complex structure, an additive structure of acoustic feature components which have different context dependencies. Since the output probability distribution is composed of additive component distributions, a number of different distributions can be efficiently represented by a combination of fewer distributions. To automatically extract additive components, this paper presents a context clustering algorithm for the additive structure model in which multiple decision trees are constructed simultaneously. Experimental results show that the proposed technique improves phoneme recognition accuracy with fewer number of distributions than the conventional triphone HMMs.
Yoshihiko Nankaku, Kazuhiro Nakamura, Heiga Zen, Keiichi Tokuda
ICASSP3
2008 Performance evaluation of the speaker-independent HMM-based speech synthesis system "HTS 2007" for the Blizzard Challenge 2007
abstract
This paper describes a speaker-independent/adaptive HMM-based speech synthesis system developed for the Blizzard Challenge 2007. The new system, named “HTS-2007”, employs speaker adaptation (CSMAPLR+MAP), feature-space adaptive training, mixed-gender modeling, and full-covariance modeling using CSMAPLR transforms, in addition to several other techniques that have proved effective in our previous systems. Subjective evaluation results show that the new system generates significantly better quality synthetic speech than that of speaker-dependent approaches with realistic amounts of speech data, and that it bears comparison with speaker-dependent approaches even when large amounts of speech data are available.
Junichi Yamagishi, Takashi Nose, Heiga Zen, Tomoki Toda, Keiichi Tokuda
ICASSP3
2008 Bayesian context clustering using cross valid prior distribution for HMM-based speech recognition
abstract
This paper proposes a prior distribution determination tech-nique using cross validation for speech recognition based on the Bayesian approach. The Bayesian method is a statisti-cal technique for estimating reliable predictive distributions by marginalizing model parameters and its approximate version, the variational Bayesian method has been applied to HMM-based speech recognition. Since prior distributions represent-ing prior information about model parameters affect the pos-terior distributions and model selection, the determination of prior distributions is an important problem. However, it has not been thoroughly investigate in speech recognition. The pro-posed method can determine reliable prior distributions with-out tuning parameters and select an appropriate model struc-ture dependently on the amount of training data. Continu-ous phoneme recognition experiments show that the proposed method achieved a higher performance than the conventional methods. Index Terms: variational Bayes, cross validation, context clus-tering, continuous phoneme recognition
Kei Hashimoto, Heiga Zen, Yoshihiko Nankaku, Akinobu Lee, Keiichi Tokuda
INTERSPEECH2
2008 Unsupervised adaptation for HMM-based speech synthesis
abstract
It is now possible to synthesise speech using HMMs with a comparable quality to unit-selection techniques. Generating speech from a model has many potential advantages over concatenating waveforms. The most exciting is model adaptation. It has been shown that supervised speaker adaptation can yield highquality synthetic voices with an order of magnitude less data than required to train a speaker-dependent model or to build a basic unit-selection system. Such supervised methods require labelled adaptation data for the target speaker. In this paper, we introduce a method capable of unsupervised adaptation, using only speech from the target speaker without any labelling. Index Terms: speech synthesis, HMM-based speech synthesis, HTS, trajectory HMMs, speaker adaptation, MLLR
Simon King 0001, Keiichi Tokuda, Heiga Zen, Junichi Yamagishi
INTERSPEECH3
2008 Acoustic modeling based on model structure annealing for speech recognition
abstract
This paper proposes an HMM training technique using multiple phonetic decision trees and evaluates it in speech recognition. In the use of context dependent models, the decision tree based context clustering is applied to find a parameter tying structure. However, the clustering is usually performed based on statistics of HMM state sequences which are obtained by unreliable models without context clustering. To avoid this problem, we optimize the decision trees and HMM state sequences simultaneously. In the proposed method, this is performed by maximum likelihood (ML) estimation of a newly defined statistical model which includes multiple decision trees as hidden variables. Applying the deterministic annealing expectation maximization (DAEM) algorithm and using multiple decision trees in early stage of model training, state sequences are reliably estimated. In continuous phoneme recognition experiments, the proposed method can improve the recognition performance.
Sayaka Shiota, Kei Hashimoto, Heiga Zen, Yoshihiko Nankaku, Akinobu Lee, Keiichi Tokuda
INTERSPEECH3
2008 Probabilistic feature mapping based on trajectory HMMs
Heiga Zen, Yoshihiko Nankaku, Keiichi Tokuda
INTERSPEECH1
2007 Statistical Parametric Speech Synthesis
abstract
This paper gives a general overview of techniques in statistical parametric speech synthesis. One of the instances of these techniques, called HMM-based generation synthesis (or simply HMM-based synthesis), has recently been shown to be very effective in generating acceptable speech synthesis. This paper also contrasts these techniques with the more conventional unit selection technology that has dominated speech synthesis over the last ten years. Advantages and disadvantages of statistical parametric synthesis are highlighted as well as identifying where we expect the key developments to appear in the immediate future.
Alan W. Black, Heiga Zen, Keiichi Tokuda
ICASSP (4)2
2007 A trainable excitation model for HMM-based speech synthesis
abstract
This paper introduces a novel excitation approach for speech synthesizers in which the final waveform is generated through parameters directly obtained from Hidden Markov Models (HMMs). Despite the attractiveness of the HMM-based speech synthesis technique, namely utilization of small corpora and flexibility concerning the achievement of different voice styles, synthesized speech presents a characteristic buzziness caused by the simple excitation model which is employed during the speech production. This paper presents an innovative scheme where mixed excitation is modeled through closed-loop training of a set of state-dependent filters and pulse trains, with minimization of the error between excitation and residual sequences. The proposed method shows effectiveness, yielding synthesized speech with quality far superior to the simple excitation baseline and comparable to the best excitation schemes thus far reported for HMM-based speech synthesis.
Ranniery Maia, Tomoki Toda, Heiga Zen, Yoshihiko Nankaku, Keiichi Tokuda
INTERSPEECH3
2007 Model-space MLLR for trajectory HMMs
abstract
This paper proposes model-space Maximum Likelihood Linear Regression (mMLLR) based speaker adaptation technique for trajectory HMMs, which have been derived from HMMs by imposing explicit relationships between static and dynamic features. This model can alleviate two limitations of the HMM: constant statistics within a state and conditional independence assumption of state out-put probabilities without increasing the number of model parameters. Results in a continuous speech recognition experiments show that the proposed algorithm can adapt trajectory HMMs to a specific speaker and improve the performance of a trajectory HMM-based speech recogni-tion system. Index Terms: trajectory HMM, speaker adaptation, model-space MLLR
Heiga Zen, Yoshihiko Nankaku, Keiichi Tokuda
INTERSPEECH1
2007 Reformulating the HMM as a trajectory model by imposing explicit relationships between static and dynamic feature vector sequences
Heiga Zen, Keiichi Tokuda, Tadashi Kitamura
Comput. Speech Lang.1
2006 Hidden Semi-Markov Model Based Speech Recognition System using Weighted Finite-State Transducer
abstract
In hidden Markov models (HMMs), state duration probabilities decrease exponentially with time. It would be inappropriate representation of temporal structure of speech. One of the solutions for this problem is integrating state duration probability distributions explicitly into the HMM. This form is known as a hidden semi-Markov model (HSMM) [1]. Although a number of attempts to use explicit duration models in speech recognition systems have been proposed, they are not consistent because various approximations were used in both training and decoding. In the present paper, a fully consistent speech recognition system based on the HSMM framework is proposed. In a speaker-dependent continuous speech recognition experiment, HSMM-based speech recognition system achieved about 5.9% relative error reduction over the corresponding HMM-based one.
Keiichiro Oura, Heiga Zen, Yoshihiko Nankaku, Akinobu Lee, Keiichi Tokuda
ICASSP (1)2
2006 Estimating Trajectory Hmm Parameters Using Monte Carlo Em With Gibbs Sampler
abstract
In the present paper, the Monte Carlo EM (MCEM) algorithm with a Gibbs sampler is applied for estimating parameters of a trajectory HMM, which has been derived from an HMM by imposing explicit relationships between static and dynamic features. The trajectory HMM can alleviate two limitations of the HMM, which are i) constant statistics within a state, and ii) conditional independence of state output probabilities, without increasing the number of model parameters. In a speaker-dependent continuous speech recognition experiment, trajectory HMMs estimated by the MCEM algorithm achieved significant improvements over the corresponding HMMs trained by the EM (Baum-Welch) algorithm.
Heiga Zen, Yoshihiko Nankaku, Keiichi Tokuda, Tadashi Kitamura
ICASSP (1)1
2006 An HMM-based singing voice synthesis system
abstract
Abstract The present paper describes a corpus-based singing voice syn-thesis system based on hidden Markov models (HMMs). Thissystem employs the HMM-based speech synthesis to synthesizesingingvoice. Musical information such aslyrics, tones, durationsis modeled simultaneously in a unified framework of the context-dependent HMM. It can mimic the voice quality and singing styleof the original singer. Results of a singing voice synthesis exper-iment show that the proposed system can synthesize smooth andnatural-sounding singing voice. Index Terms : singing voice synthesis, HMM, time-lag model. 1. Introduction In recent years, various applications of speech synthesis systemshave been proposed and investigated. Singing voice synthesis isone of the hot topics in this area [1–5]. However, only a fewcorpus-based singing voice synthesis systems which can be con-structed automatically have been proposed.Currently, there are two main paradigms in the corpus-basedspeech synthesis area: sample-based approach and statistical ap-proach. The sample-based approach such as unit selection [6]can synthesize high-quality speech. However, it requires a hugeamountoftrainingdatatorealizevariousvoicecharacteristics. Onthe other hand, the quality of statistical approach such as HMM-basedspeechsynthesis[7]isbuzzybecauseitisbasedonavocod-ingtechnique. However,itissmoothandstable,anditsvoicechar-acteristics can easily be modified by transforming HMM parame-ters appropriately. For singing voice synthesis, applying the unitselection seems to be difficult because a huge amount of singingspeech which covers vast combinations of contextual factors thataffect singing voice has to be recorded. On the other hand, theHMM-based system can be constructed using a relatively smallamount of training data. From this point of view, the HMM-basedapproach seems to be more suitable for the singing voice synthe-sizer. In the present paper, we apply the HMM-based synthesisapproach to singing voice synthesis.Although the singing voice synthesis system proposed in thepresent paper is quite similar to the HMM-based text-to-speechsynthesissystem[7],therearetwomaindifferencesbetweenthem.In the HMM-based text-to-speech synthesis system, contextualfactors which may affect reading speech (e.g. phonemes, sylla-bles, words, phrases, etc.) are taken into account. However, con-textual factors which may affect singing voice should be different
Keijiro Saino, Heiga Zen, Yoshihiko Nankaku, Akinobu Lee, Keiichi Tokuda
INTERSPEECH2
2006 Speaker adaptation of trajectory HMMs using feature-space MLLR
abstract
Abstract Recently, a trajectory model, derived from the hiddenMarkov model (HMM) by imposing explicit relationshipsbetween static and dynamic features, has been proposed.The derived model, named trajectory HMM , can alleviatetwo limitations of the HMM: constant statistics within astate and conditional independence assumption of state out-put probabilities. In the present paper, a speaker adapta-tion algorithm for the trajectory HMM based on feature-space Maximum Likelihood Linear Regression (fMLLR)is derived and evaluated. Results of a simple continu-ous speech recognition experiment shows that adapting tra-jectory HMMs using the derived adaptation algorithm im-proves the speech recognition performance. Index Terms : trajectory HMM, adaptation, fMLLR. 1. Introduction Speech recognition technologies have achieved significantprogress with the introduction of hidden Markov models(HMMs). Their tractability and efficient implementationsare achieved by a number of assumptions, such as constantstatistics within an HMM state, conditional independenceof state output probabilities. Although these assumptionsmake the HMM practically useful, they are not realistic formodeling sequences of speech spectra, especially in spon-taneous speech. To overcome these shortcomings of theHMM, a variety of alternative models have been proposed,e.g., [1–3]. Although these models can improve the speechrecognition performance, they generally require an increaseinthenumberofmodelparametersandcomputationalcom-plexity. Alternatively, the use of dynamic features (deltaand delta-delta features) [4] also improves the performanceof HMM-based speech recognizers. It can be viewed as asimple mechanism to capture time dependencies. However,it has been thought of as an ad hoc rather than an essentialsolution. Generally, dynamic features are calculated as re-gression coefficients from their neighboring static features.Therefore, relationshipsbetweenstaticanddynamicfeaturevector sequences are deterministic. However, usually theserelationships are ignored and the static and dynamic fea-tures are modeled as independent random variables. Ignor-ing these dependencies allows inconsistency between thestatic and dynamic features when the HMM is used as agenerative model in the obvious way.Recently, a trajectory model, derived from the HMM byimposing the explicit relationships between static and dy-namic features, has been proposed [5]. The derived model,named
Heiga Zen, Yoshihiko Nankaku, Keiichi Tokuda, Tadashi Kitamura
INTERSPEECH1
2005 Sparse KPCA for Feature Extraction in Speech Recognition
abstract
This paper presents an analysis of the applicability of sparse kernel principal component analysis (SKPCA) for feature extraction in speech recognition, as well as a proposed approach to make the SKPCA technique realizable for a large amount of training data, which is a usual context in speech recognition systems. Although the KPCA (kernel principal component analysis) has proved to be an efficient technique for being applied to speech recognition, it has the disadvantage of requiring training data reduction, when its amount is excessively large. The standard approach to perform this data reduction is to randomly choose frames from the original data set, which does not necessarily provide a good statistical representation of the original data set. In order to solve this problem a likelihood related re-estimation procedure was applied to the KPCA framework, thus creating the SKPCA. The experimental results show the efficiency of SKPCA technique with the proposed approach over the KPCA with the standard sparse solution using randomly chosen frames and the standard feature extraction techniques.
Amaro A. de Lima, Heiga Zen, Yoshihiko Nankaku, Keiichi Tokuda, Tadashi Kitamura, Fernando Gil Vianna Resende Jr.
ICASSP (1)2
2005 On building a concatenative speech synthesis system from the blizzard challenge speech databases
abstract
In this paper, we compare two methods of building a concatenative speech synthesis system from the relatively small, “Blizzard Challenge ” speech databases. In the first method we build a system directly from the Blizzard databases using the IBM Concatenetative Speech Synthesis System originally designed for very large speech databases. In the second method, a larger database is used to build the synthesis system and the output is “morphed ” to match the speakers in the Blizzard databases. The second method outperformed the first while maintaining the identity of the Blizzard target speakers. 1.
Wael Hamza, Raimo Bakis, Zhiwei Shuang, Heiga Zen
INTERSPEECH4
2005 An overview of nitech HMM-based speech synthesis system for blizzard challenge 2005
abstract
In the present paper, hidden Markov model (HMM) based speech synthesis system developed in Nagoya Institute of Technology (Nitech-HTS) for a competition of text-to-speech synthesis systems using the same speech databases, named Blizzard Challenge 2005, is described. We show an overview of the basic HMM-based speech synthesis system and then recent developments to the latest one such as STRAIGHT-based vocoding, hidden semi-Markov model (HSMM) based acoustic modeling, and parameter generation considering global variance are illustrated. Constructed voices can synthesize speech around 0.3 xRT (real time ratio) and their footprints are less than 2 MB. The listening test results show that performances of our systems are much better than we expected. 1.
Heiga Zen, Tomoki Toda
INTERSPEECH1
2004 A Viterbi algorithm for a trajectory model derived from HMM with explicit relationship between static and dynamic features
abstract
This paper introduces a Viterbi algorithm to obtain a sub-optimal state sequence for trajectory-HMM, which is derived from HMM with explicit relationship between static and dynamic features. The trajectory-HMM can alleviate some limitations of HMM, which are (i) constant statistics within HMM state and (ii) conditional independence of observations given the state sequence, without increasing the number of model parameters. The proposed algorithm was applied to state-boundary optimization for Viterbi training and N-best rescoring. In a speaker-dependent continuous speech recognition experiment, trajectory-HMM with the proposed algorithm achieved about 14% error reduction over the standard HMM with the conventional Viterbi algorithm.
Heiga Zen, Keiichi Tokuda, Tadashi Kitamura
ICASSP (1)1
2004 Deterministic annealing EM algorithm in parameter estimation for acoustic model
abstract
ABSTRACT This paper investigates the effectiveness of the DAEM (Determin-istic Annealing EM) algorithm in acoustic modeling for speakerand speech recognition. Although the EM algorithm has beenwidely used to approximate the ML estimates, it has the problemof initialization dependence. To relax this problem, the DAEMalgorithm has been proposed and confirmed the effectiveness insmall tasks. In this paper, we applied the DAEM algorithm tospeakerrecognitionbasedonGMMsandcontinuousspeechrecog-nitionbasedonHMMs. ExperimentalresultsshowthattheDAEMalgorithm can improve the recognition performance as comparedtotheordinaryEMalgorithmwithconventionalinitializationmeth-ods,especiallyintheflatstarttrainingforcontinuousspeechrecog-nition. 1. INTRODUCTION The EM (Expectation-Maximization) algorithm [1] is widelyused for parameter estimation of statistical models with hiddenvariables. This algorithm provides a simple iterative proceduretoobtainapproximateML(maximumlikelihood)estimates. How-ever, since the EM algorithm is a hill-climbing approach, it suffersfrom the local maxima problem.On the other hand, GMMs (Gaussian mixture models) [2] andHMMs (hidden Markov models) [3] have been commonly usedin acoustic modeling for speaker and speech recognition, respec-tively. In conventional approaches, the LBG algorithm for GMMsand the segmental k-means algorithm for HMMs have been em-ployed to obtain initial model parameters before applying the EMalgorithm. However these initial values are not guaranteed to benear the true maximum likelihood point, and the posterior den-sity becomes unreliable at an early stage of training. Especiallyin continuous speech recognition, it is difficult to obtain accuratephoneme boundaries for all training data. Hence, the embeddedtraininghasbeenusedinwhichphonemeboundariesarealsodealtas hidden variables, and estimated based on the EM algorithm.Furthermore, in the worse case that the boundary information isnot available, a method called the flat start training is often ap-plied. In this method, initial parameters of HMMs are given bymaking all states of all models equal, and then carry out the em-bedded training. In these situations, we do not have enough priorknowledge to obtain a good initial values for the EM algorithm,and it would converge to one of the local maxima or saddle points
Yohei Itaya, Heiga Zen, Yoshihiko Nankaku, Chiyomi Miyajima, Keiichi Tokuda, Tadashi Kitamura
INTERSPEECH2
2004 Constructing emotional speech synthesizers with limited speech database
abstract
This paper describes an emotional speech synthesis system based on HMMs and related modeling techniques. For concatenative speech synthesis, we require all of the concatenation units that will be used to be recorded beforehand and made available at synthesis time. To adopt this approach for synthesizing the wide variety of human emotions possible in speech, implies that this process should be repeated for every targeted emotion making this task challenging and time consuming. In this paper, we propose an emotional speech synthesis technique based on HMMs, especially for the case where only limited amount of training data is available, directly incorporating subjective evaluation results performed on the training data. Listening results performed on the synthesized speech suggest that the proposed technique helps to improve the emotional content of synthesized speech.
Heiga Zen, Tadashi Kitamura, Murtaza Bulut, Shri Narayanan, Ryosuke Tsuzuki, Keiichi Tokuda
INTERSPEECH1
2004 Hidden semi-Markov model based speech synthesis
abstract
In the present paper, a hidden-semi Markov model (HSMM) based speech synthesis system is proposed. In a hidden Markov model (HMM) based speech synthesis system which we have proposed, rhythm and tempo are controlled by state duration probability distributions modeled by single Gaussian distributions. To synthesis speech, it constructs a sentence HMM corresponding to an arbitralily given text and determine state durations maximizing their probabilities, then a speech parameter vector sequence is generated for the given state sequence. However, there is an inconsistency: although the speech is synthesized from HMMs with explicit state duration probability distributions, HMMs are trained without them. In the present paper, we introduce an HSMM, which is an HMM with explicit state duration probability distributions, into the HMM-based speech synthesis system. Experimental results show that the use of HSMM training improves the naturalness of the synthesized speech.
Heiga Zen, Keiichi Tokuda, Takashi Masuko, Takao Kobayashi, Tadashi Kitamura
INTERSPEECH1
2003 Improving the performance of HMM-based very low bit rate speech coding
abstract
In this paper, we define an F0 quantization scheme for a very low bit rate speech coder based on HMM (hidden Markov model). In the coding system, the encoder carries out phoneme recognition, and transmits phoneme indices, state durations and F0 information to the decoder. In the decoder, phoneme HMM are concatenated according to the phoneme indices, and a sequence of mel-cepstral coefficient vectors is generated from the concatenated HMM. Finally we obtain synthetic speech by using the MLSA (mel log spectrum approximation) filter according to the mel-cepstral coefficients and F0 information. In addition to the F0 quantization, we investigate encoding methods for other parameters to reduce the bit rate, yet keeping the subjective speech quality. A subjective listening test shows that the performance of the proposed coder at about 100/spl sim/150 bit/s is superior to a VQ-based vocoder at 600 bit/s (mel-cepstrum: 6 bit/frame/spl times/50 frame/s, F0: 6 bit/frame/spl times/50 frame/s).
Takahiro Hoshiya, Shinji Sako, Heiga Zen, Keiichi Tokuda, Takashi Masuko, Takao Kobayashi, Tadashi Kitamura
ICASSP (1)3
2003 Speech recognition using voice-characteristic-dependent acoustic models
abstract
This paper proposes a speech recognition technique based on acoustic models considering voice characteristic variations. Context-dependent acoustic models, which are typically triphone HMM, are often used in continuous speech recognition systems. This work hypothesizes that the speaker voice characteristics that humans can perceive by listening are also factors in acoustic variation for construction of acoustic models, and a tree-based clustering technique is also applied to speaker voice characteristics to construct voice-characteristic-dependent acoustic models. In speech recognition using triphone models, the neighboring phonetic context is given from the linguistic-phonetic knowledge in advance; in contrast, the voice characteristics of input speech are unknown in recognition using voice-characteristic-dependent acoustic models. This paper proposes a method of recognizing speech even under conditions where the voice characteristics of the input speech are unknown. The result of a gender-dependent speech recognition experiment shows that the proposed method achieves higher recognition performance in comparison to conventional methods.
Hiroyuki Suzuki, Heiga Zen, Yoshihiko Nankaku, Chiyomi Miyajima, Keiichi Tokuda, Tadashi Kitamura
ICASSP (1)2
2003 On the use of kernel PCA for feature extraction in speech recognition
abstract
This paper describes an approachfor feature extraction in speech recognition systems using kernel principal componentanalysis (KPCA). This approachconsists in representing speech features as the projection of the extracted speech features mapped into a feature space via a nonlinear mapping onto the principal components. The nonlinear mapping is implicitly performed using the kerneltrick, which is an useful way of not mapping the input space into a featurespace explicitly,makingthis mapping computationally feasible. Better results were obtained by using this approach when compared to the standard technique.
Amaro A. de Lima, Heiga Zen, Yoshihiko Nankaku, Chiyomi Miyajima, Keiichi Tokuda, Tadashi Kitamura
INTERSPEECH2
2003 Towards the development of a brazilian portuguese text-to-speech system based on HMM
abstract
This paper describes the development of a Brazilian Portuguese text-to-speech system which applies a technique wherein speech is directly synthesized from hidden Markov models. In order to build the synthesizer a speech database was recorded and phonetically segmented. Furthermore, contextual informations about syllables, words, phrases, and utterances were determined, as well as questions for decision tree-based context clustering algorithms. The resulting system presents a fair reproduction of the prosody even when a small database is used for training.
Ranniery Maia, Heiga Zen, Keiichi Tokuda, Tadashi Kitamura, Fernando Gil Vianna Resende Jr.
INTERSPEECH2
2003 Trajectory modeling based on HMMs with the explicit relationship between static and dynamic features
abstract
This paper shows that the HMM whose state output vector includes static and dynamic feature parameters can be reformulated as a trajectory model by imposing the explicit relationship between the static and dynamic features. The derived model, named trajectory HMM, can alleviate the limitations of HMMs: i) constant statistics within an HMM state and ii) independence assumption of state output probabilities. We also derive a Viterbi-type training algorithm for the trajectory HMM. A preliminary speech recognition experiment based on N-best rescoring demonstrates that the training algorithm can improve the recognition performance significantly even though the trajectory HMM has the same parameterization as the standard HMM.
Keiichi Tokuda, Heiga Zen, Tadashi Kitamura
INTERSPEECH2
2003 Decision tree-based simultaneous clustering of phonetic contexts, dimensions, and state positions for acoustic modeling
abstract
In this paper, a new decision tree-based clustering technique called Phonetic, Dimensional and State Positional Decision Tree (PDS-DT) is proposed. In PDS-DT, phonetic contexts, dimensions and state positions are grouped simultaneously during decision tree construction. PDS-DT provides a complicate distribution sharing structure without any external control parameters. In speaker-independent continuous speech recognition experiments, PDS-DT achieved about 13%--15% error reduction over the phonetic decision tree-based state-tying technique.
Heiga Zen, Keiichi Tokuda, Tadashi Kitamura
INTERSPEECH1
2002 Decision tree distribution tying based on a dimensional split technique
abstract
Split Phonetic Decision Tree (DS-PDT) is proposed. In DSPDT, state distributions are split dimensionally when applying phonetic question. This technique is an extension of the decision tree based acoustic modeling. It gives a proper context-dependent sharing structure of each dimension automatically while maintaining the correlations among the dimensions. In speaker-independent continuous speech recognition experiments, DS-PDT achieved about 8% error reduction over the phonetic decision tree clustering.
Heiga Zen, Keiichi Tokuda, Tadashi Kitamura
INTERSPEECH1