VLDB 2026 Research / reviewers in the wild / expert
Lei He 0005
dblp:75/5673-5
· DBLP profile ↗
54ranked-venue papers
0as first author
34since 2021 · last 2025
0000-0003-2106-740XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 45 · 26 since 2021Artificial intelligence and machine learning · 32 · 21 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Drop the Beat! Freestyler for Accompaniment Conditioned Rapping Voice GenerationabstractRap, a prominent genre of vocal performance, remains underexplored in vocal generation. General vocal synthesis depends on precise note and duration inputs, requiring users to have related musical knowledge, which limits flexibility. In contrast, rap typically features simpler melodies, with a core focus on a strong rhythmic sense that harmonizes with accompanying beats. In this paper, we propose Freestyler, the first system that generates rapping vocals directly from lyrics and accompaniment inputs. Freestyler utilizes language model-based token generation, followed by a conditional flow matching model to produce spectrograms and a neural vocoder to restore audio. It allows a 3-second prompt to enable zero-shot timbre control. Due to the scarcity of publicly available rap datasets, we also present RapBank, a rap song dataset collected from the internet, alongside a meticulously designed processing pipeline. Experimental results show that Freestyler produces high-quality rapping voice generation with enhanced naturalness and strong alignment with accompanying beats, both stylistically and rhythmically. Ziqian Ning, Shuai Wang 0016, Yuepeng Jiang, Jixun Yao, Lei He 0005, Shifeng Pan, Lei Xie 0001 |
AAAI | 5 |
| 2025 | Efficient Speech Watermarking for Speech Synthesis via Progressive Knowledge DistillationabstractWith the rapid advancement of speech generative models, unauthorized voice cloning poses significant privacy and security risks. Speech watermarking offers a viable solution for tracing sources and preventing misuse. Current watermarking technologies fall mainly into two categories: DSP-based methods and deep learning-based methods. DSP-based methods are efficient but vulnerable to attacks, whereas deep learning-based methods offer robust protection at the expense of significantly higher computational cost. To improve the computational efficiency and enhance the robustness, we propose PKDMark, a lightweight deep learning-based speech watermarking method that leverages progressive knowledge distillation (PKD). Our approach proceeds in two stages: (1) training a high-performance teacher model using an invertible neural network-based architecture, and (2) transferring the teacher’s capabilities to a compact student model through progressive knowledge distillation. This process reduces computational costs by 93.6% while maintaining high level of robust performance and imperceptibility. Experimental results demonstrate that our distilled model achieves an average detection F1 score of 99.6% with a PESQ of 4.30 in advanced distortions, enabling efficient speech watermarking for real-time speech synthesis applications. Peter Pan, Lei He 0005, Sheng Zhao 0002 |
ASRU | 3 |
| 2025 | ZSVC: Zero-shot Style Voice Conversion with Disentangled Latent Diffusion Models and Adversarial TrainingabstractStyle voice conversion aims to transform the speaking style of source speech into a desired style while keeping the original speaker’s identity. However, previous style voice conversion approaches primarily focus on well-defined domains such as emotional aspects, limiting their practical applications. In this study, we present ZSVC, a novel Zero-shot Style Voice Conversion approach that utilizes a speech codec and a latent diffusion model with speech prompting mechanism to facilitate in-context learning for speaking style conversion. To disentangle speaking style and speaker timbre, we introduce information bottleneck to filter speaking style in the source speech and employ Uncertainty Modeling Adaptive Instance Normalization (UMAdaIN) to perturb the speaker timbre in the style prompt. Moreover, we propose a novel adversarial training strategy to enhance in-context learning and improve style similarity. Experiments conducted on 44,000 hours of speech data demonstrate the superior performance of ZSVC in generating speech with diverse speaking styles in zero-shot scenarios. Xinfa Zhu, Lei He 0005, Yujia Xiao, Xi Wang 0016, Xu Tan 0003, Sheng Zhao 0002, Lei Xie 0001 |
ICASSP | 2 |
| 2024 | Stylespeech: Self-Supervised Style Enhancing with VQ-VAE-Based Pre-Training for Expressive Audiobook Speech SynthesisabstractThe expressive quality of synthesized speech for audiobooks is limited by generalized model architecture and unbalanced style distribution in the training data. To address these issues, in this paper, we propose a self-supervised style enhancing method with VQ-VAE-based pre-training for expressive audiobook speech synthesis. Firstly, a text style encoder is pre-trained with a large amount of unlabeled text-only data. Secondly, a spectrogram style extractor based on VQ-VAE is pre-trained in a self-supervised manner, with plenty of audio data that covers complex style variations. Then a novel architecture with two encoder-decoder paths is specially designed to model the pronunciation and high-level style expressiveness respectively, with the guidance of the style extractor. Both objective and subjective evaluations demonstrate that our proposed method can effectively improve the naturalness and expressiveness of the synthesized speech in audiobook synthesis especially for the role and out-of-domain scenarios.1 Xueyuan Chen, Xi Wang 0016, Shaofei Zhang, Lei He 0005, Zhiyong Wu 0001, Xixin Wu, Helen M. Meng |
ICASSP | 4 |
| 2024 | PromptTTS 2: Describing and Generating Voices with Text PromptabstractSpeech conveys more information than text, as the same word can be uttered in various voices to convey diverse information. Compared to traditional text-to-speech (TTS) methods relying on speech prompts (reference speech) for voice variability, using text prompts (descriptions) is more user-friendly since speech prompts can be hard to find or may not exist at all. TTS approaches based on the text prompt face two main challenges: 1) the one-to-many problem, where not all details about voice variability can be described in the text prompt, and 2) the limited availability of text prompt datasets, where vendors and large cost of data labeling are required to write text prompts for speech. In this work, we introduce PromptTTS 2 to address these challenges with a variation network to provide variability information of voice not captured by text prompts, and a prompt generation pipeline to utilize the large language models (LLM) to compose high quality text prompts. Specifically, the variation network predicts the representation extracted from the reference speech (which contains full information about voice variability) based on the text prompt representation. For the prompt generation pipeline, it generates text prompts for speech with a speech language understanding model to recognize voice attributes (e.g., gender, speed) from speech and a large language model to formulate text prompts based on the recognition results. Experiments on a large-scale (44K hours) speech dataset demonstrate that compared to the previous works, PromptTTS 2 generates voices more consistent with text prompts and supports the sampling of diverse voice variability, thereby offering users more choices on voice generation. Additionally, the prompt generation pipeline produces high-quality text prompts, eliminating the large labeling cost. The demo page of PromptTTS 2 is available (https://speechresearch.github.io/prompttts2). Yichong Leng, Zhifang Guo, Zeqian Ju, Xu Tan 0003, Eric Liu 0006, Dongchao Yang, Leying Zhang, Kaitao Song, Lei He 0005, Xiang-Yang Li 0001, Sheng Zhao 0002, Tao Qin 0001, Jiang Bian 0002 |
ICLR | 11 |
| 2024 | NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing SynthesizersabstractScaling text-to-speech (TTS) to large-scale, multi-speaker, and in-the-wild datasets is important to capture the diversity in human speech such as speaker identities, prosodies, and styles (e.g., singing). Current large TTS systems usually quantize speech into discrete tokens and use language models to generate these tokens one by one, which suffer from unstable prosody, word skipping/repeating issue, and poor voice quality. In this paper, we develop NaturalSpeech 2, a TTS system that leverages a neural audio codec with residual vector quantizers to get the quantized latent vectors and uses a diffusion model to generate these latent vectors conditioned on text input. To enhance the zero-shot capability that is important to achieve diverse speech synthesis, we design a speech prompting mechanism to facilitate in-context learning in the diffusion model and the duration/pitch predictor. We scale NaturalSpeech 2 to large-scale datasets with 44K hours of speech and singing data and evaluate its voice quality on unseen speakers. NaturalSpeech 2 outperforms previous TTS systems by a large margin in terms of prosody/timbre similarity, robustness, and voice quality in a zero-shot setting, and performs novel zero-shot singing synthesis with only a speech prompt. Audio samples are available at https://naturalspeech2.github.io/. Zeqian Ju, Xu Tan 0003, Eric Liu 0006, Yichong Leng, Lei He 0005, Tao Qin 0001, Sheng Zhao 0002, Jiang Bian 0002 |
ICLR | 6 |
| 2024 | NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion ModelsabstractWhile recent large-scale text-to-speech (TTS) models have achieved significant progress, they still fall shorts in speech quality, similarity, and prosody. Considering that speech intricately encompasses various attributes (e.g., content, prosody, timbre, and acoustic details) that pose significant challenges for generation, a natural idea is to factorize speech into individual subspaces representing different attributes and generate them individually. Motivated by it, we propose a TTS system with novel factorized diffusion models to generate natural speech in a zero-shot way. Specifically, 1) we design a neural codec with factorized vector quantization (FVQ) to disentangle speech waveform into subspaces of content, prosody, timbre, and acoustic details; 2) we propose a factorized diffusion model, which generates attributes in each subspace following its corresponding prompt. With this factorization design, our method can effectively and efficiently model the intricate speech with disentangled subspaces in a divide-and-conquer way. Experimental results show that our method outperforms the state-of-the-art TTS systems on quality, similarity, prosody, and intelligibility. Zeqian Ju, Yuancheng Wang, Xu Tan 0003, Detai Xin, Dongchao Yang, Eric Liu 0006, Yichong Leng, Kaitao Song, Siliang Tang, Zhizheng Wu 0001, Tao Qin 0001, Xiang-Yang Li 0001, Wei Ye 0004, Shikun Zhang, Jiang Bian 0002, Lei He 0005, Jinyu Li 0001, Sheng Zhao 0002 |
ICML | 17 |
| 2024 | Contrastive Context-Speech Pretraining for Expressive Text-to-Speech SynthesisabstractThe latest Text-to-Speech (TTS) systems can produce speech with voice quality and naturalness comparable to human speech. Yet the demand for large amount of high-quality data from target speakers remains a significant challenge. Particularly for long-form expressive reading, target speaker's training speech that covers rich contextual information are needed. In this paper a novel design of context-aware speech pre-trained model is developed for expressive TTS based on contrastive learning. The model can be trained with abundant speech data without explicitly labelled speaker identities. It captures the intricate relationship between the speech expression of a spoken sentence and the contextual text information. By incorporating cross-modal text and speech features into the TTS model, it enables the generation of coherent and expressive speech, which is especially beneficial when there is a scarcity of target speaker data. The pre-trained model is evaluated first in the task of Context-Speech retrieval and then as the integral part of a zero-shot TTS system. Experimental results demonstrate that the pretraining framework effectively learns Context-Speech representations and significantly enhances the expressiveness of synthesized speech. Audio demos are available at: https://ccsp2024.github.io/demo/. Yujia Xiao, Xi Wang 0016, Xu Tan 0003, Lei He 0005, Xinfa Zhu, Sheng Zhao 0002, Tan Lee |
ACM Multimedia | 4 |
| 2024 | UniStyle: Unified Style Modeling for Speaking Style Captioning and Stylistic Speech SynthesisabstractUnderstanding the speaking style, such as the emotion of the interlocutor's speech, and responding with speech in an appropriate style is a natural occurrence in human conversations. However, technically, existing research on speech synthesis and speaking style captioning typically proceeds independently. In this work, an innovative framework, referred to as UniStyle, is proposed to incorporate both the capabilities of speaking style captioning and style-controllable speech synthesizing. Specifically, UniStyle consists of a UniConnector and a style prompt-based speech generator. The role of the UniConnector is to bridge the gap between different modalities, namely speech audio and text descriptions. It enables the generation of text descriptions with speech as input and the creation of style representations from text descriptions for speech synthesis with the speech generator. Besides, to overcome the issue of data scarcity, we propose a two-stage and semi-supervised training strategy, which reduces data requirements while boosting performance. Extensive experiments conducted on open-source corpora demonstrate that UniStyle achieves state-of-the-art performance in speaking style captioning and synthesizes expressive speech with various speaker timbres and speaking styles in a zero-shot manner. Xinfa Zhu, Lei He 0005, Yujia Xiao, Xi Wang 0016, Xu Tan 0003, Sheng Zhao 0002, Lei Xie 0001 |
ACM Multimedia | 4 |
| 2024 | CoVoMix: Advancing Zero-Shot Speech Generation for Human-like Multi-talker ConversationsabstractRecent advancements in zero-shot text-to-speech (TTS) modeling have led to significant strides in generating high-fidelity and diverse speech. However, dialogue generation, along with achieving human-like naturalness in speech, continues to be a challenge. In this paper, we introduce CoVoMix: Conversational Voice Mixture Generation, a novel model for zero-shot, human-like, multi-speaker, multi-round dialogue speech generation. CoVoMix first converts dialogue text into multiple streams of discrete tokens, with each token stream representing semantic information for individual talkers. These token streams are then fed into a flow-matching based acoustic model to generate mixed mel-spectrograms. Finally, the speech waveforms are produced using a HiFi-GAN model. Furthermore, we devise a comprehensive set of metrics for measuring the effectiveness of dialogue modeling and generation. Our experimental results show that CoVoMix can generate dialogues that are not only human-like in their naturalness and coherence but also involve multiple talkers engaging in multiple rounds of conversation. This is exemplified by instances generated in a single channel where one speaker's utterance is seamlessly mixed with another's interjections or laughter, indicating the latter's role as an attentive listener. Audio samples are enclosed in the supplementary. Leying Zhang, Yao Qian, Shujie Liu 0001, Dongmei Wang, Xiaofei Wang 0009, Midia Yousefi, Yanmin Qian, Jinyu Li 0001, Lei He 0005, Sheng Zhao 0002, Michael Zeng 0001 |
NeurIPS | 10 |
| 2024 | NaturalSpeech: End-to-End Text-to-Speech Synthesis With Human-Level QualityabstractText-to-speech (TTS) has made rapid progress in both academia and industry in recent years. Some questions naturally arise that whether a TTS system can achieve human-level quality, how to define/judge that quality, and how to achieve it. In this paper, we answer these questions by first defining the human-level quality based on the statistical significance of subjective measure and introducing appropriate guidelines to judge it, and then developing a TTS system called NaturalSpeech that achieves human-level quality on benchmark datasets. Specifically, we leverage a variational auto-encoder (VAE) for end-to-end text-to-waveform generation, with several key modules to enhance the capacity of the prior from text and reduce the complexity of the posterior from speech, including phoneme pre-training, differentiable duration modeling, bidirectional prior/posterior modeling, and a memory mechanism in VAE. Experimental evaluations on the popular LJSpeech dataset show that our proposed NaturalSpeech achieves -0.01 CMOS (comparative mean opinion score) to human recordings at the sentence level, with Wilcoxon signed rank test at p-level p >> 0.05, which demonstrates no statistically significant difference from human recordings for the first time. Xu Tan 0003, Jiawei Chen 0008, Haohe Liu, Jian Cong, Chen Zhang 0020, Xi Wang 0016, Yichong Leng, Yuanhao Yi, Lei He 0005, Sheng Zhao 0002, Tao Qin 0001, Frank K. Soong, Tie-Yan Liu |
IEEE Trans. Pattern Anal. Mach. Intell. | 10 |
| 2023 | VideoDubber: Machine Translation with Speech-Aware Length Control for Video DubbingabstractVideo dubbing aims to translate the original speech in a film or television program into the speech in a target language, which can be achieved with a cascaded system consisting of speech recognition, machine translation and speech synthesis. To ensure the translated speech to be well aligned with the corresponding video, the length/duration of the translated speech should be as close as possible to that of the original speech, which requires strict length control. Previous works usually control the number of words or characters generated by the machine translation model to be similar to the source sentence, without considering the isochronicity of speech as the speech duration of words/characters in different languages varies. In this paper, we propose VideoDubber, a machine translation system tailored for the task of video dubbing, which directly considers the speech duration of each token in translation, to match the length of source and target speech. Specifically, we control the speech length of generated sentence by guiding the prediction of each word with the duration information, including the speech duration of itself as well as how much duration is left for the remaining words. We design experiments on four language directions (German -> English, Spanish -> English, Chinese English), and the results show that VideoDubber achieves better length control ability on the generated speech than baseline methods. To make up the lack of real-world datasets, we also construct a real-world test set collected from films to provide comprehensive evaluations on the video dubbing task. Yihan Wu 0008, Junliang Guo, Xu Tan 0003, Chen Zhang 0020, Bohan Li 0003, Ruihua Song, Lei He 0005, Sheng Zhao 0002, Arul Menezes, Jiang Bian 0002 |
AAAI | 7 |
| 2023 | Prosody-Aware Speecht5 for Expressive Neural TTSabstractSpeechT5, a multimodal learning framework which explores encoder-decoder pre-training by leveraging both unlabeled speech and text, has been proven to be effective on a wide variety of speech processing tasks. In this paper, we enhance SpeechT5 by adding a new sub-task on prosody modeling (prosody-aware SpeechT5) for neural text-to-speech (TTS), which can improve the model capability to learn richer contextual representations through multi-task learning. In the prosody-aware SpeechT5 training framework, most modules in neural TTS can be pre-trained with large-scale unlabeled speech and text corpus, including encoder, decoder, and variance adaptor. Experimental results show that the proposed prosody-aware SpeechT5 is effective at improving the expressiveness of neural TTS1: 1) the CMOS (comparison mean opinion score) gain is 0.154 for texts from news domain and 0.114 for texts from audiobook domain; 2) the prosody related issues in synthetic speech are reduced by 19.02% in subjective evaluation. Yuanhao Yi, Shujie Liu 0001, Lei He 0005 |
ICASSP | 5 |
| 2023 | Joint Pre-Training with Speech and Bilingual Text for Direct Speech to Speech TranslationabstractDirect speech-to-speech translation (S2ST) is an attractive research topic with many advantages compared to cascaded S2ST. However, direct S2ST suffers from the data scarcity problem because the corpora from the speech of the source language to the speech of the target language are very rare. To address this issue, we propose in this paper a Speech2S model, which is jointly pre-trained with unpaired speech and bilingual text data for direct speech-to-speech translation tasks. By effectively leveraging the paired text data, Speech2S is capable of modeling the cross-lingual speech conversion from source to target language. We verify the performance of the proposed Speech2S on Europarl-ST and VoxPopuli datasets. Experimental results demonstrate that Speech2S gets an improvement of about 5 BLEU scores compared to encoder-only pre-training models, and achieves a competitive or even better performance than existing state-of-the-art models1. Shujie Liu 0001, Lei He 0005, Jinyu Li 0001, Furu Wei |
ICASSP | 6 |
| 2023 | LeanSpeech: The Microsoft Lightweight Speech Synthesis System for Limmits Challenge 2023abstractThis paper describes the Microsoft Text-to-Speech (TTS) system: LeanSpeech for LIMMITS (Lightweight, Multi-speaker, Multi-lingual Indic TTS) Challenge 20231, which is part of ICASSP2023 to encourage the advance of TTS in Indian Languages. We propose a lightweight encoder-decoder acoustic model composed of 1-D convolution and LSTM blocks, which is trained with knowledge distillation from a multi-speaker multi-lingual teacher model, DelightfulTTS [1]. The speech corpus is reprocessed and used in both AM training and vocoder fine-tuning. In Track-2 of the challenge, our system achieves MOS 4.56 and SMOS 3.98, which indicates the efficiency of the proposed model and training strategy. Chen Zhang 0020, Shubham Bansal, Aakash Lakhera, Jinzhu Li, Gang Wang 0001, Sandeepkumar Satpal, Sheng Zhao 0002, Lei He 0005 |
ICASSP | 8 |
| 2023 | Large-Scale Automatic Audiobook Creation
Brendan Walsh, Mark Hamilton, Greg Newby, Xi Wang 0016, Serena Ruan, Sheng Zhao 0002, Lei He 0005, Shaofei Zhang, Eric Dettinger, William T. Freeman, Markus Weimer |
INTERSPEECH | 7 |
| 2023 | ContextSpeech: Expressive and Efficient Text-to-Speech for Paragraph ReadingabstractWhile state-of-the-art Text-to-Speech systems can generate natural speech of very high quality at sentence level, they still meet great challenges in speech generation for paragraph / long-form reading.Such deficiencies are due to i) ignorance of cross-sentence contextual information, and ii) high computation and memory cost for long-form synthesis.To address these issues, this work develops a lightweight yet effective TTS system, ContextSpeech.Specifically, we first design a memory-cached recurrence mechanism to incorporate global text and speech context into sentence encoding.Then we construct hierarchically-structured textual semantics to broaden the scope for global context enhancement.Additionally, we integrate linearized self-attention to improve model efficiency.Experiments show that ContextSpeech significantly improves the voice quality and prosody expressiveness in paragraph reading with competitive model efficiency. Yujia Xiao, Shaofei Zhang, Xi Wang 0016, Xu Tan 0003, Lei He 0005, Sheng Zhao 0002, Frank K. Soong, Tan Lee |
INTERSPEECH | 5 |
| 2023 | AUDIT: Audio Editing by Following Instructions with Latent Diffusion ModelsabstractAudio editing is applicable for various purposes, such as adding background sound effects, replacing a musical instrument, and repairing damaged audio. Recently, some diffusion-based methods achieved zero-shot audio editing by using a diffusion and denoising process conditioned on the text description of the output audio. However, these methods still have some problems: 1) they have not been trained on editing tasks and cannot ensure good editing effects; 2) they can erroneously modify audio segments that do not require editing; 3) they need a complete description of the output audio, which is not always available or necessary in practical scenarios. In this work, we propose AUDIT, an instruction-guided audio editing model based on latent diffusion models. Specifically, \textbf{AUDIT} has three main design features: 1) we construct triplet training data (instruction, input audio, output audio) for different audio editing tasks and train a diffusion model using instruction and input (to be edited) audio as conditions and generating output (edited) audio; 2) it can automatically learn to only modify segments that need to be edited by comparing the difference between the input and output audio; 3) it only needs edit instructions instead of full target audio descriptions as text input. AUDIT achieves state-of-the-art results in both objective and subjective metrics for several audio editing tasks (e.g., adding, dropping, replacement, inpainting, super-resolution). Demo samples are available at https://audit-demopage.github.io/. Yuancheng Wang, Zeqian Ju, Xu Tan 0003, Lei He 0005, Zhizheng Wu 0001, Jiang Bian 0002, Sheng Zhao 0002 |
NeurIPS | 4 |
| 2022 | Infergrad: Improving Diffusion Models for Vocoder by Considering Inference in TrainingabstractDenoising diffusion probabilistic models (diffusion models for short) require a large number of iterations in inference to achieve the generation quality that matches or surpasses the state-of-the-art generative models, which invariably results in slow inference speed. Previous approaches aim to optimize the choice of inference schedule over a few iterations to speed up inference. However, this results in reduced generation quality, mainly because the inference process is optimized separately, without jointly optimizing with the training process. In this paper, we propose InferGrad, a diffusion model for vocoder that incorporates inference process into training, to reduce the inference iterations while maintaining high generation quality. More specifically, during training, we generate data from random noise through a reverse process under inference schedules with a few iterations, and impose a loss to minimize the gap between the generated and ground-truth data samples. Then, unlike existing approaches, the training of InferGrad considers the inference process. The advantages of InferGrad are demonstrated through experiments on the LJSpeech dataset showing that InferGrad achieves better voice quality than the baseline WaveGrad under same conditions while maintaining the same voice quality as the baseline but with 3x speedup (2 iterations for InferGrad vs 6 iterations for WaveGrad). Zehua Chen 0005, Xu Tan 0003, Shifeng Pan, Danilo P. Mandic, Lei He 0005, Sheng Zhao 0002 |
ICASSP | 6 |
| 2022 | Improving Fastspeech TTS with Efficient Self-Attention and Compact Feed-Forward NetworkabstractFastSpeech, as a feed-forward transformer based TTS, can avoid the slow serial, autoregressive inference to generate the target mel-spectrogram in a parallel way. As a non-autoregressive TTS, the latency and computation load in inference is shifted from vocoder to transformer where the efficiency is limited by the quadratic time and memory complexity in the self-attention mechanism, particularly for a long text sequence. To tackle this challenges, We propose two models, ProbSparseFS and LinearizedFS, which have efficient self-attention arrangements to improve the inference speed and memory complexity. LinearizedFS has achieved 3.4x memory savings and 2.1x inference speedup, compared with the those of the baseline FastSpeech. A further optimized LinearizedFS with a lightweight FFN can accelerate the inference speed by 3.6x more. We do subjective voice quality evaluations in MOS and CMOS of News report and Audiobook applications, for multi-speaker and multi-style scenarios. Test results verified that the proposed models yield a TTS quality which is on-par with that of the baseline system but with much better memory efficiency and inference speed. Yujia Xiao, Xi Wang 0016, Lei He 0005, Frank K. Soong |
ICASSP | 3 |
| 2022 | Prosodyspeech: Towards Advanced Prosody Model for Neural Text-to-SpeechabstractThis paper proposes ProsodySpeech, a novel prosody model to enhance encoder-decoder neural Text-To-Speech (TTS), to generate high expressive and personalized speech even with very limited training data. First, we use a Prosody Extractor built from a large speech corpus with various speakers to generate a set of prosody exemplars from multiple reference speeches, in which Mutual Information based Style content separation (MIST) is adopted to alleviate "content leakage" problem. Second, we use a Prosody Distributor to make a soft selection of appropriate prosody exemplars in phone-level with the help of an attention mechanism. The resulting prosody feature is then aggregated into the output of text encoder, together with additional phone-level pitch feature to enrich the prosody. We apply this method into two tasks: highly expressive multi style/emotion TTS and few-shot personalized TTS. The experiments show the proposed model outperforms baseline FastSpeech 2 + GST with significant improvements in terms of similarity and style expression. Yuanhao Yi, Lei He 0005, Shifeng Pan, Xi Wang 0016, Yujia Xiao |
ICASSP | 2 |
| 2022 | Exploring Machine Speech Chain For Domain AdaptationabstractMachine Speech Chain integrates both end-to-end (E2E) automatic speech recognition (ASR) and neural text-to-speech (TTS) into one circle for joint training. It has been proven that it can effectively leverage a large amount of unpaired data in the spirit of data augmentation. In this paper, we explore the TTS→ASR pipeline in machine speech chain to perform domain adaptation for both E2E ASR and neural TTS models with only text data from the target domain. We conduct experiments by adapting from audiobook domain (i.e., LibriSpeech) to presentation domain (i.e., TED-LIUM). There is a relative word error rate (WER) reduction of 19.7% for the E2E ASR model on the TED-LIUM test set, and a relative WER reduction of 29.4% in synthetic speech generated by neural TTS in the presentation domain. Moreover, we observe that the gains from the proposed method and conventional adaptation methods of language models are additive. Fengpeng Yue, Lei He 0005, Tom Ko, Yu Zhang 0006 |
ICASSP | 3 |
| 2022 | Neural Lexicon Reader: Reduce Pronunciation Errors in End-to-end TTS by Leveraging External Textual KnowledgeabstractEnd-to-end TTS requires a large amount of speech/text paired data to cover all necessary knowledge, particularly how to pronounce different words in diverse contexts, so that a neural model may learn such knowledge accordingly. But in real applications, such high demand of training data is hard to be satisfied and additional knowledge often needs to be injected manually. For example, to capture pronunciation knowledge on languages without regular orthography, a complicated grapheme-to-phoneme pipeline needs to be built based on a large structured pronunciation lexicon, leading to extra, sometimes high, costs to extend neural TTS to such languages. In this paper, we propose a framework to learn to automatically extract knowledge from unstructured external resources using a novel Token2Knowledge attention module. The framework is applied to build a TTS model named Neural Lexicon Reader that extracts pronunciations from raw lexicon texts in an end-to-end manner. Experiments show the proposed model significantly reduces pronunciation errors in low-resource, end-to-end Chinese TTS, and the lexicon-reading capability can be transferred to other languages with a smaller amount of data. Mutian He 0001, Lei He 0005, Frank K. Soong |
INTERSPEECH | 3 |
| 2022 | DelightfulTTS 2: End-to-End Speech Synthesis with Adversarial Vector-Quantized Auto-Encoders
Ruiqing Xue, Lei He 0005, Xu Tan 0003, Sheng Zhao 0002 |
INTERSPEECH | 3 |
| 2022 | AdaSpeech 4: Adaptive Text to Speech in Zero-Shot ScenariosabstractAdaptive text to speech (TTS) can synthesize new voices in zero-shot scenarios efficiently, by using a well-trained source TTS model without adapting it on the speech data of new speakers. Considering seen and unseen speakers have diverse characteristics, zero-shot adaptive TTS requires strong generalization ability on speaker characteristics, which brings modeling challenges. In this paper, we develop AdaSpeech 4, a zero-shot adaptive TTS system for high-quality speech synthesis. We model the speaker characteristics systematically to improve the generalization on new speakers. Generally, the modeling of speaker characteristics can be categorized into three steps: extracting speaker representation, taking this speaker representation as condition, and synthesizing speech/mel-spectrogram given this speaker representation. Accordingly, we improve the modeling in three steps: 1) To extract speaker representation with better generalization, we factorize the speaker characteristics into basis vectors and extract speaker representation by weighted combining of these basis vectors through attention. 2) We leverage conditional layer normalization to integrate the extracted speaker representation to TTS model. 3) We propose a novel supervision loss based on the distribution of basis vectors to maintain the corresponding speaker characteristics in generated mel-spectrograms. Without any fine-tuning, AdaSpeech 4 achieves better voice quality and similarity than baselines in multiple datasets. Yihan Wu 0008, Xu Tan 0003, Bohan Li 0003, Lei He 0005, Sheng Zhao 0002, Ruihua Song, Tao Qin 0001, Tie-Yan Liu |
INTERSPEECH | 4 |
| 2022 | Self-supervised Context-aware Style Representation for Expressive Speech SynthesisabstractExpressive speech synthesis, like audiobook synthesis, is still challenging for style representation learning and prediction.Deriving from reference audio or predicting style tags from text requires a huge amount of labeled data, which is costly to acquire and difficult to define and annotate accurately.In this paper, we propose a novel framework for learning style representation from abundant plain text in a self-supervised manner.It leverages an emotion lexicon and uses contrastive learning and deep clustering.We further integrate the style representation as a conditioned embedding in a multi-style Transformer TTS.Comparing with multi-style TTS by predicting style tags trained on the same dataset but with human annotations, our method achieves improved results according to subjective evaluations on both in-domain and out-of-domain test sets in audiobook speech.Moreover, with implicit context-aware style representation, the emotion transition of synthesized audio in a long paragraph appears more natural.The audio samples are available on the demo website. Yihan Wu 0008, Xi Wang 0016, Shaofei Zhang, Lei He 0005, Ruihua Song, Jian-Yun Nie |
INTERSPEECH | 4 |
| 2022 | SoftSpeech: Unsupervised Duration Model in FastSpeech 2
Yuanhao Yi, Lei He 0005, Shifeng Pan, Xi Wang 0016 |
INTERSPEECH | 2 |
| 2022 | BinauralGrad: A Two-Stage Conditional Diffusion Probabilistic Model for Binaural Audio SynthesisabstractBinaural audio plays a significant role in constructing immersive augmented and virtual realities. As it is expensive to record binaural audio from the real world, synthesizing them from mono audio has attracted increasing attention. This synthesis process involves not only the basic physical warping of the mono audio, but also room reverberations and head/ear related filtration, which, however, are difficult to accurately simulate in traditional digital signal processing. In this paper, we formulate the synthesis process from a different perspective by decomposing the binaural audio into a common part that shared by the left and right channels as well as a specific part that differs in each channel. Accordingly, we propose BinauralGrad, a novel two-stage framework equipped with diffusion models to synthesize them respectively. Specifically, in the first stage, the common information of the binaural audio is generated with a single-channel diffusion model conditioned on the mono audio, based on which the binaural audio is generated by a two-channel diffusion model in the second stage. Combining this novel perspective of two-stage synthesis with advanced generative models (i.e., the diffusion models), the proposed BinauralGrad is able to generate accurate and high-fidelity binaural audio samples. Experiment results show that on a benchmark dataset, BinauralGrad outperforms the existing baselines by a large margin in terms of both object and subject evaluation metrics (Wave L2: $0.128$ vs. $0.157$, MOS: $3.80$ vs. $3.61$). The generated audio samples\footnote{\url{https://speechresearch.github.io/binauralgrad}} and code\footnote{\url{https://github.com/microsoft/NeuralSpeech/tree/master/BinauralGrad}} are available online. Yichong Leng, Zehua Chen 0005, Junliang Guo, Haohe Liu, Jiawei Chen 0008, Xu Tan 0003, Danilo P. Mandic, Lei He 0005, Xiang-Yang Li 0001, Tao Qin 0001, Sheng Zhao 0002, Tie-Yan Liu |
NeurIPS | 8 |
| 2021 | On Addressing Practical Challenges for RNN-TransducerabstractIn this paper, several works are proposed to address practi-cal challenges for deploying RNN Transducer (RNN-T) based speech recognition systems. These challenges are adapting a well-trained RNN-T model to a new domain without col-lecting the audio data, obtaining time stamps and confidence scores at word level. We solve the first challenge with a splicing data method which concatenates the speech segments ex-tracted from the source domain data. To get time stamps, a phone prediction branch is added to the RNN-T model by sharing the encoder for the purpose of forced alignment. Fi-nally, we obtain word level confidence scores by utilizing sev-eral types of features calculated during decoding and from a confusion network. Evaluated with Microsoft production data, the splicing data adaptation method improves the base-line and adaptation with the text to speech method by 58.03% and 15.25% relative word error rate reduction, respectively. The proposed time stamping method can get less than 50 mil-lisecond word timing difference from the ground truth align-ment on average while maintaining the recognition accuracy. We also obtain high confidence annotation performance with limited computation cost. Rui Zhao 0017, Jinyu Li 0001, Wenning Wei, Lei He 0005, Yifan Gong 0001 |
ASRU | 5 |
| 2021 | Speech Bert Embedding for Improving Prosody in Neural TTSabstractThis paper presents a speech BERT model to extract embedded prosody information in speech segments for improving the prosody of synthesized speech in neural text-to-speech (TTS). As a pre-trained model, it can learn prosody attributes from a large amount of speech data, which can utilize more data than the original training data used by the target TTS. The embedding is extracted from the previous segment of a fixed length in the proposed BERT. The extracted embedding is then used together with the mel-spectrogram to predict the following segment in the TTS decoder. Experimental results obtained by the Transformer TTS show that the proposed BERT can extract fine-grained, segment-level prosody, which is complementary to utterance-level prosody to improve the final prosody of the TTS speech. The objective distortions measured on a single speaker TTS are reduced between the generated speech and original recordings. Subjective listening tests also show that the proposed approach is favorably preferred over the TTS without the BERT prosody embedding module, for both in-domain and out-of-domain applications. For Microsoft professional, single/multiple speakers and the LJ Speaker in the public database, subjective preference is similarly confirmed with the new BERT prosody embedding. TTS demo audio samples are in https://judy44chen.github.io/TTSSpeechBERT/. Xi Wang 0016, Frank K. Soong, Lei He 0005 |
ICASSP | 5 |
| 2021 | Improving RNN-T for Domain Scaling Using Semi-Supervised Training with Neural TTS
Rui Zhao 0017, Zhong Meng, Xie Chen 0001, Jinyu Li 0001, Yifan Gong 0001, Lei He 0005 |
Interspeech | 8 |
| 2021 | Cross-Speaker Style Transfer with Prosody Bottleneck in Neural Speech SynthesisabstractCross-speaker style transfer is crucial to the applications of multi-style and expressive speech synthesis at scale. It does not require the target speakers to be experts in expressing all styles and to collect corresponding recordings for model training. However, the performances of existing style transfer methods are still far behind real application needs. The root causes are mainly twofold. Firstly, the style embedding extracted from single reference speech can hardly provide fine-grained and appropriate prosody information for arbitrary text to synthesize. Secondly, in these models the content/text, prosody, and speaker timbre are usually highly entangled, it's therefore not realistic to expect a satisfied result when freely combining these components, such as to transfer speaking style between speakers. In this paper, we propose a cross-speaker style transfer text-to-speech (TTS) model with explicit prosody bottleneck. The prosody bottleneck builds up the kernels accounting for speaking style robustly, and disentangles the prosody from content and speaker timbre, therefore guarantees high quality cross-speaker style transfer. Evaluation result shows the proposed method even achieves on-par performance with source speaker's speaker-dependent (SD) model in objective measurement of prosody, and significantly outperforms the cycle consistency and GMVAE-based baselines in objective and subjective evaluations. Shifeng Pan, Lei He 0005 |
Interspeech | 2 |
| 2021 | Conversational End-to-End TTS for Voice AgentsabstractEnd-to-end neural TTS has achieved excellent performance on reading style speech synthesis. However, it is still a challenge to build a high-quality conversational TTS due to the limitations of corpus and modeling capability. This study aims at building a conversational TTS for a voice agent under sequence to sequence modeling framework. We firstly construct a spontaneous conversational speech corpus well designed for the voice agent with a new recording scheme ensuring both recording quality and conversational speaking style. Secondly, we propose a conversation context-aware end-to-end TTS approach that employs an auxiliary encoder and a conversational context encoder to specifically reinforce the information about the current utterance and its context in a conversation as well. Experimental results show that the proposed approach produces more natural prosody in accordance with the conversational context, with significant preference gains at both utterance-level and conversation-level. Moreover, we find that the model has the ability to express some spontaneous behaviors like fillers and repeated words, which makes the conversational speaking style more realistic. Haohan Guo, Shaofei Zhang, Frank K. Soong, Lei He 0005, Lei Xie 0001 |
SLT | 4 |
| 2021 | Cycle consistent network for end-to-end style transfer TTS training
Liumeng Xue, Shifeng Pan, Lei He 0005, Lei Xie 0001, Frank K. Soong |
Neural Networks | 3 |
| 2020 | Using Personalized Speech Synthesis and Neural Language Generator for Rapid Speaker AdaptationabstractWe propose to use the personalized speech synthesis and the neural language generator to synthesize content relevant personalized speech for rapid speaker adaptation. It has two distinct aspects: First, it relieves the general data sparsity issue in rapid adaptation via making use of additional synthesized personalized speech; Second, it circumvents the obstacle of the explicit labeling error in unsupervised adaptation by converting it to pseudo-supervised adaptation. In this setup, the labeling error is implicitly rendered as less damaging speech distortion in the personalized synthesized speech. This results in significant performance breakthrough in the rapid unsupervised speaker adaptation. We apply the proposed methodology to a speaker adaptation task in a state-of-art speech transcription system. With 1 minute (min) adaptation data, our proposed approach yields 9.19 % or 5.98 % relative word error rate (WER) reduction for the supervised and the unsupervised adaptation, comparing to the negligible gain when adapting only with 1 min original speech. With 10 min adaptation data, it yields 12.53 % or 7.89 % relative WER reduction, doubling the gain of the baseline adaptation. The proposed approach is particularly suitable for unsupervised adaptation. Yan Huang 0028, Lei He 0005, Wenning Wei, William Gale, Jinyu Li 0001, Yifan Gong 0001 |
ICASSP | 2 |
| 2020 | Adaptation of RNN Transducer with Text-To-Speech Technology for Keyword SpottingabstractWith the advent of recurrent neural network transducer (RNN-T) model, the performance of keyword spotting (KWS) systems has greatly improved. However, the KWS systems, employed for wake-word detection, still rely on the availability of keyword specific training data for achieving reasonable performance on each keyword. With a goal to improve the KWS performance for these keywords without having to collect additional natural speech data, we explore Text-To-Speech (TTS) technology to synthetically generate training data for such keywords. Employing an RNN-T based KWS model, already well trained on large keyword-independent natural speech dataset, as a seed model, we run adaptation experiments using the generated keyword-specific TTS data. Besides observing a considerable improvement in the overall performance for the low-resource keywords, we find that the performance improvement with TTS-generated training data, similar to natural speech data, depends on speaker diversity, amount of data per speaker and data simulation. We get additional improvement in performance by selectively adapting specific parts of the RNN-T model and gain key insights into different architectural constructs of RNN-T model. Eva Sharma, Guoli Ye, Wenning Wei, Rui Zhao 0017, Jian Wu 0027, Lei He 0005, Ed Lin, Yifan Gong 0001 |
ICASSP | 7 |
| 2020 | Improving Prosody with Linguistic and Bert Derived Features in Multi-Speaker Based Mandarin Chinese Neural TTSabstractRecent advances of neural TTS have made "human parity" synthesized speech possible when a large amount of studio-quality training data from a voice talent is available. However, with only limited, casual recordings from an ordinary speaker, human-like TTS is still a big challenge, in addition to other artifacts like incomplete sentences, repetition of words, etc. Chinese, a language, of which the text is different from that of other roman-letter based languages like English, has no blank space between adjacent words, hence word segmentation errors can cause serious semantic confusions and unnatural prosody. In this study, with a multi-speaker TTS to accommodate the insufficient training data of a target speaker, we investigate linguistic features and Bert-derived information to improve the prosody of our Mandarin Chinese TTS. Three factors are studied: phone-related and prosody-related linguistic features; better predicted breaks with a refined Bert-CRF model; augmented phoneme sequence with character embedding derived from a Bert model. Subjective tests on in- and out-domain tasks of News, Chat and Audiobook, have shown that all factors are effective for improving prosody of our Mandarin TTS. The model with additional character embeddings from Bert is the best one, which outperforms the baseline by 0.17 MOS gain. Yujia Xiao, Lei He 0005, Huaiping Ming, Frank K. Soong |
ICASSP | 2 |
| 2020 | Rapid RNN-T Adaptation Using Personalized Speech Synthesis and Neural Language GeneratorabstractRapid unsupervised speaker adaptation in an E2E system posits us new challenges due to its end-to-end unified structure in addition to its intrinsic difficulty of data sparsity and imperfect label [1]. Previously we proposed utilizing the content relevant personalized speech synthesis for rapid speaker adaptation and achieved significant performance breakthrough in a hybrid system [2]. In this paper, we answer the following two questions: First, how to effectively perform rapid speaker adaptation in an RNN-T. Second, whether our previously proposed approach is still beneficial for the RNN-T and what are the modification and distinct observations. We apply the proposed methodology to a speaker adaptation task in a state-of-art presentation transcription RNN-T system. In the 1 min setup, it yields 11.58 % or 7.95 % relative word error rate (WER) reduction for the sup/unsup adaptation, comparing to the negligible gain when adapting with 1 min source speech. In the 10 min setup, it yields 15.71 % or 8.00 % relative WER reduction, doubling the gain of the source speech adaptation. We further apply various data filtering techniques and significantly bridge the gap between sup/unsup adaptation. Yan Huang 0028, Jinyu Li 0001, Lei He 0005, Wenning Wei, William Gale, Yifan Gong 0001 |
INTERSPEECH | 3 |
| 2020 | An Efficient Subband Linear Prediction for LPCNet-Based Neural Synthesis
Xi Wang 0016, Lei He 0005, Frank K. Soong |
INTERSPEECH | 3 |
| 2020 | Developing RNN-T Models Surpassing High-Performance Hybrid Models with Customization CapabilityabstractBecause of its streaming nature, recurrent neural network transducer (RNN-T) is a very promising end-to-end (E2E) model that may replace the popular hybrid model for automatic speech recognition.In this paper, we describe our recent development of RNN-T models with reduced GPU memory consumption during training, better initialization strategy, and advanced encoder modeling with future lookahead.When trained with Microsoft's 65 thousand hours of anonymized training data, the developed RNN-T model surpasses a very well trained hybrid model with both better recognition accuracy and lower latency.We further study how to customize RNN-T models to a new domain, which is important for deploying E2E models to practical scenarios.By comparing several methods leveraging text-only data in the new domain, we found that updating RNN-T's prediction and joint networks using text-to-speech generated from domain-specific text is the most effective. Jinyu Li 0001, Rui Zhao 0017, Zhong Meng, Wenning Wei, Sarangarajan Parthasarathy, Vadim Mazalov, Lei He 0005, Sheng Zhao 0002, Yifan Gong 0001 |
INTERSPEECH | 9 |
| 2019 | Learning Latent Representations for Style Control and Transfer in End-to-end Speech SynthesisabstractIn this paper, we introduce the Variational Autoencoder (VAE) to an end-to-end speech synthesis model, to learn the latent representation of speaking styles in an unsupervised manner. The style representation learned through VAE shows good properties such as disentangling, scaling, and combination, which makes it easy for style control. Style transfer can be achieved in this framework by first inferring style representation through the recognition network of VAE, then feeding it into TTS network to guide the style in synthesizing speech. To avoid Kullback-Leibler (KL) divergence collapse in training, several techniques are adopted. Finally, the proposed model shows good performance of style control and outperforms Global Style Token (GST) model in ABX preference tests on style transfer. Shifeng Pan, Lei He 0005, Zhen-Hua Ling |
ICASSP | 3 |
| 2019 | A New GAN-Based End-to-End TTS Training AlgorithmabstractEnd-to-end, autoregressive model-based TTS has shown significant performance improvements over the conventional one.However, the autoregressive module training is affected by the exposure bias, or the mismatch between the different distributions of real and predicted data.While real data is available in training, but in testing, only predicted data is available to feed the autoregressive module.By introducing both real and generated data sequences in training, we can alleviate the effects of the exposure bias.We propose to use Generative Adversarial Network (GAN) along with the key idea of Professor Forcing in training.A discriminator in GAN is jointly trained to equalize the difference between real and predicted data.In AB subjective listening test, the results show that the new approach is preferred over the standard transfer learning with a CMOS improvement of 0.1.Sentence level intelligibility tests show significant improvement in a pathological test set.The GAN-trained new model is also more stable than the baseline to produce better alignments for the Tacotron output. Haohan Guo, Frank K. Soong, Lei He 0005, Lei Xie 0001 |
INTERSPEECH | 3 |
| 2019 | Exploiting Syntactic Features in a Parsed Tree to Improve End-to-End TTSabstractThe end-to-end TTS, which can predict speech directly from a given sequence of graphemes or phonemes, has shown improved performance over the conventional TTS.However, its predicting capability is still limited by the acoustic/phonetic coverage of the training data, usually constrained by the training set size.To further improve the TTS quality in pronunciation, prosody and perceived naturalness, we propose to exploit the information embedded in a syntactically parsed tree where the inter-phrase/word information of a sentence is organized in a multilevel tree structure.Specifically, two key features: phrase structure and relations between adjacent words are investigated.Experimental results in subjective listening, measured on three test sets, show that the proposed approach is effective to improve the pronunciation clarity, prosody and naturalness of the synthesized speech of the baseline system. Haohan Guo, Frank K. Soong, Lei He 0005, Lei Xie 0001 |
INTERSPEECH | 3 |
| 2019 | Robust Sequence-to-Sequence Acoustic Modeling with Stepwise Monotonic Attention for Neural TTSabstractNeural TTS has demonstrated strong capabilities to generate human-like speech with high quality and naturalness, while its generalization to out-of-domain texts is still a challenging task, with regard to the design of attention-based sequence-tosequence acoustic modeling.Various errors occur in those inputs with unseen context, including attention collapse, skipping, repeating, etc., which limits the broader applications.In this paper, we propose a novel stepwise monotonic attention method in sequence-to-sequence acoustic modeling to improve the robustness on out-of-domain inputs.The method utilizes the strict monotonic property in TTS with constraints on monotonic hard attention that the alignments between inputs and outputs sequence must be not only monotonic but allowing no skipping on inputs.Soft attention could be used to evade mismatch between training and inference.The experimental results show that the proposed method could achieve significant improvements in robustness on out-of-domain scenarios for phoneme-based models, without any regression on the in-domain naturalness test. Mutian He 0001, Lei He 0005 |
INTERSPEECH | 3 |
| 2019 | Forward-Backward Decoding for Regularizing End-to-End TTSabstractNeural end-to-end TTS can generate very high-quality synthesized speech, and even close to human recording within similar domain text. However, it performs unsatisfactory when scaling it to challenging test sets. One concern is that the encoder-decoder with attention-based network adopts autoregressive generative sequence model with the limitation of exposure bias To address this issue, we propose two novel methods, which learn to predict future by improving agreement between forward and backward decoding sequence. The first one is achieved by introducing divergence regularization terms into model training objective to reduce the mismatch between two directional models, namely L2R and R2L (which generates targets from left-to-right and right-to-left, respectively). While the second one operates on decoder-level and exploits the future information during decoding. In addition, we employ a joint training strategy to allow forward and backward decoding to improve each other in an interactive process. Experimental results show our proposed methods especially the second one (bidirectional decoder regularization), leads a significantly improvement on both robustness and overall naturalness, as outperforming baseline (the revised version of Tacotron2) with a MOS gap of 0.14 in a challenging test, and achieving close to human quality (4.42 vs. 4.49 in MOS) on general test. Yibin Zheng, Xi Wang 0016, Lei He 0005, Shifeng Pan, Frank K. Soong, Zhengqi Wen, Jianhua Tao 0001 |
INTERSPEECH | 3 |
| 2018 | A New Glottal Neural Vocoder for Speech Synthesis
Xi Wang 0016, Lei He 0005, Frank K. Soong |
INTERSPEECH | 3 |
| 2016 | Unsupervised speaker adaptation for DNN-based TTS synthesisabstractMulti-speaker TTS trained with a general DNN has outperformed individually modelled baseline [1]. Multi-speaker DNN takes advantages of larger amount of training data from multiple speakers to find robust transformations in the hidden layers and covers more speaker variability in the output regression layer. In this paper, we propose a new approach to unsupervised speaker adaptation with multi-speaker DNN. It takes advantage of shared hidden transformation to search for the labels of unlabelled acoustic frames and the found labels are used for speaker adaption. Experimental results show that the new approach of unsupervised adaptation can achieve comparable performance with supervised adaptation both objectively and subjectively. We further extend it to cross-lingual adaptation. It can remove non-native accent and improve the naturalness while keep the same speaker's characteristics. Yuchen Fan 0001, Yao Qian, Frank K. Soong, Lei He 0005 |
ICASSP | 4 |
| 2016 | Speaker and language factorization in DNN-based TTS synthesisabstractWe have successfully proposed to use multi-speaker modelling in DNN-based TTS synthesis for improved voice quality with limited available data from a speaker. In this paper, we propose a new speaker and language factorized DNN, where speaker-specific layers are used for multi-speaker modelling, and shared layers and language-specific layers are employed for multi-language, linguistic feature transformation. Experimental results on a speech corpus of multiple speakers in both Mandarin and English show that the proposed factorized DNN can not only achieve a similar voice quality as that of a multi-speaker DNN, but also perform polyglot synthesis with a monolingual speaker's voice. Yuchen Fan 0001, Yao Qian, Frank K. Soong, Lei He 0005 |
ICASSP | 4 |
| 2016 | Learning Distributed Word Representations For Bidirectional LSTM Recurrent Neural NetworkabstractPeilu Wang, Yao Qian, Frank K. Soong, Lei He, Hai Zhao. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016. Peilu Wang, Yao Qian, Frank K. Soong, Lei He 0005 |
HLT-NAACL | 4 |
| 2016 | Modeling F0 trajectories in hierarchically structured deep neural networks
Xiang Yin 0002, Yao Qian, Frank K. Soong, Lei He 0005, Zhen-Hua Ling, Li-Rong Dai 0001 |
Speech Commun. | 5 |
| 2015 | Multi-speaker modeling and speaker adaptation for DNN-based TTS synthesisabstractIn DNN-based TTS synthesis, DNNs hidden layers can be viewed as deep transformation for linguistic features and the output layers as representation of acoustic space to regress the transformed linguistic features to acoustic parameters. The deep-layered architectures of DNN can not only represent highly-complex transformation compactly, but also take advantage of huge amount of training data. In this paper, we propose an approach to model multiple speakers TTS with a general DNN, where the same hidden layers are shared among different speakers while the output layers are composed of speaker-dependent nodes explaining the target of each speaker. The experimental results show that our approach can significantly improve the quality of synthesized speech objectively and subjectively, comparing with speech synthesized from the individual, speaker-dependent DNN-based TTS. We further transfer the hidden layers for a new speaker with limited training data and the resultant synthesized speech of the new speaker can also achieve a good quality in term of naturalness and speaker similarity. Yuchen Fan 0001, Yao Qian, Frank K. Soong, Lei He 0005 |
ICASSP | 4 |
| 2015 | Word embedding for recurrent neural network based TTS synthesisabstractThe current state of the art TTS synthesis can produce synthesized speech with highly decent quality if rich segmental and suprasegmental information are given. However, some suprasegmental features, e.g., Tone and Break (TOBI), are time consuming due to being manually labeled with a high inconsistency among different annotators. In this paper, we investigate the use of word embedding, which represents word with low dimensional continuous-valued vector and being assumed to carry a certain syntactic and semantic information, for bidirectional long short term memory (BLSTM), recurrent neural network (RNN) based TTS synthesis. Experimental results show that word embedding can significantly improve the performance of BLSTM-RNN based TTS synthesis without using features of TOBI and Part of Speech (POS). Peilu Wang, Yao Qian, Frank K. Soong, Lei He 0005 |
ICASSP | 4 |
| 2015 | Sequence generation error (SGE) minimization based deep neural networks training for text-to-speech synthesis
Yuchen Fan 0001, Yao Qian, Frank K. Soong, Lei He 0005 |
INTERSPEECH | 4 |
| 2014 | Modeling DCT parameterized F0 trajectory at intonation phrase level with DNN or decision tree
Xiang Yin 0002, Yao Qian, Frank K. Soong, Lei He 0005, Zhen-Hua Ling, Li-Rong Dai 0001 |
INTERSPEECH | 5 |