EDBT 2026 Demo / reviewers in the wild / expert
Shan Yang 0001
dblp:72/8479-1
· DBLP profile ↗
32ranked-venue papers
4as first author
22since 2021 · last 2026
0000-0003-4464-146XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 29 · 3 first-author · 20 since 2021Artificial intelligence and machine learning · 19 · 2 first-author · 12 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | UniCUE: Unified Recognition and Generation Framework for Chinese Cued Speech Video-to-Speech GenerationabstractCued Speech (CS) enhances lipreading via hand coding, offering visual phonemic cues that support precise speech perception for the hearing-impaired. The task of CS Video-to-Speech generation (CSV2S) aims to convert CS videos into intelligible speech signals. Most existing research focuses on CS Recognition (CSR), which transcribes video content into text. Consequently, a common solution for CSV2S is to integrate CSR with a text-to-speech (TTS) system. However, this pipeline relies on text as an intermediate medium, which may lead to error propagation and temporal misalignment between speech and CS video dynamics. In contrast, directly generating audio speech from CS video (direct CSV2S) often suffer from the inherent multimodal complexity and the limited availability of CS data. To address these challenges, we propose UniCUE, the first unified framework for CSV2S that directly generates speech from CS videos without relying on intermediate text. The core innovation of UniCUE lies in integrating a understanding task (CSR) that provides fine-grained CS visual-semantic cues to to guide the speech generation. Specifically, UniCUE incorporates a pose-aware visual processor, a semantic alignment pool that enables precise visual–semantic mapping, and a VisioPhonetic adapter to bridge the understanding and generation tasks within a unified architecture. To support this framework, we construct UniCUE-HI, a large-scale Mandarin CS dataset containing 11,282 videos from 14 cuers, including both hearing-impaired and normal-hearing individuals. Extensive experiments conducted on this dataset demonstrate that UniCUE achieves state-of-the-art (SOTA) performance across multiple evaluation metrics. Jinting Wang, Shan Yang 0001, Chenxing Li, Dong Yu 0001, Li Liu 0036 |
AAAI | 2 |
| 2025 | Sinba: Singing-To-Accompaniment Generation With Pitch Guidance Via Mamba-Based Language ModelabstractIn this paper, we propose Sinba, a system that can directly generate corresponding background accompaniment music from vocal input, allowing users to create complete songs using only sung vocals. Sinba adopts a decoder-only backbone network architecture. We utilize the Mamba model, which is a linear-time sequence modeling method with selective state spaces and has been proven to achieve more advanced performance than Transformers as a foundation model in long-sequence modeling tasks. However, the Mamba was initially applied to audio tasks by pre-training directly on raw audio waveform samples as the backbone model. In this paper, we convert both the training targets and inputs into discretized tokens for direct training. We also extract pitch information from the vocal input as an additional feature for the model. The proposed model is trained using source-separated data pairs. Subjective and objective experimental results demonstrate that the proposed model can generate high-quality accompaniment that matches the style and rhythm of the vocal input, outperforming the Transformerbased baseline. Synthesized audio samples are available at: https://sounddemos.github.io/sinba. Jianwei Cui 0003, Shihao Chen, Jie Zhang 0042, Chengxing Li, Shan Yang 0001, Li-Rong Dai 0001 |
ASRU | 8 |
| 2025 | DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control ConditionsabstractControlling text-to-speech (TTS) systems to synthesize speech with the prosodic characteristics expected by users has attracted much attention. To achieve controllability, current studies focus on two main directions: (1) using reference speech as prosody prompt to guide speech synthesis, and (2) using natural language descriptions to control the generation process. However, finding reference speech that exactly contains the prosody that users want to synthesize takes a lot of effort. Description-based guidance in TTS systems can only determine the overall prosody, which has difficulty in achieving fine-grained prosody control over the synthesized speech. In this paper, we propose DrawSpeech, a sketch-conditioned diffusion model capable of generating speech based on any prosody sketches drawn by users. Specifically, the prosody sketches are fed to DrawSpeech to provide a rough indication of the expected prosody trends. DrawSpeech then recovers the detailed pitch and energy contours based on the coarse sketches and synthesizes the desired speech. Experimental results show that DrawSpeech can generate speech with a wide variety of prosody and can precisely control the fine-grained prosody in a user-friendly manner. Our implementation and audio samples are publicly available1. Shan Yang 0001, Guangzhi Li, Xixin Wu |
ICASSP | 2 |
| 2025 | UniSep: Universal Target Audio Separation with Language Models at ScaleabstractWe propose Universal target audio Separation (UniSep), addressing the separation task on arbitrary mixtures of different types of audio. Distinguished from previous studies, UniSep is performed on unlimited source domains and unlimited source numbers. We formulate the separation task as a sequence-to-sequence problem, and a large language model (LLM) is used to model the audio sequence in the discrete latent space, leveraging the power of LLM in handling complex mixture audios with large-scale data. Moreover, a novel pre-training strategy is proposed to utilize audio-only data, which reduces the efforts of large-scale data simulation and enhances the ability of LLMs to understand the consistency and correlation of information within audio sequences. We also demonstrate the effectiveness of scaling datasets in an audio separation task: we use large-scale data (36.5k hours), including speech, music, and sound, to train a universal target audio separation model that is not limited to a specific domain. Experiments show that UniSep achieves competitive subjective and objective evaluation results compared with single-task models. Hangting Chen, Dongchao Yang, Guangzhi Li, Shan Yang 0001, Zhiyong Wu 0001, Helen M. Meng, Xixin Wu |
ICME | 7 |
| 2025 | Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model
Yong Ren 0006, Chenxing Li, Duzhen Zhang, Yujie Chen 0006, Manjie Xu, Ruibo Fu, Shan Yang 0001, Dong Yu 0001 |
INTERSPEECH | 9 |
| 2025 | Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning
Chenxing Li, Yong Ren 0006, Yujie Chen 0006, Ruibo Fu, Shan Yang 0001, Dong Yu 0001 |
INTERSPEECH | 7 |
| 2025 | Cued-Agent: A Collaborative Multi-Agent System for Automatic Cued Speech RecognitionabstractCued Speech (CS) is a visual communication system that combines lip-reading with hand coding to facilitate communication for individuals with hearing impairments. Automatic CS Recognition (ACSR) aims to convert CS hand gestures and lip movements into text via AI-driven methods. Traditionally, the temporal asynchrony between hand and lip movements requires the design of complex modules to facilitate effective multimodal fusion. However, constrained by limited data availability, current methods demonstrate insufficient capacity for adequately training these fusion mechanisms, resulting in suboptimal performance. Recently, multi-agent systems have shown promising capabilities in handling complex tasks with limited data availability. To this end, we propose the first collaborative multi-agent system for ACSR, named Cued-Agent. It integrates four specialized sub-agents: a Multimodal Large Language Model-based Hand Recognition agent that employs keyframe screening and CS expert prompt strategies to decode hand movements, a pretrained Transformer-based Lip Recognition agent that extracts lip features from the input video, a Hand Prompt Decoding agent that dynamically integrates hand prompts with lip features during inference in a training-free manner, and a Self-Correction Phoneme-to-Word agent that enables post-processing and end-to-end conversion from phoneme sequences to natural language sentences for the first time through semantic refinement. To support this study, we expand the existing Mandarin CS dataset by collecting data from eight hearing-impaired cuers, establishing a mixed dataset of fourteen subjects. Extensive experiments demonstrate that our Cued-Agent performs superbly in both normal and hearing-impaired scenarios compared with state-of-the-art methods. The implementation is available at https://github.com/DennisHgj/Cued-Agent. Guanjie Huang, Danny H. K. Tsang, Shan Yang 0001, Guangzhi Lei, Li Liu 0036 |
ACM Multimedia | 3 |
| 2025 | AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation
Yan Rong, Jinting Wang, Guangzhi Lei, Shan Yang 0001, Li Liu 0036 |
ACM Multimedia | 4 |
| 2023 | UniSyn: An End-to-End Unified Model for Text-to-Speech and Singing Voice SynthesisabstractText-to-speech (TTS) and singing voice synthesis (SVS) aim at generating high-quality speaking and singing voice according to textual input and music scores, respectively. Unifying TTS and SVS into a single system is crucial to the applications requiring both of them. Existing methods usually suffer from some limitations, which rely on either both singing and speaking data from the same person or cascaded models of multiple tasks. To address these problems, a simplified elegant framework for TTS and SVS, named UniSyn, is proposed in this paper. It is an end-to-end unified model that can make a voice speak and sing with only singing or speaking data from this person. To be specific, a multi-conditional variational autoencoder (MC-VAE), which constructs two independent latent sub-spaces with the speaker- and style-related (i.e. speak or sing) conditions for flexible control, is proposed in UniSyn. Moreover, supervised guided-VAE and timbre perturbation with the Wasserstein distance constraint are leveraged to further disentangle the speaker timbre and style. Experiments conducted on two speakers and two singers demonstrate that UniSyn can generate natural speaking and singing voice without corresponding training data. The proposed approach outperforms the state-of-the-art end-to-end voice generation work, which proves the effectiveness and advantages of UniSyn. Shan Yang 0001, Qicong Xie, Jixun Yao, Lei Xie 0001, Dan Su 0002 |
AAAI | 2 |
| 2023 | Multi-mode Neural Speech Coding Based on Deep Generative Networks
Shan Yang 0001, Yupeng Shi, Yuyong Kang, Dan Su 0002, Shidong Shang, Dong Yu 0001 |
INTERSPEECH | 4 |
| 2022 | Referee: Towards Reference-Free Cross-Speaker Style Transfer with Low-Quality Data for Expressive Speech SynthesisabstractCross-speaker style transfer (CSST) in text-to-speech (TTS) synthesis aims at transferring a speaking style to the synthesised speech in a target speaker’s voice. Most previous CSST approaches rely on expensive high-quality data carrying desired speaking style during training and require a reference utterance to obtain speaking style descriptors as conditioning on the generation of a new sentence. This work presents Referee, a robust reference-free CSST approach for expressive TTS, which fully leverages low-quality data to learn speaking styles from text. Referee is built by cascading a text-to-style (T2S) model with a style-to-wave (S2W) model. Phonetic PosteriorGram (PPG), phoneme-level pitch and energy contours are adopted as fine-grained speaking style descriptors, which are predicted from text using the T2S model. A novel pretrain-refinement method is adopted to learn a robust T2S model by only using readily accessible low-quality data. The S2W model is trained with high-quality target data, which is adopted to effectively aggregate style descriptors and generate high-fidelity speech in the target speaker’s voice. Experimental results are presented, showing that Referee outperforms a global-style-token (GST)-based baseline approach in CSST. Songxiang Liu, Shan Yang 0001, Dan Su 0002, Dong Yu 0001 |
ICASSP | 2 |
| 2022 | VCVTS: Multi-Speaker Video-to-Speech Synthesis Via Cross-Modal Knowledge Transfer from Voice ConversionabstractThough significant progress has been made for speaker-dependent Video-to-Speech (VTS) synthesis, little attention is devoted to multi-speaker VTS that can map silent video to speech, while allowing flexible control of speaker identity, all in a single system. This paper proposes a novel multi-speaker VTS system based on cross-modal knowledge transfer from voice conversion (VC), where vector quantization with contrastive predictive coding (VQCPC) is used for the content encoder of VC to derive discrete phoneme-like acoustic units, which are transferred to a Lip-to-Index (Lip2Ind) network to infer the index sequence of acoustic units. The Lip2Ind network can then substitute the content encoder of VC to form a multi-speaker VTS system to convert silent video to acoustic units for reconstructing accurate spoken content. The VTS system also inherits the advantages of VC by using a speaker encoder to produce speaker representations to effectively control the speaker identity of generated speech. Extensive evaluations verify the effectiveness of proposed approach, which can be applied in both constrained vocabulary and open vocabulary conditions, achieving state-of-the-art performance in generating high-quality speech with high naturalness, intelligibility and speaker similarity. Our demo page is released here1. Disong Wang, Shan Yang 0001, Dan Su 0002, Xunying Liu, Dong Yu 0001, Helen M. Meng |
ICASSP | 2 |
| 2022 | Glow-WaveGAN 2: High-quality Zero-shot Text-to-speech Synthesis and Any-to-any Voice ConversionabstractThe zero-shot scenario for speech generation aims at synthesizing a novel unseen voice with only one utterance of the target speaker.Although the challenges of adapting new voices in zero-shot scenario exist in both stages -acoustic modeling and vocoder, previous works usually consider the problem from only one stage.In this paper, we extend our previous Glow-WaveGAN to Glow-WaveGAN 2, aiming to solve the problem from both stages for high-quality zero-shot text-to-speech and any-to-any voice conversion.We first build a universal Wave-GAN model for extracting latent distribution p(z) of speech and reconstructing waveform from it.Then a flow-based acoustic model only needs to learn the same p(z) from texts, which naturally avoids the mismatch between the acoustic model and the vocoder, resulting in high-quality generated speech without model fine-tuning.Based on a continuous speaker space and the reversible property of flows, the conditional distribution can be obtained for any speaker, and thus we can further conduct highquality zero-shot speech generation for new speakers.We particularly investigate two methods to construct the speaker space, namely pre-trained speaker encoder and jointly-trained speaker encoder.The superiority of Glow-WaveGAN 2 has been proved through TTS and VC experiments conducted on LibriTTS corpus and VTCK corpus. Shan Yang 0001, Jian Cong, Lei Xie 0001, Dan Su 0002 |
INTERSPEECH | 2 |
| 2022 | Learning Noise-independent Speech Representation for High-quality Voice Conversion for Noisy Target SpeakersabstractBuilding a voice conversion system for noisy target speakers, such as users providing noisy samples or Internet found data, is a challenging task since the use of contaminated speech in model training will apparently degrade the conversion performance.In this paper, we leverage the advances of our recently proposed Glow-WaveGAN [1] and propose a noise-independent speech representation learning approach for high-quality voice conversion for noisy target speakers.Specifically, we learn a latent feature space where we ensure that the target distribution modeled by the conversion model is exactly from the modeled distribution of the waveform generator.With this premise, we further manage to make the latent feature to be noise-invariant.Specifically, we introduce a noise-controllable WaveGAN, which directly learns the noise-independent acoustic representation from waveform by the encoder and conducts noise control in the hidden space through a FiLM [2] module in the decoder.As for the conversion model, importantly, we use a flow-based model to learn the distribution of noiseindependent but speaker-related latent features from phoneme posteriorgrams.Experimental results demonstrate that the proposed model achieves high speech quality and speaker similarity in the voice conversion for noisy target speakers. Liumeng Xue, Shan Yang 0001, Na Hu, Dan Su 0002, Lei Xie 0001 |
INTERSPEECH | 2 |
| 2022 | Cross-Speaker Emotion Transfer Through Information Perturbation in Emotional Speech SynthesisabstractThrough borrowing emotional expressions from an emotional speaker, cross-speaker emotion transfer is an effective way to produce emotional speech for target speakers without emotional training data. Since emotion and timbre of the source speaker are heavily entangled in speech, existing approaches often struggle to trade off between speaker similarity and emotional expression in the synthetic speech of the target speaker. In this letter, we propose to disentangle timbre and emotion through information perturbation to conduct cross-speaker emotion transfer, which effectively learns the emotional expression of the source speaker and maintains the timbre of the target speaker. Specifically, we separately perturb the timbre and emotion-related features (e.g., formant and pitch) of source speech to obtain and model the timbre- and emotion-independent signals, based on which the proposed model can deliver the emotional expression for target speakers. Experimental results demonstrate the proposed approach significantly outperforms the baselines in terms of naturalness and similarity, indicating the effectiveness of information perturbation for cross-speaker emotion transfer. Shan Yang 0001, Xinfa Zhu, Lei Xie 0001, Dan Su 0002 |
IEEE Signal Process. Lett. | 2 |
| 2022 | MsEmoTTS: Multi-Scale Emotion Transfer, Prediction, and Control for Emotional Speech SynthesisabstractExpressive synthetic speech is essential for many human-computer interaction and audio broadcast scenarios, and thus synthesizing expressive speech has attracted much attention in recent years. Previous methods performed the expressive speech synthesis either with explicit labels or with a fixed-length style embedding extracted from reference audio, both of which can only learn an average style and thus ignores the multi-scale nature of speech prosody. In this paper, we propose MsEmoTTS, a multi-scale emotional speech synthesis framework, to model the emotion from different levels. Specifically, the proposed method is a typical attention-based sequence-to-sequence model and with proposed three modules, including global-level emotion presenting module (GM), utterance-level emotion presenting module (UM), and local-level emotion presenting module (LM), to model the global emotion category, utterance-level emotion variation, and syllable-level emotion strength, respectively. In addition to modeling the emotion from different levels, the proposed method also allows us to synthesize emotional speech in different ways, i.e., transferring the emotion from reference audio, predicting the emotion from input text, and controlling the emotion strength manually. Extensive experiments conducted on a Chinese emotional speech corpus demonstrate that the proposed method outperforms the compared reference audio-based and text-based emotional speech synthesis methods on the emotion transfer speech synthesis and text-based emotion prediction speech synthesis respectively. Besides, the experiments also show that the proposed method can control the emotion expressions flexibly. Detailed analysis shows the effectiveness of each module and the good design of the proposed method. Shan Yang 0001, Lei Xie 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2021 | Controllable Context-Aware Conversational Speech SynthesisabstractIn spoken conversations, spontaneous behaviors like filled pause and prolongations always happen.Conversational partner tends to align features of their speech with their interlocutor which is known as entrainment.To produce human-like conversations, we propose a unified controllable spontaneous conversational speech synthesis framework to model the above two phenomena.Specifically, we use explicit labels to represent two typical spontaneous behaviors filled-pause and prolongation in the acoustic model and develop a neural network based predictor to predict the occurrences of the two behaviors from text.We subsequently develop an algorithm based on the predictor to control the occurrence frequency of the behaviors, making the synthesized speech vary from less disfluent to more disfluent.To model the speech entrainment at acoustic level, we utilize a context acoustic encoder to extract a global style embedding from the previous speech conditioning on the synthesizing of current speech.Furthermore, since the current and previous utterances belong to the different speakers in a conversation, we add a domain adversarial training module to eliminate the speaker-related information in the acoustic encoder while maintaining the style-related information.Experiments show that our proposed approach can synthesize realistic conversations and control the occurrences of the spontaneous behaviors naturally. Jian Cong, Shan Yang 0001, Na Hu, Guangzhi Li, Lei Xie 0001, Dan Su 0002 |
Interspeech | 2 |
| 2021 | Glow-WaveGAN: Learning Speech Representations from GAN-Based Variational Auto-Encoder for High Fidelity Flow-Based Speech SynthesisabstractCurrent two-stage TTS framework typically integrates an acoustic model with a vocoder -- the acoustic model predicts a low resolution intermediate representation such as Mel-spectrum while the vocoder generates waveform from the intermediate representation. Although the intermediate representation is served as a bridge, there still exists critical mismatch between the acoustic model and the vocoder as they are commonly separately learned and work on different distributions of representation, leading to inevitable artifacts in the synthesized speech. In this work, different from using pre-designed intermediate representation in most previous studies, we propose to use VAE combining with GAN to learn a latent representation directly from speech and then utilize a flow-based acoustic model to model the distribution of the latent representation from text. In this way, the mismatch problem is migrated as the two stages work on the same distribution. Results demonstrate that the flow-based acoustic model can exactly model the distribution of our learned speech representation and the proposed TTS framework, namely Glow-WaveGAN, can produce high fidelity speech outperforming the state-of-the-art GAN-based model. Jian Cong, Shan Yang 0001, Lei Xie 0001, Dan Su 0002 |
Interspeech | 2 |
| 2021 | Fine-Grained Emotion Strength Transfer, Control and Prediction for Emotional Speech SynthesisabstractThis paper proposes a unified model to conduct emotion transfer, control and prediction for sequence-to-sequence based fine-grained emotional speech synthesis. Conventional emotional speech synthesis often needs manual labels or reference audio to determine the emotional expressions of synthesized speech. Such coarse labels cannot control the details of speech emotion, often resulting in an averaged emotion expression delivery, and it is also hard to choose suitable reference audio during inference. To conduct fine-grained emotion expression generation, we introduce phoneme-level emotion strength representations through a learned ranking function to describe the local emotion details, and the sentence-level emotion category is adopted to render the global emotions of synthesized speech. With the global render and local descriptors of emotions, we can obtain fine-grained emotion expressions from reference audio via its emotion descriptors (for transfer) or directly from phoneme-level manual labels (for control). As for the emotional speech synthesis with arbitrary text inputs, the proposed model can also predict phoneme-level emotion expressions from texts, which does not require any reference audio or manual label. Shan Yang 0001, Lei Xie 0001 |
SLT | 2 |
| 2021 | Learn2Sing: Target Speaker Singing Voice Synthesis by Learning from a Singing TeacherabstractSinging voice synthesis has been paid rising attention with the rapid development of speech synthesis area. In general, a studio-level singing corpus is usually necessary to produce a natural singing voice from lyrics and music-related transcription. However, such a corpus is difficult to collect since it's hard for many of us to sing like a professional singer. In this paper, we propose an approach - Learn2Sing that only needs a singing teacher to generate the target speakers' singing voice without their singing voice data. In our approach, a teacher's singing corpus and speech from multiple target speakers are trained in a frame-level auto-regressive acoustic model where singing and speaking share the common speaker embedding and style tag embedding. Meanwhile, since there is no music-related transcription for the target speaker, we use log-scale fundamental frequency (LF0) as an auxiliary feature as the inputs of the acoustic model for building a unified input representation. In order to enable the target speaker to sing without singing reference audio in the inference stage, a duration model and an LF0 prediction model are also trained. Particularly, we employ domain adversarial training (DAT) in the acoustic model, which aims to enhance the singing performance of target speakers by disentangling style from acoustic features of singing and speaking data. Our experiments indicate that the proposed approach is capable of synthesizing singing voice for target speaker given only their speech samples. Heyang Xue, Shan Yang 0001, Lei Xie 0001, Xiulin Li |
SLT | 2 |
| 2021 | Multi-Band Melgan: Faster Waveform Generation For High-Quality Text-To-SpeechabstractIn this paper, we propose multi-band MelGAN, a much faster waveform generation model targeting to high-quality text-to-speech. Specifically, we improve the original MelGAN by the following aspects. First, we increase the receptive field of the generator, which is proven to be beneficial to speech generation. Second, we substitute the feature matching loss with the multi-resolution STFT loss to better measure the difference between fake and real speech. Together with pre-training, this improvement leads to both better quality and better training stability. More importantly, we extend MelGAN with multi-band processing: the generator takes mel-spectrograms as input and produces sub-band signals which are subsequently summed back to full-band signals as discriminator input. The proposed multi-band MelGAN has achieved high MOS of 4.34 and 4.22 in waveform generation and TTS, respectively. With only 1.91M parameters, our model effectively reduces the total computational complexity of the original MelGAN from 5.85 to 0.95 GFLOPS. Our Pytorch implementation can achieve a real-time factor of 0.03 on CPU without hardware specific optimization. Shan Yang 0001, Kai Liu 0053, Wei Chen 0071, Lei Xie 0001 |
SLT | 2 |
| 2021 | Effective and direct control of neural TTS prosody by removing interactions between different attributes
Xiaochun An, Frank K. Soong, Shan Yang 0001, Lei Xie 0001 |
Neural Networks | 3 |
| 2020 | Data Efficient Voice Cloning from Noisy Samples with Domain Adversarial TrainingabstractData efficient voice cloning aims at synthesizing target speaker's voice with only a few enrollment samples at hand.To this end, speaker adaptation and speaker encoding are two typical methods based on base model trained from multiple speakers.The former uses a small set of target speaker data to transfer the multi-speaker model to target speaker's voice through direct model update, while in the latter, only a few seconds of target speaker's audio directly goes through an extra speaker encoding model along with the multi-speaker model to synthesize target speaker's voice without model update.Nevertheless, the two methods need clean target speaker data.However, the samples provided by user may inevitably contain acoustic noise in real applications.It's still challenging to generating target voice with noisy data.In this paper, we study the data efficient voice cloning problem from noisy samples under the sequenceto-sequence based TTS paradigm.Specifically, we introduce domain adversarial training (DAT) to speaker adaptation and speaker encoding, which aims to disentangle noise from speechnoise mixture.Experiments show that for both speaker adaptation and encoding, the proposed approaches can consistently synthesize clean speech from noisy speaker samples, apparently outperforming the method adopting state-of-the-art speech enhancement module. Jian Cong, Shan Yang 0001, Lei Xie 0001, Guoqiao Yu, Guanglu Wan |
INTERSPEECH | 2 |
| 2020 | Exploiting Deep Sentential Context for Expressive End-to-End Speech SynthesisabstractAttention-based seq2seq text-to-speech systems, especially those use self-attention networks (SAN), have achieved state-of-art performance. But an expressive corpus with rich prosody is still challenging to model as 1) prosodic aspects, which span across different sentential granularities and mainly determine acoustic expressiveness, are difficult to quantize and label and 2) the current seq2seq framework extracts prosodic information solely from a text encoder, which is easily collapsed to an averaged expression for expressive contents. In this paper, we propose a context extractor, which is built upon SAN-based text encoder, to sufficiently exploit the sentential context over an expressive corpus for seq2seq-based TTS. Our context extractor first collects prosodic-related sentential context information from different SAN layers and then aggregates them to learn a comprehensive sentence representation to enhance the expressiveness of the final generated speech. Specifically, we investigate two methods of context aggregation: 1) direct aggregation which directly concatenates the outputs of different SAN layers, and 2) weighted aggregation which uses multi-head attention to automatically learn contributions for different SAN layers. Experiments on two expressive corpora show that our approach can produce more natural speech with much richer prosodic variations, and weighted aggregation is more superior in modeling expressivity. Fengyu Yang 0002, Shan Yang 0001, Lei Xie 0001 |
INTERSPEECH | 2 |
| 2020 | On the localness modeling for the self-attention based end-to-end speech synthesis
Shan Yang 0001, Heng Lu 0004, Shiyin Kang, Liumeng Xue, Jinba Xiao, Dan Su 0002, Lei Xie 0001, Dong Yu 0001 |
Neural Networks | 1 |
| 2020 | Adversarial Feature Learning and Unsupervised Clustering Based Speech Synthesis for Found Data With Acoustic and Textual NoiseabstractAttention-based sequence-to-sequence (seq2seq) speech synthesis has achieved extraordinary performance. But a studio-quality corpus with manual transcription is necessary to train such seq2seq systems. In this letter, we propose an approach to build high-quality and stable seq2seq based speech synthesis system using challenging found data, where training speech contains noisy interferences (acoustic noise) and texts are imperfect speech recognition transcripts (textual noise). To deal with text-side noise, we propose a VQVAE based heuristic method to compensate erroneous linguistic feature with phonetic information learned directly from speech. As for the speech-side noise, we propose to learn a noise-independent feature in the auto-regressive decoder through adversarial training and data augmentation, which does not need an extra speech enhancement model. Experiments show the effectiveness of the proposed approach in dealing with text-side and speech-side noise. Surpassing the denoising approach based on a state-of-the-art speech enhancement model, our system built on noisy found data can synthesize clean and high-quality speech with MOS close to the system built on the clean counterpart. Shan Yang 0001, Yuxuan Wang 0002, Lei Xie 0001 |
IEEE Signal Process. Lett. | 1 |
| 2019 | Learning Hierarchical Representations for Expressive Speaking Style in End-to-End Speech SynthesisabstractAlthough Global Style Tokens (GSTs) are a recently-proposed method to uncover expressive factors of variation in speaking style, they are a mixture of style attributes without explicitly considering the factorization of multiple-level speaking styles. In this work, we introduce a hierarchical GST architecture with residuals to Tacotron, which learns multiple-level disentangled representations to model and control different style granularities in synthesized speech. We make hierarchical evaluations conditioned on individual tokens from different GST layers. As the number of layers increases, we tend to observe a coarse to fine style decomposition. For example, the first GST layer learns a good representation of speaker IDs while finer speaking style or emotion variations can be found in higher-level layers. Meanwhile, the proposed model shows good performance of style transfer. Xiaochun An, Yuxuan Wang 0002, Shan Yang 0001, Zejun Ma 0001, Lei Xie 0001 |
ASRU | 3 |
| 2019 | Improving Mandarin End-to-End Speech Synthesis by Self-Attention and Learnable Gaussian BiasabstractCompared to conventional speech synthesis, end-to-end speech synthesis has achieved much better naturalness with more simplified system building pipeline. End-to-end framework can generate natural speech directly from characters for English. But for other languages like Chinese, recent studies have indicated that extra engineering features are still needed for model robustness and naturalness, e.g, word boundaries and prosody boundaries, which makes the front-end pipeline as complicated as the traditional approach. To maintain the naturalness of generated speech and discard language-specific expertise as much as possible, in Mandarin TTS, we introduce a novel self-attention based encoder with learnable Gaussian bias in Tacotron. We evaluate different systems with and without complex prosody information and results show that the proposed approach has the ability to generate stable and natural speech with minimum language-dependent front-end modules. Fengyu Yang 0001, Shan Yang 0001, Pengcheng Zhu 0004, Pengju Yan, Lei Xie 0001 |
ASRU | 2 |
| 2019 | Controlling Emotion Strength with Relative Attribute for End-to-End Speech SynthesisabstractRecently, attention-based end-to-end speech synthesis has achieved superior performance compared to traditional speech synthesis models, and several approaches like global style tokens are proposed to explore the style controllability of the end-to-end model. Although the existing methods show good performance in style disentanglement and transfer, it is still unable to control the explicit emotion of generated speech. In this paper, we mainly focus on the subtle control of expressive speech synthesis, where the emotion category and strength can be easily controlled with a discrete emotional vector and a continuous simple scalar, respectively. The continuous strength controller is learned by a ranking function according to the relative attribute measured on an emotion dataset. Our method automatically learns the relationship between low-level acoustic features and high-level subtle emotion strength. Experiments show that our method can effectively improve the controllability for an expressive end-to-end model. Xiaolian Zhu, Shan Yang 0001, Lei Xie 0001 |
ASRU | 2 |
| 2019 | Enhancing Hybrid Self-attention Structure with Relative-position-aware Bias for Speech SynthesisabstractCompared with the conventional "front-end"-"back-end"- "vocoder" structure, based on the attention mechanism, end-to-end speech synthesis systems directly train and synthesize from text sequence to the acoustic feature sequence as a whole. Recently, a more calculation efficient end-to-end architecture named transformer, which is solely based on self-attention, was proposed to model global dependencies between the input and output sequences. However, although with many advantages, transformer lacks position information in its structure. Moreover, the weighted sum form in self-attention may disperse the attention to the whole input sequence other than focusing on the more important neighbouring positions. In order to solve the above problems, this paper introduces a hybrid self-attention structure which combines self-attention with the recurrent neural networks (RNNs). We further enhance the proposed structure with relative-position-aware biases. Mean opinion score (MOS) test results indicate that by enhancing hybrid self-attention structure with relative-position-aware biases, the proposed system achieves the best performance with only 0.11 MOS score lower than natural recording. Shan Yang 0001, Heng Lu 0004, Shiying Kang, Lei Xie 0001, Dong Yu 0001 |
ICASSP | 1 |
| 2017 | Statistical parametric speech synthesis using generative adversarial networks under a multi-task learning frameworkabstractIn this paper, we aim at improving the performance of synthesized speech in statistical parametric speech synthesis (SPSS) based on a generative adversarial network (GAN). In particular, we propose a novel architecture combining the traditional acoustic loss function and the GAN's discriminative loss under a multi-task learning (MTL) framework. The mean squared error (MSE) is usually used to estimate the parameters of deep neural networks, which only considers the numerical difference between the raw audio and the synthesized one. To mitigate this problem, we introduce the GAN as a second task to determine if the input is a natural speech with specific conditions. In this MTL framework, the MSE optimization improves the stability of GAN, and at the same time GAN produces samples with a distribution closer to natural speech. Listening tests show that the multi-task architecture can generate more natural speech that satisfies human perception than the conventional methods. Shan Yang 0001, Lei Xie 0001, Xiaoyan Lou, Dong-Yan Huang, Haizhou Li 0001 |
ASRU | 1 |
| 2016 | A deep bidirectional LSTM approach for video-realistic talking head
Lei Xie 0001, Shan Yang 0001, Frank K. Soong |
Multim. Tools Appl. | 3 |