Ya-Jun Hu

dblp:176/4075 · DBLP profile ↗
← Back
14ranked-venue papers
5as first author
7since 2021 · last 2025
0000-0003-0624-6119ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2025 Self-supervised Prosody Learning at Phoneme-level with Momentum Contrast for Speech Synthesis
abstract
This paper investigates leveraging large-scale speech data to enhance prosodic modeling in speech synthesis, and introduces a model named SP2MC which achieves self-supervised prosody learning at phoneme-level with momentum contrast. This model incorporates dual convolutional encoders for speech and linear predictive coding (LPC) residual inputs to generate phoneme-level embeddings, which are masked and processed by a Transformer model to produce prosody representations. Two supervision modules are employed to generate phoneme-level supervision from speech waveforms and residuals. Momentum contrast is utilized to manage negative sample selection in contrastive learning. Finally, the SP2MC representations are integrated into a Fastspeech2-based acoustic model for speech synthesis. Experimental results indicate that the naturalness of speech synthesized by the proposed method is significantly better than that of baselines.
Zhaoci Liu, Ya-Jun Hu, Zhen-Hua Ling
ICASSP2
2025 Anchored Monotonic Alignment and Representation Substitution for Rare Spontaneous Behaviors in Spontaneous Speech Synthesis
abstract
Spontaneous behaviors in speech pose significant challenges for speech synthesis. Existing research has not adequately addressed these behaviors, with most studies relying on specially recorded datasets. In contrast, real-world data more accurately reflects the natural, spontaneous speaking styles in everyday life and encompasses a wider range of spontaneous behaviors. However, such data is often of lower quality, and the distribution of spontaneous behaviors is highly imbalanced. In this study, we explore spontaneous speech synthesis using real-world data within the VITS2 framework. To overcome these challenges, we introduce two techniques: anchored monotonic alignment and spontaneous hidden representation substitution. Experimental results demonstrate that these methods enhance model alignment and improve the naturalness of the generated speech. Our proposed approach successfully addresses the challenge of synthesizing rare spontaneous behaviors and offers users flexible control over the synthesized speech.
Ning-Qian Wu, Ya-Jun Hu, Zhen-Hua Ling
ICASSP2
2024 Language-Independent Prosody-Enhanced Speech Representations For Multilingual Speech Synthesis
abstract
This paper proposes language-independent prosody-enhanced speech representations to improve the naturalness of speech synthesis for the target languages that lack prosodic labels. To build text-to-speech (TTS) systems for low-resource languages, recent studies have employed the representations extracted from self-supervised learning (SSL) speech models, such as wav2vec 2.0, as intermediate representations in TTS models. However, they have generally focused only on the linguistic and phonetic information in SSL representations, disregarding the prosodic information. This paper investigates the prosodic information contained in the multilingual wav2vec 2.0 model through layer-wise probing tests utilizing acoustic prosodic features and prosodic labels. Furthermore, we propose a language-independent prosody enhancement approach to improve the prosodic properties of SSL models. The proposed method introduces a prosodic label prediction loss to fine-tune wav2vec 2.0 model with multilingual prosody-annotated corpora. From the fine-tuned wav 2 vec 2.0 model, the language-independent prosody-enhanced speech representations are extracted and serve as intermediate representations of our acoustic model in the downstream TTS task. The experimental results on six target languages demonstrate that our proposed prosody-enhanced speech representations outperform the original wav2vec 2.0 representations without enhancement.
Chang Liu 0140, Zhen-Hua Ling, Ya-Jun Hu
SLT3
2024 PE-Wav2vec: A Prosody-Enhanced Speech Model for Self-Supervised Prosody Learning in TTS
abstract
This paper investigates leveraging large-scale untranscribed speech data to enhance the prosody modelling capability oftext-to-speech(TTS) models. On the basis of the self-supervised speech model wav2vec 2.0,Prosody-Enhanced wav2vec(PE-wav2vec) is proposed by introducing prosody learning. Specifically, prosody learning is achieved by applying supervision from thelinear predictive coding(LPC) residual signals on the initial Transformer blocks in the wav2vec 2.0 architecture. The embedding vectors extracted with the initial Transformer blocks of the PE-wav2vec model are utilised as prosodic representations for the corresponding frames in a speech utterance. To apply the PE-wav2vec representations in TTS, an acoustic model namedSpeech Synthesis model conditioned on Self-Supervisedly Learned Prosodic Representations(S4LPR) is designed on the basis of FastSpeech 2. The experimental results demonstrate that the proposed PE-wav2vec model can provide richer prosody descriptions of speech than the vanilla wav2vec 2.0 model can. Furthermore, the S4LPR model using PE-wav2vec representations can effectively improve the subjective naturalness and reduce the objective distortions of synthetic speech compared with baseline models.
Zhaoci Liu, Ya-Jun Hu, Zhen-Hua Ling
IEEE ACM Trans. Audio Speech Lang. Process.3
2023 Speech Synthesis with Self-Supervisedly Learnt Prosodic Representations
Zhaoci Liu, Zhen-Hua Ling, Ya-Jun Hu, Jin-Wei Wang, Yun-Di Wu
INTERSPEECH3
2022 Improving Recognition-Synthesis Based any-to-one Voice Conversion with Cyclic Training
abstract
In recognition-synthesis based any-to-one voice conversion (VC), an automatic speech recognition (ASR) model is employed to extract content-related features and a synthesizer is built to predict the acoustic features of the target speaker from the content-related features of any source speakers at the conversion stage. Since source speakers are unknown at the training stage, we have to use the content-related features of the target speaker to estimate the parameters of the synthesizer. This inconsistency between conversion and training stages constrains the speaker similarity of converted speech. To address this issue, a cyclic training method is proposed in this paper. This method designs pseudo-source acoustic features, which are generated by converting the training data of the target speaker towards multiple speakers in a reference corpus. Then, these pseudo-source acoustic features are used as the input of the synthesizer at the training stage to predict the acoustic features of the target speaker and a cyclic reconstruction loss is derived. Experimental results show that our proposed method achieved more consistent accuracy of acoustic feature prediction for various source speakers than the baseline method. It also achieved better similarity of converted speech, especially for the pairs of source and target speakers with distant speaker characteristics.
Yan-Nian Chen, Li-Juan Liu, Ya-Jun Hu, Yuan Jiang 0006, Zhen-Hua Ling
ICASSP3
2022 Neural Grapheme-To-Phoneme Conversion with Pre-Trained Grapheme Models
abstract
Neural network models have achieved state-of-the-art performance on grapheme-to-phoneme (G2P) conversion. However, their performance relies on large-scale pronunciation dictionaries, which may not be available for a lot of languages. Inspired by the success of the pre-trained language model BERT, this paper proposes a pre-trained grapheme model called grapheme BERT (GBERT), which is built by self-supervised training on a large, language-specific word list with only grapheme information. Furthermore, two approaches are developed to incorporate GBERT into the state-of-the-art Transformer-based G2P model, i.e., fine-tuning GBERT or fusing GBERT into the Transformer model by attention. Experimental results on the Dutch, Serbo-Croatian, Bulgarian and Korean datasets of the SIGMORPHON 2021 G2P task confirm the effectiveness of our GBERT-based G2P models under both medium-resource and low-resource data conditions.
Lu Dong 0005, Zhiqiang Guo, Chao-Hong Tan, Ya-Jun Hu, Yuan Jiang 0006, Zhen-Hua Ling
ICASSP4
2018 Extracting Spectral Features Using Deep Autoencoders With Binary Distributed Hidden Units for Statistical Parametric Speech Synthesis
abstract
This paper presents a spectral feature extraction method using deep autoencoders (DAEs) with binary distributed hidden units (BDAE) for statistical parametric speech synthesis (SPSS). Conventional DAEs are trained to minimize the error of reconstructing raw features. In this paper, we investigate another important property of DAEs that may influence their performances as feature extractors for regression tasks, i.e., the degree of binarization of hidden units. Our analysis shows that making the hidden units of DAEs to be binary may help alleviate the over-smoothing effect caused by acoustic modeling and parameter generation, which are one of the main deficiencies of current SPSS systems. This paper further proposes an effective BDAE training method by adding noise to the input of hidden units during model training and applying DBN-based pretraining strategies. Our experiments adopt feedforward deep neural networks as acoustic models for SPSS and compare the performances of different spectral feature extractors. Experimental results show that when extracting low-dimensional spectral features by BDAEs, the predicted spectral features can reconstruct spectral envelopes closer to natural samples than using conventional DAEs. Subjective evaluations on the synthetic voices of a Chinese speaker and an English speaker demonstrate that BDAEs achieve better naturalness of synthetic speech than conventional mel-cepstra and other neural network based feature extractors, such as DAEs and DBNs.
Ya-Jun Hu, Zhen-Hua Ling
IEEE ACM Trans. Audio Speech Lang. Process.1
2017 The USTC system for blizzard machine learning challenge 2017-ES2
abstract
The Blizzard Machine Learning Challenge (BMLC) aims to liberate participants from speech-specific processing when building speech synthesis systems. This paper describes the USTC system for the ES2 sub-task in BMLC2017, which requires participants to train a model to directly predict waveforms from linguistic features. We investigate three aspects of waveform modeling when preparing our system for this task. First, two different model structures for waveform modeling, i.e., WaveNet and SampleRNN, are compared on this task. Second, a strategy of using features extracted from waveforms as intermediate representations for waveform modeling is studied. Experimental results show that using low-level features (STFT amplitude spectra) as intermediate representations can achieve similar performance as using high-level features (mel-cepstra and F0). Third, the feasibility of applying WaveNet to wideband speech signals with more than 256 quantization levels is verified by experiments. Finally, a system which adopts STFT amplitude spectra as intermediate representations to model 24kHz speech waveforms with 1024 mu-law quantization levels is submitted for evaluation. The evaluation results of BMLC2017 demonstrate the effectiveness of our proposed methods.
Ya-Jun Hu, Li-Juan Liu, Chuang Ding, Zhen-Hua Ling, Li-Rong Dai 0001
ASRU1
2017 The iFLYTEK system for blizzard machine learning challenge 2017-ES1
abstract
This paper introduces the speech synthesis system submitted by IFLYTEK for the Blizzard Machine Learning Challenge 2017-ES1. Linguistic and acoustic features from a 4hour corpus were released for this task. Participants are expected to build a speech synthesis system on the given linguist and acoustic features without using any external data. Our system is composed of a long short term memory (LSTM) recurrent neural network (RNN)-based acoustic model and a generative adversarial network (GAN)-based post-filter for mel-cepstra. Two approaches to build GAN-based post-filter are implemented and compared in our experiments. The first one is to predict the residuals of mel-cepstra given the mel-cepstra predicted by the LSTM-based acoustic model. However, this method leads to unstable synthetic speech sounds in our experiments, which may be due to the poor quality of analysis-synthesis speech using the natural acoustic features given by this corpus. The other approach is to ignore the detailed components of natural mel-cepstra by dimension reduction using principal component analysis (PCA) and then recover them back using GAN given the main PCA components. At synthesis time, mel-cepstra predicted by the RNN acoustic model are first projected to the main PCA components, which are then sent to the GAN for detail recovering. Finally, the second approach is used in the final submitted system. The evaluation results show the effectiveness of our submitted system.
Li-Juan Liu, Chuang Ding, Ya-Jun Hu, Zhen-Hua Ling, Yuan Jiang 0006, Si Wei
ASRU3
2017 Extracting structural spectral features using what-where auto-encoders for statistical parametric speech synthesis
abstract
This paper presents a method to extract structural spectral features from spectral envelopes using what-where autoencoders (WWAE) for statistical parametric speech synthesis (SPSS). A WWAE is constructed by concatenating a convolutional net for input encoding and a deconvolutional net for reconstruction. The output values of the max-pooling layer in the encoder and the positions of the max-pooling switches are utilized as the what and where features respectively. Considering the intrinsic formant structures in the spectral envelopes of voiced speech frames, the WWAE model is adopted in this paper to detect, locate, and reconstruct the formants and other local structures in spectral envelopes. Here, the what and where features describe the prominences and positions of specific local spectral structures within a pooling frequency window. Then, the extracted what and where features are modeled as separate streams under the hidden Markov model (HMM)-based SPSS framework. Experimental results show that the speech synthesis system built using our proposed spectral features can produce synthetic speech with sharper formant structures and better naturalness than the systems using mel-cepstra and conventional auto-encoder-based spectral features.
Ya-Jun Hu, Zhen-Hua Ling, Li-Rong Dai 0001
ICASSP1
2016 Deep belief network-based post-filtering for statistical parametric speech synthesis
abstract
The speech synthesized by statistical parametric speech synthesis (SPSS) always sounds muffled. One important reason is that the generated spectral envelopes are over-smoothed and many detailed spectral structures in natural speech are lost. This paper presents a deep belief network (DBN)-based post-filtering method for hidden Markov model (HMM)-based SPSS to address this issue. At training time, a DBN is estimated using the spectral envelopes extracted from natural speech. This DBN serves as a generatively trained postfilter which processes the spectral envelopes recovered from the predicted spectral features at synthesis time. Experimental results show that the effectiveness of this method depends on the sampling strategy used to generate the training data of the restricted Boltzmann machines (RBM) which forms the higher layers of the DBN. When binary samples are adopted instead of mean-filed approximation, the DBN post-filter can alleviate the over-smoothing effect of parameter generation and improve the naturalness of synthetic speech significantly when either mel-cepstra or line spectral pairs (LSP) are used as spectral features. Its performance is comparative with the parameter generation method with global variance (GV) modeling for mel-cepstra and better than the LSP-based formant enhancement method used in previous work.
Ya-Jun Hu, Zhen-Hua Ling, Li-Rong Dai 0001
ICASSP1
2016 Modeling spectral envelopes using deep conditional restricted Boltzmann machines for statistical parametric speech synthesis
abstract
This paper proposes a spectral modeling method using a deep conditional restricted Boltzmann machine (DCRBM) for statistical parametric speech synthesis. In this method, a DCRBM, which combines a deep neural network (DNN) with a conditional restricted Boltzmann machine (CRBM), is utilized to describe the conditional distribution of spectral envelopes given linguistic features. Compared with DNN and deep mixture density network (DMDN), DCRBM is better at describing the multimodal distribution of high-dimensional acoustic features with cross-dimension correlations. At training stage, the DNN part and the CRBM part of the DCRBM are pre-trained successively and then a unified fine-tuning of all model parameters is conducted. At synthesis time, spectral envelopes are generated from the estimated DCRBM model by iterative sampling and dynamic-feature-constrained parameter generation given linguistic features of input text. Experimental results show that our proposed method can produce more natural speech sounds than the hidden Markov model (HMM)-based, DNN-based, and DMDN-based synthesis methods. This method also outperforms previous work which adopts restricted Boltzmann machines (RBM) to model the distributions of spectral envelopes at HMM states.
Xiang Yin 0002, Zhen-Hua Ling, Ya-Jun Hu, Li-Rong Dai 0001
ICASSP3
2016 DBN-based Spectral Feature Representation for Statistical Parametric Speech Synthesis
abstract
This letter presents a method of deriving spectral features using a deep belief network (DBN) for hidden Markov model (HMM)-based parametric speech synthesis. At training time, a DBN is estimated to represent the high-dimensional spectral envelopes and then transforms them into binary codes. These DBN-based binary codes (DBCs) are used as spectral features for HMM modeling. At synthesis time, spectral envelopes are recovered from the predicted DBC sequences and then used for waveform reconstruction. Experimental results show that our proposed method can achieve better naturalness than the conventional method using mel-cepstra as spectral features and considering global variance (GV) during parameter generation.
Ya-Jun Hu, Zhen-Hua Ling
IEEE Signal Process. Lett.1