EDBT 2026 Demo / reviewers in the wild / expert
Zhaoci Liu
dblp:260/4199
· DBLP profile ↗
11ranked-venue papers
6as first author
10since 2021 · last 2025
0009-0007-0102-6527ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 5 first-author · 8 since 2021Artificial intelligence and machine learning · 6 · 3 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DiffStyleTTS: Diffusion-based Hierarchical Prosody Modeling for Text-to-Speech with Diverse and Controllable StylesabstractHuman speech exhibits rich and flexible prosodic variations. To address the one-to-many mapping problem from text to prosody in a reasonable and flexible manner, we propose DiffStyleTTS, a multi-speaker acoustic model based on a conditional diffusion module and an improved classifier-free guidance, which hierarchically models speech prosodic features, and controls different prosodic styles to guide prosody prediction. Experiments show that our method outperforms all baselines in naturalness and achieves superior synthesis speed compared to three diffusion-based baselines. Additionally, by adjusting the guiding scale, DiffStyleTTS effectively controls the guidance intensity of the synthetic prosody. Zhaoci Liu, Yajun Hu, Yingying Gao, Shilei Zhang, Zhen-Hua Ling |
COLING | 2 |
| 2025 | Self-supervised Prosody Learning at Phoneme-level with Momentum Contrast for Speech SynthesisabstractThis paper investigates leveraging large-scale speech data to enhance prosodic modeling in speech synthesis, and introduces a model named SP2MC which achieves self-supervised prosody learning at phoneme-level with momentum contrast. This model incorporates dual convolutional encoders for speech and linear predictive coding (LPC) residual inputs to generate phoneme-level embeddings, which are masked and processed by a Transformer model to produce prosody representations. Two supervision modules are employed to generate phoneme-level supervision from speech waveforms and residuals. Momentum contrast is utilized to manage negative sample selection in contrastive learning. Finally, the SP2MC representations are integrated into a Fastspeech2-based acoustic model for speech synthesis. Experimental results indicate that the naturalness of speech synthesized by the proposed method is significantly better than that of baselines. Zhaoci Liu, Ya-Jun Hu, Zhen-Hua Ling |
ICASSP | 1 |
| 2025 | CMCNet: enhancing face image super-resolution through CNN-Mamba collaboration
Zhaoci Liu, Huayong Liu |
Vis. Comput. | 1 |
| 2024 | Considering Temporal Connection between Turns for Conversational Speech SynthesisabstractConversational speech synthesis aims to synthesize speech of an individual speaker based on history conversation. However, most studies in conversational speech synthesis only focus on the synthesis performance of the current speaker’s turn and neglect the temporal relationship between turns of interlocutors. Therefore, we consider the temporal connection between turns for conversational speech synthesis, which is crucial for the naturalness and coherence of conversations. Specifically, this paper formulates a task in which there is no overlap between turns and only one history turn is considered. To complete this task, an acoustic model is proposed which leverages multi-modal (including text and speech) information from previous turn to predict the acoustic features of not only current turn but also the inter-turn gap. The model is designed based on MQTTS and incorporates the global acoustic representation and BERT-based local semantic representation of previous turn when predicting the acoustic features of each frame. Experimental results demonstrate that with the introduction of global acoustic information and local semantic information, our model achieves better performance on the temporal connection between turns and the quality of synthetic speech. Audio samples can be found in https://mkd-mkd.github.io/icassp2024. Kangdi Mei, Zhaoci Liu, Hui-Peng Du, Yang Ai, Zhen-Hua Ling |
ICASSP | 2 |
| 2024 | Refining Self-supervised Learnt Speech Representation using Brain Activations
Kangdi Mei, Zhaoci Liu, Yang Ai, Jie Zhang 0042, Zhen-Hua Ling |
INTERSPEECH | 3 |
| 2024 | PE-Wav2vec: A Prosody-Enhanced Speech Model for Self-Supervised Prosody Learning in TTSabstractThis paper investigates leveraging large-scale untranscribed speech data to enhance the prosody modelling capability oftext-to-speech(TTS) models. On the basis of the self-supervised speech model wav2vec 2.0,Prosody-Enhanced wav2vec(PE-wav2vec) is proposed by introducing prosody learning. Specifically, prosody learning is achieved by applying supervision from thelinear predictive coding(LPC) residual signals on the initial Transformer blocks in the wav2vec 2.0 architecture. The embedding vectors extracted with the initial Transformer blocks of the PE-wav2vec model are utilised as prosodic representations for the corresponding frames in a speech utterance. To apply the PE-wav2vec representations in TTS, an acoustic model namedSpeech Synthesis model conditioned on Self-Supervisedly Learned Prosodic Representations(S4LPR) is designed on the basis of FastSpeech 2. The experimental results demonstrate that the proposed PE-wav2vec model can provide richer prosody descriptions of speech than the vanilla wav2vec 2.0 model can. Furthermore, the S4LPR model using PE-wav2vec representations can effectively improve the subjective naturalness and reduce the objective distortions of synthetic speech compared with baseline models. Zhaoci Liu, Ya-Jun Hu, Zhen-Hua Ling |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2023 | Speech Synthesis with Self-Supervisedly Learnt Prosodic Representations
Zhaoci Liu, Zhen-Hua Ling, Ya-Jun Hu, Jin-Wei Wang, Yun-Di Wu |
INTERSPEECH | 1 |
| 2022 | Discourse-Level Prosody Modeling with a Variational Autoencoder for Non-Autoregressive Expressive Speech SynthesisabstractTo address the issue of one-to-many mapping from phoneme sequences to acoustic features in expressive speech synthesis, this paper proposes a method of discourse-level prosody modeling with a variational autoencoder (VAE) based on the non-autoregressive architecture of FastSpeech. In this method, phone-level prosody codes are extracted from prosody features by combining VAE with FastSpeech, and are predicted using discourse-level text features together with BERT embeddings. The continuous wavelet transform (CWT) in FastSpeech2 for F0 representation is not necessary anymore. Experimental results on a Chinese audiobook dataset show that our proposed method can effectively take advantage of discourse-level linguistic information and has outperformed FastSpeech2 on the naturalness and expressiveness of synthetic speech. Ning-Qian Wu, Zhaoci Liu, Zhen-Hua Ling |
ICASSP | 2 |
| 2022 | Integrating Discrete Word-Level Style Variations into Non-Autoregressive Acoustic Models for Speech Synthesis
Zhaoci Liu, Ning-Qian Wu, Zhen-Hua Ling |
INTERSPEECH | 1 |
| 2021 | Detecting Alzheimer's Disease from Speech Using Neural Networks with Bottleneck Features and Data AugmentationabstractThis paper presents a method of detecting Alzheimer’s disease (AD) from the spontaneous speech of subjects in a picture description task using neural networks. This method does not rely on the manual transcriptions and annotations of a subject’s speech, but utilizes the bottleneck features extracted from audio using an ASR model. The neural network contains convolutional neural network (CNN) layers for local context modeling, bidirectional long shortterm memory (BiLSTM) layers for global context modeling and an attention pooling layer for classification. Furthermore, a masking- based data augmentation method is designed to deal with the data scarcity problem. Experiments on the DementiaBank dataset show that the detection accuracy of our proposed method is 82.59%, which is better than the baseline method based on manually-designed acoustic features and support vector machines (SVM), and achieves the state-of-the-art performance of detecting AD using only audio data on this dataset. Zhaoci Liu, Zhiqiang Guo, Zhen-Hua Ling, Yunxia Li |
ICASSP | 1 |
| 2020 | Text Classification by Contrastive Learning and Cross-lingual Data Augmentation for Alzheimer's Disease DetectionabstractData scarcity is always a constraint on analyzing speech transcriptions for automatic Alzheimer's disease (AD) detection, especially when the subjects are non-English speakers.To deal with this issue, this paper first proposes a contrastive learning method to obtain effective representations for text classification based on monolingual embeddings of BERT.Furthermore, a cross-lingual data augmentation method is designed by building autoencoders to learn the text representations shared by both languages.Experiments on a Mandarin AD corpus show that the contrastive learning method can achieve better detection accuracy than conventional CNN-based and BERTbased methods.Our cross-lingual data augmentation method also outperforms other compared methods when using another English AD corpus for augmentation.Finally, a best detection accuracy of 81.6% is obtained by our proposed methods on the Mandarin AD corpus. Zhiqiang Guo, Zhaoci Liu, Zhen-Hua Ling, Shijin Wang 0001, Lingjing Jin, Yunxia Li |
COLING | 2 |