Zhengchen Zhang

dblp:98/8688 · DBLP profile ↗
← Back
15ranked-venue papers
1as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 7 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2026 IRIS: Interleaved Reinforcement with Incremental Staged Curriculum for Cross-Lingual Mathematical Reasoning
abstract
Navya Gupta, Rishitej Reddy Vyalla, Avinash Anand, Chhavi Kirtani, Erik Cambria, Zhengchen Zhang, Zhengkui Wang, Timothy Liu, Aik Beng Ng, Simon See, Rajiv Ratn Shah. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Navya Gupta, Rishitej Reddy Vyalla, Avinash Anand, Chhavi Kirtani, Erik Cambria, Zhengchen Zhang, Zhengkui Wang, Timothy Liu, Aik Beng Ng, Simon See, Rajiv Ratn Shah
ACL (1)6
2023 MaskedSpeech: Context-aware Speech Synthesis with Masking Strategy
Ya-Jie Zhang, Yanghao Yue, Zhengchen Zhang, Youzheng Wu, Xiaodong He 0001
INTERSPEECH4
2023 Prosody Modelling With Pre-Trained Cross-Utterance Representations for Improved Speech Synthesis
abstract
When humans speak multiple utterances in a continuous manner, the prosodic features generated in each utterance are related to those in its neighbouring utterances. Such cross-utterance (CU) dependencies are often ignored by the current neural text-to-speech (TTS) systems, which reduces the naturalness and expressiveness of the synthesized speeches. In this paper, we propose to improve the prosody modelling ability of neural TTS systems using pre-trained CU acoustic and text representations. Such CU acoustic representations are derived using the Wav2Vec 2.0 model (W2V2) from the synthesized audios of the past utterances, while the CU text representations are extracted using the Bidirectional Encoder Representation from Transformers (BERT) model from the scripts of the future utterances. Experimental results on a Mandarin audiobook and an English audiobook showed the naturalness and expressiveness of the synthesized audios were significantly improved by incorporating such pre-trained W2V2 and BERT CU representations into the Fastspeech2 TTS framework.
Ya-Jie Zhang, Chao Zhang 0031, Zhengchen Zhang, Youzheng Wu, Xiaodong He 0001
IEEE ACM Trans. Audio Speech Lang. Process.4
2022 SCaLa: Supervised Contrastive Learning for End-to-End Speech Recognition
abstract
End-to-end Automatic Speech Recognition (ASR) models are usually trained to optimize the loss of the whole token sequence, while neglecting explicit phonemic-granularity supervision.This could result in recognition errors due to similarphoneme confusion or phoneme reduction.To alleviate this problem, we propose a novel framework based on Supervised Contrastive Learning (SCaLa) to enhance phonemic representation learning for end-to-end ASR systems.Specifically, we extend the self-supervised Masked Contrastive Predictive Coding (MCPC) to a fully-supervised setting, where the supervision is applied in the following way.First, SCaLa masks variablelength encoder features according to phoneme boundaries given phoneme forced-alignment extracted from a pre-trained acoustic model; it then predicts the masked features via contrastive learning.The forced-alignment can provide phoneme labels to mitigate the noise introduced by positive-negative pairs in selfsupervised MCPC.Experiments on reading and spontaneous speech datasets show that our proposed approach achieves 2.8 and 1.4 points Character Error Rate (CER) absolute reductions compared to the baseline, respectively.
Runyu Wang, Fan Lu 0003, Zhengchen Zhang, Meng Chen 0006, Youzheng Wu, Xiaodong He 0001
INTERSPEECH5
2022 Singing Voice Synthesis with Vibrato Modeling and Latent Energy Representation
abstract
This paper proposes an expressive singing voice synthesis system by introducing explicit vibrato modeling and latent energy representation. Vibrato is essential to the naturalness of synthesized sound, due to the inherent characteristics of human singing. Hence, a deep learning-based vibrato model is introduced in this paper to control the vibrato's likeliness, rate, depth and phase in singing, where the vibrato likeliness represents the existence probability of vibrato and it would help improve the singing voice's naturalness. Actually, there is no annotated label about vibrato likeliness in existing singing corpus. We adopt a novel vibrato likeliness labeling method to label the vibrato likeliness automatically. Meanwhile, the power spectrogram of audio contains rich information that can improve the expressiveness of singing. An autoencoder-based latent energy bottleneck feature is proposed for expressive singing voice synthesis. Experimental results on the open dataset NUS48E show that both the vibrato modeling and the latent energy representation could significantly improve the expressiveness of singing voice. The audio samples are shown in the demo website11https://mango321321.github.io/ExpressiveSing/.
Wei Zhang 0031, Zhengchen Zhang, Dan Zeng 0001, Zhi Liu 0003
MMSP4
2021 Incremental Learning for End-to-End Automatic Speech Recognition
abstract
In this paper, we propose an incremental learning method for end-to-end Automatic Speech Recognition (ASR) which enables an ASR system to perform well on new tasks while maintaining the performance on its originally learned ones. To mitigate catastrophic forgetting during incremental learning, we design a novel explainability-based knowledge distillation for ASR models, which is combined with a response-based knowledge distillation to maintain the original model's predictions and the “reason” for the predictions. Our method works without access to the training data of original tasks, which addresses the cases where the previous data is no longer available or joint training is costly. Results on a multi-stage sequential training task show that our method outperforms existing ones in mitigating forgetting. Furthermore, in two practical scenarios, compared to the target-reference joint training method, the performance drop of our method is 0.02% Character Error Rate (CER), which is 97% smaller than the drops of the baseline methods.
Libo Zi, Zhengchen Zhang, Youzheng Wu, Xiaodong He 0001, Bowen Zhou 0001
ASRU4
2021 Dian: Duration Informed Auto-Regressive Network for Voice Cloning
abstract
In this paper, we propose a novel end-to-end speech synthesis approach, Duration Informed Auto-regressive Network (DIAN), which consists of an acoustic model and a separate duration model. Un-like other auto-regressive TTS methods, the duration information of phonemes is provided as part of the input to the acoustic model, which enables the removal of the attention mechanism between its encoder and decoder parts. This eliminates the common seen skipping and repeating issues and improves speech intelligibility while ensuring high speech quality. A Transformer-based duration model is used to predict the duration of each phoneme for the attention-free acoustic model. We developed our TTS systems for the multi-speaker multi-style voice cloning challenge (M2VoC) using the proposed DIAN approach. In our procedure, a multi-speaker attention-free acoustic model and its Transformer-based duration model are first separately trained based on the training data released by M2VoC. Next, the multi-speaker models are adapted to form the speaker-specific models with the speaker-dependent data and transfer learning. At last, a speaker-specific LPCNet is estimated and used to synthesize the speech of the corresponding speaker. The M2VoC results showed that our proposed approach achieved the 3rd-place in the speech quality ranking and the 4th-place in the speaker similarity and style similarity ranking in the Track1-a task.
Zhengchen Zhang, Chao Zhang 0031, Youzheng Wu, Xiaodong He 0001, Bowen Zhou 0001
ICASSP3
2021 Improving Prosody Modelling with Cross-Utterance Bert Embeddings for End-to-End Speech Synthesis
abstract
Although speech prosody is related to the linguistic information up to the discourse structure, most text-to-speech (TTS) systems only take into account the information within each sentence. This makes it challenging when converting a paragraph of text into natural and expressive speech. In this paper, we propose to use the text embeddings of the neighboring sentences to improve the prosody generation for each utterance of a paragraph in an end-to-end fashion without using any explicit prosody features. More specifically, cross-utterance (CU) context vectors, which are produced by an additional CU encoder based on the sentence embeddings extracted by a pretrained BERT model, are used to augment the input of the Tacotron2 decoder. Two types of BERT embeddings are investigated, which leads to the use of different CU encoder structures. Experimental results on a Mandarin audiobook dataset and the LJ-Speech English audiobook dataset demonstrate the use of CU information can improve the naturalness and expressiveness of the synthesized speech. Subjective listening testing shows most of the participants prefer the voice generated using the CU encoder over that generated using standard Tacotron2. It is also found that the prosody can be controlled indirectly by changing the neighbouring sentences.
Zhengchen Zhang, Chao Zhang 0031, Xiaodong He 0001, Bowen Zhou 0001
ICASSP3
2021 ViDA-MAN: Visual Dialog with Digital Humans
abstract
We demonstrate ViDA-MAN, a digital-human agent for multi-modal interaction, which offers realtime audio-visual responses to instant speech inquiries. Compared to traditional text or voice-based system, ViDA-MAN offers human-like interactions (e.g, vivid voice, natural facial expression and body gestures). Given a speech request, the demonstration is able to response with high quality videos in sub-second latency. To deliver immersive user experience, ViDA-MAN seamlessly integrates multi-modal techniques including Acoustic Speech Recognition (ASR), multi-turn dialog, Text To Speech (TTS), talking heads video generation. Backed with large knowledge base, ViDA-MAN is able to chat with users on a number of topics including chit-chat, weather, device control, News recommendations, booking hotels, as well as answering questions via structured knowledge.
Jiawei Zuo, Liqin Jiang, Meng Chen 0006, Zhengchen Zhang, Wei Zhang 0031, Xiaodong He 0001, Tao Mei 0001
ACM Multimedia7
2020 Efficient WaveGlow: An Improved WaveGlow Vocoder with Enhanced Speed
Zhengchen Zhang, Chao Zhang 0031, Xiaodong He 0001, Bowen Zhou 0001
INTERSPEECH3
2017 Node-level parallelization for deep neural networks with conditional independent graph
Fugen Zhou, Fuxiang Wu, Zhengchen Zhang, Minghui Dong
Neurocomputing3
2014 Speaker state classification based on fusion of asymmetric simple partial least squares (SIMPLS) and support vector machines
Dong-Yan Huang, Zhengchen Zhang, Shuzhi Sam Ge
Comput. Speech Lang.2
2013 Bottom-up saliency detection for attention determination
Shuzhi Sam Ge, Hongsheng He, Zhengchen Zhang
Mach. Vis. Appl.3
2012 Mutual-reinforcement document summarization using embedded graph based sentence clustering for storytelling
Zhengchen Zhang, Shuzhi Sam Ge, Hongsheng He
Inf. Process. Manag.1
2011 Speaker State Classification Based on Fusion of Asymmetric SIMPLS and Support Vector Machines
abstract
Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH
Dong-Yan Huang, Shuzhi Sam Ge, Zhengchen Zhang
INTERSPEECH3