Zengqiang Shang

dblp:296/4686 · DBLP profile ↗
← Back
12ranked-venue papers
2as first author
12since 2021 · last 2025
0000-0002-9225-2138ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 11 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 7 since 2021
YearPublicationVenuePosition
2025 Pitch-Assistant Harmonic Recovery for Efficient Speech Enhancement
abstract
With the rapid development of low-resource online speech enhancement models, noise suppression can now be achieved with significantly fewer model parameters. However, these models often suffer from limited effectiveness in preserving speech quality. In particular, most existing online speech enhancement methods tend to distort the harmonic structure of speech while performing noise reduction, leading to noticeable degradation in perceptual quality. In this paper, we propose a novel model architecture called Pitch-Assistant Harmonic Recovery for Efficient Speech Enhancement (PHRSE). The model operates on low-dimensional Bark-scale spectral features to perform noise suppression, while leveraging estimated fundamental frequency information to guide the reconstruction of harmonic components. This pitch-guided strategy enables the model to preserve the speech’s natural harmonic structure more effectively. Experimental results demonstrate that PHRSE not only achieves higher perceptual speech quality compared to existing benchmarks, but also maintains real-time performance with significantly lower computational overhead, making it suitable for online and resource-constrained scenarios.
Zengqiang Shang, Haoyuan Xie, Mou Wang, Pengyuan Zhang
ASRU2
2025 Restoring Harmonics: Enhancing Speech Quality with Deep Mask and Harmonic Restoration Network
Zengqiang Shang, Mou Wang, Pengyuan Zhang
INTERSPEECH2
2025 Leveraging distance information for generalized spoofing speech detection
Jingze Lu, Zengqiang Shang, Pengyuan Zhang
Comput. Speech Lang.4
2025 Flexpéro: Flexible Expressive Zero-Shot Speech Refinement via In-Context Learning
abstract
Controlling speech expressiveness has emerged as a critical research frontier in speech generation, focusing on synthesizing natural, human-like speech that accurately conveys intended psychological and emotional states. While many large-scale models have demonstrated sufficient zero-shot capability by conditioning the acoustic model on reference speech—which provides cues on speaker identity and style—they often fall short of meeting desired emotional or prosodic targets at fine-grained levels. To address this challenge, we propose a novel speech refinement method based on a zero-shot voice synthesis model that can flexibly and interactively enhance expressiveness on unsatisfactory speech segments. It supports emotion modulation through chunk-wise valence/arousal and flexible keyframe-based prediction of pitch and energy, allowing for the creation of any prosodic patterns. Experimental results show that our method achieves fine-grained control, thus enriching the expressiveness of zero-shot synthetic speech.
Hua Hua, Zengqiang Shang, Xuyuan Li, Pengyuan Zhang
IEEE Signal Process. Lett.2
2024 One-Class Knowledge Distillation for Spoofing Speech Detection
abstract
The detection of spoofing speech generated by unseen algorithms remains an unresolved challenge. One reason for the lack of generalization ability is that traditional detecting systems follow the binary classification paradigm, which inherently assumes the possession of prior knowledge of spoofing speech. One-class methods attempt to learn the distribution of bonafide speech and are inherently suited to the task where spoofing speech exhibits significant differences. However, training a one-class system using only bonafide speech is challenging. In this paper, we introduce a teacher-student framework to provide guidance for the training of a one-class model. The proposed one-class knowledge distillation method outperforms other state-of-the-art methods on the ASVspoof 21DF and InTheWild datasets, demonstrating its superior generalization ability.
Jingze Lu, Zengqiang Shang, Pengyuan Zhang
ICASSP4
2024 Improving Short Utterance Anti-Spoofing with Aasist2
abstract
The wav2vec 2.0 and integrated spectro-temporal graph attention network (AASIST) based countermeasure achieves great performance in speech anti-spoofing. However, current spoof speech detection systems have fixed training and evaluation durations, while the performance degrades significantly during short utterance evaluation. To solve this problem, AASIST can be improved to AASIST2 by modifying the residual blocks to Res2Net blocks. The modified Res2Net blocks can extract multi-scale features and improve the detection performance for speech of different durations, thus improving the short utterance evaluation performance. On the other hand, adaptive large margin fine-tuning (ALMFT) has achieved performance improvement in short utterance speaker verification. Therefore, we apply Dynamic Chunk Size (DCS) and ALMFT training strategies in speech anti-spoofing to further improve the performance of short utterance evaluation. Experiments demonstrate that the proposed AASIST2 improves the performance of short utterance evaluation while maintaining the performance of regular evaluation on different datasets.
Jingze Lu, Zengqiang Shang, Pengyuan Zhang
ICASSP3
2024 Expressive paragraph text-to-speech synthesis with multi-step variational autoencoder
Xuyuan Li, Zengqiang Shang, Peiyang Shi, Hua Hua, Ta Li, Pengyuan Zhang
INTERSPEECH2
2024 Improving Copy-Synthesis Anti-Spoofing Training Method with Rhythm and Speaker Perturbation
Jingze Lu, Zengqiang Shang, Pengyuan Zhang
INTERSPEECH4
2024 Emilia: An Extensive, Multilingual, and Diverse Speech Dataset For Large-Scale Speech Generation
abstract
Recent advancements in speech generation models have been significantly driven by the use of large-scale training data. However, producing highly spontaneous, human-like speech remains a challenge due to the scarcity of large, diverse, and spontaneous speech datasets. In response, we introduce Emilia, the first large-scale, multilingual, and diverse speech generation dataset. Emilia starts with over 101k hours of speech across six languages, covering a wide range of speaking styles to enable more natural and spontaneous speech generation. To facilitate the scale-up of Emilia, we also present Emilia-Pipe, the first open-source preprocessing pipeline designed to efficiently transform raw, in-the-wild speech data into high-quality training data with speech annotations. Experimental results demonstrate the effectiveness of both Emilia and Emilia-Pipe. Demos are available at: https://emilia-dataset.github.io/Emilia-Demo-Page/.
Haorui He, Zengqiang Shang, Chaoren Wang, Xuyuan Li, Yicheng Gu, Hua Hua, Liwei Liu 0008, Jiaqi Li 0030, Peiyang Shi, Yuancheng Wang, Kai Chen 0026, Pengyuan Zhang, Zhizheng Wu 0001
SLT2
2021 The Thinkit System for Icassp2021 M2voc Challenge
abstract
In this paper, we introduce the low resource text-to-speech system from the ThinkIT team submitted to Multi-Speaker Multi-Style Voice Cloning Challenge (M2VoC). The challenge has two tasks: few-shot track1 provides 100 samples for each person and one-shot track2 offers 5 samples only. Each track contains two sub-tracks A and B. Instead of sub-track A, sub-track B can use extra public data besides the released data. But we participate in the sub-track A only. We choose the finetune as our backbone strategy. Our submitted systems include BERT based prosody boundary prediction module, FastSpeech based acoustic model to generate acoustic features from text input, and HIFIGAN based vocoder to generate waveform from acoustic features. Among them, acoustic models are susceptible to low resource speakers. To prevent over-fitting, we modified the acoustic model and split out validation set to assist the manual model selection. Evaluation results provided by the challenges organizers demonstrate the effectiveness of our system.
Zengqiang Shang, Bolin Zhou, Pengyuan Zhang
ICASSP1
2021 Incorporating Cross-Speaker Style Transfer for Multi-Language Text-to-Speech
Zengqiang Shang, Pengyuan Zhang, Yonghong Yan 0002
Interspeech1
2021 LinearSpeech: Parallel Text-to-Speech with Linear Complexity
Zengqiang Shang, Pengyuan Zhang, Yonghong Yan 0002
Interspeech3