Leying Zhang

dblp:278/7751 · DBLP profile ↗
← Back
12ranked-venue papers
8as first author
12since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 5 first-author · 8 since 2021Artificial intelligence and machine learning · 6 · 4 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 SLIDE: Integrating Speech Language Model with LLM for Spontaneous Spoken Dialogue Generation
abstract
Recently, "textless" speech language models (SLMs) based on speech units have made huge progress in generating naturalistic speech, including non-verbal vocalizations. However, the generated speech samples often lack semantic coherence. In this paper, we propose SLM and LLM Integration for spontaneous spoken Dialogue gEneration (SLIDE). Specifically, we first utilize an LLM to generate the textual content of spoken dialogue. Next, we convert the textual dialogues into phoneme sequences and use a two-tower transformer-based duration predictor to predict the duration of each phoneme. Finally, an SLM conditioned on the spoken phoneme sequences is used to vocalize the textual dialogue. Experimental results on the Fisher dataset demonstrate that our system can generate naturalistic spoken dialogue while maintaining high semantic coherence.
Haitian Lu, Gaofeng Cheng, Liuping Luo, Leying Zhang, Yanmin Qian, Pengyuan Zhang
ICASSP4
2025 Advanced Zero-Shot Text-to-Speech for Background Removal and Preservation with Controllable Masked Speech Prediction
abstract
The acoustic background plays a crucial role in natural conversation. It provides context and helps listeners understand the environment, but a strong background makes it difficult for listeners to understand spoken words. The appropriate handling of these backgrounds is situation-dependent: Although it may be necessary to remove background to ensure speech clarity, preserving the background is sometimes crucial to maintaining the contextual integrity of the speech. Despite recent advancements in zero-shot Text-to-Speech technologies, current systems often struggle with speech prompts containing backgrounds. To address these challenges, we propose a Controllable Masked Speech Prediction strategy coupled with a dual-speaker encoder, utilizing a task-related control signal to guide the prediction of dual background removal and preservation targets. Experimental results demonstrate that our approach enables precise control over the removal or preservation of background across various acoustic conditions and exhibits strong generalization capabilities in unseen scenarios.
Leying Zhang, Wangyou Zhang, Zhengyang Chen, Yanmin Qian
ICASSP1
2025 E2E-BPVC: End-to-End Background-Preserving Voice Conversion via In-Context Learning
Zhengyang Chen, Leying Zhang, Yanmin Qian
INTERSPEECH3
2025 CoVoMix2: Advancing Zero-Shot Dialogue Generation with Fully Non-Autoregressive Flow Matching
abstract
Generating natural-sounding, multi-speaker dialogue is crucial for applications such as podcast creation, virtual agents, and multimedia content generation. However, existing systems struggle to maintain speaker consistency, model overlapping speech, and synthesize coherent conversations efficiently. In this paper, we introduce CoVoMix2, a fully non-autoregressive framework for zero-shot multi-talker dialogue generation. CoVoMix2 directly predicts mel-spectrograms from multi-stream transcriptions using a flow-matching-based generative model, eliminating the reliance on intermediate token representations. To better capture realistic conversational dynamics, we propose transcription-level speaker disentanglement, sentence-level alignment, and prompt-level random masking strategies. Our approach achieves state-of-the-art performance, outperforming strong baselines like MoonCast and Sesame in speech quality, speaker consistency, and inference speed. Notably, CoVoMix2 operates without requiring transcriptions for the prompt and supports controllable dialogue generation, including overlapping speech and precise timing control, demonstrating strong generalizability to real-world speech generation scenarios. Audio samples are available at https://www.microsoft.com/en-us/research/project/covomix/covomix2.
Leying Zhang, Yao Qian, Xiaofei Wang 0009, Manthan Thakker, Dongmei Wang, Jianwei Yu 0001, Yuxuan Hu 0003, Jinyu Li 0001, Yanmin Qian, Sheng Zhao 0002
NeurIPS1
2024 Generation-Based Target Speech Extraction with Speech Discretization and Vocoder
abstract
Target speech extraction (TSE) is a task aiming at isolating the speech of a specific target speaker from an audio mixture, with the help of an auxiliary recording of that target speaker. Most existing TSE methods employ discrimination-based models to estimate the target speaker’s proportion in the mixture, but they often fail to compensate for the missing or highly corrupted frequency components in the speech signal. In contrast, the generation-based methods can naturally handle such scenarios via speech resynthesis. In this paper, we propose a novel discrete token based TSE approach by combining state-of-the-art speech discretization and vocoder techniques. By predicting a sequence of discrete tokens with the auxiliary audio and employing a vocoder that takes discrete tokens as input, the target speech can be effectively re-synthesized while eliminating interference. Our experiments conducted on the WSJ0-2mix and Libri2mix datasets demonstrate that our proposed method yields high-quality target speech without interference.
Linfeng Yu, Wangyou Zhang, Chenpeng Du, Leying Zhang, Yanmin Qian
ICASSP4
2024 PromptTTS 2: Describing and Generating Voices with Text Prompt
abstract
Speech conveys more information than text, as the same word can be uttered in various voices to convey diverse information. Compared to traditional text-to-speech (TTS) methods relying on speech prompts (reference speech) for voice variability, using text prompts (descriptions) is more user-friendly since speech prompts can be hard to find or may not exist at all. TTS approaches based on the text prompt face two main challenges: 1) the one-to-many problem, where not all details about voice variability can be described in the text prompt, and 2) the limited availability of text prompt datasets, where vendors and large cost of data labeling are required to write text prompts for speech. In this work, we introduce PromptTTS 2 to address these challenges with a variation network to provide variability information of voice not captured by text prompts, and a prompt generation pipeline to utilize the large language models (LLM) to compose high quality text prompts. Specifically, the variation network predicts the representation extracted from the reference speech (which contains full information about voice variability) based on the text prompt representation. For the prompt generation pipeline, it generates text prompts for speech with a speech language understanding model to recognize voice attributes (e.g., gender, speed) from speech and a large language model to formulate text prompts based on the recognition results. Experiments on a large-scale (44K hours) speech dataset demonstrate that compared to the previous works, PromptTTS 2 generates voices more consistent with text prompts and supports the sampling of diverse voice variability, thereby offering users more choices on voice generation. Additionally, the prompt generation pipeline produces high-quality text prompts, eliminating the large labeling cost. The demo page of PromptTTS 2 is available (https://speechresearch.github.io/prompttts2).
Yichong Leng, Zhifang Guo, Zeqian Ju, Xu Tan 0003, Eric Liu 0006, Dongchao Yang, Leying Zhang, Kaitao Song, Lei He 0005, Xiang-Yang Li 0001, Sheng Zhao 0002, Tao Qin 0001, Jiang Bian 0002
ICLR9
2024 CoVoMix: Advancing Zero-Shot Speech Generation for Human-like Multi-talker Conversations
abstract
Recent advancements in zero-shot text-to-speech (TTS) modeling have led to significant strides in generating high-fidelity and diverse speech. However, dialogue generation, along with achieving human-like naturalness in speech, continues to be a challenge. In this paper, we introduce CoVoMix: Conversational Voice Mixture Generation, a novel model for zero-shot, human-like, multi-speaker, multi-round dialogue speech generation. CoVoMix first converts dialogue text into multiple streams of discrete tokens, with each token stream representing semantic information for individual talkers. These token streams are then fed into a flow-matching based acoustic model to generate mixed mel-spectrograms. Finally, the speech waveforms are produced using a HiFi-GAN model. Furthermore, we devise a comprehensive set of metrics for measuring the effectiveness of dialogue modeling and generation. Our experimental results show that CoVoMix can generate dialogues that are not only human-like in their naturalness and coherence but also involve multiple talkers engaging in multiple rounds of conversation. This is exemplified by instances generated in a single channel where one speaker's utterance is seamlessly mixed with another's interjections or laughter, indicating the latter's role as an attentive listener. Audio samples are enclosed in the supplementary.
Leying Zhang, Yao Qian, Shujie Liu 0001, Dongmei Wang, Xiaofei Wang 0009, Midia Yousefi, Yanmin Qian, Jinyu Li 0001, Lei He 0005, Sheng Zhao 0002, Michael Zeng 0001
NeurIPS1
2024 DDTSE: Discriminative Diffusion Model for Target Speech Extraction
abstract
Diffusion models have gained attention in speech enhancement tasks, providing an alternative to conventional discriminative methods. However, research on target speech extraction under multispeaker noisy conditions remains relatively unexplored. Moreover, the superior quality of diffusion methods typically comes at the cost of slower inference speed. In this paper, we introduce the Discriminative Diffusion model for Target Speech Extraction (DDTSE). We apply the same forward process as diffusion models and utilize the reconstruction loss similar to discriminative methods. Furthermore, we devise a two-stage training strategy to emulate the inference process during model training. DDTSE not only works as a standalone system, but also can further improve the performance of discriminative models without additional retraining. Experimental results demonstrate that DDTSE not only achieves higher perceptual quality but also accelerates the inference process by 3 times compared to the conventional diffusion model.
Leying Zhang, Yao Qian, Linfeng Yu, Heming Wang, Hemin Yang, Shujie Liu 0001, Yanmin Qian
SLT1
2023 Segment Anything Model (SAM) for Medical Image Segmentation: A Preliminary Review
abstract
Medical image segmentation is a critical component in a variety of clinical applications, facilitating accurate diagnosis and treatment planning. The Segment Anything Model (SAM), a deep learning architecture, has emerged as a promising solution to the challenges inherent in medical image segmentation. SAM’s superior zero-shot capability allows it to generalize effectively, even in the absence of task-specific segmentation samples. This unique characteristic broadens its application potential across various medical image modalities. This paper provides an in-depth review of SAM, focusing on its application in medical image segmentation. The review discusses the advantages of deep learning image segmentation over traditional methods, emphasizing the superior accuracy, efficiency, and automation that deep learning models offer. The paper also highlights the applications of SAM across various medical imaging modalities, demonstrating its versatility and adaptability. A taxonomy of SAM approaches in medical image segmentation is presented, categorizing them based on modality, dimension, organ, dataset, prompt, and performance. Despite the promising results of SAM, challenges remain in the field of medical image segmentation. The paper identifies these challenges and suggests potential directions for future research. In conclusion, this review aims to provide a comprehensive understanding of SAM and its potential to revolutionize medical image analysis and contribute to advancements in healthcare.
Leying Zhang, Xiaokang Deng, Yu Lu 0001
BIBM1
2023 Adaptive Large Margin Fine-Tuning For Robust Speaker Verification
abstract
Large margin fine-tuning (LMFT) is an effective strategy to improve the speaker verification system’s performance and is widely used in speaker verification challenge systems. Because the large margin in the loss function could make the training task too difficult, people usually use longer training segments to alleviate this problem in LMFT. However, the LMFT model could have a duration mismatch with the real scenario verification, where the verification speech may be very short. In our experiments, we also find that LMFT fails in short duration and other verification scenarios. To solve this problem, we propose the duration-based and similarity-based adaptive large margin fine-tuning (ALMFT) strategy. To verify its effectiveness, we constructed fixed, variable length, and asymmetric verification trials based on VoxCeleb1. Experimental results demonstrate that ALMFT algorithms are very effective and robust, which not only achieve comparable improvement with LMFT in official VoxCeleb evaluation trials but also overcome performance degradation problems in short-duration and asymmetric scenarios respectively.
Leying Zhang, Zhengyang Chen, Yanmin Qian
ICASSP1
2022 Enroll-Aware Attentive Statistics Pooling for Target Speaker Verification
Leying Zhang, Zhengyang Chen, Yanmin Qian
INTERSPEECH1
2021 Knowledge Distillation from Multi-Modality to Single-Modality for Person Verification
Leying Zhang, Zhengyang Chen, Yanmin Qian
Interspeech1