Yuanjun Lv

dblp:357/5573 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
9since 2021 · last 2025
0009-0002-2272-9153ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 2 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021
YearPublicationVenuePosition
2025 Analyzing and Mitigating Inconsistency in Discrete Speech Tokens for Neural Codec Language Models
abstract
Building upon advancements in Large Language Models (LLMs), the field of audio processing has seen increased interest in training speech generation tasks with discrete speech token sequences.However, directly discretizing speech by neural audio codecs often results in sequences that fundamentally differ from text sequences.Unlike text, where text token sequences are deterministic, discrete speech tokens can exhibit significant variability based on contextual factors, while still producing perceptually identical audio segments.We refer to this phenomenon as Discrete Representation Inconsistency (DRI).This inconsistency can lead to a single speech segment being represented by multiple divergent sequences, which creates confusion in neural codec language models and results in poor generated speech.In this paper, we quantitatively analyze the DRI phenomenon within popular audio tokenizers such as En-Codec.Our approach effectively mitigates the DRI phenomenon of the neural audio codec.Furthermore, extensive experiments on the neural codec language model over LibriTTS and large-scale MLS dataset (44,000 hours) demonstrate the effectiveness and generality of our method.The demo of audio samples is available online 1 .
Wenrui Liu 0003, Zhifang Guo, Jin Xu 0010, Yuanjun Lv, Yunfei Chu, Junyang Lin
ACL (1)4
2024 AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension
abstract
Qian Yang, Jin Xu, Wenrui Liu, Yunfei Chu, Ziyue Jiang, Xiaohuan Zhou, Yichong Leng, Yuanjun Lv, Zhou Zhao, Chang Zhou, Jingren Zhou. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Qian Yang 0006, Jin Xu 0010, Wenrui Liu 0003, Yunfei Chu, Ziyue Jiang 0001, Xiaohuan Zhou, Yichong Leng, Yuanjun Lv, Zhou Zhao 0001, Chang Zhou 0005, Jingren Zhou 0001
ACL (1)8
2024 SELM: Speech Enhancement using Discrete Tokens and Language Models
abstract
Language models (LMs) have recently shown superior performances in various speech generation tasks, demonstrating their powerful ability for semantic context modeling. Given the intrinsic similarity between speech generation and speech enhancement, harnessing semantic information is advantageous for speech enhancement tasks. In light of this, we propose SELM, a novel speech enhancement paradigm that integrates discrete tokens and leverages language models. SELM comprises three stages: encoding, modeling, and decoding. We transform continuous waveform signals into discrete tokens using pre-trained self-supervised learning (SSL) models and a k-means tokenizer. Language models then capture comprehensive contextual information within these tokens. Finally, a de-tokenizer and HiFi-GAN restore them into enhanced speech. Experimental results demonstrate that SELM achieves comparable performance in objective metrics and superior subjective perception results. Our demos are available1.
Xinfa Zhu, Yuanjun Lv, Lei Xie 0001
ICASSP4
2024 Single-Codec: Single-Codebook Speech Codec towards High-Performance Speech Generation
Hanzhao Li, Liumeng Xue, Haohan Guo, Xinfa Zhu, Yuanjun Lv, Lei Xie 0001, Yunlin Chen
INTERSPEECH5
2024 RaD-Net 2: A causal two-stage repairing and denoising speech enhancement network with knowledge distillation and complex axial self-attention
Mingshuai Liu, Zhuangqi Chen, Xiaopeng Yan, Yuanjun Lv, Xianjun Xia, Chuanzeng Huang, Yijian Xiao, Lei Xie 0001
INTERSPEECH4
2024 FreeV: Free Lunch For Vocoders Through Pseudo Inversed Mel Filter
Yuanjun Lv, Danming Xie
INTERSPEECH1
2024 Vec-Tok-VC+: Residual-enhanced Robust Zero-shot Voice Conversion with Progressive Constraints in a Dual-mode Training Strategy
Linhan Ma, Xinfa Zhu, Yuanjun Lv, Zhichao Wang 0002, Wendi He, Hongbin Zhou, Lei Xie 0001
INTERSPEECH3
2023 HIGNN-TTS: Hierarchical Prosody Modeling With Graph Neural Networks for Expressive Long-Form TTS
abstract
Recent advances in text-to-speech, particularly those based on Graph Neural Networks (GNNs), have significantly improved the expressiveness of short-form synthetic speech. However, generating human-parity long-form speech with high dynamic prosodic variations is still challenging. To address this problem, we expand the capabilities of GNNs with a hierarchical prosody modeling approach, named HiGNNTTS. Specifically, we add a virtual global node in the graph to strengthen the interconnection of word nodes and introduce a contextual attention mechanism to broaden the prosody modeling scope of GNNs from intra-sentence to inter-sentence. Additionally, we perform hierarchical supervision from acoustic prosody on each node of the graph to capture the prosodic variations with a high dynamic range. Ablation studies show the effectiveness of HiGNN-TTS in learning hierarchical prosody. Both objective and subjective evaluations demonstrate that HiGNN-TTS significantly improves the naturalness and expressiveness of long-form synthetic speech1.1Speech samples: https://dukguo.github.io/HiGNN-TTS/
Dake Guo, Xinfa Zhu, Liumeng Xue, Tao Li 0051, Yuanjun Lv, Yuepeng Jiang, Lei Xie 0001
ASRU5
2023 Salt: Distinguishable Speaker Anonymization Through Latent Space Transformation
abstract
Speaker anonymization aims to conceal a speaker’s identity without degrading speech quality and intelligibility. Most speaker anonymization systems disentangle the speaker representation from the original speech and achieve anonymization by averaging or modifying the speaker representation. However, the anonymized speech is subject to reduction in pseudo speaker distinctiveness, speech quality and intelligibility for out-of-distribution speaker. To solve this issue, we propose SALT, a Speaker Anonymization system based on Latent space Transformation. Specifically, we extract latent features by a self-supervised feature extractor and randomly sample multiple speakers and their weights, and then interpolate the latent vectors to achieve speaker anonymization. Meanwhile, we explore the extrapolation method to further extend the diversity of pseudo speakers. Experiments on Voice Privacy Challenge dataset show our system achieves a state-of-the-art distinctiveness metric while preserving speech quality and intelligibility. Our code and demo is availible at github1.1https://github.com/BakerBunker/SALT
Yuanjun Lv, Jixun Yao, Peikun Chen, Hongbin Zhou, Heng Lu 0004, Lei Xie 0001
ASRU1