Liqun Deng

dblp:01/8280 · DBLP profile ↗
← Back
21ranked-venue papers
5as first author
16since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 15 · 3 first-author · 12 since 2021Artificial intelligence and machine learning · 12 · 1 first-author · 11 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 The GaoYao Benchmark: A Comprehensive Framework for Evaluating Multilingual and Multicultural Abilities of Large Language Models
abstract
Yilun Liu, Chunguang Zhao, Mengyao Piao, Lingqi Miao, Shimin Tao, Minggui HE, Chenxin Liu, Zhang Li, Mahongxia, Jiaxin Guo, Chen Liu, Liqun Deng, Jiansheng Wei, Xiaojun Meng, Fanyi Du, Daimeng Wei, Yanghua Xiao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yilun Liu 0001, Chunguang Zhao, Mengyao Piao, Lingqi Miao, Shimin Tao, Minggui He, Chenxin Liu, Hongxia Ma, Liqun Deng, Jiansheng Wei, Xiaojun Meng, Fanyi Du, Daimeng Wei, Yanghua Xiao
ACL (1)12
2025 Sparsing Law: Towards Large Language Models with Greater Activation Sparsity
abstract
Activation sparsity denotes the existence of substantial weakly-contributed neurons within feed-forward networks of large language models (LLMs), providing wide potential benefits such as computation acceleration. However, existing works lack thorough quantitative studies on this useful property, in terms of both its measurement and influential factors. In this paper, we address three underexplored research questions: (1) How can activation sparsity be measured more accurately? (2) How is activation sparsity affected by the model architecture and training process? (3) How can we build a more sparsely activated and efficient LLM? Specifically, we develop a generalizable and performance-friendly metric, named CETT-PPL-1%, to measure activation sparsity. Based on CETT-PPL-1%, we quantitatively study the influence of various factors and observe several important phenomena, such as the convergent power-law relationship between sparsity and training data amount, the higher competence of ReLU activation than mainstream SiLU activation, the potential sparsity merit of a small width-depth ratio, and the scale insensitivity of activation sparsity. Finally, we provide implications for building sparse and effective LLMs, and demonstrate the reliability of our findings by training a 2.4B model with a sparsity ratio of 93.52%, showing 4.1$\times$ speedup compared with its dense version. The codes and checkpoints are available at https://github.com/thunlp/SparsingLaw/.
Yuqi Luo, Xu Han 0007, Yingfa Chen, Chaojun Xiao, Xiaojun Meng, Liqun Deng, Jiansheng Wei, Zhiyuan Liu 0001, Maosong Sun 0001
ICML7
2024 Prompt-Driven Target Speech Diarization
abstract
We introduce a novel task named ‘target speech diarization’, which seeks to determine ‘when target event occurred’ within an audio signal. We devise a neural architecture called Prompt-driven Target Speech Diarization (PTSD), that works with diverse prompts that specify the target speech events of interest. We train and evaluate PTSD using sim2spk, sim3spk and sim4spk datasets, which are derived from the Librispeech. We show that the proposed framework accurately localizes target speech events. Furthermore, our framework exhibits versatility through its impressive performance in three diarization-related tasks: target speaker voice activity detection, overlapped speech detection and gender diarization. In particular, PTSD achieves comparable performance to specialized models across these tasks on both real and simulated data. This work serves as a reference benchmark and provides valuable insights into prompt-driven target speech processing.
Yidi Jiang, Zhengyang Chen, Ruijie Tao, Liqun Deng, Yanmin Qian, Haizhou Li 0001
ICASSP4
2024 SA-WavLM: Speaker-Aware Self-Supervised Pre-training for Mixture Speech
abstract
It was shown that pre-trained models with self-supervised learning (SSL) techniques are effective in various downstream speech tasks.However, most such models are trained on singlespeaker speech data, limiting their effectiveness in mixture speech.This motivates us to explore pre-training on mixture speech.This work presents SA-WavLM, a novel pre-trained model for mixture speech.Specifically, SA-WavLM follows an "extract-merge-predict" pipeline in which the representations of each speaker in the input mixture are first extracted individually and then merged before the final prediction.In this pipeline, SA-WavLM performs speaker-informed extractions with the consideration of the interactions between different speakers.Furthermore, a speaker shuffling strategy is proposed to enhance the robustness towards the speaker absence.Experiments show that SA-WavLM either matches or improves upon the state-of-the-art pre-trained models.
Jingru Lin, Meng Ge, Junyi Ao, Liqun Deng, Haizhou Li 0001
INTERSPEECH4
2023 DisCover: Disentangled Music Representation Learning for Cover Song Identification
abstract
In the field of music information retrieval (MIR), cover song identification (CSI) is a challenging task that aims to identify cover versions of a query song from a massive collection. Existing works still suffer from high intra-song variances and inter-song correlations, due to the entangled nature of version-specific and version-invariant factors in their modeling. In this work, we set the goal of disentangling version-specific and version-invariant factors, which could make it easier for the model to learn invariant music representations for unseen query songs. We analyze the CSI task in a disentanglement view with the causal graph technique, and identify the intra-version and inter-version effects biasing the invariant learning. To block these effects, we propose the disentangled music representation learning framework (DisCover) for CSI. DisCover consists of two critical components: (1) Knowledge-guided Disentanglement Module (KDM) and (2) Gradient-based Adversarial Disentanglement Module (GADM), which block intra-version and inter-version biased effects, respectively. KDM minimizes the mutual information between the learned representations and version-variant factors that are identified with prior domain knowledge. GADM identifies version-variant factors by simulating the representation transitions between intra-song versions, and exploits adversarial distillation for effect blocking. Extensive comparisons with best-performing methods and in-depth analysis demonstrate the effectiveness of DisCover and the and necessity of disentanglement for CSI.
Jiahao Xun, Shengyu Zhang 0001, Yanting Yang, Jieming Zhu, Liqun Deng, Zhou Zhao 0001, Zhenhua Dong, Ruiqi Li 0002, Fei Wu 0001
SIGIR5
2022 Context-Aware Mask Prediction Network for End-to-End Text-Based Speech Editing
abstract
The text-based speech editor allows the editing of speech through intuitive cutting, copying, and pasting operations to speed up the process of editing speech. However, the major drawback of current systems is that edited speech often sounds unnatural and it is not obvious how to synthesize records according to a new word not appearing in the transcript. This paper proposes a novel end-to-end text-based speech editing method called context-aware mask prediction network (CampNet), which avoids the unnatural phenomenon caused by cut-copy-paste operation in the traditional method and can synthesize a new word not appearing in the transcript. Besides, three text-based speech editing operations based on CampNet are designed: deletion, replacement, and insertion. These operations can comprehensively cover different kinds of situations that text-based speech editing can face. The subjective and objective experiments on VCTK and LibriTTS data sets show that the speech editing results based on CampNet are better than TTS technology, manual editing, and VoCo method (the combination of speech synthesis and speech conversion). We also conducted detailed ablation experiments to explore the effect of the CampNet structure on its performance. Examples of generated speech can be found at https://hairuo55.github.io/CampNet-demo.
Tao Wang 0074, Jiangyan Yi, Liqun Deng, Ruibo Fu, Jianhua Tao 0001, Zhengqi Wen
ICASSP3
2022 HiFiDenoise: High-Fidelity Denoising Text to Speech with Adversarial Networks
abstract
Building a high-fidelity speech synthesis system with noisy speech data is a challenging but valuable task, which could significantly reduce the cost of data collection. Existing methods usually train speech synthesis systems based on the speech denoised with an enhancement model or feed noise information as a condition into the system. These methods certainly have some effect on inhibiting noise, but the quality and the prosody of their synthesized speech are still far away from natural speech. In this paper, we propose HiFiDenoise, a speech synthesis system with adversarial networks that can synthesize high-fidelity speech with low-quality and noisy speech data. Specifically, 1) to tackle the difficulty of noise modeling, we introduce multi-length adversarial training in the noise condition module. 2) To handle the problem of inaccurate pitch extraction caused by noise, we remove the pitch predictor in the acoustic model and also add discriminators on the mel-spectrogram generator. 3) In addition, we also apply HiFiDenoise to singing voice synthesis with a noisy singing dataset. Experiments show that our model outperforms the baseline by 0.36 and 0.44 in terms of MOS on speech and singing respectively.
Yi Ren 0006, Liqun Deng, Zhou Zhao 0001
ICASSP3
2022 EditSinger: Zero-Shot Text-Based Singing Voice Editing System with Diverse Prosody Modeling
abstract
Zero-shot text-based singing editing enables singing voice modification based on the given edited lyrics without any additional data from the target singer. However, due to the different demands, challenges occur when applying existing speech editing methods to singing voice editing task, mainly including the lack of systematic consideration concerning prosody in insertion and deletion, as well as the trade-off between the naturalness of pronunciation and the preservation of prosody in replacement. In this paper we propose EditSinger, which is a novel singing voice editing model with specially designed diverse prosody modules to overcome the challenges above. Specifically, 1) a general masked variance adaptor is introduced for the comprehensive prosody modeling of the inserted lyrics and the transition of deletion boundary; and 2) we further design a fusion pitch predictor for replacement. By disentangling the reference pitch and fusing the predicted pronunciation, the edited pitch can be reconstructed, which could ensure a natural pronunciation while preserving the prosody of the original audio. In addition, to the best of our knowledge, it is the first zero-shot text-based singing voice editing system. Our experiments conducted on the OpenSinger prove that EditSinger can synthesize high-quality edited singing voices with natural prosody according to the corresponding operations.
Zhou Zhao 0001, Yi Ren 0006, Liqun Deng
IJCAI4
2022 reducing multilingual context confusion for end-to-end code-switching automatic speech recognition
Shuai Zhang 0014, Jiangyan Yi, Zhengkun Tian, Jianhua Tao 0001, Yu Ting Yeung, Liqun Deng
INTERSPEECH6
2022 Streamable Speech Representation Disentanglement and Multi-Level Prosody Modeling for Live One-Shot Voice Conversion
Haoquan Yang, Liqun Deng, Yu Ting Yeung, Nianzu Zheng
INTERSPEECH2
2022 CoCA-MDD: A Coupled Cross-Attention based Framework for Streaming Mispronunciation Detection and Diagnosis
abstract
Mispronunciation detection and diagnosis (MDD) is a popular research focus in computer-aided pronunciation training (CAPT) systems.End-to-end (e2e) approaches are becoming dominant in MDD.However an e2e MDD model usually requires entire speech utterances as input context, which leads to significant time latency especially for long paragraphs.We propose a streaming e2e MDD model called CoCA-MDD.We utilize conv-transformer structure to encode input speech in a streaming manner.A coupled cross-attention (CoCA) mechanism is proposed to integrate frame-level acoustic features with encoded reference linguistic features.CoCA also enables our model to perform mispronunciation classification with whole utterances.The proposed model allows system fusion between the streaming output and mispronunciation classification output for further performance enhancement.We evaluate CoCA-MDD on publicly available corpora.CoCA-MDD achieves F1 scores of 57.03% and 60.78% for streaming and fusion modes respectively on L2-ARCTIC.For phone-level pronunciation scoring, CoCA-MDD achieves 0.58 Pearson correlation coefficient (PCC) value on SpeechOcean762.
Nianzu Zheng, Liqun Deng, Wenyong Huang, Yu Ting Yeung, Baohua Xu, Yasheng Wang, Xiao Chen 0012, Xin Jiang 0002, Qun Liu 0001
INTERSPEECH2
2022 M4Singer: A Multi-Style, Multi-Singer and Musical Score Provided Mandarin Singing Corpus
abstract
The lack of publicly available high-quality and accurately labeled datasets has long been a major bottleneck for singing voice synthesis (SVS). To tackle this problem, we present M4Singer, a free-to-use Multi-style, Multi-singer Mandarin singing collection with elaborately annotated Musical scores as well as its benchmarks. Specifically, 1) we construct and release a large high-quality Chinese singing voice corpus, which is recorded by 20 professional singers, covering 700 Chinese pop songs as well as all the four SATB types (i.e., soprano, alto, tenor, and bass); 2) we take extensive efforts to manually compose the musical scores for each recorded song, which are necessary to the study of the prosody modeling for SVS. 3) To facilitate the use and demonstrate the quality of M4Singer, we conduct four different benchmark experiments: score-based SVS, controllable singing voice (CSV), singing voice conversion (SVC) and automatic music transcription (AMT).
Ruiqi Li 0002, Shoutong Wang, Liqun Deng, Jinglin Liu, Yi Ren 0006, Jinzheng He, Rongjie Huang 0001, Jieming Zhu, Zhou Zhao 0001
NeurIPS4
2021 EditSpeech: A Text Based Speech Editing System Using Partial Inference and Bidirectional Fusion
abstract
This paper presents the design, implementation and evaluation of a speech editing system, named EditSpeech, which allows a user to perform deletion, insertion and replacement of words in a given speech utterance, without causing audible degradation in speech quality and naturalness. The EditSpeech system is developed upon a neural text-to-speech (NTTS) synthesis framework. Partial inference and bidirectional fusion are proposed to effectively incorporate the contextual information related to the edited region and achieve smooth transition at both left and right boundaries. Distortion introduced to the unmodified parts of the utterance is alleviated. The EditSpeech system is developed and evaluated on English and Chinese in multi-speaker scenarios. Objective and subjective evaluation demonstrate that EditSpeech outperforms a few baseline systems in terms of low spectral distortion and preferred speech quality. Audio samples are available online for demonstration11https://daxintan-cuhk.github.io/EditSpeech/.
Daxin Tan, Liqun Deng, Yu Ting Yeung, Xin Jiang 0002, Xiao Chen 0012, Tan Lee
ASRU2
2021 Fcl-Taco2: Towards Fast, Controllable and Lightweight Text-to-Speech Synthesis
abstract
Sequence-to-sequence (seq2seq) learning has greatly improved text-to-speech (TTS) synthesis performance, but effective implementation on resource-restricted devices remains challenging as seq2seq models are usually computationally expensive and memory intensive. To achieve fast inference speed and small model size while maintain high-quality speech, we propose FCL-taco2, a Fast, Controllable and Lightweight (FCL) TTS model based on Tacotron2. FCL-taco2 adopts a novel semi-autoregressive (SAR) mode for phoneme level based parallel mel-spectrograms generation conditioned on prosody features, leading to faster inference speed and higher prosody controllability than Tacotron2. Besides, knowledge distillation (KD) is leveraged to compress a relatively large FCL-taco2 model to its small version with minor loss of speech quality. Experimental results on English (EN) and Chinese (CN) datasets show that the small version of FCL-taco2 achieves comparable performance with Tacotron2 in terms of speech quality, while it has a 4.8× smaller footprint with 17.7× and 18.5× faster inference speeds on average for EN and CN experiments respectively. Besides, execution on mobile devices shows that the proposed model can achieve faster than real-time speech synthesis. Our code and audio samples are released1.
Disong Wang, Liqun Deng, Yang Zhang 0025, Nianzu Zheng, Yu Ting Yeung, Xiao Chen 0012, Xunying Liu, Helen M. Meng
ICASSP2
2021 VQMIVC: Vector Quantization and Mutual Information-Based Unsupervised Speech Representation Disentanglement for One-Shot Voice Conversion
abstract
One-shot voice conversion (VC), which performs conversion across arbitrary speakers with only a single target-speaker utterance for reference, can be effectively achieved by speech representation disentanglement.Existing work generally ignores the correlation between different speech representations during training, which causes leakage of content information into the speaker representation and thus degrades VC performance.To alleviate this issue, we employ vector quantization (VQ) for content encoding and introduce mutual information (MI) as the correlation metric during training, to achieve proper disentanglement of content, speaker and pitch representations, by reducing their inter-dependencies in an unsupervised manner.Experimental results reflect the superiority of the proposed method in learning effective disentangled speech representations for retaining source linguistic content and intonation variations, while capturing target speaker characteristics.In doing so, the proposed approach achieves higher speech naturalness and speaker similarity than current state-of-the-art one-shot VC systems.Our code, pre-trained models and demo are available at https://github.com/Wendison/VQMIVC.
Disong Wang, Liqun Deng, Yu Ting Yeung, Xiao Chen 0012, Xunying Liu, Helen M. Meng
Interspeech2
2021 Unsupervised Domain Adaptation for Dysarthric Speech Detection via Domain Adversarial Training and Mutual Information Minimization
abstract
Dysarthric speech detection (DSD) systems aim to detect characteristics of the neuromotor disorder from speech.Such systems are particularly susceptible to domain mismatch where the training and testing data come from the source and target domains respectively, but the two domains may differ in terms of speech stimuli, disease etiology, etc.It is hard to acquire labelled data in the target domain, due to high costs of annotating sizeable datasets.This paper makes a first attempt to formulate cross-domain DSD as an unsupervised domain adaptation (UDA) problem.We use labelled source-domain data and unlabelled target-domain data, and propose a multi-task learning strategy, including dysarthria presence classification (DPC), domain adversarial training (DAT) and mutual information minimization (MIM), which aim to learn dysarthriadiscriminative and domain-invariant biomarker embeddings.Specifically, DPC helps biomarker embeddings capture critical indicators of dysarthria; DAT forces biomarker embeddings to be indistinguishable in source and target domains; and MIM further reduces the correlation between biomarker embeddings and domain-related cues.By treating the UASPEECH and TORGO corpora respectively as the source and target domains, experiments show that the incorporation of UDA attains absolute increases of 22.2% and 20.0% respectively in utterancelevel weighted average recall and speaker-level accuracy.
Disong Wang, Liqun Deng, Yu Ting Yeung, Xiao Chen 0012, Xunying Liu, Helen M. Meng
Interspeech2
2016 HiGene: A high-performance platform for genomic data analysis
abstract
Post-sequencing genomic data analysis becomes a major challenge while next-generation sequencing technologies evolve by leaps and bounds. The data-intensive and compute-intensive nature of genome analysis makes cluster computing an attractive choice for building efficient solutions. This paper presents HiGene, a high-performance genome analysis platform that exploits big data technology to revolutionize genomics data crunching power. HiGene reconstructs the genome analysis pipeline by exploiting both multi-core and multi-node parallelization using Apache Spark, and employs two key techniques to further boost the performance. First, a dynamic computing resource re-allocator is implemented, which allows flexible on-demand resource allocation for operations inside tasks. Second, an efficient skew mitigation approach is proposed, which automatically identifies and resolves data skew and computation skew through task repartitioning and resource reallocating respectively. HiGene has been evaluated with a whole human genome dataset on a 10-node Huawei 5885 cluster. Experimental results show that HiGene achieves remarkable high performance that reduces the total running time on a whole genome sequence dataset from days to nearly one hour. Furthermore, it is two times faster than state-of-the-art cluster based approaches.
Liqun Deng, Guowei Huang 0002, Yuzheng Zhuang, Jiansheng Wei, Youliang Yan
BIBM1
2012 Generalized Model-Based Human Motion Recognition with Body Partition Index Maps
abstract
Abstract Content‐based human motion analysis has captured extensive concerns of researchers from the domains of computer animation, human‐machine interaction, entertainment, etc. However, it is a non‐trivial task due to the spatial and temporal variations in the motion data. In this paper, we propose a generalized model (GM)‐based approach to model the variations and accurately recognize motion patterns. We partition the human character model into five parts, and extract the features of the submotions of each specific body part using clustering techniques. These features from the training trials in each class are combined to build the GM. We propose a new penalty based similarity measure for DTW to be used with the GMs for isolated motion recognition. On the other hand, from the GMs five body partition index maps are constructed and used for matching together with a flexible end point detection scheme during continuous motion recognition. In the experiments, we examine the effectiveness and efficiency of the approach in both isolated motion and continuous motion recognition. The results show that our proposed method has good performance compared with other state‐of‐the‐art methods in recognition accuracy and processing speed.
Liqun Deng, Howard Leung, Naijie Gu, Yang Yang 0046
Comput. Graph. Forum1
2011 Real-time mocap dance recognition for an interactive dancing game
abstract
Abstract In this paper, we present an interactive dancing game based on motion capture technology. We address the problem of real‐time recognition of the user's live dance performance in order to determine the interactive motion to be rendered by a virtual dance partner. The real‐time recognition algorithm is based on a human body partition indexing scheme with flexible matching to determine the end of a move as well as to detect unwanted motion. We show that the system can recognize the live dance motions of users with good accuracy and render the interactive dance move of the virtual partner. Copyright © 2011 John Wiley & Sons, Ltd.
Liqun Deng, Howard Leung, Naijie Gu, Yang Yang 0046
Comput. Animat. Virtual Worlds1
2010 Recognizing Dance Motions with Segmental SVD
abstract
In this paper, a novel concept of segmental singular value decomposition (SegSVD) is proposed to represent a motion pattern with a hierarchical structure. The similarity measure based on the SegSVD representation is also proposed. SegSVD is capable of capturing the temporal information of the time series. It is effective in matching patterns in a time series in which the start and end points of the patterns are not known in advance. We evaluate the performance of our method on both isolated motion classification and continuous motion recognition for dance movements. Experiments show that our method outperforms existing work in terms of higher recognition accuracy.
Liqun Deng, Howard Leung, Naijie Gu, Yang Yang 0046
ICPR1
2010 Automated Recognition of Sequential Patterns in Captured Motion Streams
Liqun Deng, Howard Leung, Naijie Gu, Yang Yang 0046
WAIM1