EDBT 2026 Demo / reviewers in the wild / expert
Zhaoqing Li
dblp:240/4302
· DBLP profile ↗
18ranked-venue papers
6as first author
17since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 6 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 4 first-author · 12 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A heterogeneous agent reinforcement learning approach with curriculum learning for variable speed limit control
Zhaoqing Li, Silai Chen, Guosheng Xiao, Yangsheng Jiang, Zhihong Yao, Puxin Yang |
Expert Syst. Appl. | 1 |
| 2025 | ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow MatchingabstractExisting large-scale zero-shot text-to-speech (TTS) models deliver high speech quality but suffer from slow inference speeds due to massive parameters. To address this issue, this paper introduces ZipVoice, a high-quality flow-matching-based zero-shot TTS model with a compact model size and fast inference speed. Key designs include: 1) a Zipformer-based vector field estimator to maintain adequate modeling capabilities under constrained size; 2) Average upsampling-based initial speech-text alignment and Zipformer-based text encoder to improve speech intelligibility; 3) A flow distillation method to reduce sampling steps and eliminate the inference overhead associated with classifier-free guidance. Experiments on 100 k hours multilingual datasets show that ZipVoice matches state-of-the-art models in speech quality, while being 3 times smaller and up to 30 times faster than a DiT-based flow-matching baseline. Codes, model checkpoints and demo samples are publicly available.11https://github.com/k2-fsa/ZipVoice Zhu Han 0001, Wei Kang 0006, Zengwei Yao, Liyong Guo, Zhaoqing Li, Weiji Zhuang, Long Lin, Daniel Povey |
ASRU | 6 |
| 2025 | Large Language Model Can Transcribe Speech in Multi-Talker Scenarios with Versatile InstructionsabstractRecent advancements in large language models (LLMs) have revolutionized various domains, bringing significant progress and new opportunities. Despite progress in speech-related tasks, LLMs have not been sufficiently explored in multi-talker scenarios. In this work, we present a pioneering effort to investigate the capability of LLMs in transcribing speech in multi-talker environments, following versatile instructions related to multi-talker automatic speech recognition (ASR), target talker ASR, and ASR based on specific talker attributes such as sex, occurrence order, language, and keyword spoken. Our approach utilizes WavLM and Whisper encoder to extract multi-faceted speech representations that are sensitive to speaker characteristics and semantic context. These representations are then fed into an LLM fine-tuned using LoRA, enabling the capabilities for speech comprehension and transcription. Comprehensive experiments reveal the promising performance of our proposed system, MT-LLM, in cocktail party scenarios, highlighting the potential of LLM to handle speech-related tasks based on user instructions in such complex settings1. Lingwei Meng, Shujie Hu, Jiawen Kang 0002, Zhaoqing Li, Yuejiao Wang, Xixin Wu, Xunying Liu, Helen M. Meng |
ICASSP | 4 |
| 2025 | Phone-purity Guided Discrete Tokens for Dysarthric Speech RecognitionabstractDiscrete tokens provide compact and domain-adaptable representations of speech features. However, their application to disordered speech, characterized by articulation imprecision and significant mismatch with normal voice, remains unexplored. To this end, this paper proposes novel phone-purity guided (PPG) discrete tokens to address the weakened phonetic discrimination arising during unsupervised K-means clustering or vector quantization of continuous features. Phonetic label supervision is incorporated to regularize the maximum likelihood and reconstruction error costs in standard K-means and VAE-VQ-based token extraction. Experiments on the UASpeech corpus show that PPG-based discrete tokens extracted from HuBERT consistently outperform hybrid TDNN and End-to-End (E2E) Conformer systems using non-PPG tokens. Statistically significant word error rate (WER) reductions of up to 0.99% and 1.77% absolute (3.21% and 4.82% relative) are achieved across varying codebook sizes for the 16 UASpeech test dysarthric speakers. The lowest WER of 23.25% is obtained by combining systems using complementary token features. Consistent improvements are also observed in phone purity, and t-SNE visualizations demonstrate sharper decision boundaries between K-means/VAE-VQ clusters with the introduction of phone-purity guidance. Huimeng Wang, Xurong Xie, Mengzhe Geng, Shujie Hu, Haoning Xu, Youjun Chen, Zhaoqing Li, Jiajun Deng, Xunying Liu |
ICASSP | 7 |
| 2025 | Effective and Efficient Mixed Precision Quantization of Speech Foundation ModelsabstractThis paper presents a novel mixed-precision quantization approach for speech foundation models that tightly integrates mixed-precision learning and quantized model parameter estimation into one single model compression stage. Experiments conducted on LibriSpeech dataset with fine-tuned wav2vec2.0-base and HuBERT-large models suggest the resulting mixed-precision quantized models increased the lossless compression ratio by factors up to 1.7x and 1.9x over the respective uniform-precision and two-stage mixed-precision quantized baselines that perform precision learning and model parameters quantization in separate and disjointed stages, while incurring no statistically word error rate (WER) increase over the 32-bit full-precision models. The system compression time of wav2vec2.0-base and HuBERT-large models is reduced by up to 1.9 and 1.5 times over the two-stage mixed-precision baselines, while both produce lower WERs. The best-performing 3.5-bit mixed-precision quantized HuBERT-large model produces a lossless compression ratio of 8.6x over the 32-bit full-precision system. Haoning Xu, Zhaoqing Li, Zengrui Jin, Huimeng Wang, Youjun Chen, Guinan Li, Mengzhe Geng, Shujie Hu, Jiajun Deng, Xunying Liu |
ICASSP | 2 |
| 2025 | CR-CTC: Consistency regularization on CTC for improved speech recognitionabstractConnectionist Temporal Classification (CTC) is a widely used method for automatic speech recognition (ASR), renowned for its simplicity and computational efficiency. However, it often falls short in recognition performance. In this work, we propose the Consistency-Regularized CTC (CR-CTC), which enforces consistency between two CTC distributions obtained from different augmented views of the input speech mel-spectrogram. We provide in-depth insights into its essential behaviors from three perspectives: 1) it conducts self-distillation between random pairs of sub-models that process different augmented views; 2) it learns contextual representation through masked prediction for positions within time-masked regions, especially when we increase the amount of time masking; 3) it suppresses the extremely peaky CTC distributions, thereby reducing overfitting and improving the generalization ability. Extensive experiments on LibriSpeech, Aishell-1, and GigaSpeech datasets demonstrate the effectiveness of our CR-CTC. It significantly improves the CTC performance, achieving state-of-the-art results comparable to those attained by transducer or systems combining CTC and attention-based encoder-decoder (CTC/AED). We release our code at \url{https://github.com/k2-fsa/icefall}. Zengwei Yao, Wei Kang 0006, Xiaoyu Yang 0005, Liyong Guo, Han Zhu 0004, Zengrui Jin, Zhaoqing Li, Long Lin, Daniel Povey |
ICLR | 8 |
| 2025 | Exploring SSL Discrete Speech Features for Zipformer-based Contextual ASR
Yifan Yang 0005, Jiajun Deng, Jiawen Kang 0002, Shujie Hu, Tianzi Wang, Zhaoqing Li, Shiliang Zhang, Xie Chen 0001, Xunying Liu |
INTERSPEECH | 7 |
| 2025 | Towards One-bit ASR: Extremely Low-bit Conformer Quantization Using Co-training and Stochastic Precision
Zhaoqing Li, Haoning Xu, Zengrui Jin, Lingwei Meng, Tianzi Wang, Huimeng Wang, Youjun Chen, Shujie Hu, Xunying Liu |
INTERSPEECH | 1 |
| 2025 | Unfolding A Few Structures for The Many: Memory-Efficient Compression of Conformer and Speech Foundation Models
Zhaoqing Li, Haoning Xu, Xurong Xie, Zengrui Jin, Tianzi Wang, Xunying Liu |
INTERSPEECH | 1 |
| 2025 | Effective and Efficient One-pass Compression of Speech Foundation Models Using Sparsity-aware Self-pinching Gates
Haoning Xu, Zhaoqing Li, Youjun Chen, Huimeng Wang, Guinan Li, Mengzhe Geng, Chengxi Deng, Xunying Liu |
INTERSPEECH | 2 |
| 2024 | Towards High-Performance and Low-Latency Feature-Based Speaker Adaptation of Conformer Speech Recognition SystemsabstractPractical application of model-based speaker adaptation techniques to end-to-end ASR systems is hindered by speaker-level data scarcity and latency in speaker-dependent (SD) parameters update. To this end, data-efficient and low-latency rapid feature-based speaker adaptation approaches are proposed in this paper for state-of-the-art Conformer ASR systems. Compact subspace projection of training data estimated SD hidden layer output scaling or bias parameters is used to represent the most distinctive speaker "bases". A feature-driven prediction network containing purpose-built speaker-aware memory is designed to on-the-fly produce homogeneous SD basis interpolation, and facilitate rapid speaker adaptation. Experimental results on the 300-hr Switchboard corpus suggest that the proposed adaptation approach produces statistically significant word error rate (WER) reductions of up to 1.0% absolute (8.4% relative) over the baseline speaker-independent and i-vector adapted Conformers before and after external LM rescoring. Consistent WER reductions of up to 2.0% absolute (16.3% relative) and real-time factor speeding up ratios of up to 10.9 times are also obtained over offline model-based adaptation across different speaker-level data quantity operating points. T-SNE visualization reveals the on-the-fly predicted SD basis weights present intuitively more consistent speaker features than i-vectors. Jiajun Deng, Xurong Xie, Guinan Li, Mengzhe Geng, Zengrui Jin, Tianzi Wang, Shujie Hu, Zhaoqing Li, Xunying Liu |
ICASSP | 9 |
| 2024 | One-pass Multiple Conformer and Foundation Speech Systems Compression and Quantization Using An All-in-one Neural ModelabstractWe propose a novel one-pass multiple ASR systems joint compression and quantization approach using an all-in-one neural model. A single compression cycle allows multiple nested systems with varying Encoder depths, widths, and quantization precision settings to be simultaneously constructed without the need to train and store individual target systems separately. Experiments consistently demonstrate the multiple ASR systems compressed in a single all-in-one model produced a word error rate (WER) comparable to, or lower by up to 1.01% absolute (6.98% relative) than individually trained systems of equal complexity. A 3.4x overall system compression and training time speed-up was achieved. Maximum model size compression ratios of 12.8x and 3.93x were obtained over the baseline Switchboard-300hr Conformer and LibriSpeech-100hr fine-tuned wav2vec2.0 models, respectively, incurring no statistically significant WER increase. Zhaoqing Li, Haoning Xu, Tianzi Wang, Shoukang Hu, Zengrui Jin, Shujie Hu, Jiajun Deng, Mengzhe Geng, Xunying Liu |
INTERSPEECH | 1 |
| 2024 | Towards Effective and Efficient Non-autoregressive Decoding Using Block-based Attention MaskabstractThis paper proposes a novel non-autoregressive (NAR) block-based Attention Mask Decoder (AMD) that flexibly balances performance-efficiency trade-offs for Conformer ASR systems. AMD performs parallel NAR inference within contiguous blocks of output labels that are concealed using attention masks, while conducting left-to-right AR prediction and history context amalgamation between blocks. A beam search algorithm is designed to leverage a dynamic fusion of CTC, AR Decoder, and AMD probabilities. Experiments on the LibriSpeech-100hr corpus suggest the tripartite Decoder incorporating the AMD module produces a maximum decoding speed-up ratio of 1.73x over the baseline CTC+AR decoding, while incurring no statistically significant word error rate (WER) increase on the test sets. When operating with the same decoding real time factors, statistically significant WER reductions of up to 0.7% and 0.3% absolute (5.3% and 6.1% relative) were obtained over the CTC+AR baseline. Tianzi Wang, Xurong Xie, Zhaoqing Li, Shoukang Hu, Zengrui Jin, Jiajun Deng, Shujie Hu, Mengzhe Geng, Guinan Li, Helen M. Meng, Xunying Liu |
INTERSPEECH | 3 |
| 2024 | A Survey on Privacy in Graph Neural Networks: Attacks, Preservation, and ApplicationsabstractGraph Neural Networks (GNNs) have gained significant attention owing to their ability to handle graph-structured data and the improvement in practical applications. However, many of these models prioritize high utility performance, such as accuracy, with a lack of privacy consideration, which is a major concern in modern society where privacy attacks are rampant. To address this issue, researchers have started to develop privacy-preserving GNNs. Despite this progress, there is a lack of a comprehensive overview of the attacks and the techniques for preserving privacy in the graph domain. In this survey, we aim to address this gap by summarizing the attacks on graph data according to the targeted information, categorizing the privacy preservation techniques in GNNs, and reviewing the datasets and applications that could be used for analyzing/solving privacy issues in GNNs. We also outline potential directions for future research in order to build better privacy-preserving GNNs. Yuying Zhao, Zhaoqing Li, Xueqi Cheng 0002, Yu Wang 0160, Olivera Kotevska, Philip S. Yu, Tyler Derr |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2023 | Lossless 4-bit Quantization of Architecture Compressed Conformer ASR Systems on the 300-hr Switchboard Corpus
Zhaoqing Li, Tianzi Wang, Jiajun Deng, Shoukang Hu, Xunying Liu |
INTERSPEECH | 1 |
| 2023 | Hate Speech Detection via Dual Contrastive LearningabstractThe fast spread of hate speech on social media impacts the Internet environment and our society by increasing prejudice and hurting people. Detecting hate speech has aroused broad attention in the field of natural language processing. Although hate speech detection has been addressed in recent work, this task still faces two inherent unsolved challenges. The first challenge lies in the complex semantic information conveyed in hate speech, particularly the interference of insulting words in hate speech detection. The second challenge is the imbalanced distribution of hate speech and non-hate speech, which may significantly deteriorate the performance of models. To tackle these challenges, we propose a novel dual contrastive learning (DCL) framework for hate speech detection. Our framework jointly optimizes the self-supervised and the supervised contrastive learning loss for capturing span-level information beyond the token-level emotional semantics used in existing models, particularly detecting speech containing abusive and insulting words. Moreover, we integrate the focal loss into the dual contrastive learning framework to alleviate the problem of data imbalance. We conduct experiments on two publicly available English datasets, and experimental results show that the proposed model outperforms the state-of-the-art models and precisely detects hate speeches. Junyu Lu 0001, Hongfei Lin, Xiaokun Zhang 0001, Zhaoqing Li, Tongyue Zhang, Linlin Zong, Fenglong Ma, Bo Xu 0009 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2021 | Large Graph Clustering With Simultaneous Spectral Embedding and DiscretizationabstractSpectral clustering methods are gaining more and more interests and successfully applied in many fields because of their superior performance. However, there still exist two main problems to be solved: 1) spectral clustering methods consist of two successive optimization stages-spectral embedding and spectral rotation, which may not lead to globally optimal solutions, 2) and it is known that spectral methods are time-consuming with very high computational complexity. There are methods proposed to reduce the complexity for data vectors but not for graphs that only have information about similarity matrices. In this paper, we propose a new method to solve these two challenging problems for graph clustering. In the new method, a new framework is established to perform spectral embedding and spectral rotation simultaneously. The newly designed objective function consists of both terms of embedding and rotation, and we use an improved spectral rotation method to make it mathematically rigorous for the optimization. To further accelerate the algorithm, we derive a low-dimensional representation matrix from a graph by using label propagation, with which, in return, we can reconstruct a double-stochastic and positive semidefinite similarity matrix. Experimental results demonstrate that our method has excellent performance in time cost and accuracy. Zhen Wang 0004, Zhaoqing Li, Rong Wang 0001, Feiping Nie 0001, Xuelong Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2020 | Pedestrian detection via deep segmentation and context network
Zhaoqing Li, Zhenxue Chen, Q. M. Jonathan Wu, Chengyun Liu |
Neural Comput. Appl. | 1 |