Bei Liu 0003

dblp:39/3711-3 · DBLP profile ↗
← Back
19ranked-venue papers
9as first author
19since 2021 · last 2025
0000-0002-6208-003XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 16 · 7 first-author · 16 since 2021Artificial intelligence and machine learning · 14 · 8 first-author · 14 since 2021
YearPublicationVenuePosition
2025 OOQ: Outlier-Oriented Quantization for Efficient Large Language Models
abstract
Parameter quantization for Large Language Models (LLMs) has gained significant attention for reducing memory costs and improving computational efficiency. However, existing methods struggle with performance degradation in low-bit scenarios. In this paper, we propose Outlier-Oriented Quantization (OOQ), a novel framework designed to address these challenges through three key innovations. First, we design an outlieroriented metric to determine quantization precision based on outlier percentages in channel parameters. Second, we dynamically allocate varying quantization precision to different parts of the model according to the outlier distribution. Finally, guided by the outlier-oriented metric, we preserve some high-precision outliers during the quantization process. Experiments on LLaMA models demonstrate that OOQ achieves state-of-the-art results across various bit settings, particularly in extremely low-bit regimes. Additionally, OOQ improves inference speed by up to 24%.
Haoyu Wang 0007, Bei Liu 0003, Hang Shao 0005, Guanglu Wan, Yanmin Qian
ASRU2
2025 Efficient Pruning for Large-Scale Seq2Seq Speech Models without Back-Propagation
abstract
Large-scale Seq2Seq speech models like Whisper excel in speech recognition but are limited by their high computational demands, making them difficult to be deployed on resource-constrained devices. This paper introduces a novel and efficient pruning method for compressing these models without retraining and back-propagation, focusing on models with encoder-decoder architectures. We adapt layer-wise pruning to large speech models and introduce a mixed sparsity allocation strategy that uses only forward propagation. This approach effectively reduces model size while maintaining high performance. Evaluated on the Whisper-large-v3 across various datasets, our method can almost maintain Whisper’s performance and robustness with about 60% reduction in parameters. It could also be combined with other model compression methods such as distillation to further reducing model size.
Tianteng Gu, Bei Liu 0003, Yanmin Qian
ICASSP2
2025 Ultra-Low Bit Post-Training Quantization of Large Speech Models via K-Means Clustering and Mixed Precision Allocation
Tianteng Gu, Bei Liu 0003, Haoyu Wang 0007, Yanmin Qian
INTERSPEECH2
2025 DenoiseRotator: Enhance Pruning Robustness for LLMs via Importance Concentration
abstract
Pruning is a widely used technique to compress large language models (LLMs) by removing unimportant weights, but it often suffers from significant performance degradation—especially under semi-structured sparsity constraints. Existing pruning methods primarily focus on estimating the importance of individual weights, which limits their ability to preserve critical capabilities of the model. In this work, we propose a new perspective: rather than merely selecting which weights to prune, we first redistribute parameter importance to make the model inherently more amenable to pruning. By minimizing the information entropy of normalized importance scores, our approach concentrates importance onto a smaller subset of weights, thereby enhancing pruning robustness. We instantiate this idea through DenoiseRotator, which applies learnable orthogonal transformations to the model’s weight matrices. Our method is model-agnostic and can be seamlessly integrated with existing pruning techniques such as Magnitude, SparseGPT, and Wanda. Evaluated on LLaMA3, Qwen2.5, and Mistral models under 50% unstructured and 2:4 semi-structured sparsity, DenoiseRotator consistently improves perplexity and zero-shot accuracy. For instance, on LLaMA3-70B pruned with SparseGPT at 2:4 semi-structured sparsity, DenoiseRotator reduces the perplexity gap to the dense model by 58%, narrowing the degradation from 8.1 to 3.4 points.
Tianteng Gu, Bei Liu 0003, Yanmin Qian
NeurIPS2
2024 One-Shot Sensitivity-Aware Mixed Sparsity Pruning for Large Language Models
abstract
Various Large Language Models (LLMs) from the Generative Pretrained Transformer (GPT) family have achieved outstanding performances in a wide range of text generation tasks. However, the enormous model sizes have hindered their practical use in real-world applications due to high inference latency. Therefore, improving the efficiencies of LLMs through quantization, pruning, and other means has been a key issue in LLM studies. In this work, we propose a method based on Hessian sensitivity-aware mixed sparsity pruning to prune LLMs to at least 50% sparsity without the need of any retraining. It allocates sparsity adaptively based on sensitivity, allowing us to reduce pruning-induced error while maintaining the overall sparsity level. The advantages of the proposed method exhibit even more when the sparsity is extremely high. Furthermore, our method is compatible with quantization, enabling further compression of LLMs.
Hang Shao 0005, Bei Liu 0003, Yanmin Qian
ICASSP2
2024 SparseWAV: Fast and Accurate One-Shot Unstructured Pruning for Large Speech Foundation Models
Tianteng Gu, Bei Liu 0003, Hang Shao 0005, Yanmin Qian
INTERSPEECH2
2024 DQ-Whisper: Joint Distillation and Quantization for Efficient Multilingual Speech Recognition
abstract
As a popular multilingual and multitask pre-trained speech model, Whisper has the problem of curse of multilinguality. To enhance multilingual capabilities in small Whisper models, we propose DQ-Whisper, a novel joint distillation and quantization framework to compress Whisper for efficient inference. Firstly, we propose a novel dynamic matching distillation strategy. Then, a quantization-aware distillation framework is introduced to integrate quantization with distillation. Experimental results on various multilingual datasets show that our suggested distillation approach can effectively enhance the multilingual capabilities of small Whisper models without increasing computational costs. Up to 5.18x reduction in model size is achieved with marginal performance degradation. In addition, quantization is compatible with distillation, which can result in a higher compression rate.
Hang Shao 0005, Bei Liu 0003, Wei Wang 0010, Xun Gong 0005, Yanmin Qian
SLT2
2024 Towards Lightweight Speaker Verification via Adaptive Neural Network Quantization
abstract
Modern speaker verification (SV) systems typically demand expensive storage and computing resources, thereby hindering their deployment on mobile devices. In this paper, we explore adaptive neural network quantization for lightweight speaker verification. Firstly, we propose a novel adaptive uniform precision quantization method which enables the dynamic generation of quantization centroids customized for each network layer based on k-means clustering. By applying it to the pre-trained SV systems, we obtain a series of quantized variants with different bit widths. To enhance low-bit quantized models, a mixed precision quantization algorithm along with a multi-stage fine-tuning (MSFT) strategy is further introduced. This approach assigns varying bit widths to different network layers. When bit combinations are determined, MSFT progressively quantizes and fine-tunes the network in a specific order. Finally, we design two distinct binary quantization schemes to mitigate performance degradation of 1-bit quantized models: the static and adaptive quantizers. Experiments on VoxCeleb demonstrate that lossless 4-bit uniform precision quantization is achieved on both ResNets and DF-ResNets, yielding a promising compression ratio of$\sim$8. Moreover, compared to uniform precision approach, mixed precision quantization not only obtains additional performance improvements with a similar model size but also offers the flexibility to generate bit combination for any desirable model size. In addition, our suggested 1-bit quantization schemes remarkably boost the performance of binarized models. Finally, a thorough comparison with existing lightweight SV systems reveals that our proposed models outperform all previous methods by a large margin across various model size ranges.
Bei Liu 0003, Haoyu Wang 0007, Yanmin Qian
IEEE ACM Trans. Audio Speech Lang. Process.1
2023 Build a SRE Challenge System: Lessons from VoxSRC 2022 and CNSRC 2022
Zhengyang Chen, Bing Han 0008, Xu Xiang, Houjun Huang, Bei Liu 0003, Yanmin Qian
INTERSPEECH5
2023 Reversible Neural Networks for Memory-Efficient Speaker Verification
Bei Liu 0003, Yanmin Qian
INTERSPEECH1
2023 ECAPA++: Fine-grained Deep Embedding Learning for TDNN Based Speaker Verification
Bei Liu 0003, Yanmin Qian
INTERSPEECH1
2023 Extremely Low Bit Quantization for Mobile Speaker Verification Systems Under 1MB Memory
Bei Liu 0003, Haoyu Wang 0007, Yanmin Qian
INTERSPEECH1
2023 Adaptive Neural Network Quantization For Lightweight Speaker Verification
Haoyu Wang 0007, Bei Liu 0003, Yanmin Qian
INTERSPEECH2
2023 Depth-First Neural Architecture With Attentive Feature Fusion for Efficient Speaker Verification
abstract
Deep speaker embedding learning based on neural networks has become the predominant approach in speaker verification (SV) currently. In prior studies, researchers have investigated various network architectures. However, rare works pay attention to the question of how to design and scale up networks in a principled way to achieve a better trade-off on model performance and computational complexity. In this paper, we focus on efficient architecture design for speaker verification. Firstly, we systematically study the effect of the network depth and width on performance and empirically discover thatdepth is more important than the width of networks for speaker verification task. Based on this observation, we propose a novel depth-first (DF) architecture design rule. By applying it to ResNet and ECAPA-TDNN, two new families of much deeper models, namely DF-ResNets and DF-ECAPAs, are constructed. In addition, to further boost the performance of small models in the low computation regime, a novel attentive feature fusion (AFF) scheme is proposed to replace the conventional feature fusion methods. Specifically, we design two different fusion strategies, including sequential AFF (S-AFF) and parallel AFF (P-AFF), which can dynamically fuse features in a learnable way. Experimental results on the VoxCeleb dataset show that the newly proposed DF-ResNets and DF-ECAPAs can achieve a much better trade-off on performance and complexity than the original ResNet and ECAPA-TDNN. Moreover, small models can further obtain up to 40% relative improvement in EER by adopting AFF scheme with negligible computational cost. Finally, a comprehensive comparison with various other published SV systems illustrates that our proposed models achieve the best trade-off on performance and complexity in both low and high computation scenarios.
Bei Liu 0003, Zhengyang Chen, Yanmin Qian
IEEE ACM Trans. Audio Speech Lang. Process.1
2022 MLP-SVNET: A Multi-Layer Perceptrons Based Network for Speaker Verification
abstract
Convolution and self-attention based neural networks have both obtained excellent performance in automatic speaker verification. However, the convolution model often lacks the ability of long-term dependency modeling due to the limitation of receptive field, while the self-attention model is insufficient to model local information. To tackle this limitation, we propose a new multi-layer perceptrons based speaker verification network (MLP-SVNet) which can apply MLPs across temporal and frequency dimensions to capture the local and global information at the same time. The experimental results conducted on Voxceleb show that the proposed model is very competitive when compared to other systems based on convolution or self-attention. In addition, we demonstrate that MLP-SVNet based on multi-layer per-ceptrons can produce complementary embeddings, which can be fused with the state-of-the-art system to further improve the performance.
Bing Han 0008, Zhengyang Chen, Bei Liu 0003, Yanmin Qian
ICASSP3
2022 Self-Knowledge Distillation via Feature Enhancement for Speaker Verification
abstract
As the most widely used technique, deep speaker embedding learning has become predominant in speaker verification task recently. Very large neural networks such as ECAPA-TDNN and ResNet can achieve the state-of-the-art performance. However, large models are computationally unfriendly in general, which require massive storage and computation resources. Model compression has been a hot research topic. Parameter quantization usually results in significant performance degradation. Knowledge distillation demands a pretrained complex teacher model. In this paper, we introduce a novel self-knowledge distillation method, namely Self-Knowledge Distillation via Feature Enhancement (SKDFE). It utilizes an auxiliary self-teacher network to distill its own refined knowledge without the need of a pretrained teacher network. Additionally, we apply the self-knowledge distillation at two different levels: label level and feature level. Experiments on Voxceleb dataset show that our proposed self-knowledge distillation method can make small models have comparable or even better performance than large ones. Large models can also be further improved when applying our method.
Bei Liu 0003, Haoyu Wang 0007, Zhengyang Chen, Shuai Wang 0016, Yanmin Qian
ICASSP1
2022 Attentive Feature Fusion for Robust Speaker Verification
Bei Liu 0003, Zhengyang Chen, Yanmin Qian
INTERSPEECH1
2022 Dual Path Embedding Learning for Speaker Verification with Triplet Attention
Bei Liu 0003, Zhengyang Chen, Yanmin Qian
INTERSPEECH1
2022 DF-ResNet: Boosting Speaker Verification Performance with Depth-First Design
Bei Liu 0003, Zhengyang Chen, Shuai Wang 0016, Haoyu Wang 0007, Bing Han 0008, Yanmin Qian
INTERSPEECH1