Hang Shao 0005

dblp:33/11442-5 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
7since 2021 · last 2025
0000-0002-1322-4789ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021
YearPublicationVenuePosition
2025 OOQ: Outlier-Oriented Quantization for Efficient Large Language Models
abstract
Parameter quantization for Large Language Models (LLMs) has gained significant attention for reducing memory costs and improving computational efficiency. However, existing methods struggle with performance degradation in low-bit scenarios. In this paper, we propose Outlier-Oriented Quantization (OOQ), a novel framework designed to address these challenges through three key innovations. First, we design an outlieroriented metric to determine quantization precision based on outlier percentages in channel parameters. Second, we dynamically allocate varying quantization precision to different parts of the model according to the outlier distribution. Finally, guided by the outlier-oriented metric, we preserve some high-precision outliers during the quantization process. Experiments on LLaMA models demonstrate that OOQ achieves state-of-the-art results across various bit settings, particularly in extremely low-bit regimes. Additionally, OOQ improves inference speed by up to 24%.
Haoyu Wang 0007, Bei Liu 0003, Hang Shao 0005, Guanglu Wan, Yanmin Qian
ASRU3
2024 One-Shot Sensitivity-Aware Mixed Sparsity Pruning for Large Language Models
abstract
Various Large Language Models (LLMs) from the Generative Pretrained Transformer (GPT) family have achieved outstanding performances in a wide range of text generation tasks. However, the enormous model sizes have hindered their practical use in real-world applications due to high inference latency. Therefore, improving the efficiencies of LLMs through quantization, pruning, and other means has been a key issue in LLM studies. In this work, we propose a method based on Hessian sensitivity-aware mixed sparsity pruning to prune LLMs to at least 50% sparsity without the need of any retraining. It allocates sparsity adaptively based on sensitivity, allowing us to reduce pruning-induced error while maintaining the overall sparsity level. The advantages of the proposed method exhibit even more when the sparsity is extremely high. Furthermore, our method is compatible with quantization, enabling further compression of LLMs.
Hang Shao 0005, Bei Liu 0003, Yanmin Qian
ICASSP1
2024 SparseWAV: Fast and Accurate One-Shot Unstructured Pruning for Large Speech Foundation Models
Tianteng Gu, Bei Liu 0003, Hang Shao 0005, Yanmin Qian
INTERSPEECH3
2024 DQ-Whisper: Joint Distillation and Quantization for Efficient Multilingual Speech Recognition
abstract
As a popular multilingual and multitask pre-trained speech model, Whisper has the problem of curse of multilinguality. To enhance multilingual capabilities in small Whisper models, we propose DQ-Whisper, a novel joint distillation and quantization framework to compress Whisper for efficient inference. Firstly, we propose a novel dynamic matching distillation strategy. Then, a quantization-aware distillation framework is introduced to integrate quantization with distillation. Experimental results on various multilingual datasets show that our suggested distillation approach can effectively enhance the multilingual capabilities of small Whisper models without increasing computational costs. Up to 5.18x reduction in model size is achieved with marginal performance degradation. In addition, quantization is compatible with distillation, which can result in a higher compression rate.
Hang Shao 0005, Bei Liu 0003, Wei Wang 0010, Xun Gong 0005, Yanmin Qian
SLT1
2023 Factorized AED: Factorized Attention-Based Encoder-Decoder for Text-Only Domain Adaptive ASR
abstract
End-to-end automatic speech recognition (ASR) systems have gained popularity given their simplified architecture and promising results. However, text-only domain adaptation remains a big challenge for E2E systems. Text-to-speech (TTS) based approaches fine-tune ASR models by synthesized speech with an auxiliary TTS model, thus increase deployment costs. Language model (LM) fusion based approaches can achieve good performance but are sensitive to interpolation parameters. In order to factorize out the language component in the AED model, we propose the factorized attention-based encoder-decoder (Factorized AED) model whose decoder takes as input the posterior probabilities of a jointly trained LM. Moreover, in the context of domain adaptation, the domain specific LM serves as a plug-and-play component for a well-trained factorized AED model. In-domain experiments on LibriSpeech and out-of-domain experiments adapting from LibriSpeech to a variety of domains in GigaSpeech are conducted to validate the effectiveness of our proposed methods. Results show 20% / 24% relative word error rate (WER) reduction for LibriSpeech test sets and 8 ∼34% relative WER reduction for 8 GigaSpeech target domains test sets compared to the AED baseline.
Xun Gong 0005, Wei Wang 0010, Hang Shao 0005, Xie Chen 0001, Yanmin Qian
ICASSP3
2023 Joint Discriminator and Transfer Based Fast Domain Adaptation For End-To-End Speech Recognition
abstract
Adapting End-to-End (E2E) models to unseen domains is still a big challenge since training E2E models requires lots of paired audio and text training data. We propose a novel domain adaptation framework for the E2E model, which only uses the text of the target domain. Moreover, the proposed methods can keep the performance on the source domain intact while greatly improving the performance on the target domain. The proposed framework consists of two parts: the discriminator and the transfer which were optimized separately. Finally, optimized discriminator and transfer were combined and evaluated on two domain adaption tasks. In the experiments of adapting the English Librispeech to Gigaspeech, we obtained an average relative 11.6% and 11.8% on word error rate (WER) reduction for the target domain dev and test sets, respectively, while almost without WER degradation on the source domain. For the inhouse Chinese corpus aviation and TV, the character error rate (CER) of the source domain increased within 5%, while the CER on the target domain achieved around relative 85% and 42% improvement, respectively. In addition, our approach is also more effective in the mixed domain scenarios in the evaluation.
Hang Shao 0005, Tian Tan 0002, Wei Wang 0010, Xun Gong 0005, Yanmin Qian
ICASSP1
2023 Text Only Domain Adaptation with Phoneme Guided Data Splicing for End-to-End Speech Recognition
Wei Wang 0010, Xun Gong 0005, Hang Shao 0005, Dongning Yang, Yanmin Qian
INTERSPEECH3