Zhuqing Yuan

dblp:267/1733 · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
3since 2021 · last 2025
0000-0002-1042-9711ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 3 since 2021Artificial intelligence and machine learning · 1Graphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 An Integer-only Quantization Framework for Edge Deployment of Large Language Models
abstract
The rapid growth in the parameter size of large language models (LLMs) has introduced significant challenges for deployment on edge devices. To address these challenges, this paper focuses on developing a post-training quantization (PTQ) framework tailored for LLM deployment on edge devices. Our framework introduces an enhanced channel smoothing technique based on channel value ranges, combined with channel reordering to mitigate quantization errors associated with large channel value range differences and activation outliers, reducing the need for quantization and dequantization steps during inference, making it more suitable for edge deployment. Our approach achieves full integer quantization for all model operations, reducing model size by 4×. Through extensive experiments on various language tasks using the OPT model, we demonstrate that our framework surpasses state-of-the-art methods under the W4A4 configuration.
Yaqi Hu, Zhuqing Yuan, Weichen Gao, Yongpan Liu
ISCAS2
2022 Efficient Neural Networks with Spatial Wise Sparsity Using Unified Importance Map
abstract
Exploiting neural network sparsity is one of the most important directions to accelerate CNN executions. Plenty of techniques are proposed to exploit neural network sparsity, where spatial-wise pruning is quite effective for input image. However, previous spatial-wise pruning methods need nontrivial hardware overhead for dynamic execution, due to layer-by-layer binary sampling and online scheduling. This paper proposes a structured configured, spatial-wise pruning technique. Numerous computation will be saved by skipping unimportant region. By using a unified importance map, the computing graph could be compiled in advance to make it more hardware friendly. Additionally, due to multi-level measurement of importance for each region, our method can have a better performance on various tasks. On image classification task, the method can have around 50% fewer top-1 accuracy drop than previous spatialwise pruning methods at similar sparse level. On super resolution and image deraining task, the method can bring $5 \times$ to $19 \times$ acceleration while causing neglectable effect on reconstruction quality. Hardware implementation is also included.
Wenyu Sun, Wenxun Wang, Zhuqing Yuan, Yongpan Liu
ISCAS4
2022 PACA: A Pattern Pruning Algorithm and Channel-Fused High PE Utilization Accelerator for CNNs
abstract
In recent years, convolutional neural networks (CNNs) have achieved significant advancements in various fields. However, the computation and storage overheads of CNNs are overwhelming for Internet-of-Things devices. Both network pruning algorithms and hardware accelerators have been introduced to empower CNN inference at the edge. Network pruning algorithms reduce the size and computational cost of CNNs by regularizing unimportant weights to zeros. However, existing works lack intrakernel structured types to tradeoff between sparsity and hardware efficiency, and the index storage for irregularly pruned networks is significant. Hardware accelerators leverage the sparsity of pruned CNNs to improve energy efficiency. However, their process element (PE) utilization rate is low because of uneven sparsity among input convolutional kernels. To overcome these problems, we propose PACA: a Pattern pruning Algorithm and Channel-fused high PE utilization Accelerator for CNNs. It includes three parts: a pattern pruning algorithm to explore the intrakernel sparsity type and reduce the index storage, a channel-fused hardware architecture to reduce the PEs’ idle rate and improve the performance, and a heuristic and taboo search-based smart fusion scheduler to analyze the idle PE problem and schedule the channel fusion in hardware. To demonstrate the effectiveness of PACA, we have implemented the software parts by Python and the hardware architecture by RTL codes. Experimental results on various datasets show that compared with an existing work, PACA can reduce the index storage overhead by$3.47\times $–$5.63\times $with 3.85–9.12 average patterns, and it can improve the hardware performance by$2.01\times $–$5.53\times $because of PEs’ idle rate reduction.
Jingyu Wang 0004, Songming Yu, Zhuqing Yuan, Jinshan Yue, Ruoyang Liu, Yanzhi Wang 0001, Huazhong Yang, Xueqing Li 0002, Yongpan Liu
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2020 High PE Utilization CNN Accelerator with Channel Fusion Supporting Pattern-Compressed Sparse Neural Networks
abstract
Recently CNN-based methods have made remarkable progress in broad fields. Both network pruning algorithms and hardware accelerators have been introduced to accelerate CNN. However, existing pruning algorithms have not fully studied the pattern pruning method, and current index storage scheme of sparse CNN is not efficient. Furthermore, the performance of existing accelerators suffers from no-load PEs on sparse networks. This work proposes a software-hardware co-design to address these problems. The software includes an ADMM-based method which compresses the patterns of convolution kernels with acceptable accuracy loss, and a Huffman encoding method which reduces index storage overhead. The hardware is a fusion-enabled systolic architecture, which can reduce PEs' no-load rate and improve performance by supporting the channel fusion. On CIFAR-10, this work achieves 5.63× index storage reduction with 2-7 patterns among different layers with 0.87% top-1 accuracy loss. Compared with the state-of-art accelerator, this work achieves 1.54×-1.79× performance and 25%-34% reduction of no-load rate with reasonable area and power overheads.
Jingyu Wang 0004, Songming Yu, Jinshan Yue, Zhuqing Yuan, Huazhong Yang, Xueqing Li 0002, Yongpan Liu
DAC5
2020 High-Quality Single-Model Deep Video Compression with Frame-Conv3D and Multi-frame Differential Modulation
Wenyu Sun, Weigui Li, Zhuqing Yuan, Huazhong Yang, Yongpan Liu
ECCV (30)4
2020 Multi-channel precision-sparsity-adapted inter-frame differential data codec for video neural network processor
abstract
Activation I/O traffic is a critical bottleneck of video neural network processor. Recent works adopted an inter-frame difference method to reduce activation size. However, current methods can't fully adapt to the various precision and sparsity in differential data. In this paper, we propose the multi-channel precision-sparsity-adapted codec, which will separate the differential activation and encode activation in multiple channels. We analyze the most adapted encoding of each channel, and select the optimal channel number with the best performance. A two-channel codec hardware has been implemented in the ASIC accelerator, which can encode/decode activations in parallel. Experiment results show that our coding achieves 2.2x-18.2x compression rate in three scenarios with no accuracy loss, and the hardware has 42x/174x improvement on speed and energy-efficiency compared with the software codec.
Yixiong Yang, Fang Su, Fanyang Cheng, Zhuqing Yuan, Huazhong Yang, Yongpan Liu
ISLPED5