Gang Wang 0063

dblp:71/4292-63 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
11since 2021 · last 2025
0009-0003-6944-2958ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 3 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 KVO-LLM: Boosting Long-Context Generation Throughput for Batched LLM Inference
abstract
With the widespread deployment of long-context large language models (LLMs), efficient and high-quality generation is becoming increasingly important. Modern LLMs employ batching and key-value (KV) cache to improve generation throughput and quality. However, as the context length and batch size rise drastically, the KV cache incurs extreme external memory access (EMA) issues. Recent LLM accelerators face substantial processing element (PE) under-utilization due to the low arithmetic intensity of attention with KV cache, while existing KV cache compression algorithms struggle with hardware inefficiency or significant accuracy degradation. To address these issues, an algorithm-architecture co-optimization, KVO-LLM, is proposed for long-context batched LLM generation. At the algorithm level, we propose a KV cache quantization-aware pruning method that first adopts salient-token-aware quantization and then prunes KV channels and tokens by attention guided pruning based on salient tokens identified during quantization. Achieving substantial savings on hardware overhead, our algorithm reduces the EMA of KV cache over 91% with significant accuracy advantages compared to previous KV cache compression algorithms. At the architecture level, we propose a multi-core jointly optimized accelerator that adopts operator fusion and cross-batch interleaving strategy, maximizing PE and DRAM bandwidth utilization. Compared to the state-of-the-art LLM accelerators, KVO-LLM improves generation throughput by up to $7.32 \times$, and attains $5.52 \sim 8.38 \times$ better energy efficiency.
Dongxu Lyu, Gang Wang 0063, Wenjie Li 0003, Jianfei Jiang 0001, Yanan Sun 0003, Guanghui He 0002
DAC3
2025 BitPattern: Enabling Efficient Bit-Serial Acceleration of Deep Neural Networks through Bit-Pattern Pruning
abstract
Bit-serial computation shows promise for accelerating deep neural networks (DNNs) by exploiting inherent bit sparsity. However, the original unstructured bit sparsity poses two major challenges for existing bit-serial accelerators (BSA): (1) workload imbalance from irregular bit distribution, and (2) inefficient memory access due to unpredictable non-zero bit locations. To address these issues, this paper proposes BitPattern, an algorithm/hardware co-design to efficiently accelerate bitserial computation through bit-pattern pruning. At the algorithm level, we employ bit-pattern pruning to identify optimal combinations of predefined patterns and apply compression encoding to minimize weight storage. We further devise a pattern-similaritybased merging method to balance the bit-serial workload. At the hardware level, we co-design a bit-serial accelerator with a dedicated bit-pattern decoder and PE to leverage the potential of structured bit-pattern sparsity. The evaluation on several deep learning benchmarks shows that BitPattern can achieve $1.72 \times$ memory reduction with negligible accuracy loss, and up to $2.11 \times$ speedup and $1.86 \times$ energy saving compared to state-of-the-art bit-serial accelerators.
Gang Wang 0063, Wenjie Li 0003, Dongxu Lyu, Yanan Sun 0003, Jianfei Jiang 0001, Guanghui He 0002
DAC1
2025 VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator
abstract
Large Language Models (LLMs) excel in natural language processing tasks but pose significant computational and memory challenges for edge deployment due to their intensive resource demands. This work addresses the efficiency of LLM inference by algorithm-hardwaredataflow tri-optimizations. We propose a novel voting-based KV cache eviction algorithm, balancing hardware efficiency and algorithm accuracy by adaptively identifying unimportant kv vectors. From a dataflow perspective, we introduce a flexible-product dataflow and a runtime reconfigurable PE array for matrix-vector multiplication. The proposed approach effectively handles the diverse dimensional requirements and solves the challenges of incrementally varying sequence lengths. Additionally, an element-serial scheduling scheme is proposed for nonlinear operations, such as softmax and layer normalization (layernorm). Results demonstrate a substantial reduction in latency, accompanied by a significant decrease in hardware complexity, from $O(N)$ to $O(1)$. The proposed solution is realized in a custom-designed accelerator, VEDA, which outperforms existing hardware platforms. This research represents a significant advancement in LLM inference on resource-constrained edge devices, facilitating real-time processing, enhancing data privacy, and enabling model customization.
Zhican Wang, Hongxiang Fan, Haroon Waris, Gang Wang 0063, Jianfei Jiang 0001, Yanan Sun 0003, Guanghui He 0002
DAC4
2025 COSA Plus: Enhanced Co-Operative Systolic Arrays for Attention Mechanism in Transformers
abstract
The attention mechanism is becoming a vital building block across various modern neural networks, e.g., Transformers. However, it encounters low efficiency when deployed on the general-purpose GPU/CPU platform, which motivates the dedicated accelerator design. Existing accelerators are commonly devised by exploring the potential sparsity in attention mechanism using a hardware-software codesign scheme, which suffers from complicated training, fine-tuning processes, and possible accuracy degradation. More importantly, the sparse pattern only focuses on certain datasets with less generality, and the fine-grained sparse pattern could also bring hardware inefficiency. Instead, we try to solve these issues from another perspective: by systematically analysing the inherent dataflow characteristics of the attention mechanism, we propose the co-operative systolic arrays (COSAs) with an optimized dataflow to support the general purpose attention mechanism and pursue higher computational efficiency. COSA system exploits the high parallelism from the inherent model and leverages run-time configurable hybrid dataflows, i.e., weight and output stationary (OS) for a systolic array (SA) to support the varying matrix multiplication in the attention mechanism. Regarding the cascaded matrix multiplications, COSA proposes levels of fusion methodologies to reduce the off-chip access and enhance processing element (PE) utilization, such as directly using the result of OS as the weight of weight stationary SA by deep fusion. Additionally, the COSA system also provides the solution to hide the latency and radically save the buffer size related to the softmax. Experiment results show that, across various benchmarks, COSA can achieve$2.29-2.60\times $throughput improvement over the traditional SA of the same MAC number, with up to 94.7% PE utilization rate and$8.2\times $less off-chip memory access. Compared with the general-purpose platforms,$7.6-12.4\times $energy efficiency over NVIDIA GeForce 3090 GPU and$35.2-80.9\times $energy efficiency over Intel 6226R server CPU.
Zhican Wang, Gang Wang 0063, Guanghui He 0002
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2025 An Efficient Multi-View Cross-Attention Accelerator for Vision-Centric 3D Perception in Autonomous Driving
abstract
Vision-centric 3D perception has become a key mechanism in autonomous driving. It achieves exceptional perceptual performance mainly by introducing a novel attention,multi-view cross-attention(MVCA), for learnable feature extraction and fusion from surround-view cameras. Despite its superiority, MVCA encounters severe inefficiencies in sample, processing elements (PE), and pipelined processing, owing to the redundant and non-uniform sampling-aggregation and rigorous inter-operator dependencies. To address these issues, this article proposes a dedicated MVCA accelerator, MVAtor, with algorithm-architecture co-optimization for vision-centric 3D perception based on multi-view inputs flexibly. For sample inefficiency, a 3-tier hybrid static-dynamic sample and a sensitivity-aware feature pruning approach are proposed to eliminate the 86.03% sample overhead and 24.48% memory requirement, only incuring <1% accuracy loss with no need of fine-tuning. For PE inefficiency, a spatial pruner and sequential sampler collaboration strategy is proposed to improve the sampler utilization without compromising pruner’s throughput, which outperforms the previous design by 53.7~96.1% energy-delay product reduction. For pipeline inefficiency, a fine-grained-tiling assisted highly-pipelined architecture is constructed in MVAtor by exploiting the decoupling opportunities on inter-view sparsity, thereby saving 61.03% external memory access while boosting the overall throughputs by 1.83×. Extensively evaluated on representative benchmarks, MVAtor attains 1.38~7.67× and 1.67~11.15× improvement on energy and area efficiency respectively, compared to the state-of-the-art related accelerators.
Dongxu Lyu, Gang Wang 0063, Wenjie Li 0003, Weifeng He, Guanghui He 0002
IEEE Trans. Circuits Syst. I Regul. Pap.4
2025 OFQ-LLM: Outlier-Flexing Quantization for Efficient Low-Bit Large Language Model Acceleration
abstract
Large Language Models (LLMs) have achieved significant success in various Natural Language Processing (NLP) tasks, becoming essential to modern intelligent computing. Their large memory footprint and high computational cost hinder efficient deployment. Post-Training Quantization (PTQ) is a promising technique to alleviate this issue and accelerate LLM inference. However, the presence of outliers impedes the advancement of LLM quantization to lower bit levels. In this paper, we introduce OFQ-LLM, an algorithm-hardware co-design solution that adopts outlier-flexing quantization to efficiently accelerate LLM at low-bit levels. The key insight of OFQ-LLM is that normal data can be efficiently quantized in a slightly reduced data encoding space, while the rest encoding space can be used for flexible outlier values. During quantization, we use rescale-based clipping (RBC) to optimize accuracy for normal data and group outlier clustering (GOC) to flexibly represent outlier values. At the hardware level, we introduce a memory-aligned outlier-flexing encoding scheme to encode activations and weights in LLMs at a low bit level. The outlier-normal mixed hardware architecture is devised to leverage the encoding scheme and accelerate LLMs with high speed and high energy efficiency. Our experiments show that OFQ-LLM achieves better accuracy compared to state-of-the-art (SOTA) low-bit LLM PTQ works. OFQ-LLM-based accelerator surpasses the SOTA outlier-aware accelerators by up to$2.69\times $core energy efficiency, up to$3.83\times $speed up and$2.44\times $energy reduction in LLM prefilling phase, and up to$2.01\times $speed up and$2.88\times $energy reduction in LLM decoding phase, with superior accuracy.
Gang Wang 0063, Wenjie Li 0003, Dongxu Lyu, Guanghui He 0002
IEEE Trans. Circuits Syst. I Regul. Pap.1
2024 DEFA: Efficient Deformable Attention Acceleration via Pruning-Assisted Grid-Sampling and Multi-Scale Parallel Processing
abstract
Multi-scale deformable attention (MSDeformAttn) has emerged as a key mechanism in various vision tasks, demonstrating explicit superiority attributed to multi-scale grid-sampling. However, this newly introduced operator incurs irregular data access and enormous memory requirement, leading to severe PE under-utilization. Meanwhile, existing approaches for attention acceleration cannot be directly applied to MSDeformAttn due to lack of support for this distinct procedure. Therefore, we propose a dedicated algorithm-architecture co-design dubbed DEFA, the first-of-its-kind method for MSDeformAttn acceleration. At the algorithm level, DEFA adopts frequency-weighted pruning and probability-aware pruning for feature maps and sampling points respectively, alleviating the memory footprint by over 80%. At the architecture level, it explores the multi-scale parallelism to boost the throughput significantly and further reduces the memory access via fine-grained layer fusion and feature map reusing. Extensively evaluated on representative benchmarks, DEFA achieves 10.1-31.9X speedup and 20.3-37.7X energy efficiency boost compared to powerful GPU platforms. It also rivals the related accelerators by 2.2-3.7X energy efficiency improvement while providing pioneering support of MSDeformAttn.
Dongxu Lyu, Zilong Wang 0030, Gang Wang 0063, Zhican Wang, Haomin Li 0002, Guanghui He 0002
DAC6
2024 A High-Throughput Lossless Image Compression Engine Optimized for Compression Ratio
abstract
Image compression is an essential technique for graphics processing units to support high-resolution video and high-quality 3D rendering. Many works sacrifice compression ratio (CR) or image quality in order to achieve a higher throughput or lower latency. In this paper, a lossless high-throughput compression-decompression engine optimized for CR is proposed. At the algorithm level, the engine utilizes pixel inter-channel correlation and positional correlation to jointly improve CR. At the hardware level, the diagonal-level parallel decoding (DLPD) is used to obtain a high throughput (44.1 Gbytes/sec). The proposed engine is designed and validated on the JPEG AIC-3 dataset. The average CR on the dataset is 2.23, an improvement of 0.35 over the state-of-the-art high-throughput work with a 24.2 % increase in throughput.
Zeyuan Jin, Gang Wang 0063, Guanghui He 0002
ISCAS5
2024 Hardware-oriented algorithms for softmax and layer normalization of large language models
Wenjie Li 0003, Dongxu Lyu, Gang Wang 0063, Aokun Hu, Ningyi Xu, Guanghui He 0002
Sci. China Inf. Sci.3
2024 BSViT: A Bit-Serial Vision Transformer Accelerator Exploiting Dynamic Patch and Weight Bit-Group Quantization
abstract
Vision Transformers (ViTs) have achieved remarkable success in computer vision (CV) and are increasingly recognized as the new backbone for vision-language multi-modal tasks. Despite their success, the high computational cost associated with ViTs hinders their inference efficiency. In this paper, we introduce BSViT, a bit-serial Vision Transformer accelerator enhanced by algorithm-hardware co-design. BSViT can efficiently accelerate both plain and hierarchical Vision Transformer inference. At the algorithm level, we propose a post-training quantization scheme named dynamic patch and weight bit-group quantization. We first introduce a dynamic patch quantization (DPQ) scheme to dynamically allocate bit-width to different image patches based on their importance, thus reducing bit width and saving computation without significantly impacting accuracy. Second, we propose a weight bit-group quantization (BGQ) scheme to evenly distribute bits within groups and achieve workload balance across processing elements (PEs). At the hardware level, we propose a term-separate bit-serial accelerator to efficiently support DPQ and BGQ. We introduce dense and sparse bit-serial PEs to manipulate the dense least significant term (LST) and sparse most significant term (MST) workloads. A dense-sparse hybrid dataflow is devised to efficiently balance the two kinds of workloads. Our experiments show that BSViT can achieve up to$1.95\times $speedup and$2.72\times $energy efficiency compared to state-of-the-art (SOTA) bit-serial accelerators and achieve up to$3.69\times $energy efficiency compared to SOTA Transformer accelerators.
Gang Wang 0063, Wenjie Li 0003, Dongxu Lyu, Guanghui He 0002
IEEE Trans. Circuits Syst. I Regul. Pap.1
2023 COSA:Co-Operative Systolic Arrays for Multi-head Attention Mechanism in Neural Network using Hybrid Data Reuse and Fusion Methodologies
abstract
Attention mechanism acceleration is becoming increasingly vital to achieve superior performance in deep learning tasks. Existing accelerators are commonly devised dedicatedly by exploring the potential sparsity in neural network (NN) models, which suffer from complicated training, tuning processes, and accuracy degradation. By systematically analyzing the inherent dataflow characteristics of attention mechanism, we propose the Co-Operative Systolic Array (COSA) to pursue higher computational efficiency for its acceleration. In COSA, two systolic arrays that can be dynamically configured into weight or output stationary modes are cascaded to enable efficient attention operation. Thus, hybrid dataflows are simultaneously supported in COSA. Furthermore, various fusion methodologies and an advanced softmax unit are designed. Experimental results show that the COSA-based accelerator can achieve 2.95-28.82× speedup compared with the existing designs, with up to 97.4% PE utilization rate and less memory access.
Zhican Wang, Gang Wang 0063, Honglan Jiang, Ningyi Xu, Guanghui He 0002
DAC2