Guang Yang 0036

dblp:25/5712-36 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
5since 2021 · last 2025
0009-0001-0000-1447ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021
YearPublicationVenuePosition
2025 TWDP: A Vision Transformer Accelerator with Token-Weight Dual-Pruning Strategy for Edge Device Deployment
abstract
Vision Transformers (ViTs) have attracted significant attention due to their superior accuracy compared to convolutional neural networks (CNNs) in various computer vision tasks. However, their substantial computational load and significant memory footprint lead to excessive delay and considerable data storage overhead, posing challenges for resource-limited edge device deployment. To address these issues, we present TWDP, a vision transformer accelerator employing a Token-Weight Dual-Pruning strategy to enhance the efficiency of the inference process. Firstly, we propose a parameter-free self-adaptive token pruning method to skip redundant computations in an image-dependent manner. Secondly, we apply a Hessian-aware layer-wise N:M weight pruning approach to minimize storage overhead, memory access, and computational power consumption. Additionally, to manage the complex computing patterns in ViTs, an overlapping dataflow is utilized to further reduce temporal storage and inference latency. Implemented and evaluated under an industrial 28nm technology, the proposed TWDP framework reduces 66.1% weight storage requirements and achieves an energy efficiency of 2070.9 FPS/W. Compared to state-of-the-art architectures, TWDP obtains a 1.6× energy efficiency improvement with negligible accuracy loss, demonstrating the superiority of TWDP in edge device deployment scenarios.
Guang Yang 0036, Xinming Yan, Hui Kou, Zihan Zou, Qingwen Wei, Hao Cai 0001, Bo Liu 0019
ASP-DAC1
2025 S-DMA: Sparse Diffusion Models Acceleration via Spatiality-Aware Prediction and Dimension-Adaptive Dataflow
abstract
Diffusion Models (DMs) have demonstrated remarkable performance in a variety of image generation tasks.However, their complex architectures and intensive computations result in significant overhead and latency, posing challenges for hardware deployment.To address these issues, researchers have explored the sparsity in DMs to reduce computational workloads, including semantic sparsity in image generation and spatial sparsity in local editing.Unfortunately, existing sparsity prediction methods face critical limitations in deployment: 1) additional prediction overheads offset the benefits of sparsity; 2) convolution and general matrix multiplication (GEMM) exhibit distinct sparsity patterns, which current co-design frameworks struggle to process.In this paper, we introduce S-DMA, a software-hardware co-design framework that unifies efficient sparsity prediction while supporting various sparse operators.First, we propose a spatiality-aware similarity computation method that leverages the local similarity of images, reducing the computational complexity of sparsity prediction from O(𝑁 2 ) to O(N ).Second, we implement NAND-based similarity for sparsity prediction, which minimizes the computational overheads and ensures adaptability to different sparsity schemes.Finally, a dedicated hardware architecture is designed to efficiently leverage the algorithm optimizations.A NAND-based sparsity prediction processing unit is designed to adaptively handle the sparsity patterns.Additionally, a sparsity-aware reduction network and a dimension-adaptive Bo Liu is the corresponding author.
Zihan Zou, Xinming Yan, Guang Yang 0036, Hao Cai 0001, Bo Liu 0019
MICRO5
2025 PF²A-ViT: Parameter-Free and Feature-Aware Dynamic Token Pruning Accelerator With Complementary Quantization-Encoding for Vision Transformer
abstract
Vision Transformers (ViTs) have achieved outstanding performance in visual applications. However, ViT’s increasing parameters and computation overhead limit its deployment on hardware. Previous ViT accelerators focus on optimizing the core attention mechanism of ViTs due to the high overheads of language transformer-based neural networks. Nevertheless, linear layers are the actual bottleneck of ViT inference, accounting for larger than 90% FLOPs on numerous ViTs because of the short and fixed token length of ViTs. To this end, we propose PF2A-ViT, an algorithm and accelerator co-design framework, comprehensively accelerating both core attention and linear layers. At the algorithm level, a parameter-free and feature-aware dynamic token pruning (PF2ATP) strategy is proposed to reduce the dimension of feature maps and dynamically adjust the pruning ratio according to the complexity of the feature without complex execution of subnetworks. Meanwhile, a mixed-precision quantization strategy combines Hessian trace and parameter-aware signal-to-quantization-noise ratio to boost the deployment efficiency of PF2ATP. In addition, a cluster-regroup-based weight encoding strategy is proposed to compensate for the bit-wise redundant information of the quantization strategy. At the hardware level, a token pruning module based on bitonic sorters is designed to fully leverage PF2ATP. Simultaneously, a 3D-PE array with reconfigurable 4/8-bit processing elements (PE) is designed to implement the mixed-precision quantization strategy and equipped with greedy bit-wise compensation decoders to exploit the encoding strategy. Extensive experiments on multiple ViTs demonstrate the achievements of PF2A-ViT: (1) Maximally realize 3.89× speedup, 5.54× energy efficiency compared to state-of-the-art ViT accelerators. (2) Occupying a 2.25 mm area and consuming 76 mW power in 28-nm technology.
Zihan Zou, Xinming Yan, Chen Zhang 0025, Shikuang Chen, Guang Yang 0036, Han Yan 0014, Hao Cai 0001, Bo Liu 0019
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2024 Timing Error Tolerant CNN Accelerator With Layerwise Approximate Multiplication
abstract
Exploiting the error tolerance in computation, approximate circuits become an emerging computing paradigm to increase the energy efficiency in digital systems, which is crucial in high-performance and low-power systems for the edge Internet-of-Things (EIoT) devices. Inspired by the state-of-the-art high-efficiency NN accelerators, three techniques are proposed for effectively integrating the approximate computing unit into CNN accelerator to achieve a dynamic energy-accuracy trade-off: (1) An approximate multiplier that can be configured to three precision modes is proposed. A weight pre-encoding method is used to save hardware overhead. (2) For hybrid-accuracy layer-wise mapping, the hessian-aware layer-wise accuracy scaling is proposed, which concerns inference accuracy and hardware overhead simultaneously. A progressive re-training approach is proposed to enable an aggressive approximation configuration and higher power reduction. (3) A tensor multiplication unit (TMU) with timing error detection and correction (TEDC) approach is proposed, enabling an aggressive voltage scaling and a 41.5% power reduction is obtained. An energy-efficient CNN accelerator is proposed and shows how deep learning can be brought to EIoT devices by running each layer at its appropriate computational accuracy. Implemented under 28-nm CMOS technology, the CNN accelerator achieves the energy efficiency of 14.4 TOPS/W. The proposed accelerator and method are conducted on the applications of keyword spotting of GSCD, CIFAR10 and CIFAR100, 44.5%~46.7% multiplication energy is saved while reducing the accuracy by less than 0.6%.
Bo Liu 0019, Na Xie, Qingwen Wei, Guang Yang 0036, Chonghang Xie, Weiqiang Liu 0001, Hao Cai 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2024 Layer-Wise Mixed-Modes CNN Processing Architecture With Double-Stationary Dataflow and Dimension-Reshape Strategy
abstract
With the development of convolutional neural networks (CNN) across various domains, the growth in network structure complexity and computational load has increasingly become a research focus in the deployment of neural networks. The key to current research on neural network accelerators lies in striking a balance between computational accuracy and energy efficiency. This paper proposes a software-hardware co-design to strike the balance for CNN edge applications. On the hardware side, a 3-dimensional tensor engine (3D-TE), achieved with reconfigurable Tensor Processing Units (TPUs), is introduced for efficient convolution computation. We optimize the CNN dataflow on 3D-TE using a dimension reshaping method for feature maps rearrangement, and a double stationary dataflow scheduling to reduce memory access. This paper adopts a configurable approximate multiplier design based on Boolean Matrix Factorization (BMF) based logic synthesis applied in the architecture of TPU. The proposed 3D-TE, characterized by its configurable precision, enables the TPUs to dynamically adapt the bitwidth of features and weights in response to varying precision requirements. On the software side, a hessian-guided layer precision mapping is adopted to reduce unnecessary computational overhead, and a progressive re-training approach is proposed to enable a better approximation configuration and higher power reduction. Fabricated on 28-nm CMOS, this work achieves an optimized energy efficiency of 14.9 TOPS/W and 12.1 TOPS/W for ResNet56 and MobileNetV2 respectively, with 0.6V supply voltage and 150MHz clock frequency, representing an improvement of$1.33\times \sim 8.28\times $over the state-of-the-art works.
Bo Liu 0019, Xinxiang Huang, Yang Zhang 0132, Guang Yang 0036, Han Yan 0014, Chen Zhang 0025, Zejv Li, Yuanhao Wang 0009, Hao Cai 0001
IEEE Trans. Circuits Syst. I Regul. Pap.4