Han Yan 0014

dblp:63/49-14 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2025
0009-0004-6738-8510ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Hardware accelerators and domain-specific architectures · 100%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.912025
PF²A-ViT: Parameter-Free and Feature-Aware Dynamic Token Pruning Accelerator With Complementary Quantization-Encoding for Vision Transformer · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2025
Hardware accelerators and domain-specific architectures › quantization
mixed-precision quantization
0.912025
PF²A-ViT: Parameter-Free and Feature-Aware Dynamic Token Pruning Accelerator With Complementary Quantization-Encoding for Vision Transformer · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2025
Hardware accelerators and domain-specific architectures
quantization
0.912025
PF²A-ViT: Parameter-Free and Feature-Aware Dynamic Token Pruning Accelerator With Complementary Quantization-Encoding for Vision Transformer · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2025
Hardware accelerators and domain-specific architectures › machine learning accelerator › DNN inference
token pruning
0.912025
PF²A-ViT: Parameter-Free and Feature-Aware Dynamic Token Pruning Accelerator With Complementary Quantization-Encoding for Vision Transformer · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2025
Hardware accelerators and domain-specific architectures › machine learning accelerator › transformer accelerator
vision transformer accelerator
0.912025
PF²A-ViT: Parameter-Free and Feature-Aware Dynamic Token Pruning Accelerator With Complementary Quantization-Encoding for Vision Transformer · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2025

Methods — techniques the papers use, named apart from their topics

weight encoding · 0.9bitonic sort · 0.9
YearPublicationVenuePosition
2025 PF²A-ViT: Parameter-Free and Feature-Aware Dynamic Token Pruning Accelerator With Complementary Quantization-Encoding for Vision Transformer
abstract
Vision Transformers (ViTs) have achieved outstanding performance in visual applications. However, ViT’s increasing parameters and computation overhead limit its deployment on hardware. Previous ViT accelerators focus on optimizing the core attention mechanism of ViTs due to the high overheads of language transformer-based neural networks. Nevertheless, linear layers are the actual bottleneck of ViT inference, accounting for larger than 90% FLOPs on numerous ViTs because of the short and fixed token length of ViTs. To this end, we propose PF2A-ViT, an algorithm and accelerator co-design framework, comprehensively accelerating both core attention and linear layers. At the algorithm level, a parameter-free and feature-aware dynamic token pruning (PF2ATP) strategy is proposed to reduce the dimension of feature maps and dynamically adjust the pruning ratio according to the complexity of the feature without complex execution of subnetworks. Meanwhile, a mixed-precision quantization strategy combines Hessian trace and parameter-aware signal-to-quantization-noise ratio to boost the deployment efficiency of PF2ATP. In addition, a cluster-regroup-based weight encoding strategy is proposed to compensate for the bit-wise redundant information of the quantization strategy. At the hardware level, a token pruning module based on bitonic sorters is designed to fully leverage PF2ATP. Simultaneously, a 3D-PE array with reconfigurable 4/8-bit processing elements (PE) is designed to implement the mixed-precision quantization strategy and equipped with greedy bit-wise compensation decoders to exploit the encoding strategy. Extensive experiments on multiple ViTs demonstrate the achievements of PF2A-ViT: (1) Maximally realize 3.89× speedup, 5.54× energy efficiency compared to state-of-the-art ViT accelerators. (2) Occupying a 2.25 mm area and consuming 76 mW power in 28-nm technology.
Zihan Zou, Xinming Yan, Chen Zhang 0025, Shikuang Chen, Guang Yang 0036, Han Yan 0014, Hao Cai 0001, Bo Liu 0019
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2024 Layer-Wise Mixed-Modes CNN Processing Architecture With Double-Stationary Dataflow and Dimension-Reshape Strategy
abstract
With the development of convolutional neural networks (CNN) across various domains, the growth in network structure complexity and computational load has increasingly become a research focus in the deployment of neural networks. The key to current research on neural network accelerators lies in striking a balance between computational accuracy and energy efficiency. This paper proposes a software-hardware co-design to strike the balance for CNN edge applications. On the hardware side, a 3-dimensional tensor engine (3D-TE), achieved with reconfigurable Tensor Processing Units (TPUs), is introduced for efficient convolution computation. We optimize the CNN dataflow on 3D-TE using a dimension reshaping method for feature maps rearrangement, and a double stationary dataflow scheduling to reduce memory access. This paper adopts a configurable approximate multiplier design based on Boolean Matrix Factorization (BMF) based logic synthesis applied in the architecture of TPU. The proposed 3D-TE, characterized by its configurable precision, enables the TPUs to dynamically adapt the bitwidth of features and weights in response to varying precision requirements. On the software side, a hessian-guided layer precision mapping is adopted to reduce unnecessary computational overhead, and a progressive re-training approach is proposed to enable a better approximation configuration and higher power reduction. Fabricated on 28-nm CMOS, this work achieves an optimized energy efficiency of 14.9 TOPS/W and 12.1 TOPS/W for ResNet56 and MobileNetV2 respectively, with 0.6V supply voltage and 150MHz clock frequency, representing an improvement of$1.33\times \sim 8.28\times $over the state-of-the-art works.
Bo Liu 0019, Xinxiang Huang, Yang Zhang 0132, Guang Yang 0036, Han Yan 0014, Chen Zhang 0025, Zejv Li, Yuanhao Wang 0009, Hao Cai 0001
IEEE Trans. Circuits Syst. I Regul. Pap.5