Aokun Hu

dblp:344/4608 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
4since 2021 · last 2024
0000-0001-5412-1340ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Hardware accelerators and domain-specific architectures · 100%
Artificial intelligence
2 papers
Efficient and distributed learning · 62% Deep learning architectures and training · 38%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware accelerators and domain-specific architectures
machine learning accelerator
1.522024
A Precision-Scalable Deep Neural Network Accelerator With Activation Sparsity Exploitation · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2024
Hardware-oriented algorithms for softmax and layer normalization of large language models · Sci. China Inf. Sci. 2024
Machine learning › Efficient and distributed learning
model compression
0.812024
A Precision-Scalable Deep Neural Network Accelerator With Activation Sparsity Exploitation · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2024
Hardware accelerators and domain-specific architectures
algorithm-hardware co-design
0.812024
CoDA: A Co-Design Framework for Versatile and Efficient Attention Accelerators · IEEE Trans. Computers 2024
Hardware accelerators and domain-specific architectures › machine learning accelerator › transformer accelerator
attention accelerator
0.812024
CoDA: A Co-Design Framework for Versatile and Efficient Attention Accelerators · IEEE Trans. Computers 2024
Hardware accelerators and domain-specific architectures › machine learning accelerator
transformer accelerator
0.812024
Hardware-oriented algorithms for softmax and layer normalization of large language models · Sci. China Inf. Sci. 2024
Machine learning › Deep learning architectures and training
attention mechanism
0.212024
CoDA: A Co-Design Framework for Versatile and Efficient Attention Accelerators · IEEE Trans. Computers 2024
Machine learning › Deep learning architectures and training › neural network inference
DNN inference
0.212024
A Precision-Scalable Deep Neural Network Accelerator With Activation Sparsity Exploitation · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2024

Methods — techniques the papers use, named apart from their topics

zero-skipping · 1.5sparsity-aware mapping · 1.5sparse attention · 1.5softmax sharing · 1.5low-rank attention · 1.5decomposable multiplier · 1.5block-wise dataflow · 1.5hardware-oriented algorithm · 0.8bit-splitting · 0.8bit splitting · 0.8
YearPublicationVenuePosition
2024 Hardware-oriented algorithms for softmax and layer normalization of large language models
Wenjie Li 0003, Dongxu Lyu, Gang Wang 0063, Aokun Hu, Ningyi Xu, Guanghui He 0002
Sci. China Inf. Sci.4
2024 CoDA: A Co-Design Framework for Versatile and Efficient Attention Accelerators
abstract
As a primary component of Transformers, attention mechanism suffers from quadratic computational complexity. To achieve efficient implementations, its hardware accelerator designs have aroused great research interest. However, most existing accelerators only support a single type of application and a single type of attention, making it difficult to meet the demands of diverse application scenarios. Additionally, they mainly focus on the dynamic pruning of attention matrices, which requires the deployment of pre-processing units, thereby reducing overall hardware efficiency. This paper presents CoDA which is an algorithm, dataflow and architecture co-design framework for versatile and efficient attention accelerators. The designed accelerator supports both NLP and CV applications, and can be configured into the mode supporting low-rank attention or low-rank plus sparse attention. We apply algorithmic transformations to low-rank attention to significantly reduce computational complexity. To prevent an increase in storage overhead resulting from the proposed algorithmic transformations, we carefully design the dataflows and adopt a block-wise fashion. Down-scaling softmax is further supported by architecture and dataflow co-design. Moreover, we propose a softmax sharing strategy to reduce the area cost. Our experiment results demonstrate that the proposed accelerator outperforms the state-of-the-art designs in terms of throughput, area efficiency and energy efficiency.
Wenjie Li 0003, Aokun Hu, Ningyi Xu, Guanghui He 0002
IEEE Trans. Computers2
2024 A Precision-Scalable Deep Neural Network Accelerator With Activation Sparsity Exploitation
abstract
To meet the demand in a wide range of practical applications, precision-scalable deep neural network (DNN) accelerators are becoming an unavoidable trend. On the other hand, it has been demonstrated that a DNN accelerator may achieve better computation efficiency through exploiting the sparsity. Therefore, DNN accelerators with both precision scalability and sparsity exploitation are expected to have better performance. In this article, we propose an efficient precision-scalable DNN accelerator that can exploit the sparsity of activations. The precision scalability is obtained from the decomposable multiplier which is inspired by the well-known design, Bit Fusion. Besides, a zero-skipping scheme is adopted to leverage the inherent sparsity of activations. We first modify the architecture of the conventional fusion unit (FU) to make it amenable to the zero-skipping scheme. Then, a segmentation approach is devised to tackle the memory access conflict. Furthermore, a sparsity-aware mapping method is proposed to balance the workload of processing elements (PEs). Moreover, we present a bit-splitting strategy which can take advantage of the sparsity in the bit level. Compared with the state-of-the-art precision-scalable designs, our proposed accelerator can provide speedups of$4.12\times $,$4.07\times $, and$6.62\times $in the precision modes$8b\times 8b$,$4b\times 4b$, and$2b\times 2b$, respectively. Meanwhile, it also achieves$3.92\times $peak area efficiency and competitive peak energy efficiency.
Wenjie Li 0003, Aokun Hu, Ningyi Xu, Guanghui He 0002
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2024 Quantization and Hardware Architecture Co-Design for Matrix-Vector Multiplications of Large Language Models
abstract
Large language models (LLMs) have sparked a new revolution in the field of natural language processing (NLP), and have garnered tremendous attention in both academic research and everyday life, thanks to their unprecedented performance in a wide range of applications. However, their deployment remains a significant challenge, primarily due to their intensive computational and memory requirements. Hardware acceleration and efficient quantization are promising solutions to address the two issues. In this paper, a quantization and hardware architecture co-design is presented for matrix-vector multiplications (MVMs) of LLMs. During quantization, we uniformly group weights and activations to ensure workload balance for hardware. To enhance the performance of quantization, we further propose two approaches called channel sorting and channel selection, which can be applied simultaneously. To support the proposed quantization scheme, we develop two precision-scalable MVM hardware architectures. They are specifically designed for high speed and high energy efficiency, respectively. Experimental results show that our proposed quantization scheme achieves state-of-the-art performance among all the reported post-training schemes that quantize both weights and activations into integers. Compared to MVM architecture of the state-of-the-art LLM accelerator OliVe, our design exhibits significant advantages in terms of area efficiency and energy efficiency.
Wenjie Li 0003, Aokun Hu, Ningyi Xu, Guanghui He 0002
IEEE Trans. Circuits Syst. I Regul. Pap.2