EDBT 2026 Demo / reviewers in the wild / expert
Aokun Hu
dblp:344/4608
· DBLP profile ↗
4ranked-venue papers
0as first author
4since 2021 · last 2024
0000-0001-5412-1340ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Hardware accelerators and domain-specific architectures · 100% | |
| Artificial intelligence
2 papers |
Efficient and distributed learning · 62% Deep learning architectures and training · 38% |
Topics — the 7 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Hardware accelerators and domain-specific architectures
machine learning accelerator |
1.5 | 2 | 2024 | A Precision-Scalable Deep Neural Network Accelerator With Activation Sparsity Exploitation · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2024 Hardware-oriented algorithms for softmax and layer normalization of large language models · Sci. China Inf. Sci. 2024 |
Machine learning › Efficient and distributed learning
model compression |
0.8 | 1 | 2024 | A Precision-Scalable Deep Neural Network Accelerator With Activation Sparsity Exploitation · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2024 |
Hardware accelerators and domain-specific architectures
algorithm-hardware co-design |
0.8 | 1 | 2024 | CoDA: A Co-Design Framework for Versatile and Efficient Attention Accelerators · IEEE Trans. Computers 2024 |
Hardware accelerators and domain-specific architectures › machine learning accelerator › transformer accelerator
attention accelerator |
0.8 | 1 | 2024 | CoDA: A Co-Design Framework for Versatile and Efficient Attention Accelerators · IEEE Trans. Computers 2024 |
Hardware accelerators and domain-specific architectures › machine learning accelerator
transformer accelerator |
0.8 | 1 | 2024 | Hardware-oriented algorithms for softmax and layer normalization of large language models · Sci. China Inf. Sci. 2024 |
Machine learning › Deep learning architectures and training
attention mechanism |
0.2 | 1 | 2024 | CoDA: A Co-Design Framework for Versatile and Efficient Attention Accelerators · IEEE Trans. Computers 2024 |
Machine learning › Deep learning architectures and training › neural network inference
DNN inference |
0.2 | 1 | 2024 | A Precision-Scalable Deep Neural Network Accelerator With Activation Sparsity Exploitation · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2024 |
Methods — techniques the papers use, named apart from their topics
zero-skipping · 1.5sparsity-aware mapping · 1.5sparse attention · 1.5softmax sharing · 1.5low-rank attention · 1.5decomposable multiplier · 1.5block-wise dataflow · 1.5hardware-oriented algorithm · 0.8bit-splitting · 0.8bit splitting · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Hardware-oriented algorithms for softmax and layer normalization of large language models
Wenjie Li 0003, Dongxu Lyu, Gang Wang 0063, Aokun Hu, Ningyi Xu, Guanghui He 0002 |
Sci. China Inf. Sci. | 4 |
| 2024 | CoDA: A Co-Design Framework for Versatile and Efficient Attention AcceleratorsabstractAs a primary component of Transformers, attention mechanism suffers from quadratic computational complexity. To achieve efficient implementations, its hardware accelerator designs have aroused great research interest. However, most existing accelerators only support a single type of application and a single type of attention, making it difficult to meet the demands of diverse application scenarios. Additionally, they mainly focus on the dynamic pruning of attention matrices, which requires the deployment of pre-processing units, thereby reducing overall hardware efficiency. This paper presents CoDA which is an algorithm, dataflow and architecture co-design framework for versatile and efficient attention accelerators. The designed accelerator supports both NLP and CV applications, and can be configured into the mode supporting low-rank attention or low-rank plus sparse attention. We apply algorithmic transformations to low-rank attention to significantly reduce computational complexity. To prevent an increase in storage overhead resulting from the proposed algorithmic transformations, we carefully design the dataflows and adopt a block-wise fashion. Down-scaling softmax is further supported by architecture and dataflow co-design. Moreover, we propose a softmax sharing strategy to reduce the area cost. Our experiment results demonstrate that the proposed accelerator outperforms the state-of-the-art designs in terms of throughput, area efficiency and energy efficiency. Wenjie Li 0003, Aokun Hu, Ningyi Xu, Guanghui He 0002 |
IEEE Trans. Computers | 2 |
| 2024 | A Precision-Scalable Deep Neural Network Accelerator With Activation Sparsity ExploitationabstractTo meet the demand in a wide range of practical applications, precision-scalable deep neural network (DNN) accelerators are becoming an unavoidable trend. On the other hand, it has been demonstrated that a DNN accelerator may achieve better computation efficiency through exploiting the sparsity. Therefore, DNN accelerators with both precision scalability and sparsity exploitation are expected to have better performance. In this article, we propose an efficient precision-scalable DNN accelerator that can exploit the sparsity of activations. The precision scalability is obtained from the decomposable multiplier which is inspired by the well-known design, Bit Fusion. Besides, a zero-skipping scheme is adopted to leverage the inherent sparsity of activations. We first modify the architecture of the conventional fusion unit (FU) to make it amenable to the zero-skipping scheme. Then, a segmentation approach is devised to tackle the memory access conflict. Furthermore, a sparsity-aware mapping method is proposed to balance the workload of processing elements (PEs). Moreover, we present a bit-splitting strategy which can take advantage of the sparsity in the bit level. Compared with the state-of-the-art precision-scalable designs, our proposed accelerator can provide speedups of$4.12\times $,$4.07\times $, and$6.62\times $in the precision modes$8b\times 8b$,$4b\times 4b$, and$2b\times 2b$, respectively. Meanwhile, it also achieves$3.92\times $peak area efficiency and competitive peak energy efficiency. Wenjie Li 0003, Aokun Hu, Ningyi Xu, Guanghui He 0002 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2024 | Quantization and Hardware Architecture Co-Design for Matrix-Vector Multiplications of Large Language ModelsabstractLarge language models (LLMs) have sparked a new revolution in the field of natural language processing (NLP), and have garnered tremendous attention in both academic research and everyday life, thanks to their unprecedented performance in a wide range of applications. However, their deployment remains a significant challenge, primarily due to their intensive computational and memory requirements. Hardware acceleration and efficient quantization are promising solutions to address the two issues. In this paper, a quantization and hardware architecture co-design is presented for matrix-vector multiplications (MVMs) of LLMs. During quantization, we uniformly group weights and activations to ensure workload balance for hardware. To enhance the performance of quantization, we further propose two approaches called channel sorting and channel selection, which can be applied simultaneously. To support the proposed quantization scheme, we develop two precision-scalable MVM hardware architectures. They are specifically designed for high speed and high energy efficiency, respectively. Experimental results show that our proposed quantization scheme achieves state-of-the-art performance among all the reported post-training schemes that quantize both weights and activations into integers. Compared to MVM architecture of the state-of-the-art LLM accelerator OliVe, our design exhibits significant advantages in terms of area efficiency and energy efficiency. Wenjie Li 0003, Aokun Hu, Ningyi Xu, Guanghui He 0002 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |