Haonan Du

dblp:378/0114 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
6since 2021 · last 2026
0009-0003-6345-8991ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 2 first-author · 6 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 QUNF+: A Quadratic Approximation Framework With Hardware Co-Design for Universal Nonlinear Function Acceleration in Neural Networks
abstract
Modern neural networks have undergone extensive hardware optimization to address the increasing computational demands. While most existing acceleration strategies concentrate on linear operations, the relative cost of these nonlinear operations has become a critical efficiency bottleneck. This work presents QUNF+, a hardware-centric, quadratic-based approximation framework that offers a universal, scalable, and accurate approach for accelerating a wide range of nonlinear functions. Unlike conventional piecewise linear methods, QUNF+ segments functions uniformly and applies second-order Taylor expansions, yielding superior accuracy with fewer segments. QUNF+ also introduces a hardware-efficient approximation scheme with adjustable precision and an optional remainder compensation mechanism. In addition, we propose a set of novel function remapping techniques further reducing approximation error with minimal overhead. We design a fully integrated hardware architecture incorporating Canonical Signed Digit encoding and logic pruning to minimize resource usage without accuracy loss. Experimental results demonstrate that QUNF+ achieves up to 4.1x, 1.5x and 1.7x improvements in power, area, and latency, respectively, over state-of-the-art PWL methods. Application to real-world Transformer models results in an accuracy degradation of less than 0.72% and 0.10% on NLP and CV tasks, with system-level evaluations showing up to 24.38% reduction in total energy consumption and 92.39% in nonlinear operation energy. These results establish QUNF+ as a robust and scalable solution for nonlinear acceleration in modern AI hardware.
Haonan Du, Chenyi Wen, Zhengrui Chen, Li Zhang 0021, Qi Sun 0002, Zheyu Yan, Cheng Zhuo
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2026 A High-Parallelism Softmax Hardware-Software Co-Design for Fast and Efficient LLM Inference
abstract
Large language models (LLMs) have been the driving force behind significant advancements in artificial intelligence. However, their unique self-attention mechanism leads to difficulties in accelerating the inference. Softmax, with its complex nonlinear operations and low parallelism, significantly limits the efficiency of LLMs for long sequences. This work proposes a novel high-parallelism hardware/software co-design Softmax solution. By incorporating a sum estimation algorithm, we eliminate the need for complex exponential on all elements. A high-speed low-energy hardware architecture is introduced by applying high-parallelism statistical module and simplified division module. Experimental results demonstrate that our approach achieves a 9.90–44.75% reduction in latency and up to a 17.36% reduction in energy consumption compared with state-of-the-art Softmax hardware, with minimal impact on model inference accuracy.
Chenyi Wen, Haonan Du, Xuyang He, Zheyu Yan, Qi Sun 0002, Cheng Zhuo
ACM Trans. Design Autom. Electr. Syst.2
2025 VQT-CiM: Accelerating Vector Quantization Enhanced Transformer with Ferroelectric Compute-in-Memory
abstract
Transformer models have achieved state-of-the-art performance in various natural language processing (NLP) and computer vision (CV) tasks. To meet their substantial computational demands, the compute-in-memory (CiM) architectures, which alleviate the memory wall problem and enable efficient vector-matrix multiplication (VMM), have been adopted for transformer accelerators. However, the dynamic VMM involved in the attention mechanism, which necessitates runtime write operations, presents significant challenges for non-volatile memory (NVM)-based CiM designs. High write overhead, complex compute-write-compute (CWC) dependencies, and limited endurance reduce their effectiveness. In this paper, we propose VQT-CiM, a ferroelectric FET (FeFET)-based CiM design that accelerates vector quantization (VQ) enhanced transformers by eliminating the runtime write operations. VQT-CiM quantizes keys and values in self-attention to convert dynamic VMMs in inner-product and weighted-sum into static VMMs with the codebooks, enabling efficient calculations with CiM crossbars. However, directly applying VQ hinders the accuracy of transformer model due to its limited representation capability. To address this, we introduce a vector quantization scheme that integrates residual VQ (RVQ) and product VQ (PVQ) for enhanced representation space. We present an efficient hardware implementation for the proposed VQT-CiM with optimized dataflow in RVQ, which incorporates the FeFET-based CiM crossbars and peripheral digital circuits. Evaluation results suggest that VQT-CiM achieves the $3.54 \times$ and $4.53 \times$ improvements in energy efficiency and throughput, respectively, compared to state-of-the-art NVM-based CiM transformer designs.
Xuchu Huang, Haonan Du, Zheyu Yan, Cheng Zhuo, Xunzhao Yin
DAC2
2025 Algorithm-Hardware Co-Design of a Unified Accelerator for Non-Linear Functions in Transformers
abstract
Nonlinear functions (NFs) in Transformers require high-precision computation consuming significant time and energy, despite the aggressive quantization schemes for other components. Piece-wise Linear (PWL) approximation-based methods offer more efficient processing schemes for NFs but fall short in dealing with functions with high nonlinearities. Moreover, PWL-based methods still suffer from inevitably high latency introduced by the Multiply-And-Add (MADD) unit. To address these issues, this paper proposes a novel quadratic approximation scheme and a highly integrated, multiplier-less hardware structure, as a unified method to accelerate any unary nonlinear function. We also demonstrate implementation examples for GELU, Softmax, and LayerNorm. The experimental results show that the proposed method achieves up to 5.41% higher inference accuracy and 60.12% lower area-delay product.
Haonan Du, Chenyi Wen, Zhengrui Chen, Li Zhang 0021, Qi Sun 0002, Zheyu Yan, Cheng Zhuo
DATE1
2025 PACE: A Piece-Wise Approximate Floating-Point Divider with Runtime Configurability and High Energy Efficiency
abstract
Approximate computing emerges as a viable solution to enhance energy efficiency in applications sensitive to human perception, particularly on edge devices. This work introduces a novel piece-wise approximate floating-point divider that boasts resource efficiency and runtime configurability. Our method leverages a piece-wise approximation algorithm for computing 1/ y by exploiting powers of 2, complemented by an error compensation technique grounded in thorough mathematical analysis. This approach facilitates the realization of a reciprocal-based floating-point divider devoid of multipliers, which not only mitigates hardware resource consumption but also reduces latency. Additionally, we unveil a multi-level runtime configurable hardware architecture that significantly improves flexibility across diverse application contexts. Compared to the existing state-of-the-art approximate dividers and truncated exact dividers, our proposed solution achieves a superior compromise between precision and resource efficiency. Application-level evaluations reveal that our design provides over 87.7% energy saving while maintaining a negligible impact on output quality.
Chenyi Wen, Haonan Du, Zhengrui Chen, Li Zhang 0021, Qi Sun 0002, Cheng Zhuo
ACM Trans. Design Autom. Electr. Syst.2
2024 PACE: A Piece-Wise Approximate and Configurable Floating - Point Divider for Energy - Efficient Computing
abstract
Approximate computing is a promising alternative to improve energy efficiency for human perception related applications on the edge. This work proposes a piece-wise approximate floating-point divider, which is resource-efficient and run-time configurable. We provide a piece-wise approximation algorithm for 1/ y, utilizing powers of 2. This approach enables the implementation of a reciprocal-based floating-point divider that is independent of multipliers, which not only reduces hardware consumption but also results in shorter latency. Furthermore, a multi-level run-time configurable hardware structure is intro-duced, enhancing the adaptability to various application scenarios. When compared to the prior state-of-the-art approximate divider, the proposed divider strikes an advantageous balance between accuracy and resource efficiency. The application-level evaluation of the proposed dividers demonstrates manageable and minimal degradation of the output quality when compared to the exact divider.
Chenyi Wen, Haonan Du, Zhengrui Chen, Li Zhang 0021, Qi Sun 0002, Cheng Zhuo
DATE2