Li Zhang 0021

dblp:89/5992-21 · DBLP profile ↗
← Back
17ranked-venue papers
1as first author
14since 2021 · last 2026
0000-0002-8951-4969ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 15 · 1 first-author · 12 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 QUNF+: A Quadratic Approximation Framework With Hardware Co-Design for Universal Nonlinear Function Acceleration in Neural Networks
abstract
Modern neural networks have undergone extensive hardware optimization to address the increasing computational demands. While most existing acceleration strategies concentrate on linear operations, the relative cost of these nonlinear operations has become a critical efficiency bottleneck. This work presents QUNF+, a hardware-centric, quadratic-based approximation framework that offers a universal, scalable, and accurate approach for accelerating a wide range of nonlinear functions. Unlike conventional piecewise linear methods, QUNF+ segments functions uniformly and applies second-order Taylor expansions, yielding superior accuracy with fewer segments. QUNF+ also introduces a hardware-efficient approximation scheme with adjustable precision and an optional remainder compensation mechanism. In addition, we propose a set of novel function remapping techniques further reducing approximation error with minimal overhead. We design a fully integrated hardware architecture incorporating Canonical Signed Digit encoding and logic pruning to minimize resource usage without accuracy loss. Experimental results demonstrate that QUNF+ achieves up to 4.1x, 1.5x and 1.7x improvements in power, area, and latency, respectively, over state-of-the-art PWL methods. Application to real-world Transformer models results in an accuracy degradation of less than 0.72% and 0.10% on NLP and CV tasks, with system-level evaluations showing up to 24.38% reduction in total energy consumption and 92.39% in nonlinear operation energy. These results establish QUNF+ as a robust and scalable solution for nonlinear acceleration in modern AI hardware.
Haonan Du, Chenyi Wen, Zhengrui Chen, Li Zhang 0021, Qi Sun 0002, Zheyu Yan, Cheng Zhuo
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2026 ZlibBoost: An Efficient and Flexible Open-Source Framework for Standard Cell Characterization
abstract
As VLSI designs grow increasingly complex and transition to smaller process nodes, accurate and efficient library characterization has become essential for modern design workflows. Existing open-source tools are often constrained by limited functionality, efficiency, and accuracy, making them insufficient for today’s design challenges. This article reviews the shortcomings of current open-source tools and introduces ZlibBoost, a novel open-source framework designed to provide both flexibility and high performance. Its modular, front-end and back-end separated architecture, along with user-friendly interfaces, enables seamless customization, integration of machine learning models, and expanded simulator compatibility. A variety of key features are introduced to significantly enhance both accuracy and efficiency of library characterization. Experimental results demonstrate ZlibBoost’s capability to meet the demands of both academic research and practical applications, establishing it as a robust solution for advancing semiconductor design.
Zhengrui Chen, Chengjun Guo, Shizhang Wang, Guozhu Feng, Zixuan Song, Xunzhao Yin, Weiquan Song, Li Zhang 0021, Zheyu Yan, Cheng Zhuo
ACM Trans. Design Autom. Electr. Syst.9
2026 Machine Learning-Assisted VCD Processing for Accelerated Dynamic Voltage Drop Analysis
abstract
With escalating power integrity challenges in advanced technologies, acquiring accurate dynamic power supply noise through Dynamic Voltage Drop (DVD) analysis becomes increasingly demanding. As noise margins shrink, the use of Value Change Dump (VCD) files for precise DVD analysis is indispensable but computationally expensive. Furthermore, the substantial storage requirements of VCD files, which record digital waveforms from logical simulations, pose significant challenges. In this article, we propose a machine learning (ML)-assisted VCD processing framework to accelerate DVD analysis and improve data efficiency. Transitions recorded in VCD files are mapped to a Physical Design-Aware Circuit Hierarchy Tree (CHT) for efficient feature extraction. These features are leveraged by an XGBoost-based predictor to identify critical vector time windows within the VCD, significantly reducing simulation complexity. Additionally, Huffman encoding is applied to compress signal names, further optimizing storage utilization. Experimental results show that DVD analysis using our profiled VCD files achieves a speedup of approximately 3.53× with an error margin of only 3.89%.
Jingchao Hu, Yufei Chen 0007, Songyu Sun, Jianfei Song, Li Zhang 0021, Xunzhao Yin, Zhou Jin 0001, Cheng Zhuo
ACM Trans. Design Autom. Electr. Syst.5
2026 A Multimode Built-In Self-Test Circuit for Sequential Cell Timing Characteristics Based on Digital-to-Time Converter
abstract
To address the growing challenges in high-precision timing characterization, where conventional measurement methods face limitations due to increasing process variations and multicharacteristic testing demands, this work presents an innovative multimode built-in self-test (BIST) circuit based on a digital-to-time converter (DTC). This circuit comprises a multimode design under test (DUT) module, a DTC, and a controller, enabling the simultaneous measurement of critical timing characteristics (setup/hold times and CK–Q/QN delays) across multiple sequential cells with varying trigger types and clock edges. To enhance the measurement performance, we propose a DTC architecture featuring a digitally controlled multilevel tunable delay cell within a two-stage “coarse–fine” delay chain and a multilevel switching mechanism. The BIST circuit achieves a minimum resolution of 18.7 ps, with the DTC offering a 10.5-ns dynamic range and high linearity, occupying 0.03 mm2, while the total BIST area is 0.58 mm2. Experimental results in a 180-nm process show setup/hold time measurement errors of −21 to 27 and −27 to 25 ps, respectively, and a CK–Q(QN) delay error range of −22.3 to 30.7 ps.
Wenwen Cai, Yanhui Zhao, Zhengrui Chen, Xuecheng Zou, Cheng Zhuo, Li Zhang 0021
IEEE Trans. Very Large Scale Integr. Syst.9
2025 Invited Paper: Boosting Standard Cell Library Characterization with Machine Learning
abstract
As VLSI designs grow more complex and transition to smaller process nodes, accurate and efficient library characterization has become increasingly crucial within DTCO and STCO flows. Current open-source tools, however, are constrained to basic library characterization functions and fail to adequately meet modern design demands. In this paper, we review the existing open-source standard cell characterization tools, summarize their limitations, and introduce ZlibBoost---a new open-source framework designed to offer both flexibility and efficiency. We leverage ZlibBoost for LUT index optimization, dynamic power supply noise modeling, and machine learning-based prediction to enhance efficiency and accuracy in library characterization. Experimental results show that such a tool is helpful for both academia and industry to effectively navigate DTCO and STCO challenges.
Zhengrui Chen, Chengjun Guo, Zixuan Song, Guozhu Feng, Shizhang Wang, Li Zhang 0021, Xunzhao Yin, Zheyu Yan, Cheng Zhuo
ASP-DAC6
2025 Algorithm-Hardware Co-Design of a Unified Accelerator for Non-Linear Functions in Transformers
abstract
Nonlinear functions (NFs) in Transformers require high-precision computation consuming significant time and energy, despite the aggressive quantization schemes for other components. Piece-wise Linear (PWL) approximation-based methods offer more efficient processing schemes for NFs but fall short in dealing with functions with high nonlinearities. Moreover, PWL-based methods still suffer from inevitably high latency introduced by the Multiply-And-Add (MADD) unit. To address these issues, this paper proposes a novel quadratic approximation scheme and a highly integrated, multiplier-less hardware structure, as a unified method to accelerate any unary nonlinear function. We also demonstrate implementation examples for GELU, Softmax, and LayerNorm. The experimental results show that the proposed method achieves up to 5.41% higher inference accuracy and 60.12% lower area-delay product.
Haonan Du, Chenyi Wen, Zhengrui Chen, Li Zhang 0021, Qi Sun 0002, Zheyu Yan, Cheng Zhuo
DATE4
2025 Dynamic Multi-scale Feature Integration Network for unsupervised MR-CT synthesis
Jiuming Jiang, Tao Zhou 0002, Yizhe Zhang 0001, Bin Qiu, Li Zhang 0021
Neural Networks7
2025 PACE: A Piece-Wise Approximate Floating-Point Divider with Runtime Configurability and High Energy Efficiency
abstract
Approximate computing emerges as a viable solution to enhance energy efficiency in applications sensitive to human perception, particularly on edge devices. This work introduces a novel piece-wise approximate floating-point divider that boasts resource efficiency and runtime configurability. Our method leverages a piece-wise approximation algorithm for computing 1/ y by exploiting powers of 2, complemented by an error compensation technique grounded in thorough mathematical analysis. This approach facilitates the realization of a reciprocal-based floating-point divider devoid of multipliers, which not only mitigates hardware resource consumption but also reduces latency. Additionally, we unveil a multi-level runtime configurable hardware architecture that significantly improves flexibility across diverse application contexts. Compared to the existing state-of-the-art approximate dividers and truncated exact dividers, our proposed solution achieves a superior compromise between precision and resource efficiency. Application-level evaluations reveal that our design provides over 87.7% energy saving while maintaining a negligible impact on output quality.
Chenyi Wen, Haonan Du, Zhengrui Chen, Li Zhang 0021, Qi Sun 0002, Cheng Zhuo
ACM Trans. Design Autom. Electr. Syst.5
2024 PACE: A Piece-Wise Approximate and Configurable Floating - Point Divider for Energy - Efficient Computing
abstract
Approximate computing is a promising alternative to improve energy efficiency for human perception related applications on the edge. This work proposes a piece-wise approximate floating-point divider, which is resource-efficient and run-time configurable. We provide a piece-wise approximation algorithm for 1/ y, utilizing powers of 2. This approach enables the implementation of a reciprocal-based floating-point divider that is independent of multipliers, which not only reduces hardware consumption but also results in shorter latency. Furthermore, a multi-level run-time configurable hardware structure is intro-duced, enhancing the adaptability to various application scenarios. When compared to the prior state-of-the-art approximate divider, the proposed divider strikes an advantageous balance between accuracy and resource efficiency. The application-level evaluation of the proposed dividers demonstrates manageable and minimal degradation of the output quality when compared to the exact divider.
Chenyi Wen, Haonan Du, Zhengrui Chen, Li Zhang 0021, Qi Sun 0002, Cheng Zhuo
DATE4
2024 An Agile Framework for Efficient LLM Accelerator Development and Model Inference
abstract
Large Language Models (LLMs) have revolutionized many domains with exceptional performance while their large sizes hinder their broad applicability, especially in the edge computation scenarios. Designing large-scale LLM-specific accelerators is also challenging, suffering from the complicated, cumbersome, and time-consuming design, simulation, and optimization process. This paper meticulously proposes an agile framework for accelerator development, supporting efficient LLM inference. Firstly, we investigate the architecture of LLMs, uncover performance bottlenecks, and design an optimized binarized accelerator and a configurable RISC-V-based SoC to boost the inference of binary LLMs. Further, a novel fidelity-driven method is proposed to learn the multi-fidelity representation, solving the modeling and accuracy issues due to the lack of accurate later-stage data in the EDA flow, by capturing complex relationships among simulation metrics in and across different fidelities. Tailored strategies across model preparation, backend kernel implementations, agile accelerator and SoC design, and inference simulation are incorporated into our framework to refine the development workflow. Our method significantly accelerates the hardware design, simulation, and optimization processes. Experimental results illustrate the impressive speed and effectiveness of our framework in designing edge LLM accelerators and optimizing LLM inference.
Lvcheng Chen, Chenyi Wen, Shizhang Wang, Li Zhang 0021, Bei Yu 0001, Qi Sun 0002, Cheng Zhuo
ICCAD5
2023 A fine-grained mixed precision DNN accelerator using a two-stage big-little core RISC-V MCU
Li Zhang 0021, Qishen Lv, Wenchao Meng, Qinmin Yang, Cheng Zhuo
Integr.1
2022 Data-Driven Deep Supervision for Skin Lesion Classification
Suraj Mishra, Yizhe Zhang 0001, Li Zhang 0021, Tianyu Zhang 0001, Xiaobo Sharon Hu, Danny Ziyi Chen
MICCAI (1)3
2021 A Physical-Aware Framework for Memory Network Design Space Exploration
abstract
At the era of big data, there have been growing demands for server memory capacity and performance. Memory network is a promising alternative to provide high bandwidth and low latency through distributed memory nodes connected by high speed interconnect. However, most of them implement the design from a pure-logic-level and ignore the physical impact from network interconnect latency, processor placement and the interplay between processor and memory. In this work, we propose a Physical-Aware framework for memory network design space exploration, which facilitates the design of an energy efficient and physical-aware memory network system. Experimental results on various workloads show that the proposed framework can help customize network topology with significant improvements on various design metrics when compared to the other commonly used topologies.
Tianhao Shen, Li Zhang 0021, Jishen Zhao, Cheng Zhuo
ASP-DAC3
2021 A Reconfigurable Multiplier for Signed Multiplications with Asymmetric Bit-Widths
abstract
Multiplications have been commonly conducted in quantized CNNs, filters, and reconfigurable cores, and so on, which are widely deployed in mobile and embedded applications. Most multipliers are designed to perform multiplications with symmetric bit-widths, i.e., n - by n -bit multiplication. Such features would cause extra area overhead and performance loss when m - by n -bit multiplications ( m > n ) are deployed in the same hardware design, resulting in inefficient multiplication operations. It is highly desired and challenging to propose a reconfigurable multiplier design to accommodate operands with both symmetric and asymmetric bit-widths. In this work, we propose a reconfigurable approximate multiplier to support multiplications at various precisions, i.e., bit-widths. Unlike prior works of approximate adders assuming a uniform weight distribution with bit-wise independence, scenarios like a quantized CNN may have a centralized weight distribution and hence follow a Gaussian-like distribution with correlated adjacent bits. Thus, a new block-based approximate adder is also proposed as part of the multiplier to ensure energy-efficient operation with an awareness of the bit-wise correlation. Our experimental results show that the proposed approximate adder significantly reduces the error rate by 76% to 98% over a state-of-the-art approximate adder for Gaussian-like distribution scenarios. Evaluation results show that the proposed multiplier is 19% faster and 22% more power saving than a Xilinx multiplier IP at the same bit precision and achieves a 23.94-dB peak signal-to-noise ratio, which is comparable to the accurate one of 24.10 dB when deployed in a Gaussian filter for image processing tasks.
Chuliang Guo, Li Zhang 0021, Grace Li Zhang, Bing Li 0005, Weikang Qian, Xunzhao Yin, Cheng Zhuo
ACM J. Emerg. Technol. Comput. Syst.2
2020 A Reconfigurable Approximate Multiplier for Quantized CNN Applications
abstract
Quantized CNNs, featured with different bit-widths at different layers, have been widely deployed in mobile and embedded applications. The implementation of a quantized CNN may have multiple multipliers at different precisions with limited resource reuse or one multiplier at higher precision than needed causing area overhead. It is then highly desired to design a multiplier by accounting for the characteristics of quantized CNNs to ensure both flexibility and energy efficiency. In this work, we present a reconfigurable approximate multiplier to support multiplications at various precisions, i.e., bit-widths. Moreover, unlike prior works assuming uniform distribution with bit-wise independence, a quantized CNN may have centralized weight distribution and hence follow a Gaussian-like distribution with correlated adjacent bits. Thus, a new block-based approximate adder is also proposed as part of the multiplier to ensure energy efficient operation with awareness of bit-wise correlation. Our experimental results show that the proposed adder significantly reduces the error rate by 76-98% over a state-of-the-art approximate adder for such scenarios. Moreover, with the deployment of the proposed multiplier, which is 17% faster and 22% more power saving than a Xilinx multiplier IP at the same precision, a quantized CNN implemented in FPGA achieves 17% latency reduction and 15% power saving compared with a full precision case.
Chuliang Guo, Li Zhang 0021, Weikang Qian, Cheng Zhuo
ASP-DAC2
2020 A Convolutional Neural Network Accelerator Architecture with Fine-Granular Mixed Precision Configurability
abstract
Convolutional neural networks (CNNs) have been widely deployed in deep learning applications, especially on power hungry GP-GPUs. Recent efforts in designing CNN accelerators are considered as a promising alternative to achieve higher energy efficiency. Unfortunately, with the growing complexity of CNN, the demanded computational and storage resources for accelerators keep increasing, hindering its wider applications in mobile devices. On the other hand, many quantization algorithms have been proposed for efficient CNN training, which brings many small or zero weights. This is a unique opportunity for accelerator designers to employ much fewer bits, e.g., 4 bits, in both arithmetic core and storage, thereby saving significant design cost. However, such a single precision strategy inevitably compromises the accuracy as some key operations may demand a higher precision. Thus, this paper proposes a low power CNN accelerator architecture that can simultaneously conduct computations with mixed precisions and assign the appropriate arithmetic cores to operation with different precision demands. This proposed architecture can achieve significant area and energy savings, without accuracy compromise. The experimental results show that the proposed architecture implemented on FPGA can reduces almost half of the weight storage and MAC area, and lower the dynamic power by 12.1% when compared with a state-of-the-art CNN accelerator design.
Li Zhang 0021, Chuliang Guo, Xunzhao Yin, Cheng Zhuo
ISCAS2
2019 Optimal design of a low-power, phase-switching modulator for implantable medical applications
Dawei Li 0012, Xiaowei Xu 0004, Leibo Liu, Li Zhang 0021, Cheng Zhuo, Yiyu Shi 0001
Integr.4