Zhen Li 0059

dblp:74/2397-59 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
5since 2021 · last 2026
0000-0003-3994-9304ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021
YearPublicationVenuePosition
2026 LOFMPL: An Open-source Logic Optimization Framework with MFFC-based Hypergraph Partition and Reinforcement Learning for Large Circuits
abstract
As the size of a circuit increases, previous reinforcement learning (RL) approaches struggle to effectively explore the logic optimization sequences of large-scale Boolean networks due to the long runtime overhead with poor optimization results. This article proposes LOFMPL: an open-source logic optimization framework with Maximum Fanout-Free Cone (MFFC) based hypergraph partitioning and reinforcement learning. The novel two-stage MFFC-based hypergraph partitioning can divide the circuit into highly independent subnetworks, which can be explored by an enhanced parallel RL-based design space exploration engine with an improved objective function. The experiment is conducted based on more than 150 benchmarks with logic optimization and ASIC technology mapping tasks and compared with other ML-based and greedy methods. The different partitioning algorithms are also compared for the subsequent logic optimization. Experimental results demonstrate that the proposed partitioning algorithm significantly enhances optimization quality without greatly increasing partitioning time, outperforming the KaHypar algorithm. Additionally, for the logic optimization task, the proposed method achieves a node-level-product improvement of 13% over the RLG synthesis exploration technique, 3% over the ESE reinforcement learning framework, 14% over the Boils synthesis method, and 7% over the DRiLLS synthesis method, while delivering greater reductions in node count compared with the Bulls-Eye optimization technique. For the ASIC technology mapping task, the proposed method achieves an area-delay-product improvement of 23% over the LSOracle framework, 9% over the Boils synthesis method, and 5% over the DRiLLS synthesis method. Hence, LOFMPL can achieve better results within the same runtime constraints compared with state-of-the-art works.
Kaixiang Zhu, Zhen Li 0059, Jide Zhang, Wai-Shing Luk, Lingli Wang
ACM Trans. Design Autom. Electr. Syst.2
2022 Low Error-Rate Approximate Multiplier Design for DNNs with Hardware-Driven Co-Optimization
abstract
In this paper, two approximate 3 × 3 multipliers are proposed and the synthesis results of the ASAP-7nm process library justify that they can reduce the area by 31.38% and 36.17%, and the power consumption by 36.73% and 35.66% compared with the exact multiplier, respectively. They can be aggregated with a 2 × 2 multiplier to produce an 8 × 8 multiplier with low error-rate based on the distribution of DNN weights. We propose a hardware-driven software co-optimization method to improve the DNN accuracy by retraining. Based on the proposed two approximate 3-bit multipliers, three approximate 8-bit multipliers with low error-rate are designed for DNNs. Compared with the exact 8-bit unsigned multiplier, our design can achieve a significant advantage over other approximate multipliers on the public dataset.
Jide Zhang, Su Zheng, Zhen Li 0059, Lingli Wang
ISCAS4
2022 HEAM: High-Efficiency Approximate Multiplier optimization for Deep Neural Networks
abstract
We propose an optimization method for the automatic design of approximate multipliers, which minimizes the average error according to the operand distributions. Our multiplier achieves up to 50.24% higher accuracy than the best reproduced approximate multiplier in DNNs, with 15.76% smaller area, 25.05% less power consumption, and 3.50% shorter delay. Compared with an exact multiplier, our multiplier reduces the area, power consumption, and delay by 44.94%, 47.63%, and 16.78%, respectively, with negligible accuracy losses. The tested DNN accelerator modules with our multiplier obtain up to 18.70% smaller area and 9.99% less power consumption than the original modules.
Su Zheng, Zhen Li 0059, Jingbo Gao, Jide Zhang, Lingli Wang
ISCAS2
2022 Adaptable Approximate Multiplier Design Based on Input Distribution and Polarity
abstract
Approximate computing is an efficient approach to reduce the design complexity for error-resilient applications. Multipliers are key arithmetic units in many applications, such as deep neural networks (DNNs) and digital signal processing (DSP) systems. In this article, an open-source adaptable approximate multiplier design driven by input distribution and polarity is proposed to generate optimized approximate multipliers to trade off between the application-level performance and the hardware cost. The proposed method minimizes the average square of the absolute error of an approximate multiplier according to the probability distributions of operands extracted from the target application with consideration of input polarity, achieving low hardware cost and negligible application-level performance loss. The proposed method can generate unsigned multipliers (or signed multipliers) based on the Braun multiplier (or Baugh–Wooley multiplier). To demonstrate the effectiveness of the method, three different-scale quantized DNNs, including LeNet, AlexNet, and VGG16 with 8$\times $8 unsigned multiplication and an adaptive least mean square (LMS)-based finite impulse response (FIR) filter with 16$\times $16 fixed-point signed multiplication, are evaluated. In the DNN training process, a noise training technique is adopted to reduce the accuracy loss due to the approximation. When compared to the state-of-the-art approximate multipliers, the generated multipliers can achieve up to 26.4% and 27.1% product of power, delay, and area gains with negligible application-level performance loss in VGG16 and FIR applications, respectively.
Zhen Li 0059, Su Zheng, Jide Zhang, Jingbo Gao, Jun Tao 0001, Lingli Wang
IEEE Trans. Very Large Scale Integr. Syst.1
2021 Parallelized Technology Mapping to General PLBs by Adaptive Circuit Partitioning
abstract
Technology mapping from logic netlists to programmable logic blocks (PLB) plays an important role in FPGA EDA flow, especially for architecture exploration of PLBs. However, technology mapping becomes time-consuming due to the booming scale and complexity of IC designs as well as the growing complexity of PLB architectures. To speed up this process, a parallelized technology mapping approach based on adaptive circuit partitioning is proposed in this paper to perform fast multi-thread technology mapping. First, We choose the best of the three candidate partitioning strategies for the given netlist by circuit analysis to partition the original netlist into several independent sub-netlists. Secondly, these sub-netlists are mapped to the given PLB architecture simultaneously in their corresponding mapping threads. Finally, the complete mapped netlist is generated by merging the mapped sub-netlists. The proposed approach is implemented in ABC, independent of the detailed mapping algorithm. 13 large circuits from the Titan23 benchmark set are used as benchmarks to evaluate the proposed approach. Experimental results show that the proposed approach leads to an average of 5.76 × speedup over the single-thread version (up to 8.21 × individually) with no delay loss and less than 0.57% average area penalty.
Xiaoxi Wang, Moucheng Yang, Zhen Li 0059, Lingli Wang
FPT3