Benzheng Li

dblp:284/7892 · DBLP profile ↗
← Back
6ranked-venue papers
5as first author
6since 2021 · last 2024
0000-0002-1227-3909ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 5 first-author · 6 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2024 An Efficient Hypergraph Partitioner under Inter - Block Interconnection Constraints
abstract
Multi-FPGA systems are increasingly employed for very large scale integration circuit emulation and prototyping. Due to limited I/O resources, each FPGA often only has direct physical connections to a few other FPGAs. Therefore, if signals between FPGAs originate from a source FPGA and flow toward a target FPGA not directly connected to the source FPGA, intermediate FPGAs will be used as hops in the signal path. These FPGA-hops increase signal delays and the number of physical lines used in signal multiplexing between FPGAs, degrading system performance. To address these issues, researchers proposed partitioners that guarantees zero hop, but they lead to a considerable cut-size. In this paper, building on previous research, we introduce a new candidate block propagation theorem and optimize the initial partition process based on its corollary. Additionally, we also present a method for correcting violations during uncoarsening to improve the solver capability. Results of experiments demonstrate that our proposed method significantly reduces the cut size by 96% while retaining comparable running times.
Benzheng Li, Hailong You, Shunyang Bi
DATE1
2024 MaPart: An Efficient Multi-FPGA System-Aware Hypergraph Partitioning Framework
abstract
Multi-FPGA systems (MFSs) are increasingly important in addressing VLSI circuit emulation and prototyping. However, the limitations of I/O resources between FPGAs have driven the usage of TDM and FPGA-hop technologies, which complicate the partitioning problem. Consequently, designing a suitable partitioning process for MFS has emerged as a critical research question affecting overall system performance. This paper proposes MaPart, a novel hypergraph partitioning framework, which aims to minimize the maximum path delay in MFS. MaPart combines binary search with a non-hop partitioner, TopoPart+, to minimize the maximum hop count during the partitioning process. Compared to previous non-hop partitioner, TopoPart+ provides enhanced problem-solving capabilities and achieves a remarkable 96% reduction in cut-size. Furthermore, the framework incorporates two successive local refinement algorithms that optimize the time-division multiplexing ratio, reduce total hop count, and alleviate congestion on critical paths. Additionally, MaPart includes a system-level router based on layered graphs, enabling flexible control of the hop count based on the timing criticality of each path. Experimental results demonstrate that the proposed framework achieves a significant 37% reduction in delay compared to baseline algorithms when evaluated using publicly available benchmarks.
Benzheng Li, Shunyang Bi, Hailong You, Zhongdong Qi, Richard Sun
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2023 ASSURER: A PPA-friendly Security Closure Framework for Physical Design
abstract
Hardware security is emerging in the very large scale integration (VLSI). The seminal threats, like hardware Trojan insertion, probing attacks, and fault injection, are hard to detect and almost impossible to fix at post-design stage. The optimal solution is to prevent them at the physical design stage. Usually, defending against them may cause a lot of power, performance, and area (PPA) loss. In this paper, we propose a PPA-friendly physical layout security closure framework ASSURER. Reward-directed placement refinement and multi-threshold partition algorithm are proposed to assure Trojan threats are empty. Cleaning up probing attacks is established on a patch-based ECO routing flow. Evaluated on the ISPD'22 benchmarks, ASSURER can clean out the Trojan threat with no leakage power increase when shrinking the physical layout area. When not shrinking, ASSURER only increases 14% total power. Compared with the work of first place in the ISPD2022 Contest, ASSURE reduced 53% additional total power consumption, and probing vulnerability can be reduced by 97.6% under the premise of timing closure. We believe this work shall open up a new perspective for preventing Trojan insertion and probing attacks.
Hailong You, Zhengguang Tang, Benzheng Li, Cong Li 0023, Xiaojue Zhang
ASP-DAC4
2023 Machine Learning Based Framework for Fast Resource Estimation of RTL Designs Targeting FPGAs
abstract
Field-programmable gate arrays (FPGAs) have grown to be an important platform for integrated circuit design and hardware emulation. However, with the dramatic increase in design scale, it has become a key challenge to partition very large scale integration into multi-FPGA systems. Fast estimation of FPGA on-chip resource usage for individual sub-circuit blocks early in the circuit design flow will provide an essential basis for reasonable circuit partition. It will also help FPGA designers to tune the circuits in hardware description language. In this article, we propose a framework for fast estimation of the on-chip resources consumed by register transfer level (RTL) designs with machine learning methods. We extensively collect RTL designs as a dataset, extract features from the result of a parser tool and analyze their roles, and train a targeted three-stage ensemble learning model. A 5,513× speedup is achieved while having 27% relative absolute error. Although the effect is sufficient to support RTL circuit partition, we discuss how the estimation quality continues to be improved.
Benzheng Li, Hailong You, Zhongdong Qi
ACM Trans. Design Autom. Electr. Syst.1
2022 High quality hypergraph partitioning for logic emulation
Benzheng Li, Zhongdong Qi, Zhengguang Tang, Xiyi He, Hailong You
Integr.1
2021 Placement for Wafer-Scale Deep Learning Accelerator
abstract
To meet the growing demand from deep learning applications for computing resources, accelerators by ASIC are necessary. A wafer-scale engine (WSE) is recently proposed [1], which is able to simultaneously accelerate multiple layers from a neural network (NN). However, without a high-quality placement that properly maps NN layers onto the WSE, the acceleration efficiency cannot be achieved. Here, the WSE placement resembles the traditional ASIC floor plan problem of placing blocks onto a chip region, but they are fundamentally different. Since the slowest layer determines the compute time of the whole NN on WSE, a layer with a heavier workload needs more computing resources. Besides, locations of layers and protocol adapter cost of internal 10 connections will influence inter-layer communication overhead. In this paper, we propose GigaPlacer to handle this new challenge. A binary-search-based framework is developed to obtain a minimum compute time of the NN. Two dynamic-programming-based algorithms with different optimizing strategies are integrated to produce legal placement. The distance and adapter cost between connected layers will be further minimized by some refinements. Compared with the first place of the ISPD2020 Contest, GigaPlacer reduces the contest metric by up to 6.89% and on average 2.09%, while runs 7.23X faster.
Benzheng Li, Qi Du, Dingcheng Liu, Jingchong Zhang, Gengjie Chen, Hailong You
ASP-DAC1