Huizhen Kuang

dblp:369/4911 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
4since 2021 · last 2025
0009-0006-9808-8179ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 4 since 2021
YearPublicationVenuePosition
2025 Two-Phase Transistor Sizing for FPGAs via Bayesian Optimization
abstract
Transistor-level design is pivotal for the accurate evaluation of FPGA architectures. Since traditional linear models are increasingly inadequate in advanced technology nodes, simulation-driven approaches have become the standard for FPGA architecture exploration. However, due to the non-analytical nature of delay measurements from simulations, the transistor sizing process becomes a black-box optimization problem. COFFE2 [1], [2], the state-of-the-art academic sizing tool employs a division-based brute-force approach, which is relatively time-consuming and may lose optimal solutions. In this paper, we propose a two-phase transistor sizing methodology and enhance the COFFE2 framework with Bayesian Optimization, which is well-suited for black-box optimization problems. Our proposed approach not only achieves a 11.7% improvement in the quality of results but also reduces runtime by ~40%-70%. Through extensive benchmarking on a complete design flow, from transistor-level sizing to routing with VTR benchmarks, we demonstrate that FPGA architectures optimized by our approach offer an 11.0% reduction in the area-delay product, proving the efficacy of our method.
Xianfeng Cao, Huizhen Kuang, Yuanqi Wang, Lingli Wang
FPGA2
2025 FLAIC: A Novel FPGA Logic Architecture via Fine-Grained Cut Topology Analysis
abstract
Look-up table (LUT)-based programmable logic blocks (PLBs) serve as the foundation for FPGAs. As increasing the input number of LUTs to improve logic capacity will introduce exponential area overhead, substantial research has focused on designing more efficient alternatives. Previous approaches primarily design dedicated hardware by analyzing the distribution of Boolean functions and implementing those with high frequency. However, these approaches face scalability challenges due to the explosive growth in the function space. In this paper, we consider the topology of cuts rather than Boolean functions they represent. By identifying topologies that occur commonly in cuts and integrating them with LUTs, we propose a new 8-input PLB architecture, named FLAIC. This architecture incurs only a slight area overhead compared to a 6-LUT while achieving logic capacity comparable to that of an 8LUT. Post-synthesis results demonstrate that FLAIC reduces the logic levels by over 20 % and the number of PLBs by more than$\mathbf{1 0 \%}$, compared to 6-LUTs. Additionally, post-implementation results show improvement in critical path delay by 10.6 % and a reduction in the number of Configurable Logic Blocks (CLBs) by 5.3 % on MCNC and VTR benchmarks, compared to the Intel Stratix 10-like architecture.
Xianfeng Cao, Huizhen Kuang, Yuanqi Wang, Lingli Wang
FPL2
2025 GEF: A GNN-Based Evaluation Framework for FPGA Routing Architecture
abstract
The routing architecture significantly impacts the performance of modern FPGAs, motivating extensive research into its design space exploration (DSE). However, DSE efficiency is hindered by non-generalizable parametrization methods and considerable runtime overhead of FPGA architecture evaluation tools. In this paper, we propose GEF, a GNN-based FPGA Evaluation Framework that predicts routability and area-delay product (ADP) across various routing architectures. In GEF, we introduce Intra-Tile Graph, a novel intermediate representation (IR) that encodes global routing patterns in a compact form, serving as the input to predictors. The Routability Predictor (Rou-P) integrates Self-Attention Pooling (SAGPool), while the ADP Predictor (ADP-P) benefits from intermediate supervision through auxiliary node-level labels. Experimental results demonstrate the high accuracy of GEF, with Rou-P achieving 94.56% and ADP-P 94.57 %, respectively. We also conduct ablation studies, which further validate that GEF achieves substantial enhancements through efficient architecture modeling and timingaware analysis. Finally, a case study on routing architecture exploration with the incorporation of GEF is presented, which achieves a$\mathbf{1 5} \boldsymbol{\times}$speedup and enhanced improvements. Our codes are are available from https://github.com/RapidFlex/GEF.
Yuanqi Wang, Yunfei Dai, Kaixiang Zhu, Huizhen Kuang, Eric Ren, Xifan Tang, Weijun Qin, Lingli Wang
FPL5
2025 DEFA: Design Space Exploration for FPGA Overlay Accelerators Through Frequency Prediction and Bayesian Optimization
abstract
In edge AI inference, FPGAs demonstrate superiority in performance-area balance. FPGA Overlay Accelerators (FOAs) are programmable accelerators implemented on FPGAs, typically highly parameterized to enable flexible hardware realization. These parameters, varying across a wide design space, have a significant impact on performance and require efficient Design Space Exploration (DSE). However, current frameworks struggle to accurately predict performance metrics like maximum frequency and fail to fully explore the design space, limiting DSE's effectiveness. In this paper, we propose a DSE framework for FOA (DEFA) based on Bayesian optimization, providing more effective and comprehensive DSE. To address complex parameter interdependencies in FOA, a dependency-aware design space modeling approach (DAM) is proposed. This approach applies fine-grained pruning to the parameter space while addressing dependency constraints. Based on this pruned parameter space, we develop a custom regression predictor (CREP) for maximum frequency using LightGBM, significantly enhancing performance estimation accuracy. Furthermore, the search efficiency is improved through enhanced Latin hypercube sampling and the Tree-Structured Parzen Estimator. We use the proposed framework to optimize an FOA template, Intel FPGA AI Suite. The Pearson correlation coefficient of CREP's predictions regarding the maximum frequency of accelerator instances achieves 0.87. In the throughput optimization experiment, the proposed DSE framework improves 30.16 % compared to the architecture optimization functionality provided by Intel FPGA AI Suite across the given 10 benchmarks on average. In the areathroughput trade-off optimization experiment, compared with FPGA AI Suite, the proposed DSE framework improves 5.01 % in frequency, 18.48 % in throughput and 21.60 % in area.
Qilong Zhu, Yunfei Dai, Shiyan Bi, Huizhen Kuang, Dylan Wang, Wenbo Yin, Lingli Wang
FPL4