Yunfei Dai

dblp:119/5940 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
6since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 4 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Compacted-LUT: Fine-Grained Customizable LUT Architecture via SRAM-MUX Co-Optimization
abstract
Traditional FPGA PLB designs are constrained by the exponential increase in LUT area with the augmentation of inputs. Recent work has explored a pruned LUT based on the non-uniform distribution of Boolean functions in practical benchmarks, designing an 8-input PLB with enhanced functionality and a modest area overhead. Nonetheless, the existing LUT pruning algorithm is prone to local optima and focuses exclusively on SRAM pruning, neglecting lookahead optimization of the MUX tree. In this paper, we propose Compacted-LUT (CLUT), a fine-grained customizable LUT architecture via SRAM-MUX co-optimization. Based on the principle of LUT pruning, we design a novel representation for Boolean functions. This representation directly associates each Boolean function with the number of required SRAMs and MUX-tree transistors. On this basis, a novel evaluation model for the hardware-friendliness of Boolean functions can be formulated. We further design a beam search algorithm to identify an optimal subset of Boolean functions in target benchmarks based on evaluation results. With this subset, the customizable SRAM-MUX co-optimized CLUT architecture can be generated. Furthermore, we propose Asym-CLUT6, a function-diverse 8-input PLB composed of two variant 6-input CLUTs. We evaluate Asym-CLUT6 on VTR and Koios benchmarks. Post-route results show that, compared to the Altera Stratix 10-like architecture and Dual-RLUT6, Asym-CLUT6 reduces the area-delay product by 13.65% and 10.06% on average.
Yunfei Dai, Wai-Shing Luk, Lingli Wang
DATE2
2026 A Collaborative Framework for Multi-Level Multi-Objective Design Space Exploration
abstract
High-level synthesis (HLS) tools have drawn considerable attention in recent years because they can automatically generate hardware description code from high-level semantics under compiler-controlled configurations. However, the time-consuming design process, the inherent trade-offs among design objectives, and the often suboptimal quality of RTL produced by HLS have meant that prior studies rarely scale to or investigate the downstream stages beyond HLS.In this paper, we present COLA, an end to end design space exploration (DSE) framework that effectively automates the adaptive tuning of compiler transformation sequences and logic synthesis directives. First, we introduce MOEBO, a holistic Bayesian optimization method that builds multiple local surrogate models within trust regions while maintaining a global surrogate to correct local search bias and align decisions across regions. We further design a cooperative acquisition maximization scheme that coordinates these surrogates to propose diverse and promising candidates in parallel. Additionally, we employ reinforcement learning (RL) techniques to optimize logic synthesis by exploring the design space more effectively, improving the quality of the generated RTL and minimizing the area-delay product. The RL model dynamically adapts the logic synthesis directives to achieve better optimization outcomes over traditional methods. Experimental results show that, our framework achieves a substantial speedup across diverse accelerators for varying kernel granularities with a better trade-off between area and performance.
Kaixiang Zhu, Yuping Bai, Yunfei Dai, Lingli Wang
DATE4
2025 GEF: A GNN-Based Evaluation Framework for FPGA Routing Architecture
abstract
The routing architecture significantly impacts the performance of modern FPGAs, motivating extensive research into its design space exploration (DSE). However, DSE efficiency is hindered by non-generalizable parametrization methods and considerable runtime overhead of FPGA architecture evaluation tools. In this paper, we propose GEF, a GNN-based FPGA Evaluation Framework that predicts routability and area-delay product (ADP) across various routing architectures. In GEF, we introduce Intra-Tile Graph, a novel intermediate representation (IR) that encodes global routing patterns in a compact form, serving as the input to predictors. The Routability Predictor (Rou-P) integrates Self-Attention Pooling (SAGPool), while the ADP Predictor (ADP-P) benefits from intermediate supervision through auxiliary node-level labels. Experimental results demonstrate the high accuracy of GEF, with Rou-P achieving 94.56% and ADP-P 94.57 %, respectively. We also conduct ablation studies, which further validate that GEF achieves substantial enhancements through efficient architecture modeling and timingaware analysis. Finally, a case study on routing architecture exploration with the incorporation of GEF is presented, which achieves a$\mathbf{1 5} \boldsymbol{\times}$speedup and enhanced improvements. Our codes are are available from https://github.com/RapidFlex/GEF.
Yuanqi Wang, Yunfei Dai, Kaixiang Zhu, Huizhen Kuang, Eric Ren, Xifan Tang, Weijun Qin, Lingli Wang
FPL2
2025 DEFA: Design Space Exploration for FPGA Overlay Accelerators Through Frequency Prediction and Bayesian Optimization
abstract
In edge AI inference, FPGAs demonstrate superiority in performance-area balance. FPGA Overlay Accelerators (FOAs) are programmable accelerators implemented on FPGAs, typically highly parameterized to enable flexible hardware realization. These parameters, varying across a wide design space, have a significant impact on performance and require efficient Design Space Exploration (DSE). However, current frameworks struggle to accurately predict performance metrics like maximum frequency and fail to fully explore the design space, limiting DSE's effectiveness. In this paper, we propose a DSE framework for FOA (DEFA) based on Bayesian optimization, providing more effective and comprehensive DSE. To address complex parameter interdependencies in FOA, a dependency-aware design space modeling approach (DAM) is proposed. This approach applies fine-grained pruning to the parameter space while addressing dependency constraints. Based on this pruned parameter space, we develop a custom regression predictor (CREP) for maximum frequency using LightGBM, significantly enhancing performance estimation accuracy. Furthermore, the search efficiency is improved through enhanced Latin hypercube sampling and the Tree-Structured Parzen Estimator. We use the proposed framework to optimize an FOA template, Intel FPGA AI Suite. The Pearson correlation coefficient of CREP's predictions regarding the maximum frequency of accelerator instances achieves 0.87. In the throughput optimization experiment, the proposed DSE framework improves 30.16 % compared to the architecture optimization functionality provided by Intel FPGA AI Suite across the given 10 benchmarks on average. In the areathroughput trade-off optimization experiment, compared with FPGA AI Suite, the proposed DSE framework improves 5.01 % in frequency, 18.48 % in throughput and 21.60 % in area.
Qilong Zhu, Yunfei Dai, Shiyan Bi, Huizhen Kuang, Dylan Wang, Wenbo Yin, Lingli Wang
FPL2
2025 Exponential Tracking Control With Guaranteed Performance for Strict-Feedback Systems Under Deferred Full-State Asymmetric Constraints
abstract
This work provides an exponential tracking control solution for a class of nonlinear systems characterized by unmatched, non-vanishing uncertainties and deferred full-state constraints. The presence of both non-vanishing and unmatched uncertainties complicates the task of achieving exponential tracking rather than regulation, particularly in scenarios involving deferred full-state asymmetric constraints alongside steady-state and transient performance requirements. To address these challenges, several key techniques are employed. Firstly, a timevarying feedback gain technique is utilized to ensure exponential tracking of the strict-feedback system. Secondly, we develop an asymmetric constraint mapping function that integrates the system state, tracking error, a finite-time adjustment function (AF), and an exponential AF to tackle performance issues without violating the deferred full-state asymmetric constraints, even when the initial conditions are unknown. Thirdly, an important lemma (Lemma 2) is derived to guarantee the boundedness of virtual controller derivatives, even as the exponential AF approaches infinity. Additionally, all signals in the closed-loop system are ultimately uniformly bounded. The effectiveness of the proposed scheme is validated through two examples.
Yunfei Dai, Yujuan Wang 0001, Zhuwu Shao, Yongduan Song 0001
IEEE Trans Autom. Sci. Eng.1
2025 Adaptive Asymptotic Exponential Regulation for Nonparametric Uncertain Strict-Feedback Systems With Asymmetric Time-Varying Output Constraints
abstract
This article provides an adaptive asymptotic exponential regulation solution for nonparametric uncertain strict-feedback systems under asymmetric time-varying output constraints. Unlike most existing methods for achieving exponential regulation, where the persistent excitation (PE) condition is in need, here in this work the restriction on the PE condition is removed by introducing a novel parameter estimation error deceleration transformation technique. In addition, more general nonparametric uncertainties but not parametric uncertainties are considered in this article. Furthermore, a new nonlinear state-dependent function (NSDF) is employed to ensure that the asymmetric time-varying output constraint cannot be violated all the time, which also allows the same convergence rate between the system state and transformed function. Additionally, all signals in the closed-loop system are ensured to be bounded. Both theoretical analysis and simulation examples validate the effectiveness of the proposed control algorithm.
Yunfei Dai, Yujuan Wang 0001
IEEE Trans. Syst. Man Cybern. Syst.1