EDBT 2026 Demo / reviewers in the wild / expert
Xianfeng Cao
dblp:323/9968
· DBLP profile ↗
5ranked-venue papers
2as first author
5since 2021 · last 2026
0009-0005-8540-987XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RL-FRA: Exploration and Optimization of FPGA Routing Architectures via Model-Based Reinforcement Learning
Yuanqi Wang, Xianfeng Cao, Lingli Wang |
ISCAS | 2 |
| 2025 | Two-Phase Transistor Sizing for FPGAs via Bayesian OptimizationabstractTransistor-level design is pivotal for the accurate evaluation of FPGA architectures. Since traditional linear models are increasingly inadequate in advanced technology nodes, simulation-driven approaches have become the standard for FPGA architecture exploration. However, due to the non-analytical nature of delay measurements from simulations, the transistor sizing process becomes a black-box optimization problem. COFFE2 [1], [2], the state-of-the-art academic sizing tool employs a division-based brute-force approach, which is relatively time-consuming and may lose optimal solutions. In this paper, we propose a two-phase transistor sizing methodology and enhance the COFFE2 framework with Bayesian Optimization, which is well-suited for black-box optimization problems. Our proposed approach not only achieves a 11.7% improvement in the quality of results but also reduces runtime by ~40%-70%. Through extensive benchmarking on a complete design flow, from transistor-level sizing to routing with VTR benchmarks, we demonstrate that FPGA architectures optimized by our approach offer an 11.0% reduction in the area-delay product, proving the efficacy of our method. Xianfeng Cao, Huizhen Kuang, Yuanqi Wang, Lingli Wang |
FPGA | 1 |
| 2025 | FLAIC: A Novel FPGA Logic Architecture via Fine-Grained Cut Topology AnalysisabstractLook-up table (LUT)-based programmable logic blocks (PLBs) serve as the foundation for FPGAs. As increasing the input number of LUTs to improve logic capacity will introduce exponential area overhead, substantial research has focused on designing more efficient alternatives. Previous approaches primarily design dedicated hardware by analyzing the distribution of Boolean functions and implementing those with high frequency. However, these approaches face scalability challenges due to the explosive growth in the function space. In this paper, we consider the topology of cuts rather than Boolean functions they represent. By identifying topologies that occur commonly in cuts and integrating them with LUTs, we propose a new 8-input PLB architecture, named FLAIC. This architecture incurs only a slight area overhead compared to a 6-LUT while achieving logic capacity comparable to that of an 8LUT. Post-synthesis results demonstrate that FLAIC reduces the logic levels by over 20 % and the number of PLBs by more than$\mathbf{1 0 \%}$, compared to 6-LUTs. Additionally, post-implementation results show improvement in critical path delay by 10.6 % and a reduction in the number of Configurable Logic Blocks (CLBs) by 5.3 % on MCNC and VTR benchmarks, compared to the Intel Stratix 10-like architecture. Xianfeng Cao, Huizhen Kuang, Yuanqi Wang, Lingli Wang |
FPL | 1 |
| 2025 | DynVec: An End-to-End Framework for Efficient Vector-Dataflow ExecutionabstractHigh-performance computing (HPC) and hardware acceleration increasingly rely on dataflow architectures to achieve scalable parallelism and efficiency. High-level synthesis (HLS) facilitates accelerator design from high-level programs, but conventional tools often require intrusive source-level modifications and struggle to optimize irregular workloads. Dynamically scheduled HLS frameworks offer a promising direction for addressing control flow divergence and memory irregularity by generating dataflow accelerators. However, they lack compile-time parallelism optimizations such as vectorization and incur significant hardware overhead. Moreover, modern compilers can generate vectorized code using memory access and computational patterns. Nevertheless, in programs with irregular control flow or data-dependent behavior, such patterns are unknown until runtime, limiting the effectiveness of static vectorization strategies.To address these challenges, we propose DynVec, a unified vector-dataflow framework that integrates dynamic scheduling and vectorization to exploit runtime parallelism beyond conventional models. We address the vectorization of irregular kernels through an MLIR-based context-aware vectorizer that effectively identifies vectorizable operations and, through dataflow scheduling, generates a vector-dataflow execution graph that explicitly models control flow constructs, data and control interfaces, and memory operations. DynVec encapsulates high-level elastic units designed with built-in vectorization support, allowing customizable and adaptive execution behavior. Our compiler preserves the structural hierarchy of the kernel by combining vector and scalar operations in a bottom-up, type-safe manner. Experiments show that our approach achieves significant speedup compared to state-of-the-art HLS implementations across various regular and irregular applications. Moreover, compared to hybrid accelerators that separately support dynamic parallelism and vectorization, DynVec delivers superior performance. Xianfeng Cao, Kaixiang Zhu, Wenbo Yin, Lingli Wang |
ICCAD | 2 |
| 2024 | Label-aware Attention Network with Multi-scale Boosting for Medical Image Segmentation
Linbo Wang 0001, Peng Xu 0045, Xianfeng Cao, Michele Nappi, Shaohua Wan 0001 |
Expert Syst. Appl. | 3 |