EDBT 2026 Demo / reviewers in the wild / expert
Qilong Zhu
dblp:330/4047
· DBLP profile ↗
5ranked-venue papers
1as first author
5since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Adora Compiler: End-to-End Optimization for High-Efficiency Dataflow Acceleration and Task Pipelining on CGRAsabstractTo fully harness emerging computing architectures, compilers must provide intuitive input handling alongside powerful code optimization to unlock maximum performance. Coarse-Grained Reconfigurable Arrays (CGRAs) — highly energy-efficient for nested-loop applications — have lacked a compiler capable of meeting these objectives. This paper introduces the Adora compiler [1], which effectively bridges user-friendly, lightweight coding inputs with high-performance acceleration on the CGRA SoC. Adora utilizes CGRA-target loop transformations to achieve efficient data-flow level execution while optimizing data communication and task pipelining at the task-flow level. Additionally, it incorporates a comprehensive automated algorithm with a thoughtfully designed optimization sequence. A series of comprehensive experiments highlights the exceptional efficiency and scalability of the Adora compiler, demonstrating its transformative impact in leveraging CGRA capabilities for acceleration in edge computing. Jiahang Lou, Qilong Zhu, Yuan Dai, Zewei Zhong, Wenbo Yin, Lingli Wang |
DAC | 2 |
| 2025 | DEFA: Design Space Exploration for FPGA Overlay Accelerators Through Frequency Prediction and Bayesian OptimizationabstractIn edge AI inference, FPGAs demonstrate superiority in performance-area balance. FPGA Overlay Accelerators (FOAs) are programmable accelerators implemented on FPGAs, typically highly parameterized to enable flexible hardware realization. These parameters, varying across a wide design space, have a significant impact on performance and require efficient Design Space Exploration (DSE). However, current frameworks struggle to accurately predict performance metrics like maximum frequency and fail to fully explore the design space, limiting DSE's effectiveness. In this paper, we propose a DSE framework for FOA (DEFA) based on Bayesian optimization, providing more effective and comprehensive DSE. To address complex parameter interdependencies in FOA, a dependency-aware design space modeling approach (DAM) is proposed. This approach applies fine-grained pruning to the parameter space while addressing dependency constraints. Based on this pruned parameter space, we develop a custom regression predictor (CREP) for maximum frequency using LightGBM, significantly enhancing performance estimation accuracy. Furthermore, the search efficiency is improved through enhanced Latin hypercube sampling and the Tree-Structured Parzen Estimator. We use the proposed framework to optimize an FOA template, Intel FPGA AI Suite. The Pearson correlation coefficient of CREP's predictions regarding the maximum frequency of accelerator instances achieves 0.87. In the throughput optimization experiment, the proposed DSE framework improves 30.16 % compared to the architecture optimization functionality provided by Intel FPGA AI Suite across the given 10 benchmarks on average. In the areathroughput trade-off optimization experiment, compared with FPGA AI Suite, the proposed DSE framework improves 5.01 % in frequency, 18.48 % in throughput and 21.60 % in area. Qilong Zhu, Yunfei Dai, Shiyan Bi, Huizhen Kuang, Dylan Wang, Wenbo Yin, Lingli Wang |
FPL | 1 |
| 2024 | HETA: A Heterogeneous Temporal CGRA Modeling and Design Space Exploration via Bayesian OptimizationabstractDue to its high energy efficiency and flexibility, coarse-grained reconfigurable architecture (CGRA) has gained increasing attention. Temporal CGRA is a typical category of CGRA that supports single-cycle context switching and time-multiplexing hardware resources to perform spatial and temporal computations. Although multiple temporal CGRAs have been proposed, an architecture with rich design parameters and heterogeneous modeling is still lacking. To this end, we propose a highly parameterized heterogeneous temporal CGRA, called HETA. However, the highly parameterized and heterogeneous design introduces a challenging design space for manual exploration. To address this challenge, we introduce a Bayesian-optimization (BO)-based design space exploration (DSE) of homogeneous and heterogeneous architectures. Different from other DSE processes that require defining the heterogeneous exploration strategy, our approach adopts a searching-pruning-based method without manual intervention. To improve the efficiency of DSE, we develop a fast statistic model for area evaluation, whose error is below 1%. In addition, a pipeline mapping (PiPMap) algorithm is developed to alleviate the restrictions caused by data synchronization and unleash the potential of the proposed architecture. Experimental results show that HETA can achieve 89%, 52%, and 47% improvement in throughput, area efficiency, and energy efficiency over the neighbor-to-neighbor (N2N)-based interconnect CGRA, respectively. Compared with the Switch-based interconnect CGRA, HETA’s area efficiency is increased by 61%. Furthermore, compared with the homogeneous architecture of HETA, the optimized heterogeneous architecture improves area efficiency and energy efficiency by 14.7% and 4.8%, respectively. Yuan Dai, Jingyuan Li 0003, Qilong Zhu, Yunhui Qiu, Yihan Hu 0003, Wenbo Yin, Lingli Wang |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2023 | UPTRA: An Ultra-Parameterized Temporal CGRA Modeling and OptimizationabstractTemporal Coarse-Grained Reconfigurable Architecture (CGRA) is a typical category of CGRA that supports single-cycle context switching and time-multiplexing hardware resources to perform both spatial and temporal computations. Compared with the spatial CGRA, it can be used in area and power budget-constrained scenarios, with the sacrifice of the throughput. Therefore, achieving minimum Initialization Interval (II) for higher throughput is the main objective in many works for temporal CGRA mapping. Yuan Dai, Yunhui Qiu, Qilong Zhu, Jingyuan Li 0003, Wenbo Yin, Lingli Wang |
FCCM | 3 |
| 2023 | THRAM: A Template-based Heterogeneous CGRA Modeling Framework Supporting Fast DSEabstractCoarse-grained reconfigurable architecture (CGRA), composed of word-level processing elements (PEs) and interconnects, has emerged as a promising architecture due to its high performance, energy efficiency, and flexibility. Although multiple CGRA frameworks have been proposed, a complete heterogeneous CGRA exploration framework with tunable interconnect flexibility and fast design space exploration (DSE) is still lacking. In this paper, we propose an open-source template-based CGRA exploration framework that integrates the modeling of heterogeneous PEs and interconnects, RTL generation, DFG mapping, automatic simulation and verification, and fast DSE based on a CGRA framework TRAM. Moreover, we present a novel resource-efficient shared reconfigurable delay unit (RDU) for data synchronization, which can save the CGRA area by 7%, compared with the separated RDU. Further, the explored optimal heterogeneous architecture can reduce the area and power by 44.7% and 42.9% respectively, and improve the PE utilization by 20.4%, compared with the 8 × 8 baseline architecture in TRAM. Jingyuan Li 0003, Yunhui Qiu, Guowei Zhu, Qilong Zhu, Wenbo Yin, Lingli Wang |
ISCAS | 4 |