VLDB 2026 Research / reviewers in the wild / expert
Jingyuan Li 0003
dblp:28/3576-3
· DBLP profile ↗
10ranked-venue papers
4as first author
10since 2021 · last 2026
0009-0005-3349-6940ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 4 first-author · 10 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Toward Efficient Edge AI With Heterogeneous Computing and Multilevel OptimizationabstractThe rapid progress of artificial intelligence (AI) has brought increasing demands on hardware accelerators, particularly as modern models combine dense linear operations with a growing number of irregular, nonlinear, and control-intensive operators. While tensor cores and systolic arrays offer high throughput for regular computations, they often struggle to efficiently support the diverse operations emerging in recent model structures. Coarse-grained reconfigurable arrays (CGRAs), with their spatial parallelism and reconfigurability, may serve as a natural complement to dense accelerators in such heterogeneous workloads. In this work, we propose EUREKA, a heterogeneous acceleration framework that integrates tensor cores with CGRAs through a unified instruction set, cross-architecture data scheduling, tailored hardware support for nonlinear operators, and optimizations at the instruction, task, and operator levels to exploit parallelism. At the software level, we introduce a hierarchical compilation strategy that combines graph-level optimizations with tensor-level scheduling techniques. To address the large design space of hardware–software co-optimization, we further develop a Bayesian optimization-based exploration scheme enhanced with kernel compression methods, which provides an efficient means of identifying promising hardware configurations and scheduling strategies. Experiment results on representative AI benchmarks show that EUREKA improves execution efficiency, achieving an average$12.6\times $normalized performance gain over state-of-the-art frameworks. Jingyuan Li 0003, Xinyu Cai, Yuan Dai, Wenbo Yin, Lingli Wang |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2025 | LEMOE: LLM-Enhanced Multi-Objective Bayesian Optimization for Microarchitecture ExplorationabstractDesigning processor microarchitectures is increasingly challenging due to a vast design space and the need to balance multiple metrics. Traditional algorithm-driven design space exploration (DSE) approaches often struggle to incorporate the extensive domain knowledge of expert architects. To address this, we introduce LEMOE, a multi-objective microarchitecture optimization framework that leverages large language model (LLM) to enhance an implicit Bayesian model. LEMOE features a program-aware warm-up phase utilizing LLM and LLVM to produce an initial design set with rich prior knowledge. By harnessing LLM’s contextual learning, our approach improves surrogate modeling and sampling under sparse data conditions. Experiment results show that LEMOE achieves a $22.8 \%$ improvement in energy efficiency with the same number of iterations and a $2.9 \times$ runtime speedup for the same target compared to prior works. Jingyuan Li 0003, Jianrong Zhang, Wenbo Yin, Lingli Wang |
DAC | 1 |
| 2025 | COFFA: A Co-Design Framework for Fused-Grained Reconfigurable Architecture Towards Efficient Irregular Loop HandlingabstractCoarse-Grained Reconfigurable Architecture (CGRA) emerges as a competitive accelerator due to its high flexibility and energy efficiency. However, most CGRAs are effective for computation-intensive applications with regular loops but struggle with irregular loops containing control flows. These loops introduce fine-grained logic operations and are costly to execute by coarse-grained arithmetic units in CGRA. Efficiently handling such logic operations necessitates incorporating Boolean algebra optimization, which can improve logic density and reduce logic depth. Unfortunately, no previous research has incorporated it into the compilation flow to support irregular loops efficiently.We proposeCOFFA, an open-source framework for heterogeneous architecture with a RISC-V CPU and a fused-grained reconfigurable accelerator, which integrates coarse-grained arithmetic and fine-grained logic units, along with flexible IO units and distributed interconnects. As a software/hardware co-design framework,COFFAhas a powerful compiler that extracts and optimizes fine-grained logic operations from irregular loops, performs coarse-grained arithmetic and memory optimizations, and offloads the loops to the accelerator.Across various challenging benchmarks with irregular loops,COFFAachieves significant performance and energy efficiency improvements over an in-order, an out-of-order RISC-V CPUs, and a recent FPGA, respectively. Moreover, compared with the state-of-the-art CGRAUE-CGRAandHycube,COFFAcan achieve 2.5× and 3.5× performance gains, respectively. Yuan Dai, Xuchen Gao, Yunhui Qiu, Jingyuan Li 0003, Yuhang Cao, Yiqing Mao, Sichao Chen, Wenbo Yin, Wai-Shing Luk, Lingli Wang |
IEEE Trans. Computers | 4 |
| 2025 | MoDAF: A Multi-objective Divide-and-Conquer Parameter Tuning Framework for CGRAsabstractCoarse-grained reconfigurable architectures (CGRAs) are gaining increasing attention as domain-specific accelerators due to their high flexibility and energy efficiency. These architectures offer a compelling solution for applications that require custom hardware performance while retaining a degree of programmability. However, the design space of CGRAs is inherently vast and complex, presenting significant challenges for architects to explore design choices efficiently and systematically. Existing design space exploration (DSE) methodologies for CGRAs are often time-demanding and struggle to deliver optimal solutions when confronted with high-dimensional and multi-objective design space. Therefore, we consider constructing a CGRA parameter tuning framework called MoDAF. MoDAF initializes the design space using the most representative and diverse samples. It adopts a divide-and-conquer approach, utilizing Monte Carlo Tree Search (MCTS) and space partitioning techniques to dynamically break down the complex design space into more manageable subspaces. A hybrid model handles local fluctuations within each subspace, while a dual sampling algorithm is designed to increase sampling efficiency. MoDAF also incorporates a fast evaluation model to estimate CGRA throughput and area, significantly speeding up the exploration process. Compared with previous approaches, experiments show that our proposed framework reduces the average distance from the reference set by 53.0% and the hypervolume deviation by 64.2%, while also cutting wall time by 57.5%. Jingyuan Li 0003, Yuan Dai, Wenbo Yin, Lingli Wang |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2024 | MDCRA: A Reconfigurable Accelerator Framework for Multiple Dataflow LanesabstractCoarse-grained reconfigurable architecture (CGRA) is a type of reconfigurable computing architecture suitable for emerging applications that require dynamic compilation hardware. However, the resource utilization of existing CGRA is low due to the lack of flexibility across varied application granularity. In this paper, we propose a CGRA framework for multiple dataflow lanes (MDCRA). It supports post-silicon computational granularity adjustments. Evaluated with Polybench, Machsuite and Express, the speedup of MDCRA is$24.83\times$higher than CPU CVA6, and$2.08\times$higher than vector processor Ara. Compared with TRAM and DSAGEN, MDCRA achieves an area reduction of 27% and 47% respectively with the same speedup. Besides, compared with OpenCGRA, the average utilization of function units is improved by 20.05%. Shaoyang Sun, Boyin Jin, Jiahang Lou, Yuhang Cao, Jingyuan Li 0003, Yuan Dai, Wenbo Yin, Wai-Shing Luk, Lingli Wang |
ASAP | 6 |
| 2024 | HierCGRA: A Novel Framework for Large-scale CGRA with Hierarchical Modeling and Automated Design Space ExplorationabstractCoarse-grained reconfigurable arrays (CGRAs) are promising design choices in computation-intensive domains, since they can strike a balance between energy efficiency and flexibility. A typical CGRA comprises processing elements (PEs) that can execute operations in applications and interconnections between them. Nevertheless, most CGRAs suffer from the ineffectiveness of supporting flexible architecture design and solving large-scale mapping problems. To address these challenges, we introduce HierCGRA, a novel framework that integrates hierarchical CGRA modeling, Chisel-based Verilog generation, LLVM-based data flow graph (DFG) generation, DFG mapping, and design space exploration (DSE). With the graph homomorphism (GH) mapping algorithm, HierCGRA achieves a faster mapping speed and higher PE utilization rate compared with the existing state-of-the-art CGRA frameworks. The proposed hierarchical mapping strategy achieves 41× speedup on average compared with the ILP mapping algorithm in CGRA-ME. Furthermore, the automated DSE based on Bayesian optimization achieves a significant performance improvement by the heterogeneity of PEs and interconnections. With these features, HierCGRA enables the agile development for large-scale CGRA and accelerates the process of finding a better CGRA architecture. Sichao Chen, Su Zheng, Guowei Zhu, Jingyuan Li 0003, Yazhou Yan, Yuan Dai, Wenbo Yin, Lingli Wang |
ACM Trans. Reconfigurable Technol. Syst. | 6 |
| 2024 | HETA: A Heterogeneous Temporal CGRA Modeling and Design Space Exploration via Bayesian OptimizationabstractDue to its high energy efficiency and flexibility, coarse-grained reconfigurable architecture (CGRA) has gained increasing attention. Temporal CGRA is a typical category of CGRA that supports single-cycle context switching and time-multiplexing hardware resources to perform spatial and temporal computations. Although multiple temporal CGRAs have been proposed, an architecture with rich design parameters and heterogeneous modeling is still lacking. To this end, we propose a highly parameterized heterogeneous temporal CGRA, called HETA. However, the highly parameterized and heterogeneous design introduces a challenging design space for manual exploration. To address this challenge, we introduce a Bayesian-optimization (BO)-based design space exploration (DSE) of homogeneous and heterogeneous architectures. Different from other DSE processes that require defining the heterogeneous exploration strategy, our approach adopts a searching-pruning-based method without manual intervention. To improve the efficiency of DSE, we develop a fast statistic model for area evaluation, whose error is below 1%. In addition, a pipeline mapping (PiPMap) algorithm is developed to alleviate the restrictions caused by data synchronization and unleash the potential of the proposed architecture. Experimental results show that HETA can achieve 89%, 52%, and 47% improvement in throughput, area efficiency, and energy efficiency over the neighbor-to-neighbor (N2N)-based interconnect CGRA, respectively. Compared with the Switch-based interconnect CGRA, HETA’s area efficiency is increased by 61%. Furthermore, compared with the homogeneous architecture of HETA, the optimized heterogeneous architecture improves area efficiency and energy efficiency by 14.7% and 4.8%, respectively. Yuan Dai, Jingyuan Li 0003, Qilong Zhu, Yunhui Qiu, Yihan Hu 0003, Wenbo Yin, Lingli Wang |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2023 | UPTRA: An Ultra-Parameterized Temporal CGRA Modeling and OptimizationabstractTemporal Coarse-Grained Reconfigurable Architecture (CGRA) is a typical category of CGRA that supports single-cycle context switching and time-multiplexing hardware resources to perform both spatial and temporal computations. Compared with the spatial CGRA, it can be used in area and power budget-constrained scenarios, with the sacrifice of the throughput. Therefore, achieving minimum Initialization Interval (II) for higher throughput is the main objective in many works for temporal CGRA mapping. Yuan Dai, Yunhui Qiu, Qilong Zhu, Jingyuan Li 0003, Wenbo Yin, Lingli Wang |
FCCM | 4 |
| 2023 | PRAD: A Bayesian Optimization-based DSE Framework for Parameterized Reconfigurable Architecture DesignabstractCoarse-Grained Reconfigurable Architecture (CGRA) is a domain-specific reconfigurable architecture. Generally, the CGRA architecture consists of IO, memory, coarse-grained processing element (PE), and interconnect. Usually, ALU in PE contains a relatively complete set of operations and most of the interconnects adopt neighbor-to-neighbor (N2N) [1], switch-based [2], and combination of the connection box and switch box (CB-SB) patterns [3]. However, the complex operation sets and switch-based/CB-SB fully-connected interconnects provide sufficient reconfigurability at the cost of resource overhead. Thus, it is important to build a parameterized architecture of CGRA to achieve a balance among hardware overhead, flexibility and performance through automatic design space exploration (DSE). Bingbing Peng, Shaoyang Sun, Yuan Dai, Jingyuan Li 0003, Yunhui Qiu, Kaihang Wang, Wenbo Yin, Lingli Wang |
FCCM | 4 |
| 2023 | THRAM: A Template-based Heterogeneous CGRA Modeling Framework Supporting Fast DSEabstractCoarse-grained reconfigurable architecture (CGRA), composed of word-level processing elements (PEs) and interconnects, has emerged as a promising architecture due to its high performance, energy efficiency, and flexibility. Although multiple CGRA frameworks have been proposed, a complete heterogeneous CGRA exploration framework with tunable interconnect flexibility and fast design space exploration (DSE) is still lacking. In this paper, we propose an open-source template-based CGRA exploration framework that integrates the modeling of heterogeneous PEs and interconnects, RTL generation, DFG mapping, automatic simulation and verification, and fast DSE based on a CGRA framework TRAM. Moreover, we present a novel resource-efficient shared reconfigurable delay unit (RDU) for data synchronization, which can save the CGRA area by 7%, compared with the separated RDU. Further, the explored optimal heterogeneous architecture can reduce the area and power by 44.7% and 42.9% respectively, and improve the PE utilization by 20.4%, compared with the 8 × 8 baseline architecture in TRAM. Jingyuan Li 0003, Yunhui Qiu, Guowei Zhu, Qilong Zhu, Wenbo Yin, Lingli Wang |
ISCAS | 1 |