Xuchen Gao

dblp:369/3282 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 1 first-author · 7 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Target detection and SINR, CRB analysis for bistatic coherent FDA radar based on multichannel parallel ADMF receiving structure
Xuchen Gao, Junwei Xie 0001, Chuan Sheng, Jingwei Xu 0002, Haowei Zhang 0001
Signal Process.1
2026 Dependency-Aware Data Parallelism on Spatial CGRA via Constraint Satisfaction and Graph Coloring
abstract
Coarse-grained Reconfigurable Architecture (CGRA) is a competitive accelerator architecture for computation-intensive loop kernels. Spatial CGRA is a typical CGRA that performs all the operations spatially to reduce reconfiguration costs within a single iteration, demanding high data parallelism. To achieve this goal, one of the main challenges is the loop-carried dependency between memory accesses. Many existing CGRA compilers struggle to precisely analyze the dependency distance, especially when accesses involve complex address patterns. Consequently, these compilers often default to setting the distance to one, based on a worst-case assumption, leading to degraded performance. However, we observe that a precise distance can improve performance significantly, raising the requirement for an efficient distance calculation approach. Another challenge is the performance constraints of single-bank memory, which necessitate the designer partitioning the original data into a multi-bank memory. However, we observe that the mapping result can cause the inter-iteration conflict, thereby invalidating the memory partition scheme. Therefore, an efficient post-mapping conflict detection is required. In this paper, we develop a constraint satisfaction problem (CSP)-based approach for calculating dependency distance and detecting conflicts, which determines the maximum available dependency distance and identifies conflicts within both intra- and inter-iterations. Besides, we formulate access scheduling as a graph coloring problem, which can minimize conflicts and improve performance. Overall, we develop a comprehensive end-to-end framework with architectural and compiler support for efficient data parallelism on spatial CGRA. We conduct extensive experiments to systematically evaluate the impact of different approaches on performance and compilation. Evaluation results show that our architecture can achieve 13.16× and 1.19× (up to 1.68×) average performance improvements compared to a RISC-V CPU and a state-of-the-art CGRA SoC, respectively. Besides, our architecture has 7.38× and 1.18× (up to 1.65×) average energy efficiency gains compared to these two architectures.
Yuan Dai, Xuchen Gao, Wenbo Yin, Wai-Shing Luk, Lingli Wang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2025 Towards Efficient Data Parallelism on Spatial CGRA via Constraint Satisfaction and Graph Coloring
abstract
Coarse-Grained Reconfigurable Architecture (CGRA) is a competitive accelerator architecture for computation-intensive loop kernels. Spatial CGRA is a typical CGRA that performs all the operations spatially, demanding high data parallelism. Given the performance limitations of single-bank memory, partitioning original data into multi-bank memory within the spatial CGRA is favored. However, we observe that the mapping result can cause the inter-iteration conflict, thereby invalidating the memory partition scheme.
Yuan Dai, Xuchen Gao, Bingbing Peng, Wenbo Yin, Wai-Shing Luk, Lingli Wang
ASP-DAC2
2025 COFFA: A Co-Design Framework for Fused-Grained Reconfigurable Architecture Towards Efficient Irregular Loop Handling
abstract
Coarse-Grained Reconfigurable Architecture (CGRA) emerges as a competitive accelerator due to its high flexibility and energy efficiency. However, most CGRAs are effective for computation-intensive applications with regular loops but struggle with irregular loops containing control flows. These loops introduce fine-grained logic operations and are costly to execute by coarse-grained arithmetic units in CGRA. Efficiently handling such logic operations necessitates incorporating Boolean algebra optimization, which can improve logic density and reduce logic depth. Unfortunately, no previous research has incorporated it into the compilation flow to support irregular loops efficiently.We proposeCOFFA, an open-source framework for heterogeneous architecture with a RISC-V CPU and a fused-grained reconfigurable accelerator, which integrates coarse-grained arithmetic and fine-grained logic units, along with flexible IO units and distributed interconnects. As a software/hardware co-design framework,COFFAhas a powerful compiler that extracts and optimizes fine-grained logic operations from irregular loops, performs coarse-grained arithmetic and memory optimizations, and offloads the loops to the accelerator.Across various challenging benchmarks with irregular loops,COFFAachieves significant performance and energy efficiency improvements over an in-order, an out-of-order RISC-V CPUs, and a recent FPGA, respectively. Moreover, compared with the state-of-the-art CGRAUE-CGRAandHycube,COFFAcan achieve 2.5× and 3.5× performance gains, respectively.
Yuan Dai, Xuchen Gao, Yunhui Qiu, Jingyuan Li 0003, Yuhang Cao, Yiqing Mao, Sichao Chen, Wenbo Yin, Wai-Shing Luk, Lingli Wang
IEEE Trans. Computers2
2024 A CGRA Front-end Compiler Enabling Extraction of General Control and Dedicated Operators
abstract
Coarse-grained reconfigurable architecture (CGRA) gradually becomes an extraordinarily promising accelerator due to its flexibility and power efficiency. However, most CGRA front-end compilers focus on the innermost body of regular loops with a pure data flow. Therefore, we propose CO-Compiler, an LLVM-based CGRA front-end compiler to generate an optimized control-data flow graph (CDFG), which can handle versatile loops in C/C++, including general control flow, arbitrary nested levels, and imperfect statements. Then we extract multi-dimension memory access patterns and various dedicated operators adapting to concrete hardware functions. In addition, we analyze variable loop bounds which are settled at runtime, and realize the SoC runtime configuration of CGRA. The feasibility of our methodology is verified by a RISC-V based SoC simulation. The experimental results demonstrate that our dedicated operator extraction can reduce 43% PE resources and decrease 84% initiation interval (II) on a TRAM architecture. Furthermore, compared with state-of-the-art (SOTA) CGRA front-end compilers, CO-Compiler has the highest 88.1% success rate in CDFG generation for a wide range of benchmarks. Moreover, by using the same back-end mappers, our work can reach 78% reduction for II and $2.06\times$ PE spatio-temporal utilization in contrast with their own front-end compilers.
Xuchen Gao, Yunhui Qiu, Yuan Dai, Wenbo Yin, Lingli Wang
ASPDAC1
2024 An Agile Deploying Approach for Large-Scale Workloads on CGRA-CPU Architecture
abstract
Adopting specialized accelerators such as Coarse-Grained Reconfigurable Architectures (CGRAs) alongside CPUs to enhance performance within specific domains is an astute choice. However, the integration of heterogeneous architectures introduces complex challenges for compiler design. Simultaneously, the ever-expanding scale of workloads imposes substantial burdens on deployment. To address above challenges, this paper introduces CGRV-OPT, a user-friendly multi-level compiler designed to deploy large-scale workloads to CGRA and RISC-V CPU architecture. Built upon the MLIR framework, CGRV-OPT serves as a pivotal bridge, facilitating the seamless conversion of high-level workload descriptions into low-level intermediate representations (IRs) for different architectures. A salient feature of our approach is the automation of a comprehensive suite of optimizations and transformations, which speed up each kernel computing within the intricate SoC. Additionally, we have seamlessly integrated an automated software-hardware partitioning mechanism, guided by our multi-level optimizations, resulting in a remarkable 2.14 × speed up over large-scale workloads. The CGRV-OPT framework significantly alleviates the challenges faced by software developers, including those with limited expertise in hardware architectures.
Jiahang Lou, Xuchen Gao, Yiqing Mao, Yunhui Qiu, Yihan Hu 0003, Wenbo Yin, Lingli Wang
DATE2
2024 CFEACT: A CGRA-based Framework Enabling Agile CNN and Transformer Accelerator Design
abstract
Convolutional neural networks (CNNs) and transformer neural networks have been adopted in a wide range of applications such as natural language processing and computer vision. Coarse-grained reconfigurable architectures (CGRAs) are highly suitable for CNN and transformer applications due to their high flexibility and energy efficiency. However, current implementations of CGRA for CNNs and transformers have several limitations including the lack of System-on-Chip (SoC), insufficient support for nonlinear functions and the absence of a software toolchain. To address these challenges, we present CFEACT, a CGRA-based framework that enables agile development of CNN and transformer accelerators. CFEACT offers a broad design space of efficient CGRA accelerators through a highly flexible architecture template. The well-designed SoC, innovative mapping schemes, and comprehensive software toolchain offer a complete solution for implementing various CNN and transformer models on the generated CGRAs. Compared with the state-of-the-art works, accelerators generated by CFEACT can achieve more than $2 \times$ improvement in area-delay product for CNNs and an average of $2 \times$ higher performance for transformers.
Yiqing Mao, Xuchen Gao, Jiahang Lou, Yunhui Qiu, Wenbo Yin, Wai-Shing Luk, Lingli Wang
FPL2
2024 FDRA: A Framework for a Dynamically Reconfigurable Accelerator Supporting Multi-Level Parallelism
abstract
Coarse-grained reconfigurable architectures (CGRAs) have emerged as promising accelerators due to their high flexibility and energy efficiency. However, existing open source works often lack integration of CGRAs with CPU systems and corresponding toolchains. Moreover, there is rare support for the accelerator instruction pipelining to overlap data communication, computation, and configuration across multiple tasks. In this article, we propose FDRA, an open source exploration framework for a heterogeneous system-on-chip (SoC) with a RISC-V processor and a dynamically reconfigurable accelerator (DRA) supporting loop, instruction, and task levels of parallelism. FDRA encompasses parameterized SoC modeling, Verilog generation, source-to-source application code transformation using frontend and DRA compilers, SoC simulation, and FPGA prototyping. FDRA incorporates the extraction of periodic accumulative operators and multi-dimensional linear load/store operators from nested loops. The DRA enables accessing the shared L2 cache with virtual addresses and supports direct memory access with arbitrary start addresses and data lengths. Integrated into the RISC-V Rocket SoC, our DRA achieves a remarkable 55× acceleration for loop kernels and improves energy efficiency by 29×. Compared to state-of-the-art RISC-V vector units, our DRA demonstrates a 2.9× speed improvement and 3.5× greater energy efficiency. In contrast to previous CGRA+RISC-V SoCs, our SoC achieves a minimum speedup of 5.2×.
Yunhui Qiu, Yiqing Mao, Xuchen Gao, Sichao Chen, Wenbo Yin, Lingli Wang
ACM Trans. Reconfigurable Technol. Syst.3