Yiqing Mao

dblp:135/3576 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
6since 2021 · last 2026
0009-0005-9167-2279ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 2 first-author · 6 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 UltraMalloc: Efficient FPGA-based Memory Allocation Framework Optimized for HBM
abstract
Memory allocation efficiency remains a significant challenge in High-Level Synthesis (HLS) frameworks. Current dynamic memory management (DMM) techniques suffer from issues such as inefficiency, fragmentation, considerable hardware overhead, and difficulties in handling complex workloads. Conversely, existing static memory approaches often exhibit poor efficiency when addressing large-scale applications. Moreover, both dynamic and static methods lack sufficient support for High Bandwidth Memory (HBM), thereby limiting their effectiveness in complex neural network scenarios. To overcome these challenges, we propose an optimized static memory allocation strategy specifically designed for FPGA systems. Our approach leverages the computational characteristic of neural network applications, which typically exhibit cyclic and fixed-bound behaviors. We tightly integrate an MLIR-based compiler with an efficient static memory allocator, eliminating the need for dedicated allocator hardware while ensuring efficient runtime memory access. Furthermore, we introduce a customized AXI bus distribution mechanism and an address mapping strategy optimized for the multi-port and multi-bank architecture of HBM. This design significantly enhances bandwidth utilization and reduces latency. Experimental results confirm that our proposed methodology substantially improves allocation efficiency, spatial utilization, and effectively manages complex memory scenarios, thereby outperforming existing state-of-the-art solutions.
Yuwei Qu, Yiqing Mao, Yanxing Jin, Wai-Shing Luk, Lingli Wang
ASP-DAC2
2025 COFFA: A Co-Design Framework for Fused-Grained Reconfigurable Architecture Towards Efficient Irregular Loop Handling
abstract
Coarse-Grained Reconfigurable Architecture (CGRA) emerges as a competitive accelerator due to its high flexibility and energy efficiency. However, most CGRAs are effective for computation-intensive applications with regular loops but struggle with irregular loops containing control flows. These loops introduce fine-grained logic operations and are costly to execute by coarse-grained arithmetic units in CGRA. Efficiently handling such logic operations necessitates incorporating Boolean algebra optimization, which can improve logic density and reduce logic depth. Unfortunately, no previous research has incorporated it into the compilation flow to support irregular loops efficiently.We proposeCOFFA, an open-source framework for heterogeneous architecture with a RISC-V CPU and a fused-grained reconfigurable accelerator, which integrates coarse-grained arithmetic and fine-grained logic units, along with flexible IO units and distributed interconnects. As a software/hardware co-design framework,COFFAhas a powerful compiler that extracts and optimizes fine-grained logic operations from irregular loops, performs coarse-grained arithmetic and memory optimizations, and offloads the loops to the accelerator.Across various challenging benchmarks with irregular loops,COFFAachieves significant performance and energy efficiency improvements over an in-order, an out-of-order RISC-V CPUs, and a recent FPGA, respectively. Moreover, compared with the state-of-the-art CGRAUE-CGRAandHycube,COFFAcan achieve 2.5× and 3.5× performance gains, respectively.
Yuan Dai, Xuchen Gao, Yunhui Qiu, Jingyuan Li 0003, Yuhang Cao, Yiqing Mao, Sichao Chen, Wenbo Yin, Wai-Shing Luk, Lingli Wang
IEEE Trans. Computers6
2024 An Agile Deploying Approach for Large-Scale Workloads on CGRA-CPU Architecture
abstract
Adopting specialized accelerators such as Coarse-Grained Reconfigurable Architectures (CGRAs) alongside CPUs to enhance performance within specific domains is an astute choice. However, the integration of heterogeneous architectures introduces complex challenges for compiler design. Simultaneously, the ever-expanding scale of workloads imposes substantial burdens on deployment. To address above challenges, this paper introduces CGRV-OPT, a user-friendly multi-level compiler designed to deploy large-scale workloads to CGRA and RISC-V CPU architecture. Built upon the MLIR framework, CGRV-OPT serves as a pivotal bridge, facilitating the seamless conversion of high-level workload descriptions into low-level intermediate representations (IRs) for different architectures. A salient feature of our approach is the automation of a comprehensive suite of optimizations and transformations, which speed up each kernel computing within the intricate SoC. Additionally, we have seamlessly integrated an automated software-hardware partitioning mechanism, guided by our multi-level optimizations, resulting in a remarkable 2.14 × speed up over large-scale workloads. The CGRV-OPT framework significantly alleviates the challenges faced by software developers, including those with limited expertise in hardware architectures.
Jiahang Lou, Xuchen Gao, Yiqing Mao, Yunhui Qiu, Yihan Hu 0003, Wenbo Yin, Lingli Wang
DATE3
2024 CFEACT: A CGRA-based Framework Enabling Agile CNN and Transformer Accelerator Design
abstract
Convolutional neural networks (CNNs) and transformer neural networks have been adopted in a wide range of applications such as natural language processing and computer vision. Coarse-grained reconfigurable architectures (CGRAs) are highly suitable for CNN and transformer applications due to their high flexibility and energy efficiency. However, current implementations of CGRA for CNNs and transformers have several limitations including the lack of System-on-Chip (SoC), insufficient support for nonlinear functions and the absence of a software toolchain. To address these challenges, we present CFEACT, a CGRA-based framework that enables agile development of CNN and transformer accelerators. CFEACT offers a broad design space of efficient CGRA accelerators through a highly flexible architecture template. The well-designed SoC, innovative mapping schemes, and comprehensive software toolchain offer a complete solution for implementing various CNN and transformer models on the generated CGRAs. Compared with the state-of-the-art works, accelerators generated by CFEACT can achieve more than $2 \times$ improvement in area-delay product for CNNs and an average of $2 \times$ higher performance for transformers.
Yiqing Mao, Xuchen Gao, Jiahang Lou, Yunhui Qiu, Wenbo Yin, Wai-Shing Luk, Lingli Wang
FPL1
2024 FDRA: A Framework for a Dynamically Reconfigurable Accelerator Supporting Multi-Level Parallelism
abstract
Coarse-grained reconfigurable architectures (CGRAs) have emerged as promising accelerators due to their high flexibility and energy efficiency. However, existing open source works often lack integration of CGRAs with CPU systems and corresponding toolchains. Moreover, there is rare support for the accelerator instruction pipelining to overlap data communication, computation, and configuration across multiple tasks. In this article, we propose FDRA, an open source exploration framework for a heterogeneous system-on-chip (SoC) with a RISC-V processor and a dynamically reconfigurable accelerator (DRA) supporting loop, instruction, and task levels of parallelism. FDRA encompasses parameterized SoC modeling, Verilog generation, source-to-source application code transformation using frontend and DRA compilers, SoC simulation, and FPGA prototyping. FDRA incorporates the extraction of periodic accumulative operators and multi-dimensional linear load/store operators from nested loops. The DRA enables accessing the shared L2 cache with virtual addresses and supports direct memory access with arbitrary start addresses and data lengths. Integrated into the RISC-V Rocket SoC, our DRA achieves a remarkable 55× acceleration for loop kernels and improves energy efficiency by 29×. Compared to state-of-the-art RISC-V vector units, our DRA demonstrates a 2.9× speed improvement and 3.5× greater energy efficiency. In contrast to previous CGRA+RISC-V SoCs, our SoC achieves a minimum speedup of 5.2×.
Yunhui Qiu, Yiqing Mao, Xuchen Gao, Sichao Chen, Wenbo Yin, Lingli Wang
ACM Trans. Reconfigurable Technol. Syst.2
2022 Safety Assurance System for Electric Vehicles Based on Infrared LiDAR
abstract
With the development of electric vehicles, more and more people choose electric vehicles as their means of transportation. Microwave radar and terahertz radar can be used to build safety assurance system for electric vehicles. However, both of the two technologies have the disadvantage of low stability. Infrared LiDAR has the advantages of high resolution, low cost and high reliability, but it has not yet been widely applied in the safety assurance system for electric vehicles. In this paper, we propose a safety assurance system for electric vehicles based on infrared LiDAR. In the design, infrared sensors are used for obstacle detection, and an MCU (microprogrammed control unit) is used for control. The excellent performance of the electric vehicle equipped with the safety assurance system in avoiding obstacles proves the high efficiency and high reliability of the system proposed in this paper. This experiment can deepen students’ understanding of the circuit system and cultivate their ability to construct the system independently, which can be of important meaning in their education.
Yiqing Mao, Yun Chen 0001
ISCAS1