Yunhui Qiu

dblp:243/5124 · DBLP profile ↗
← Back
15ranked-venue papers
6as first author
12since 2021 · last 2025
0000-0003-2754-6655ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 15 · 6 first-author · 12 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2025 COFFA: A Co-Design Framework for Fused-Grained Reconfigurable Architecture Towards Efficient Irregular Loop Handling
abstract
Coarse-Grained Reconfigurable Architecture (CGRA) emerges as a competitive accelerator due to its high flexibility and energy efficiency. However, most CGRAs are effective for computation-intensive applications with regular loops but struggle with irregular loops containing control flows. These loops introduce fine-grained logic operations and are costly to execute by coarse-grained arithmetic units in CGRA. Efficiently handling such logic operations necessitates incorporating Boolean algebra optimization, which can improve logic density and reduce logic depth. Unfortunately, no previous research has incorporated it into the compilation flow to support irregular loops efficiently.We proposeCOFFA, an open-source framework for heterogeneous architecture with a RISC-V CPU and a fused-grained reconfigurable accelerator, which integrates coarse-grained arithmetic and fine-grained logic units, along with flexible IO units and distributed interconnects. As a software/hardware co-design framework,COFFAhas a powerful compiler that extracts and optimizes fine-grained logic operations from irregular loops, performs coarse-grained arithmetic and memory optimizations, and offloads the loops to the accelerator.Across various challenging benchmarks with irregular loops,COFFAachieves significant performance and energy efficiency improvements over an in-order, an out-of-order RISC-V CPUs, and a recent FPGA, respectively. Moreover, compared with the state-of-the-art CGRAUE-CGRAandHycube,COFFAcan achieve 2.5× and 3.5× performance gains, respectively.
Yuan Dai, Xuchen Gao, Yunhui Qiu, Jingyuan Li 0003, Yuhang Cao, Yiqing Mao, Sichao Chen, Wenbo Yin, Wai-Shing Luk, Lingli Wang
IEEE Trans. Computers3
2024 A CGRA Front-end Compiler Enabling Extraction of General Control and Dedicated Operators
abstract
Coarse-grained reconfigurable architecture (CGRA) gradually becomes an extraordinarily promising accelerator due to its flexibility and power efficiency. However, most CGRA front-end compilers focus on the innermost body of regular loops with a pure data flow. Therefore, we propose CO-Compiler, an LLVM-based CGRA front-end compiler to generate an optimized control-data flow graph (CDFG), which can handle versatile loops in C/C++, including general control flow, arbitrary nested levels, and imperfect statements. Then we extract multi-dimension memory access patterns and various dedicated operators adapting to concrete hardware functions. In addition, we analyze variable loop bounds which are settled at runtime, and realize the SoC runtime configuration of CGRA. The feasibility of our methodology is verified by a RISC-V based SoC simulation. The experimental results demonstrate that our dedicated operator extraction can reduce 43% PE resources and decrease 84% initiation interval (II) on a TRAM architecture. Furthermore, compared with state-of-the-art (SOTA) CGRA front-end compilers, CO-Compiler has the highest 88.1% success rate in CDFG generation for a wide range of benchmarks. Moreover, by using the same back-end mappers, our work can reach 78% reduction for II and $2.06\times$ PE spatio-temporal utilization in contrast with their own front-end compilers.
Xuchen Gao, Yunhui Qiu, Yuan Dai, Wenbo Yin, Lingli Wang
ASPDAC2
2024 An Agile Deploying Approach for Large-Scale Workloads on CGRA-CPU Architecture
abstract
Adopting specialized accelerators such as Coarse-Grained Reconfigurable Architectures (CGRAs) alongside CPUs to enhance performance within specific domains is an astute choice. However, the integration of heterogeneous architectures introduces complex challenges for compiler design. Simultaneously, the ever-expanding scale of workloads imposes substantial burdens on deployment. To address above challenges, this paper introduces CGRV-OPT, a user-friendly multi-level compiler designed to deploy large-scale workloads to CGRA and RISC-V CPU architecture. Built upon the MLIR framework, CGRV-OPT serves as a pivotal bridge, facilitating the seamless conversion of high-level workload descriptions into low-level intermediate representations (IRs) for different architectures. A salient feature of our approach is the automation of a comprehensive suite of optimizations and transformations, which speed up each kernel computing within the intricate SoC. Additionally, we have seamlessly integrated an automated software-hardware partitioning mechanism, guided by our multi-level optimizations, resulting in a remarkable 2.14 × speed up over large-scale workloads. The CGRV-OPT framework significantly alleviates the challenges faced by software developers, including those with limited expertise in hardware architectures.
Jiahang Lou, Xuchen Gao, Yiqing Mao, Yunhui Qiu, Yihan Hu 0003, Wenbo Yin, Lingli Wang
DATE4
2024 CFEACT: A CGRA-based Framework Enabling Agile CNN and Transformer Accelerator Design
abstract
Convolutional neural networks (CNNs) and transformer neural networks have been adopted in a wide range of applications such as natural language processing and computer vision. Coarse-grained reconfigurable architectures (CGRAs) are highly suitable for CNN and transformer applications due to their high flexibility and energy efficiency. However, current implementations of CGRA for CNNs and transformers have several limitations including the lack of System-on-Chip (SoC), insufficient support for nonlinear functions and the absence of a software toolchain. To address these challenges, we present CFEACT, a CGRA-based framework that enables agile development of CNN and transformer accelerators. CFEACT offers a broad design space of efficient CGRA accelerators through a highly flexible architecture template. The well-designed SoC, innovative mapping schemes, and comprehensive software toolchain offer a complete solution for implementing various CNN and transformer models on the generated CGRAs. Compared with the state-of-the-art works, accelerators generated by CFEACT can achieve more than $2 \times$ improvement in area-delay product for CNNs and an average of $2 \times$ higher performance for transformers.
Yiqing Mao, Xuchen Gao, Jiahang Lou, Yunhui Qiu, Wenbo Yin, Wai-Shing Luk, Lingli Wang
FPL4
2024 FDRA: A Framework for a Dynamically Reconfigurable Accelerator Supporting Multi-Level Parallelism
abstract
Coarse-grained reconfigurable architectures (CGRAs) have emerged as promising accelerators due to their high flexibility and energy efficiency. However, existing open source works often lack integration of CGRAs with CPU systems and corresponding toolchains. Moreover, there is rare support for the accelerator instruction pipelining to overlap data communication, computation, and configuration across multiple tasks. In this article, we propose FDRA, an open source exploration framework for a heterogeneous system-on-chip (SoC) with a RISC-V processor and a dynamically reconfigurable accelerator (DRA) supporting loop, instruction, and task levels of parallelism. FDRA encompasses parameterized SoC modeling, Verilog generation, source-to-source application code transformation using frontend and DRA compilers, SoC simulation, and FPGA prototyping. FDRA incorporates the extraction of periodic accumulative operators and multi-dimensional linear load/store operators from nested loops. The DRA enables accessing the shared L2 cache with virtual addresses and supports direct memory access with arbitrary start addresses and data lengths. Integrated into the RISC-V Rocket SoC, our DRA achieves a remarkable 55× acceleration for loop kernels and improves energy efficiency by 29×. Compared to state-of-the-art RISC-V vector units, our DRA demonstrates a 2.9× speed improvement and 3.5× greater energy efficiency. In contrast to previous CGRA+RISC-V SoCs, our SoC achieves a minimum speedup of 5.2×.
Yunhui Qiu, Yiqing Mao, Xuchen Gao, Sichao Chen, Wenbo Yin, Lingli Wang
ACM Trans. Reconfigurable Technol. Syst.1
2024 HETA: A Heterogeneous Temporal CGRA Modeling and Design Space Exploration via Bayesian Optimization
abstract
Due to its high energy efficiency and flexibility, coarse-grained reconfigurable architecture (CGRA) has gained increasing attention. Temporal CGRA is a typical category of CGRA that supports single-cycle context switching and time-multiplexing hardware resources to perform spatial and temporal computations. Although multiple temporal CGRAs have been proposed, an architecture with rich design parameters and heterogeneous modeling is still lacking. To this end, we propose a highly parameterized heterogeneous temporal CGRA, called HETA. However, the highly parameterized and heterogeneous design introduces a challenging design space for manual exploration. To address this challenge, we introduce a Bayesian-optimization (BO)-based design space exploration (DSE) of homogeneous and heterogeneous architectures. Different from other DSE processes that require defining the heterogeneous exploration strategy, our approach adopts a searching-pruning-based method without manual intervention. To improve the efficiency of DSE, we develop a fast statistic model for area evaluation, whose error is below 1%. In addition, a pipeline mapping (PiPMap) algorithm is developed to alleviate the restrictions caused by data synchronization and unleash the potential of the proposed architecture. Experimental results show that HETA can achieve 89%, 52%, and 47% improvement in throughput, area efficiency, and energy efficiency over the neighbor-to-neighbor (N2N)-based interconnect CGRA, respectively. Compared with the Switch-based interconnect CGRA, HETA’s area efficiency is increased by 61%. Furthermore, compared with the homogeneous architecture of HETA, the optimized heterogeneous architecture improves area efficiency and energy efficiency by 14.7% and 4.8%, respectively.
Yuan Dai, Jingyuan Li 0003, Qilong Zhu, Yunhui Qiu, Yihan Hu 0003, Wenbo Yin, Lingli Wang
IEEE Trans. Very Large Scale Integr. Syst.4
2023 UPTRA: An Ultra-Parameterized Temporal CGRA Modeling and Optimization
abstract
Temporal Coarse-Grained Reconfigurable Architecture (CGRA) is a typical category of CGRA that supports single-cycle context switching and time-multiplexing hardware resources to perform both spatial and temporal computations. Compared with the spatial CGRA, it can be used in area and power budget-constrained scenarios, with the sacrifice of the throughput. Therefore, achieving minimum Initialization Interval (II) for higher throughput is the main objective in many works for temporal CGRA mapping.
Yuan Dai, Yunhui Qiu, Qilong Zhu, Jingyuan Li 0003, Wenbo Yin, Lingli Wang
FCCM2
2023 PRAD: A Bayesian Optimization-based DSE Framework for Parameterized Reconfigurable Architecture Design
abstract
Coarse-Grained Reconfigurable Architecture (CGRA) is a domain-specific reconfigurable architecture. Generally, the CGRA architecture consists of IO, memory, coarse-grained processing element (PE), and interconnect. Usually, ALU in PE contains a relatively complete set of operations and most of the interconnects adopt neighbor-to-neighbor (N2N) [1], switch-based [2], and combination of the connection box and switch box (CB-SB) patterns [3]. However, the complex operation sets and switch-based/CB-SB fully-connected interconnects provide sufficient reconfigurability at the cost of resource overhead. Thus, it is important to build a parameterized architecture of CGRA to achieve a balance among hardware overhead, flexibility and performance through automatic design space exploration (DSE).
Bingbing Peng, Shaoyang Sun, Yuan Dai, Jingyuan Li 0003, Yunhui Qiu, Kaihang Wang, Wenbo Yin, Lingli Wang
FCCM5
2023 THRAM: A Template-based Heterogeneous CGRA Modeling Framework Supporting Fast DSE
abstract
Coarse-grained reconfigurable architecture (CGRA), composed of word-level processing elements (PEs) and interconnects, has emerged as a promising architecture due to its high performance, energy efficiency, and flexibility. Although multiple CGRA frameworks have been proposed, a complete heterogeneous CGRA exploration framework with tunable interconnect flexibility and fast design space exploration (DSE) is still lacking. In this paper, we propose an open-source template-based CGRA exploration framework that integrates the modeling of heterogeneous PEs and interconnects, RTL generation, DFG mapping, automatic simulation and verification, and fast DSE based on a CGRA framework TRAM. Moreover, we present a novel resource-efficient shared reconfigurable delay unit (RDU) for data synchronization, which can save the CGRA area by 7%, compared with the separated RDU. Further, the explored optimal heterogeneous architecture can reduce the area and power by 44.7% and 42.9% respectively, and improve the PE utilization by 20.4%, compared with the 8 × 8 baseline architecture in TRAM.
Jingyuan Li 0003, Yunhui Qiu, Guowei Zhu, Qilong Zhu, Wenbo Yin, Lingli Wang
ISCAS2
2022 TRAM: An Open-Source Template-based Reconfigurable Architecture Modeling Framework
abstract
Coarse-grained reconfigurable architecture (CGRA) is a promising accelerator design choice due to its high performance and power efficiency in the computation or data-intensive application domains, such as security, multimedia, digital signal processing, machine learning, and high-performance computing. CGRA consists of coarse-grained processing elements (PEs) and interconnects that determine the architecture flexibility to support different applications and also affect the performance and power efficiency significantly. Although multiple types of interconnects have been proposed, a parameterized unified model is still lacking. In this paper, we propose a flexible and scalable CGRA template with a novel interconnect model that can unify the typical neighbor-to-neighbor, switch-based, and FPGA-like interconnects. Furthermore, we present TRAM, an open-source template-based reconfigurable architecture modeling framework that integrates the Chisel-based CGRA modeling, architecture intermediate representation (IR) and Verilog generation, dataflow graph (DFG) mapping, simulation, and evaluation. The mapping flow contains graph-based placement and routing, critical-path-driven data synchronization, and simulated-annealing-based optimization. We evaluate the impacts of the rich design parameters, which demonstrate the significance of such a flexible template to facilitate architecture optimization. Compared with the related work, TRAM can achieve a 4.1× smaller DFG latency and a faster mapping speed for both the 8×8 and 16×16 CGRAs. Moreover, TRAM is able to attain an extremely high PE utilization of 94.4 % on average by architecture tuning.
Yunhui Qiu, Yuhang Cao, Yuan Dai, Wenbo Yin, Lingli Wang
FPL1
2022 A High-Performance and Scalable NVMe Controller Featuring Hardware Acceleration
abstract
Nonvolatile memory express (NVMe) is a high-performance and scalable PCI express (PCIe)-based interface for the host software communicating with NVMs, including NAND Flash and the storage class memories (SCMs). NVMe solid-state drives (SSDs) have been deployed in cloud platforms and data-centers for a variety of I/O intensive applications due to their performance benefits compared to SATA/SAS SSDs. Considering the design flexibility, firmware-based NVMe controllers are typically used in Flash-based NVMe SSDs but may occupy a significant portion of processor resources and power consumption to achieve high performance. Moreover, the firmware component can be a critical performance bottleneck for SCMs that are an order-of-magnitude faster than Flash. To address these challenges, hardware-accelerated NVMe controllers have emerged in both industry and academia. The commercial hardware controllers are confidential, whereas current academic studies still spare much room for architecture innovations. In this article, we propose an opensource ultralow-latency and high-throughput NVMe controller with a highly parallel, pipelined, and scalable architecture that accommodates one admin controller and multiple fully hardware-automated I/O controllers. We perform extensive empirical performance evaluations concerning the NVMe I/O size, queue depth, queue number, read-to-write ratio, and access pattern. The maximum read/write bandwidth can achieve 7.0 GB/s, accounting for 89% of the PCIe bandwidth. The 4-KB-sized read/write throughput can attain 1.7 million I/O operations per second (MIOPS), whereas the average latency is merely 2.4$\mu \text{s}$/3.2$\mu \text{s}$. Compared to state-of-the-art NVMe controllers in academia, the 4-KB-sized read/write bandwidth of our controller reaches$2.2 \times /2.3\times $as high and the latency is$5.1 \times /4.9\times $lower.
Yunhui Qiu, Wenbo Yin, Lingli Wang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2021 A High-performance Open-channel Open-way NAND Flash Controller Architecture
abstract
NAND-Flash-based SSDs have been widely employed in diverse computing domains and storage systems due to their higher performance and lower power consumption than HDDs. There have been various studies to explore the internal parallelism inside SSDs, including the channel-way-plane levels of interleaving and the cache mode pipelining. However, most current studies are based on simulators or focus on part of the parallelism. In this paper, we present an open-source high-performance open-channel open-way NAND Flash controller supporting all the parallelism. Several architecture innovations are proposed to improve performance and resource efficiency. Firstly, the controller exposes the multi-channel, multi-way topology with a queue-based asynchronous interface for each way. Secondly, a dual-level command scheduler is integrated to enable the fine-grained way-level interleaving, plane-level interleaving, and cache mode pipelining. Finally, four finite state machines are designed for the classified Flash command groups. Evaluated on an FPGA platform, the maximum bandwidth can reach 1.2GB/s, accounting for 93% of the theoretical bandwidth, 13% higher than the bandwidth utilization of other Flash controllers. The minimum latencies for the page reading and programming are 119ps and 2ms respectively, which can be further speeded up by 1.9x and 3.1x on average with the multi-level parallelism.
Yunhui Qiu, Wenbo Yin, Lingli Wang
FPL1
2020 FULL-KV: Flexible and Ultra-Low-Latency In-Memory Key-Value Store System Design on CPU-FPGA
abstract
In-memory key-value store (IMKVS) has gained great popularity in data centers. However, big data brings great challenges in performance and power consumption because of the general-purpose Von Neumann computer architecture. Remote direct memory access (RDMA) technology supporting zero-copy networking could partly alleviate the problem but is still not efficient for KVS. To overcome this problem, we present a flexible and ultra-low-latency IMKVS system named FULL-KV, based on a CPU-FPGA heterogeneous architecture. The FPGA serves as a KVS accelerator that can bypass the CPU and implement both the network stacks and the KVS processing with a highly parallel hardware architecture. The system latency of FULL-KV can achieve as low as 1.5μs/2.2μs for the PUT/GET operation, which is 3.0x/1.5x faster than current state-of-the-art hardware-based KVS systems. Besides, FULL-KV can support 4x larger values (up to 4M bytes). Given a total Ethernet bandwidth of 20Gbps, the peak throughput of the single-node FULL-KV can reach 26.0 million key-value operations per second (Mops). In the two-node test system with a commercial Ethernet switch, the peak throughput can reach 52Mops, manifesting the system scalability and practicability.
Yunhui Qiu, Jinyu Xie, Hankun Lv, Wenbo Yin, Wai-Shing Luk, Lingli Wang, Bowei Yu, Xianjun Ge, Zhijian Liao, Xiaozhong Shi
IEEE Trans. Parallel Distributed Syst.1
2019 A Low-Latency Multi-Version Key-Value Store Using B-Tree on an FPGA-CPU Platform
abstract
In recent years, a variety of methods for low-latency key-value store (KVS), such as remote direct memory access (RDMA), have been proposed. However, a majority of KVS systems do not support to access and query data stored in multiple versions, making them inapplicable to application scenarios like snapshot. In this paper, we present a low-latency multi-version in-memory KVS on an FPGA-CPU platform. In order to reduce latency, we store all keys into a hash table using cuckoo hashing on an FPGA board and organize each group of version-value pairs that are corresponded with the identical key into a B-tree in the host memory. The proposed architecture can perform put, get, delete, CAS, getPredecessor and range query operations within a B-tree. Every operation except range query can be completed bypassing the host CPU. The experimental results show that the average latency of get operation within a B-tree of 5 levels is less than 8?s which is at least 9x faster than other implementations with the help of CPU.
Zhijian Liao, Xiaozhong Shi, Jinyu Xie, Yunhui Qiu, Hankun Lv, Wenbo Yin, Lingli Wang, Bowei Yu, Xianjun He
FPL5
2018 Ultra-Low-Latency and Flexible In-memory Key-Value Store System Design on CPU-FPGA
abstract
In-memory key-value store (KVS) is critical infrastructure in data centers and is facing challenges in performance and power consumption with the development of the big data technology, which mainly results from the low efficiency of the multi-level memory hierarchy of the CPU-based system. Remote direct memory access (RDMA) technology partly alleviates the problems, but it is still not efficient for KVS, especially for the PUT operation. In this paper, we present an ultra-low-latency and flexible in-memory KVS system based on the CPU-FPGA heterogeneous architecture, which leverages FPGA to serve as a KVS accelerator. We design a highly parallel accelerator architecture with several novel techniques, including memory pre-allocation, fragmentation processing, and decoupling design, to achieve ultra-low latency, high flexibility, efficiency, and scalability. The system workload can scale up with the storage capacity due to the decoupling design which stores the hash table in onboard DRAM memory and values in the host memory. For each KVS operation, at most one PCIe DMA is needed, which achieves high efficiency. Compared with current hardware-based KVS systems, the proposed one is more flexible, where the supported value range is 4x wider (from 1 byte to 4M bytes). In 10Gbps Ethernet, the peak throughput of the system can reach 13.6 million key-value operations per second (Mops), achieving nearly full utilization of the Ethernet bandwidth. The system latency can achieve as low as 1.2us for the PUT operation and 1.7us for the GET operation, which is 3.8x and 2.0x faster respectively than current state-of-the-art KVS systems.
Yunhui Qiu, Hankun Lv, Jinyu Xie, Wenbo Yin, Lingli Wang
FPT1