EDBT 2026 Demo / reviewers in the wild / expert
Wenbo Yin
dblp:161/4708
· DBLP profile ↗
38ranked-venue papers
2as first author
30since 2021 · last 2026
0009-0004-1507-3242ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 36 · 28 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Lora: Towards Improved Applicability of Reconfigurable Architecture for Versatile Nonlinear Functions
Yuan Dai, Guibin Zou, Yuanda Yang, Jiahang Lou, Yiwen Luo, Xinyu Cai, Wenbo Yin, Wai-Shing Luk, Lingli Wang |
ISCA | 8 |
| 2026 | Live Demonstration: An Agile FPGA-Overlayed CGRA SoC for High-Efficiency Computing
Jiahang Lou, Jianrong Zhang, Yuan Dai, Zewei Zhong, Wenbo Yin, Lingli Wang |
ISCAS | 6 |
| 2026 | MOE: An Efficient Multicasting and One-hot Encoding Hybrid Configuration Compression Technique for CGRAs
Yuan Dai, Wenbo Yin, Lingli Wang |
ISCAS | 3 |
| 2026 | FlexEdge: A Hardware-Software Co-Design for Flexible Edge Transformer Inference
Jiewen Zheng, Zifeng Zhao, Gengsheng Chen, Wenbo Yin |
ISCAS | 5 |
| 2026 | Dependency-Aware Data Parallelism on Spatial CGRA via Constraint Satisfaction and Graph ColoringabstractCoarse-grained Reconfigurable Architecture (CGRA) is a competitive accelerator architecture for computation-intensive loop kernels. Spatial CGRA is a typical CGRA that performs all the operations spatially to reduce reconfiguration costs within a single iteration, demanding high data parallelism. To achieve this goal, one of the main challenges is the loop-carried dependency between memory accesses. Many existing CGRA compilers struggle to precisely analyze the dependency distance, especially when accesses involve complex address patterns. Consequently, these compilers often default to setting the distance to one, based on a worst-case assumption, leading to degraded performance. However, we observe that a precise distance can improve performance significantly, raising the requirement for an efficient distance calculation approach. Another challenge is the performance constraints of single-bank memory, which necessitate the designer partitioning the original data into a multi-bank memory. However, we observe that the mapping result can cause the inter-iteration conflict, thereby invalidating the memory partition scheme. Therefore, an efficient post-mapping conflict detection is required. In this paper, we develop a constraint satisfaction problem (CSP)-based approach for calculating dependency distance and detecting conflicts, which determines the maximum available dependency distance and identifies conflicts within both intra- and inter-iterations. Besides, we formulate access scheduling as a graph coloring problem, which can minimize conflicts and improve performance. Overall, we develop a comprehensive end-to-end framework with architectural and compiler support for efficient data parallelism on spatial CGRA. We conduct extensive experiments to systematically evaluate the impact of different approaches on performance and compilation. Evaluation results show that our architecture can achieve 13.16× and 1.19× (up to 1.68×) average performance improvements compared to a RISC-V CPU and a state-of-the-art CGRA SoC, respectively. Besides, our architecture has 7.38× and 1.18× (up to 1.65×) average energy efficiency gains compared to these two architectures. Yuan Dai, Xuchen Gao, Wenbo Yin, Wai-Shing Luk, Lingli Wang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2026 | Toward Efficient Edge AI With Heterogeneous Computing and Multilevel OptimizationabstractThe rapid progress of artificial intelligence (AI) has brought increasing demands on hardware accelerators, particularly as modern models combine dense linear operations with a growing number of irregular, nonlinear, and control-intensive operators. While tensor cores and systolic arrays offer high throughput for regular computations, they often struggle to efficiently support the diverse operations emerging in recent model structures. Coarse-grained reconfigurable arrays (CGRAs), with their spatial parallelism and reconfigurability, may serve as a natural complement to dense accelerators in such heterogeneous workloads. In this work, we propose EUREKA, a heterogeneous acceleration framework that integrates tensor cores with CGRAs through a unified instruction set, cross-architecture data scheduling, tailored hardware support for nonlinear operators, and optimizations at the instruction, task, and operator levels to exploit parallelism. At the software level, we introduce a hierarchical compilation strategy that combines graph-level optimizations with tensor-level scheduling techniques. To address the large design space of hardware–software co-optimization, we further develop a Bayesian optimization-based exploration scheme enhanced with kernel compression methods, which provides an efficient means of identifying promising hardware configurations and scheduling strategies. Experiment results on representative AI benchmarks show that EUREKA improves execution efficiency, achieving an average$12.6\times $normalized performance gain over state-of-the-art frameworks. Jingyuan Li 0003, Xinyu Cai, Yuan Dai, Wenbo Yin, Lingli Wang |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2025 | Towards Efficient Data Parallelism on Spatial CGRA via Constraint Satisfaction and Graph ColoringabstractCoarse-Grained Reconfigurable Architecture (CGRA) is a competitive accelerator architecture for computation-intensive loop kernels. Spatial CGRA is a typical CGRA that performs all the operations spatially, demanding high data parallelism. Given the performance limitations of single-bank memory, partitioning original data into multi-bank memory within the spatial CGRA is favored. However, we observe that the mapping result can cause the inter-iteration conflict, thereby invalidating the memory partition scheme. Yuan Dai, Xuchen Gao, Bingbing Peng, Wenbo Yin, Wai-Shing Luk, Lingli Wang |
ASP-DAC | 5 |
| 2025 | LEMOE: LLM-Enhanced Multi-Objective Bayesian Optimization for Microarchitecture ExplorationabstractDesigning processor microarchitectures is increasingly challenging due to a vast design space and the need to balance multiple metrics. Traditional algorithm-driven design space exploration (DSE) approaches often struggle to incorporate the extensive domain knowledge of expert architects. To address this, we introduce LEMOE, a multi-objective microarchitecture optimization framework that leverages large language model (LLM) to enhance an implicit Bayesian model. LEMOE features a program-aware warm-up phase utilizing LLM and LLVM to produce an initial design set with rich prior knowledge. By harnessing LLM’s contextual learning, our approach improves surrogate modeling and sampling under sparse data conditions. Experiment results show that LEMOE achieves a $22.8 \%$ improvement in energy efficiency with the same number of iterations and a $2.9 \times$ runtime speedup for the same target compared to prior works. Jingyuan Li 0003, Jianrong Zhang, Wenbo Yin, Lingli Wang |
DAC | 4 |
| 2025 | Adora Compiler: End-to-End Optimization for High-Efficiency Dataflow Acceleration and Task Pipelining on CGRAsabstractTo fully harness emerging computing architectures, compilers must provide intuitive input handling alongside powerful code optimization to unlock maximum performance. Coarse-Grained Reconfigurable Arrays (CGRAs) — highly energy-efficient for nested-loop applications — have lacked a compiler capable of meeting these objectives. This paper introduces the Adora compiler [1], which effectively bridges user-friendly, lightweight coding inputs with high-performance acceleration on the CGRA SoC. Adora utilizes CGRA-target loop transformations to achieve efficient data-flow level execution while optimizing data communication and task pipelining at the task-flow level. Additionally, it incorporates a comprehensive automated algorithm with a thoughtfully designed optimization sequence. A series of comprehensive experiments highlights the exceptional efficiency and scalability of the Adora compiler, demonstrating its transformative impact in leveraging CGRA capabilities for acceleration in edge computing. Jiahang Lou, Qilong Zhu, Yuan Dai, Zewei Zhong, Wenbo Yin, Lingli Wang |
DAC | 5 |
| 2025 | DEFA: Design Space Exploration for FPGA Overlay Accelerators Through Frequency Prediction and Bayesian OptimizationabstractIn edge AI inference, FPGAs demonstrate superiority in performance-area balance. FPGA Overlay Accelerators (FOAs) are programmable accelerators implemented on FPGAs, typically highly parameterized to enable flexible hardware realization. These parameters, varying across a wide design space, have a significant impact on performance and require efficient Design Space Exploration (DSE). However, current frameworks struggle to accurately predict performance metrics like maximum frequency and fail to fully explore the design space, limiting DSE's effectiveness. In this paper, we propose a DSE framework for FOA (DEFA) based on Bayesian optimization, providing more effective and comprehensive DSE. To address complex parameter interdependencies in FOA, a dependency-aware design space modeling approach (DAM) is proposed. This approach applies fine-grained pruning to the parameter space while addressing dependency constraints. Based on this pruned parameter space, we develop a custom regression predictor (CREP) for maximum frequency using LightGBM, significantly enhancing performance estimation accuracy. Furthermore, the search efficiency is improved through enhanced Latin hypercube sampling and the Tree-Structured Parzen Estimator. We use the proposed framework to optimize an FOA template, Intel FPGA AI Suite. The Pearson correlation coefficient of CREP's predictions regarding the maximum frequency of accelerator instances achieves 0.87. In the throughput optimization experiment, the proposed DSE framework improves 30.16 % compared to the architecture optimization functionality provided by Intel FPGA AI Suite across the given 10 benchmarks on average. In the areathroughput trade-off optimization experiment, compared with FPGA AI Suite, the proposed DSE framework improves 5.01 % in frequency, 18.48 % in throughput and 21.60 % in area. Qilong Zhu, Yunfei Dai, Shiyan Bi, Huizhen Kuang, Dylan Wang, Wenbo Yin, Lingli Wang |
FPL | 6 |
| 2025 | ICIMG-Net: Inject Context Information to Motion Generation for Optical Flow EstimationabstractAlthough the overall performance of existing optical flow estimation methods has improved rapidly, motion discontinuities caused by large displacements and occlusions remain significant challenges for accurate optical flow estimation. To address this issue, we propose a novel Inject Context Information for Motion Generation Network (ICIMG-Net). Firstly, we design a Context Injection Module (CIM), which enhances the generation of correlation volumes by injecting context information. This process supplements the semantic details needed for accurate pixel matching, improving matching precision. Then, we construct a Dual-Injection-GRU (DIG), which facilitates dual interactions between semantic and motion information to address the motion discontinuities problem. Finally, we conduct a comprehensively evaluat of our ICIMG-Net against state-of-the-art methods on the MPI-Sintel and KITTI benchmarks, demonstrating that our method achieves competitive results, particularly in complex synthetic and real-world traffic scenarios. Wenbo Yin, Congxuan Zhang, Zhen Chen 0004, Liyue Ge, Zige Wang |
ICASSP | 1 |
| 2025 | DynVec: An End-to-End Framework for Efficient Vector-Dataflow ExecutionabstractHigh-performance computing (HPC) and hardware acceleration increasingly rely on dataflow architectures to achieve scalable parallelism and efficiency. High-level synthesis (HLS) facilitates accelerator design from high-level programs, but conventional tools often require intrusive source-level modifications and struggle to optimize irregular workloads. Dynamically scheduled HLS frameworks offer a promising direction for addressing control flow divergence and memory irregularity by generating dataflow accelerators. However, they lack compile-time parallelism optimizations such as vectorization and incur significant hardware overhead. Moreover, modern compilers can generate vectorized code using memory access and computational patterns. Nevertheless, in programs with irregular control flow or data-dependent behavior, such patterns are unknown until runtime, limiting the effectiveness of static vectorization strategies.To address these challenges, we propose DynVec, a unified vector-dataflow framework that integrates dynamic scheduling and vectorization to exploit runtime parallelism beyond conventional models. We address the vectorization of irregular kernels through an MLIR-based context-aware vectorizer that effectively identifies vectorizable operations and, through dataflow scheduling, generates a vector-dataflow execution graph that explicitly models control flow constructs, data and control interfaces, and memory operations. DynVec encapsulates high-level elastic units designed with built-in vectorization support, allowing customizable and adaptive execution behavior. Our compiler preserves the structural hierarchy of the kernel by combining vector and scalar operations in a bottom-up, type-safe manner. Experiments show that our approach achieves significant speedup compared to state-of-the-art HLS implementations across various regular and irregular applications. Moreover, compared to hybrid accelerators that separately support dynamic parallelism and vectorization, DynVec delivers superior performance. Xianfeng Cao, Kaixiang Zhu, Wenbo Yin, Lingli Wang |
ICCAD | 4 |
| 2025 | COFFA: A Co-Design Framework for Fused-Grained Reconfigurable Architecture Towards Efficient Irregular Loop HandlingabstractCoarse-Grained Reconfigurable Architecture (CGRA) emerges as a competitive accelerator due to its high flexibility and energy efficiency. However, most CGRAs are effective for computation-intensive applications with regular loops but struggle with irregular loops containing control flows. These loops introduce fine-grained logic operations and are costly to execute by coarse-grained arithmetic units in CGRA. Efficiently handling such logic operations necessitates incorporating Boolean algebra optimization, which can improve logic density and reduce logic depth. Unfortunately, no previous research has incorporated it into the compilation flow to support irregular loops efficiently.We proposeCOFFA, an open-source framework for heterogeneous architecture with a RISC-V CPU and a fused-grained reconfigurable accelerator, which integrates coarse-grained arithmetic and fine-grained logic units, along with flexible IO units and distributed interconnects. As a software/hardware co-design framework,COFFAhas a powerful compiler that extracts and optimizes fine-grained logic operations from irregular loops, performs coarse-grained arithmetic and memory optimizations, and offloads the loops to the accelerator.Across various challenging benchmarks with irregular loops,COFFAachieves significant performance and energy efficiency improvements over an in-order, an out-of-order RISC-V CPUs, and a recent FPGA, respectively. Moreover, compared with the state-of-the-art CGRAUE-CGRAandHycube,COFFAcan achieve 2.5× and 3.5× performance gains, respectively. Yuan Dai, Xuchen Gao, Yunhui Qiu, Jingyuan Li 0003, Yuhang Cao, Yiqing Mao, Sichao Chen, Wenbo Yin, Wai-Shing Luk, Lingli Wang |
IEEE Trans. Computers | 8 |
| 2025 | MoDAF: A Multi-objective Divide-and-Conquer Parameter Tuning Framework for CGRAsabstractCoarse-grained reconfigurable architectures (CGRAs) are gaining increasing attention as domain-specific accelerators due to their high flexibility and energy efficiency. These architectures offer a compelling solution for applications that require custom hardware performance while retaining a degree of programmability. However, the design space of CGRAs is inherently vast and complex, presenting significant challenges for architects to explore design choices efficiently and systematically. Existing design space exploration (DSE) methodologies for CGRAs are often time-demanding and struggle to deliver optimal solutions when confronted with high-dimensional and multi-objective design space. Therefore, we consider constructing a CGRA parameter tuning framework called MoDAF. MoDAF initializes the design space using the most representative and diverse samples. It adopts a divide-and-conquer approach, utilizing Monte Carlo Tree Search (MCTS) and space partitioning techniques to dynamically break down the complex design space into more manageable subspaces. A hybrid model handles local fluctuations within each subspace, while a dual sampling algorithm is designed to increase sampling efficiency. MoDAF also incorporates a fast evaluation model to estimate CGRA throughput and area, significantly speeding up the exploration process. Compared with previous approaches, experiments show that our proposed framework reduces the average distance from the reference set by 53.0% and the hypervolume deviation by 64.2%, while also cutting wall time by 57.5%. Jingyuan Li 0003, Yuan Dai, Wenbo Yin, Lingli Wang |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2024 | An End-to-End Agile Design Framework to Improve Energy Efficiency on CGRAsabstractIn this paper, we propose a domain-specific frame-work that integrates Chisel-based Coarse-grained reconfigurable architecture (CGRA) Modeling, RTL generation, Architecture Graph Intermediate Representation (IR), dataflow graph (DFG) Mapping, interconnect exploration, and physical implementation. Within this framework, we propose an interconnect exploration flow based on a novel interconnect architecture called Matrix, which realizes a heterogeneous interconnect architecture for a set of specific applications by application mapping, design space exploration (DSE) and pruning. We design an agile mapper built on a graph-based two-level architectural IR, which better adapts to the flexible interconnect model and enables greater interconnect exploration to improve the mapping success rate. Experiments show a significant reduction in architecture area, improved energy efficiency, and high PE utilization compared to the state-of-the-art tool, with high-quality mapping results due to architecture tuning. Additionally, our pruning strategies reduce interconnect paths in the Matrix, ensuring interconnect efficiency and further improving PE utilization. Yazhou Yan, Guowei Zhu, Wenbo Yin, Lingli Wang |
ASAP | 4 |
| 2024 | MDCRA: A Reconfigurable Accelerator Framework for Multiple Dataflow LanesabstractCoarse-grained reconfigurable architecture (CGRA) is a type of reconfigurable computing architecture suitable for emerging applications that require dynamic compilation hardware. However, the resource utilization of existing CGRA is low due to the lack of flexibility across varied application granularity. In this paper, we propose a CGRA framework for multiple dataflow lanes (MDCRA). It supports post-silicon computational granularity adjustments. Evaluated with Polybench, Machsuite and Express, the speedup of MDCRA is$24.83\times$higher than CPU CVA6, and$2.08\times$higher than vector processor Ara. Compared with TRAM and DSAGEN, MDCRA achieves an area reduction of 27% and 47% respectively with the same speedup. Besides, compared with OpenCGRA, the average utilization of function units is improved by 20.05%. Shaoyang Sun, Boyin Jin, Jiahang Lou, Yuhang Cao, Jingyuan Li 0003, Yuan Dai, Wenbo Yin, Wai-Shing Luk, Lingli Wang |
ASAP | 9 |
| 2024 | A CGRA Front-end Compiler Enabling Extraction of General Control and Dedicated OperatorsabstractCoarse-grained reconfigurable architecture (CGRA) gradually becomes an extraordinarily promising accelerator due to its flexibility and power efficiency. However, most CGRA front-end compilers focus on the innermost body of regular loops with a pure data flow. Therefore, we propose CO-Compiler, an LLVM-based CGRA front-end compiler to generate an optimized control-data flow graph (CDFG), which can handle versatile loops in C/C++, including general control flow, arbitrary nested levels, and imperfect statements. Then we extract multi-dimension memory access patterns and various dedicated operators adapting to concrete hardware functions. In addition, we analyze variable loop bounds which are settled at runtime, and realize the SoC runtime configuration of CGRA. The feasibility of our methodology is verified by a RISC-V based SoC simulation. The experimental results demonstrate that our dedicated operator extraction can reduce 43% PE resources and decrease 84% initiation interval (II) on a TRAM architecture. Furthermore, compared with state-of-the-art (SOTA) CGRA front-end compilers, CO-Compiler has the highest 88.1% success rate in CDFG generation for a wide range of benchmarks. Moreover, by using the same back-end mappers, our work can reach 78% reduction for II and $2.06\times$ PE spatio-temporal utilization in contrast with their own front-end compilers. Xuchen Gao, Yunhui Qiu, Yuan Dai, Wenbo Yin, Lingli Wang |
ASPDAC | 4 |
| 2024 | An Agile Deploying Approach for Large-Scale Workloads on CGRA-CPU ArchitectureabstractAdopting specialized accelerators such as Coarse-Grained Reconfigurable Architectures (CGRAs) alongside CPUs to enhance performance within specific domains is an astute choice. However, the integration of heterogeneous architectures introduces complex challenges for compiler design. Simultaneously, the ever-expanding scale of workloads imposes substantial burdens on deployment. To address above challenges, this paper introduces CGRV-OPT, a user-friendly multi-level compiler designed to deploy large-scale workloads to CGRA and RISC-V CPU architecture. Built upon the MLIR framework, CGRV-OPT serves as a pivotal bridge, facilitating the seamless conversion of high-level workload descriptions into low-level intermediate representations (IRs) for different architectures. A salient feature of our approach is the automation of a comprehensive suite of optimizations and transformations, which speed up each kernel computing within the intricate SoC. Additionally, we have seamlessly integrated an automated software-hardware partitioning mechanism, guided by our multi-level optimizations, resulting in a remarkable 2.14 × speed up over large-scale workloads. The CGRV-OPT framework significantly alleviates the challenges faced by software developers, including those with limited expertise in hardware architectures. Jiahang Lou, Xuchen Gao, Yiqing Mao, Yunhui Qiu, Yihan Hu 0003, Wenbo Yin, Lingli Wang |
DATE | 6 |
| 2024 | CFEACT: A CGRA-based Framework Enabling Agile CNN and Transformer Accelerator DesignabstractConvolutional neural networks (CNNs) and transformer neural networks have been adopted in a wide range of applications such as natural language processing and computer vision. Coarse-grained reconfigurable architectures (CGRAs) are highly suitable for CNN and transformer applications due to their high flexibility and energy efficiency. However, current implementations of CGRA for CNNs and transformers have several limitations including the lack of System-on-Chip (SoC), insufficient support for nonlinear functions and the absence of a software toolchain. To address these challenges, we present CFEACT, a CGRA-based framework that enables agile development of CNN and transformer accelerators. CFEACT offers a broad design space of efficient CGRA accelerators through a highly flexible architecture template. The well-designed SoC, innovative mapping schemes, and comprehensive software toolchain offer a complete solution for implementing various CNN and transformer models on the generated CGRAs. Compared with the state-of-the-art works, accelerators generated by CFEACT can achieve more than $2 \times$ improvement in area-delay product for CNNs and an average of $2 \times$ higher performance for transformers. Yiqing Mao, Xuchen Gao, Jiahang Lou, Yunhui Qiu, Wenbo Yin, Wai-Shing Luk, Lingli Wang |
FPL | 5 |
| 2024 | HierCGRA: A Novel Framework for Large-scale CGRA with Hierarchical Modeling and Automated Design Space ExplorationabstractCoarse-grained reconfigurable arrays (CGRAs) are promising design choices in computation-intensive domains, since they can strike a balance between energy efficiency and flexibility. A typical CGRA comprises processing elements (PEs) that can execute operations in applications and interconnections between them. Nevertheless, most CGRAs suffer from the ineffectiveness of supporting flexible architecture design and solving large-scale mapping problems. To address these challenges, we introduce HierCGRA, a novel framework that integrates hierarchical CGRA modeling, Chisel-based Verilog generation, LLVM-based data flow graph (DFG) generation, DFG mapping, and design space exploration (DSE). With the graph homomorphism (GH) mapping algorithm, HierCGRA achieves a faster mapping speed and higher PE utilization rate compared with the existing state-of-the-art CGRA frameworks. The proposed hierarchical mapping strategy achieves 41× speedup on average compared with the ILP mapping algorithm in CGRA-ME. Furthermore, the automated DSE based on Bayesian optimization achieves a significant performance improvement by the heterogeneity of PEs and interconnections. With these features, HierCGRA enables the agile development for large-scale CGRA and accelerates the process of finding a better CGRA architecture. Sichao Chen, Su Zheng, Guowei Zhu, Jingyuan Li 0003, Yazhou Yan, Yuan Dai, Wenbo Yin, Lingli Wang |
ACM Trans. Reconfigurable Technol. Syst. | 9 |
| 2024 | FDRA: A Framework for a Dynamically Reconfigurable Accelerator Supporting Multi-Level ParallelismabstractCoarse-grained reconfigurable architectures (CGRAs) have emerged as promising accelerators due to their high flexibility and energy efficiency. However, existing open source works often lack integration of CGRAs with CPU systems and corresponding toolchains. Moreover, there is rare support for the accelerator instruction pipelining to overlap data communication, computation, and configuration across multiple tasks. In this article, we propose FDRA, an open source exploration framework for a heterogeneous system-on-chip (SoC) with a RISC-V processor and a dynamically reconfigurable accelerator (DRA) supporting loop, instruction, and task levels of parallelism. FDRA encompasses parameterized SoC modeling, Verilog generation, source-to-source application code transformation using frontend and DRA compilers, SoC simulation, and FPGA prototyping. FDRA incorporates the extraction of periodic accumulative operators and multi-dimensional linear load/store operators from nested loops. The DRA enables accessing the shared L2 cache with virtual addresses and supports direct memory access with arbitrary start addresses and data lengths. Integrated into the RISC-V Rocket SoC, our DRA achieves a remarkable 55× acceleration for loop kernels and improves energy efficiency by 29×. Compared to state-of-the-art RISC-V vector units, our DRA demonstrates a 2.9× speed improvement and 3.5× greater energy efficiency. In contrast to previous CGRA+RISC-V SoCs, our SoC achieves a minimum speedup of 5.2×. Yunhui Qiu, Yiqing Mao, Xuchen Gao, Sichao Chen, Wenbo Yin, Lingli Wang |
ACM Trans. Reconfigurable Technol. Syst. | 6 |
| 2024 | HETA: A Heterogeneous Temporal CGRA Modeling and Design Space Exploration via Bayesian OptimizationabstractDue to its high energy efficiency and flexibility, coarse-grained reconfigurable architecture (CGRA) has gained increasing attention. Temporal CGRA is a typical category of CGRA that supports single-cycle context switching and time-multiplexing hardware resources to perform spatial and temporal computations. Although multiple temporal CGRAs have been proposed, an architecture with rich design parameters and heterogeneous modeling is still lacking. To this end, we propose a highly parameterized heterogeneous temporal CGRA, called HETA. However, the highly parameterized and heterogeneous design introduces a challenging design space for manual exploration. To address this challenge, we introduce a Bayesian-optimization (BO)-based design space exploration (DSE) of homogeneous and heterogeneous architectures. Different from other DSE processes that require defining the heterogeneous exploration strategy, our approach adopts a searching-pruning-based method without manual intervention. To improve the efficiency of DSE, we develop a fast statistic model for area evaluation, whose error is below 1%. In addition, a pipeline mapping (PiPMap) algorithm is developed to alleviate the restrictions caused by data synchronization and unleash the potential of the proposed architecture. Experimental results show that HETA can achieve 89%, 52%, and 47% improvement in throughput, area efficiency, and energy efficiency over the neighbor-to-neighbor (N2N)-based interconnect CGRA, respectively. Compared with the Switch-based interconnect CGRA, HETA’s area efficiency is increased by 61%. Furthermore, compared with the homogeneous architecture of HETA, the optimized heterogeneous architecture improves area efficiency and energy efficiency by 14.7% and 4.8%, respectively. Yuan Dai, Jingyuan Li 0003, Qilong Zhu, Yunhui Qiu, Yihan Hu 0003, Wenbo Yin, Lingli Wang |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2023 | UPTRA: An Ultra-Parameterized Temporal CGRA Modeling and OptimizationabstractTemporal Coarse-Grained Reconfigurable Architecture (CGRA) is a typical category of CGRA that supports single-cycle context switching and time-multiplexing hardware resources to perform both spatial and temporal computations. Compared with the spatial CGRA, it can be used in area and power budget-constrained scenarios, with the sacrifice of the throughput. Therefore, achieving minimum Initialization Interval (II) for higher throughput is the main objective in many works for temporal CGRA mapping. Yuan Dai, Yunhui Qiu, Qilong Zhu, Jingyuan Li 0003, Wenbo Yin, Lingli Wang |
FCCM | 5 |
| 2023 | PRAD: A Bayesian Optimization-based DSE Framework for Parameterized Reconfigurable Architecture DesignabstractCoarse-Grained Reconfigurable Architecture (CGRA) is a domain-specific reconfigurable architecture. Generally, the CGRA architecture consists of IO, memory, coarse-grained processing element (PE), and interconnect. Usually, ALU in PE contains a relatively complete set of operations and most of the interconnects adopt neighbor-to-neighbor (N2N) [1], switch-based [2], and combination of the connection box and switch box (CB-SB) patterns [3]. However, the complex operation sets and switch-based/CB-SB fully-connected interconnects provide sufficient reconfigurability at the cost of resource overhead. Thus, it is important to build a parameterized architecture of CGRA to achieve a balance among hardware overhead, flexibility and performance through automatic design space exploration (DSE). Bingbing Peng, Shaoyang Sun, Yuan Dai, Jingyuan Li 0003, Yunhui Qiu, Kaihang Wang, Wenbo Yin, Lingli Wang |
FCCM | 7 |
| 2023 | THRAM: A Template-based Heterogeneous CGRA Modeling Framework Supporting Fast DSEabstractCoarse-grained reconfigurable architecture (CGRA), composed of word-level processing elements (PEs) and interconnects, has emerged as a promising architecture due to its high performance, energy efficiency, and flexibility. Although multiple CGRA frameworks have been proposed, a complete heterogeneous CGRA exploration framework with tunable interconnect flexibility and fast design space exploration (DSE) is still lacking. In this paper, we propose an open-source template-based CGRA exploration framework that integrates the modeling of heterogeneous PEs and interconnects, RTL generation, DFG mapping, automatic simulation and verification, and fast DSE based on a CGRA framework TRAM. Moreover, we present a novel resource-efficient shared reconfigurable delay unit (RDU) for data synchronization, which can save the CGRA area by 7%, compared with the separated RDU. Further, the explored optimal heterogeneous architecture can reduce the area and power by 44.7% and 42.9% respectively, and improve the PE utilization by 20.4%, compared with the 8 × 8 baseline architecture in TRAM. Jingyuan Li 0003, Yunhui Qiu, Guowei Zhu, Qilong Zhu, Wenbo Yin, Lingli Wang |
ISCAS | 5 |
| 2022 | TRAM: An Open-Source Template-based Reconfigurable Architecture Modeling FrameworkabstractCoarse-grained reconfigurable architecture (CGRA) is a promising accelerator design choice due to its high performance and power efficiency in the computation or data-intensive application domains, such as security, multimedia, digital signal processing, machine learning, and high-performance computing. CGRA consists of coarse-grained processing elements (PEs) and interconnects that determine the architecture flexibility to support different applications and also affect the performance and power efficiency significantly. Although multiple types of interconnects have been proposed, a parameterized unified model is still lacking. In this paper, we propose a flexible and scalable CGRA template with a novel interconnect model that can unify the typical neighbor-to-neighbor, switch-based, and FPGA-like interconnects. Furthermore, we present TRAM, an open-source template-based reconfigurable architecture modeling framework that integrates the Chisel-based CGRA modeling, architecture intermediate representation (IR) and Verilog generation, dataflow graph (DFG) mapping, simulation, and evaluation. The mapping flow contains graph-based placement and routing, critical-path-driven data synchronization, and simulated-annealing-based optimization. We evaluate the impacts of the rich design parameters, which demonstrate the significance of such a flexible template to facilitate architecture optimization. Compared with the related work, TRAM can achieve a 4.1× smaller DFG latency and a faster mapping speed for both the 8×8 and 16×16 CGRAs. Moreover, TRAM is able to attain an extremely high PE utilization of 94.4 % on average by architecture tuning. Yunhui Qiu, Yuhang Cao, Yuan Dai, Wenbo Yin, Lingli Wang |
FPL | 4 |
| 2022 | Knowledge Enhanced Pre-trained Language Model for Product Summarization
Wenbo Yin, Junxiang Ren, Yuejiao Wu, Ruilin Song, Sibo Wang 0010 |
NLPCC (2) | 1 |
| 2022 | A High-Performance and Scalable NVMe Controller Featuring Hardware AccelerationabstractNonvolatile memory express (NVMe) is a high-performance and scalable PCI express (PCIe)-based interface for the host software communicating with NVMs, including NAND Flash and the storage class memories (SCMs). NVMe solid-state drives (SSDs) have been deployed in cloud platforms and data-centers for a variety of I/O intensive applications due to their performance benefits compared to SATA/SAS SSDs. Considering the design flexibility, firmware-based NVMe controllers are typically used in Flash-based NVMe SSDs but may occupy a significant portion of processor resources and power consumption to achieve high performance. Moreover, the firmware component can be a critical performance bottleneck for SCMs that are an order-of-magnitude faster than Flash. To address these challenges, hardware-accelerated NVMe controllers have emerged in both industry and academia. The commercial hardware controllers are confidential, whereas current academic studies still spare much room for architecture innovations. In this article, we propose an opensource ultralow-latency and high-throughput NVMe controller with a highly parallel, pipelined, and scalable architecture that accommodates one admin controller and multiple fully hardware-automated I/O controllers. We perform extensive empirical performance evaluations concerning the NVMe I/O size, queue depth, queue number, read-to-write ratio, and access pattern. The maximum read/write bandwidth can achieve 7.0 GB/s, accounting for 89% of the PCIe bandwidth. The 4-KB-sized read/write throughput can attain 1.7 million I/O operations per second (MIOPS), whereas the average latency is merely 2.4$\mu \text{s}$/3.2$\mu \text{s}$. Compared to state-of-the-art NVMe controllers in academia, the 4-KB-sized read/write bandwidth of our controller reaches$2.2 \times /2.3\times $as high and the latency is$5.1 \times /4.9\times $lower. Yunhui Qiu, Wenbo Yin, Lingli Wang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2021 | A High-performance Open-channel Open-way NAND Flash Controller ArchitectureabstractNAND-Flash-based SSDs have been widely employed in diverse computing domains and storage systems due to their higher performance and lower power consumption than HDDs. There have been various studies to explore the internal parallelism inside SSDs, including the channel-way-plane levels of interleaving and the cache mode pipelining. However, most current studies are based on simulators or focus on part of the parallelism. In this paper, we present an open-source high-performance open-channel open-way NAND Flash controller supporting all the parallelism. Several architecture innovations are proposed to improve performance and resource efficiency. Firstly, the controller exposes the multi-channel, multi-way topology with a queue-based asynchronous interface for each way. Secondly, a dual-level command scheduler is integrated to enable the fine-grained way-level interleaving, plane-level interleaving, and cache mode pipelining. Finally, four finite state machines are designed for the classified Flash command groups. Evaluated on an FPGA platform, the maximum bandwidth can reach 1.2GB/s, accounting for 93% of the theoretical bandwidth, 13% higher than the bandwidth utilization of other Flash controllers. The minimum latencies for the page reading and programming are 119ps and 2ms respectively, which can be further speeded up by 1.9x and 3.1x on average with the multi-level parallelism. Yunhui Qiu, Wenbo Yin, Lingli Wang |
FPL | 2 |
| 2021 | FastCGRA: A Modeling, Evaluation, and Exploration Platform for Large-Scale Coarse-Grained Reconfigurable ArraysabstractCoarse-Grained Reconfigurable Arrays (CGRAs) provide sufficient flexibility in domain-specific applications with high hardware efficiency, which make CGRAs suitable for fast-evolving fields such as neural network acceleration and edge computing. To meet the requirement of the fast evolution, we propose FastCGRA, the modeling, mapping, and exploration platform for large-scale CGRAs. FastCGRA supports hierarchical architecture description and automatic switch module generation. Connectivity-aware packing and graph partition algorithms are designed to reduce the complexity of placement and routing. The graph homomorphism placement algorithm in FastCGRA enables efficient placement on large-scale CGRAs. The packing and placement algorithms cooperate with a negotiation-based routing algorithm to form an integral mapping procedure. FastCGRA can support the modeling and mapping of large-scale CGRAs with significantly higher placement and routing efficiency than existing platforms. The automatic switch module generation method can reduce the complexity of CGRA interconnection design. With these features, FastCGRA can boost the exploration of large-scale CGRAs. Su Zheng, Kaisen Zhang, Yaoguang Tian, Wenbo Yin, Lingli Wang, Xuegong Zhou |
FPT | 4 |
| 2020 | FULL-KV: Flexible and Ultra-Low-Latency In-Memory Key-Value Store System Design on CPU-FPGAabstractIn-memory key-value store (IMKVS) has gained great popularity in data centers. However, big data brings great challenges in performance and power consumption because of the general-purpose Von Neumann computer architecture. Remote direct memory access (RDMA) technology supporting zero-copy networking could partly alleviate the problem but is still not efficient for KVS. To overcome this problem, we present a flexible and ultra-low-latency IMKVS system named FULL-KV, based on a CPU-FPGA heterogeneous architecture. The FPGA serves as a KVS accelerator that can bypass the CPU and implement both the network stacks and the KVS processing with a highly parallel hardware architecture. The system latency of FULL-KV can achieve as low as 1.5μs/2.2μs for the PUT/GET operation, which is 3.0x/1.5x faster than current state-of-the-art hardware-based KVS systems. Besides, FULL-KV can support 4x larger values (up to 4M bytes). Given a total Ethernet bandwidth of 20Gbps, the peak throughput of the single-node FULL-KV can reach 26.0 million key-value operations per second (Mops). In the two-node test system with a commercial Ethernet switch, the peak throughput can reach 52Mops, manifesting the system scalability and practicability. Yunhui Qiu, Jinyu Xie, Hankun Lv, Wenbo Yin, Wai-Shing Luk, Lingli Wang, Bowei Yu, Xianjun Ge, Zhijian Liao, Xiaozhong Shi |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2019 | A Low-Latency Multi-Version Key-Value Store Using B-Tree on an FPGA-CPU PlatformabstractIn recent years, a variety of methods for low-latency key-value store (KVS), such as remote direct memory access (RDMA), have been proposed. However, a majority of KVS systems do not support to access and query data stored in multiple versions, making them inapplicable to application scenarios like snapshot. In this paper, we present a low-latency multi-version in-memory KVS on an FPGA-CPU platform. In order to reduce latency, we store all keys into a hash table using cuckoo hashing on an FPGA board and organize each group of version-value pairs that are corresponded with the identical key into a B-tree in the host memory. The proposed architecture can perform put, get, delete, CAS, getPredecessor and range query operations within a B-tree. Every operation except range query can be completed bypassing the host CPU. The experimental results show that the average latency of get operation within a B-tree of 5 levels is less than 8?s which is at least 9x faster than other implementations with the help of CPU. Zhijian Liao, Xiaozhong Shi, Jinyu Xie, Yunhui Qiu, Hankun Lv, Wenbo Yin, Lingli Wang, Bowei Yu, Xianjun He |
FPL | 7 |
| 2018 | Ultra-Low-Latency and Flexible In-memory Key-Value Store System Design on CPU-FPGAabstractIn-memory key-value store (KVS) is critical infrastructure in data centers and is facing challenges in performance and power consumption with the development of the big data technology, which mainly results from the low efficiency of the multi-level memory hierarchy of the CPU-based system. Remote direct memory access (RDMA) technology partly alleviates the problems, but it is still not efficient for KVS, especially for the PUT operation. In this paper, we present an ultra-low-latency and flexible in-memory KVS system based on the CPU-FPGA heterogeneous architecture, which leverages FPGA to serve as a KVS accelerator. We design a highly parallel accelerator architecture with several novel techniques, including memory pre-allocation, fragmentation processing, and decoupling design, to achieve ultra-low latency, high flexibility, efficiency, and scalability. The system workload can scale up with the storage capacity due to the decoupling design which stores the hash table in onboard DRAM memory and values in the host memory. For each KVS operation, at most one PCIe DMA is needed, which achieves high efficiency. Compared with current hardware-based KVS systems, the proposed one is more flexible, where the supported value range is 4x wider (from 1 byte to 4M bytes). In 10Gbps Ethernet, the peak throughput of the system can reach 13.6 million key-value operations per second (Mops), achieving nearly full utilization of the Ethernet bandwidth. The system latency can achieve as low as 1.2us for the PUT operation and 1.7us for the GET operation, which is 3.8x and 2.0x faster respectively than current state-of-the-art KVS systems. Yunhui Qiu, Hankun Lv, Jinyu Xie, Wenbo Yin, Lingli Wang |
FPT | 4 |
| 2018 | Ultra-Low Latency and High Throughput Key-Value Store Systems Over EthernetabstractKey-value store (KVS) systems are playing important roles as the caches of database to improve the data access efficiency. In this paper, we propose general architectures of KVS systems based on field programmable gate arrays (FPGAs) which can adapt to different application scenarios. Data hazards introduced by the pipeline are fully handled to improve the throughput. The batch operation is proposed to support the variable-length key-value pairs. The memory address management system reduces CPU load further by managing memory addresses on hardware. We implement and evaluate our KVS systems on VC709 evaluation board targeting throughput, latency and capacity respectively. The ultra-low latency KVS system can achieve 160 ns latency and about 155 million request per second (MRPS) to retrieve 96-bit values with 96-bit keys. The ultra-high throughput KVS system can achieve about 200 ns latency and 200 MRPS throughput to retrieve the same size key-value pairs. The ultra-capacity KVS system retrieves 64-byte values with 12-byte keys, serving a maximum request rate of more than 35 MRPS with latency less than 500 ns. Wenbo Yin, Lingli Wang |
ISCAS | 2 |
| 2017 | FPGA acceleration of the scoring process of X!TANDEM for protein identificationabstractTandem mass spectrometry has been a main method for protein identification. X!Tandem, a widely used database search engine, may spend hours or days accomplishing a certain searching task due to the increased search space, which generates urgent demands for computationally efficient database searching. Profiling analysis indicates that it takes X!Tandem about 70%-90% of the total time to conduct the scoring process. The scoring process is composed of fragment ion generation and score generation. This paper proposes a scalable hardware design to speed up the scoring process of X!Tandem that exploits the flexibility of Field Programmable Gate Arrays (FPGAs). The hardware implementation of the scoring process that instantiates 1 fragment ion generation module and 6 score generation modules running on a Xilinx Virtex-7 XC7VX690T FPGA can achieve a 26 times speedup, compared with X!Tandem software implementation running on a 2.5GHz Intel i7-4870 processor with 16 GB memory, whilst fragment ion generation can achieve a 67 times speedup and score generation can achieve a 17 times speedup. Besides, the scalability of score generation modules is linear and outperforms previous parallel approaches. Jin Qiu, Ping Kang, Yipeng Yuan, Wenbo Yin, Lingli Wang |
FPL | 5 |
| 2016 | Memory efficient and high performance key-value store on FPGA using Cuckoo hashingabstractKey-value stores (KVS) become critical in many applications because of the data explosion recently. There is a strong demand to improve the throughput and reduce the latency for KVS. FPGA-based parallel architecture can bring excellent performance and power efficiency. Cuckoo hashing has proven to be an efficient approach to implement KVS with good memory utilization and constant worst case access time. In this paper, an FPGA-based KVS implementation is proposed based on Cuckoo hashing, with a decoupled storage to achieve 81.7% memory utilization, and a pipeline scheme to achieve high performance. The latency of insert, search and delete operations is only 40 ns. And the throughput for search and delete can be 200 million requests per second (MRPS) which is 5× faster than [1]. Even when the load factor becomes 0.9, the throughput for insert can still achieve 147 MRPS. Wenbo Yin, Ping Kang, Lingli Wang |
FPL | 2 |
| 2016 | Hardware TCP Offload Engine based on 10-Gbps Ethernet for low-latency network communicationabstractThis paper introduces a hardware TCP Offload Engine (TOE) aiming at low-latency communication systems. The throughput can reach 9.99 Gbps with the Jumbo frame. The input-to-output receiving latency of a packet consists of 100 bytes payload and 64 bytes header with timestamp is close to 90 nanoseconds. The application-to-application latency between the proposed acceleration system and the native Windows socket is one tenth of the measurement result compared with the latency between the native Linux socket without TOE acceleration and the native Windows socket. Ping Kang, Wenbo Yin, Linli Wang |
FPT | 3 |
| 2014 | No zero padded sparse matrix-vector multiplication on FPGAsabstractSparse Matrix-Vector Multiplication (SpMxV) algorithms suffer heavy performance penalties due to irregular memory accesses. In this paper, we introduce a novel compressed element storage (CES) format, in which the additional data structures for indexing are abandoned, and each location associated with the non-zero element of the matrix is now indicated by the name of a variable multiplied by the corresponding element of the vector. To ensure fastest access and parallel access without data hazards, on-chip registers are used exclusively to replace the BRAM or off-chip DRAM/SRAM to hold all the SpMxV data. On-chip DSP resources are fully utilized so as to ensure a maximum number of multipliers concurrently working. Jiasen Huang, Junyan Ren, Wenbo Yin, Lingli Wang |
FPT | 3 |