EDBT 2026 Demo / reviewers in the wild / expert
Tuo Dai
dblp:220/8746
· DBLP profile ↗
7ranked-venue papers
4as first author
7since 2021 · last 2024
0000-0003-0684-7299ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 4 first-author · 7 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | G2PM: Performance Modeling for ACAP Architecture with Dual-Tiered Graph Representation LearningabstractPerformance estimation is a crucial component in the optimization processes of accelerator development on the Versal ACAP architecture. However, existing approaches present limitations - they are either too slow to facilitate efficient iterations, or they lack the necessary accuracy due to the specific AIE array architecture and two-level programming model of Versal ACAP. To tackle this challenge, we propose G2PM, a performance modeling technique based on a hierarchical graph representation centered on the AIE array. More specifically, we employ a hierarchical graph neural network to identify features of both kernel programs and dataflow programs, taking into account the hardware and software characteristics of the Versal ACAP architecture. In our evaluations, our method demonstrates significant improvements, achieving a mean error rate of less than 1.6% and providing a speed-up factor of 4165X compared to the simulation-based method. Tuo Dai, Bizhao Shi, Guojie Luo |
DAC | 1 |
| 2024 | PT-Map: Efficient Program Transformation Optimization for CGRA MappingabstractCoarse-Grained Reconfigurable Array (CGRA) is a parallel architecture providing high energy efficiency and spatial-temporal re-configurability. Beyond loop scheduling for throughput optimization, program transformation is also crucial in CGRA mapping to optimize overall performance and efficiency. However, existing studies on program transformation optimization face challenges in exploring the transformation space systematically and evaluating candidates efficiently, leading to sub-optimal results. To tackle these challenges, this paper introduces PT-Map, an efficient program transformation optimization framework for CGRA mapping. PT-Map defines a comprehensive transformation space and employs a CGRA-specialized top-down exploration approach. It also incorporates a bottom-up evaluation scheme using architectural parameters and a graph neural network-based predictive model. Experiments demonstrate that PT-Map achieves up to 2.95X/1.80X speedups and 59.0%/23.2% energy-delay-product (EDP) reductions over the state-of-the-art approaches MapZero and PBP, respectively. Bizhao Shi, Tuo Dai, Jiaxi Zhang 0001, Xuechao Wei, Guojie Luo |
DAC | 2 |
| 2024 | WideSA: A High Array Utilization Mapping Scheme for Uniform Recurrences on ACAPabstractThe Versal Adaptive Compute Acceleration Platform (ACAP) is a new architecture that combines AI Engines (AIEs) with reconfigurable fabric. This architecture offers significant acceleration potential for uniform recurrences in various domains, such as deep learning, high-performance computation, and signal processing. However, efficiently mapping these computations onto the Versal ACAP architecture while achieving high utilization of AIEs poses a challenge. To address this issue, we propose a mapping scheme called WideSA, which aims to accelerate uniform recurrences on the Versal ACAP architecture by leveraging the features of both the hardware and the computations. Considering the array architecture of AIEs, our approach utilizes space-time transformations based on the polyhedral model to generate legally optimized systolic array mappings. Concurrently, we have developed a routing-aware PLIO assignment algorithm tailored for communication on the AlE array, and the algorithm aims at successful compilation while maximizing array utilization. Furthermore, we introduce an automatic mapping framework. This framework is designed to generate the corresponding executable code for uniform recurrences, which encompasses the AlE kernel program, programmable logic bitstreams, and the host program. The experimental results validate the effectiveness of our mapping scheme. Specifically, when applying our scheme to matrix multiplication computations on the VCK5000 board, we achieve a throughput of 4.15TOPS on float data type, which is 1.11 x higher compared to the state-of-the-art accelerator on the Versal ACAP architecture. Tuo Dai, Bizhao Shi, Guojie Luo |
DATE | 1 |
| 2024 | ImageMap: Enabling Efficient Mapping from Image Processing DSL to CGRA
Bizhao Shi, Tuo Dai, Sunan Zou, Xinming Wei, Guojie Luo |
Euro-Par (1) | 2 |
| 2024 | Weave: Abstraction and Integration Flow for Accelerators of Generated ModulesabstractIn modern times, domain-specific accelerators require numerous functional components to execute complex applications in a particular domain. To ensure efficient development, the conventional approach involves decomposing, implementing, and integrating modules. Over the past decade, the generator-based method has proven to enhance the productivity of module implementation. However, current abstractions pose challenges for integrating modules implemented by generators, due to implicit interface definitions, nonunified performance modeling, and fragmented memory management. These limitations result in a lower productivity of the integration process and decreased performance of the integrated accelerators. To overcome these drawbacks, we propose Weave, an abstraction for integrating generated modules and an agile design flow for domain-specific accelerators. The Weave abstraction guides module implementation and integration with a unified performance model and memory management. Furthermore, the Weave integration flow, consisting of generation, selection, and integration phrases, enables optimization of the performance of the integrated accelerator with a design space exploration algorithm and hierarchical memory management. In the experiments, the accelerator developed by Weave achieves$1.93\times $higher performance in the deep learning domain compared to an open-source accelerator, and the integrated accelerator maintains performance for various applications with different memory access patterns. Tuo Dai, Bizhao Shi, Guojie Luo |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2023 | Weave: Abstraction for Accelerator Integration of Generated ModulesabstractAs domain-specific accelerators demand multiple functional components for complex applications in a domain, the conventional wisdom for effective development involves module decomposition, module implementation, and module integration. In the recent decade, the generator-based design methodology improves the productivity of module implementation. However, with the guidance of current abstractions, it is difficult to integrate modules implemented by generators because of implicit interface definition, non-unified performance modeling, and fragmented memory management. These disadvantages cause low productivity of the integration flow and low performance of the integrated accelerators. Tuo Dai, Bizhao Shi, Guojie Luo |
FPGA | 1 |
| 2023 | An Intermediate-Centric Dataflow for Transposed Convolution Acceleration on FPGAabstractTransposed convolution has been prevailing in convolutional neural networks (CNNs), playing an important role in multiple scenarios such as image segmentation and back-propagation process of training CNNs. This mainly benefits from the ability to up-sample the input feature maps by interpolating new information from the input feature pixels. However, the backward-stencil computation constrains its performance and hindered its wide application in diverse platforms. Moreover, in contrast to the efforts on accelerating the convolution, there is a rare investigation on the acceleration of transposed convolution that is identically compute-intensive as the former. For acceleration of transposed convolution, we propose an intermediate-centric dataflow scheme, in which we decouple the generation of the intermediate patch from its further process, aim at efficiently performing the backward-stencil computation . The intermediate-centric dataflow breaks the transposed convolution into several phases/stages, achieving feeding the input feature maps and performing the backward-stencil computation in a pipelining manner. It also provides four-degree computation parallelism and efficient data reuse of input feature maps/weights. Furthermore, we also theoretically analyze the irregular data dependence leveraging the polyhedral model, which constrains the parallel computing of transposed convolution. Additionally, we devise an optimization problem to explore the design space and automatically generate the optimal design configurations for different transposed convolutional layers and hardware platforms. By selecting the representative transposed convolutional layers from DCGAN, FSRCNN, and FCN, we generate the corresponding accelerator arrays of intermediate-centric dataflow on the Xilinx Alveo U200 platform and reach the performance of 3.92 TOPS, 2.72 TOPS, and 4.76 TOPS, respectively. Zhengzheng Ma, Tuo Dai, Xuechao Wei, Guojie Luo |
ACM Trans. Embed. Comput. Syst. | 2 |