VLDB 2026 Research / reviewers in the wild / expert
Yuanlong Xiao
dblp:198/5622
· DBLP profile ↗
11ranked-venue papers
4as first author
8since 2021 · last 2024
0000-0002-3749-2729ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 4 first-author · 8 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | RapidStream IR: Infrastructure for FPGA High-Level Physical SynthesisabstractThe increasing complexity of large-scale FPGA accelerators poses significant challenges in achieving high performance while maintaining design productivity. High-level synthesis (HLS) has been adopted as a solution, but the mismatch between the high-level description and the physical layout often leads to suboptimal operating frequency. Although existing proposals for high-level physical synthesis, which use coarse-grained design partitioning, floorplanning, and pipelining to improve frequency, have gained traction, they lack a framework enabling (1) pipelining of real-world designs at arbitrary hierarchical levels, (2) integration of HLS blocks, vendor IPs, and handcrafted RTL designs, (3) portability to emerging new target FPGA devices, and (4) extensibility for the easy implementation of new design optimization tools. Jason Lau, Yuanlong Xiao, Yutong Xie 0011, Yuze Chi, Linghao Song, Shaojie Xiang, Michael Lo, Zhiru Zhang, Jason Cong, Licheng Guo |
ICCAD | 2 |
| 2024 | ExHiPR: Extended High-Level Partial Reconfiguration for Fast Incremental FPGA CompilationabstractPartial Reconfiguration (PR) is a key technique in the application design on modern FPGAs. However, current PR tools heavily rely on the developer to manually conduct PR module definition, floorplanning, and flow control at a low level. The existing PR tools do not consider High-Level-Synthesis languages either, which are of great interest to software developers. We propose HiPR, an open-source framework, to bridge the gap between HLS and PR. HiPR allows the developer to define partially reconfigurable C/C++ functions, instead of Verilog modules, to accelerate the FPGA incremental compilation and automate the flow from C/C++ to bitstreams. We use a lightweight Simulated Annealing floorplanner and show that it can produce high-quality PR floorplans an order of magnitude faster than analytic methods. By mapping Rosetta HLS benchmarks, we demonstrate that the incremental compilation can be accelerated by 3–10× compared with state-of-the-art Xilinx Vitis flow without performance loss, at the cost of 15–67% one-time overlay set-up time. Yuanlong Xiao, Dongjoon Park, Zeyu Jason Niu, Aditya Hota, André DeHon |
ACM Trans. Reconfigurable Technol. Syst. | 1 |
| 2023 | RapidStream 2.0: Automated Parallel Implementation of Latency-Insensitive FPGA Designs Through Partial ReconfigurationabstractField-programmable gate arrays (FPGAs) require a much longer compilation cycle than conventional computing platforms such as CPUs. In this article, we shorten the overall compilation time by co-optimizing the HLS compilation (C-to-RTL) and the back-end physical implementation (RTL-to-bitstream). We propose a split compilation approach based on the pipelining flexibility at the HLS level, which allows us to partition designs for parallel placement and routing. We outline a number of technical challenges and address them by breaking the conventional boundaries between different stages of the traditional FPGA tool flow and reorganizing them to achieve a fast end-to-end compilation. Our research produces RapidStream, a parallelized and physical-integrated compilation framework that takes in a latency-insensitive program in C/C++ and generates a fully placed and routed implementation. We present two approaches. The first approach (RapidStream 1.0) resolves inter-partition routing conflicts at the end when separate partitions are stitched together. When tested on the Xilinx U250 FPGA with a set of realistic HLS designs, RapidStream achieves a 5 to 7× reduction in compile time and up to 1.3× increase in frequency when compared with a commercial off-the-shelf toolchain. In addition, we provide preliminary results using a customized open-source router to reduce the compile time up to an order of magnitude in cases with lower performance requirements. The second approach (RapidStream 2.0) prevents routing conflicts using virtual pins. Testing on Xilinx U280 FPGA, we observed 5 to 7× compile time reduction and 1.3× frequency increase. Licheng Guo, Pongstorn Maidee, Chris Lavin, Eddie Hung, Wuxi Li, Jason Lau, Weikang Qiao, Yuze Chi, Linghao Song, Yuanlong Xiao, Alireza Kaviani, Zhiru Zhang, Jason Cong |
ACM Trans. Reconfigurable Technol. Syst. | 11 |
| 2022 | PLD: fast FPGA compilation to make reconfigurable acceleration compatible with modern incremental refinement software developmentabstractFPGA-based accelerators are demonstrating significant absolute performance and energy efficiency compared with general-purpose CPUs. While FPGA computations can now be described in standard, programming languages, like C, development for FPGAs accelerators remains tedious and inaccessible to modern software engineers. Slow compiles (potentially taking tens of hours) inhibit the rapid, incremental refinement of designs that is the hallmark of modern software engineering. To address this issue, we introduce separate compilation and linkage into the FPGA design flow, providing faster design turns more familiar to software development. To realize this flow, we provide abstractions, compiler options, and compiler flow that allow the same C source code to be compiled to processor cores in seconds and to FPGA regions in minutes, providing the missing -O0 and -O1 options familiar in software development. This raises the FPGA programming level and standardizes the programming experience, bringing FPGA-based accelerators into a more familiar software platform ecosystem for software engineers. Yuanlong Xiao, Eric Micallef, Andrew Butt, Matthew Hofmann, Marc Alston, Matthew Goldsmith, Andrew Merczynski-Hait, André DeHon |
ASPLOS | 1 |
| 2022 | HiPR: Fast, Incremental Custom Partial Reconfiguration for HLS DevelopersabstractHigh-Level Synthesis can abstract away low-level circuits design and improve the coding productivity. However, it also lengthens the compilation time, exacerbating an already slow edit-compile-debug loop that discourages the development and refinement of FPGA accelerators. Partial Reconfiguration techniques can decrease the compilation time by reducing and parallelizing the size of the compilation task. But defining partial reconfigurable regions also needs expert layout-level knowledge, making this approach inaccessible to the high-level developers that HLS is intended to attract. To address the problems above, we propose HiPR, a framework that bridges the gap between HLS and PR. With HiPR, users can define a C/C++ function (rather than a Verilog module) as partially reconfigurable without considering detailed low-level constraints. HiPR automates the PR floorplan and allows the users to define elastic resource requirements for the C-level PR function for quick further tuning later. By mapping the full set of Rosetta Benchmarks, we show HiPR can find the proper floorplan solution within seconds and generate the overlay for later tuning. Significantly, the incremental compilation time can be accelerated by 3-10X with no performance loss. Yuanlong Xiao, André DeHon |
FPGA | 1 |
| 2022 | HiPR: High-level Partial Reconfiguration for Fast Incremental FPGA CompilationabstractPartial Reconfiguration (PR) is a key technique in the design of modern FPGAs. However, current PR tools heavily rely on the developers to manually conduct PR module definition, floorplanning, and flow control at a low level. The existing PR tools do not consider High-Level-Synthesis languages either, which is of great interest to software developers. We propose HiPR, an open-source framework, to bridge the gap between HLS and PR. HiPR allows the developer to define partially reconfigurable C/C++ functions instead of Verilog modules, which benefits the FPGA incremental compilation and automates the flow from C/C++ to bitstreams. By mapping Rosetta HLS benchmarks, the incremental compilation can be accelerated by 3–10× compared with Xilinx Vitis normal flow without performance loss. Yuanlong Xiao, Aditya Hota, Dongjoon Park, André DeHon |
FPL | 1 |
| 2022 | Fast and Flexible FPGA Development using Hierarchical Partial ReconfigurationabstractTo address slow FPGA compilation, researchers have proposed to run separate compilations for smaller design components in parallel. This approach provides small pages on the FPGA, allowing users to separately generate partial designs on the pages and load them together. However, this method either forces users to manually decompose a design into small components that fit in small, fixed-sized pages or to use large, fixed-sized pages, reducing the potential compilation speedup benefits. This restriction often results in suboptimal decomposition of a design or diminishes productivity. To overcome these limitations, we utilize the recently supported Hierarchical Partial Reconfiguration technology from Xilinx to generate a more flexible framework. Depending on the size of user designs, our framework provides larger pages that are hierarchically recombined from multiple smaller pages. This flexibility relieves users of the burden to decompose the original design and offers more opportunities for design-space exploration. When tested on the ZCU102 embedded platform with the Rosetta HLS benchmarks, our system achieves$1.4-4.9\times$mapped application performance improvement compared to the system with fixed-sized pages while still compiling in 2–5 minutes$(2.2-5.3\times$faster than the vendor tool). Dongjoon Park, Yuanlong Xiao, André DeHon |
FPT | 2 |
| 2021 | HLS-Compatible, Embedded-Processor Stream Links
Eric Micallef, Yuanlong Xiao, André DeHon |
FCCM | 2 |
| 2019 | Transistor-Level Optimization Methodology for GRM FPGA Interconnect CircuitsabstractDue to its dominance in the whole chip area, power and delay, the FPGA interconnect circuits are traditionally designed by full custom design method. We present an automated transistor-level sizing optimization methodology for GRM FPGA interconnect circuits. In order to get accurate and effective predicated area, the commonly used diffusion sharing, transistor folding and inputs sharing are considered. To get the accurate and effective delay value, we avoid the inaccuracy of using linear device model, and use two schemes to build wire model: the wire within a circuit and the wire between interconnect circuits. To decrease simulation time, we propose multi-thread acceleration method and the Minimum-Final-Delay (MFD) algorithm which optimizes interconnect circuit as a whole, not separated part. For switch box optimization, MFD algorithm requires 38% less number of simulations than COFFE's algorithm. We use 65nm CMOS process technology for evaluation. For different optimization strategy, we emphasize either representative critical path delay or overall layout area. Compare to full-custom design method, the global cost can be decrease by 3% ~ 17%. For different transistor sizing combinations, 10/50 threads can be ~ 9X/15X faster than single-thread. Compared with the manual design method, our optimization methodology explores larger design space, and it decreases the circuit design optimization time from months to hours. Zhengjie Li, Yuanlong Xiao, Yunbing Pang, Jian Wang 0036, Jinmei Lai 0001 |
FPGA | 2 |
| 2019 | An Automatic Transistor-Level Tool for GRM FPGA Interconnect Circuits OptimizationabstractDue to its dominance in FPGA area and delay, the interconnect circuit is traditionally designed and optimized in full customized fashion, which can be extremely time consuming. In this paper, we propose an automated transistor-level sizing optimization method for the widely-used General Routing Matrix FPGA interconnect circuits with the following three features: (1) an area model that takes into account the commonly used diffusion sharing, transistor folding and inputs sharing techniques in order to have an accurate area predication; (2) an accurate and effective non-linear delay model that treats the wire within a circuit and the wire between interconnect circuits separately; (3) a multi-thread acceleration method and the Minimum-Final-Delay algorithm to speed-up the simulation. The global optimization cost is measured by the product of the interconnect circuit area and the representative path delay based on our proposed models. The cost reduces 10.9%, when we use 65nm CMOS process chip for evaluation. The simulation time for different transistor sizing combinations is improved by 9X and 15X when 10 and 50 threads are used, respectively, faster than single-thread. Compared with the manual design method, our proposed optimization approach explores a larger design space and reduces the optimization time from months to hours. Zhengjie Li, Yuanlong Xiao, Yunbing Pang, Jian Wang 0036, Jinmei Lai 0001 |
ACM Great Lakes Symposium on VLSI | 2 |
| 2018 | Case for Fast FPGA Compilation Using Partial ReconfigurationabstractDespite the FPGA's advantages over other hardware platforms, long compilation time prevents FPGA engineers from efficiently exploring the design space and discourages new users who want to quickly iterate for debugging. To reduce compilation time, this work adopts a divide-and-conquer approach using Partial Reconfiguration with a Packet-Switched Fat-Tree network. Partially reconfigured leaves in the packet-switched network are independent from each other and can be compiled separately in parallel. Also, when a minor fix is required to a bitstream, only the corresponding leaves need to be incrementally compiled. Preliminary experimental evidence from our work-in-progress effort illustrates how a 30 minute full-chip compile time can be reduced to 7 minutes. Dongjoon Park, Yuanlong Xiao, Nevo Magnezi, André DeHon |
FPL | 2 |