Ruifan Xu

dblp:336/0854 · DBLP profile ↗
← Back
8ranked-venue papers
5as first author
8since 2021 · last 2026
0000-0001-7241-9802ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 5 first-author · 8 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 FESTAL: Dataflow Accelerator Synthesis Framework with Graph-Based Fusion for FPGA
abstract
High-Level Synthesis (HLS) provides a promising approach to design hardware at the software level. However, recent research efforts primarily focus on computational optimization while assuming perfect memory system. As a result, issues such as limited on-chip buffer capacity and high-latency off-chip memory access frequently become the performance bottlenecks. Dataflow architectures address this by enabling parallel task execution with direct on-chip communication, reducing the need for external memory access. However, dataflow implementation presents significant challenges, such as determining inter-task communication, and balancing compute and memory resources. A comprehensive modeling approach is necessary to fully leverage the benefits of dataflow for enhanced hardware performance. In this paper, we present Festal, a holistic FPGA synthesis framework that automatically generates efficient dataflow accelerators. Festal introduces a novel graph-based algorithm that systematically explores task fusion opportunities, optimizing inter-task communication patterns entirely on-chip and thereby reducing the need for off-chip memory access. By explicitly modeling memory constraints, the framework achieves a critical balance between computational workload and memory resources. Built on the MLIR infrastructure, Festal models memory management and streaming channels during code generation, providing an efficient solution for dataflow designs. Experimental results show that Festal achieves an average speedup of $2.06 \times$ on standard benchmark suites, outperforming the state-of-theart synthesis framework. For real-world applications, Festal demonstrates performance comparable to custom FPGA accelerators, underscoring its practical effectiveness.
Ruifan Xu, Yuyang Zou, Yun Liang 0001
ASP-DAC1
2026 Graph.hls: A Compiler Framework for Composable Graph Accelerator Design
Feiyang Wu, Xuxiao Yang, Zhuohang Bian, Ruifan Xu, Yun Liang 0001, Youwei Zhuo
ISCA5
2026 PipeComm: Maximizing Link Utilization Through Pipeline-Aware Collective Communication Synthesis
Ruifan Xu, Yuze Luo, Yuhao Meng, Size Zheng 0001, Meng Li 0004, Yun Liang 0001
ISCA1
2025 A Unified Synthesis Framework for Dataflow Accelerators Through Multi-level Software and Hardware Intermediate Representations
Xiaochen Hao, Ruifan Xu, Yun Liang 0001
APPT2
2025 Invited Paper: APS: Open-Source Hardware-Software Co-Design Framework for Agile Processor Specialization
abstract
APS is an open-source framework for agile hardware-software co-design of domain-specific processors. It provides both hardware synthesis and compiler infrastructure to facilitate the development of instruction extensions (ISAXs) for application acceleration. The framework proposes a unified instruction extension interface for seamless integration with diverse RISC-V SoC ecosystems. Based on the unified interface, APS introduces a cross-level architecture description language (CADL) for comprehensive instruction behavior specification, which is translated into a dynamic pipeline architecture through its synthesis flow. Besides, APS’s compiler infrastructure introduces a pattern-matching engine for the automated utilization of ISAXs in general programs. It also incorporates bitwidth-aware vectorization that leverages operand bitwidth information to reduce the overhead of calling ISAXs. We conduct case studies across multiple workloads, including cryptography, machine learning, and digital signal processing. With fewer than 175 lines of ISAX description, APS achieves 2.29× to 14.99× speedup for each case study, demonstrating APS’s practical productivity and acceleration capability. Overall, APS offers a complete, end-to-end methodology that significantly reduces the development cycle of ISAXs, making agile processor specialization practical to the research and open-source hardware communities.
Youwei Xiao, Yuyang Zou, Yitian Sun, Chenyun Yin, Ruifan Xu, Renze Chen, Yun Liang 0001
ICCAD7
2024 Hermes: Enhancing Extensibility in High-Level Synthesis through Multi-Level IRs
abstract
Field-Programmable Gate Arrays (FPGAs) have become integral components in diverse application domains due to their adaptability and reconfigurable capabilities. However, the intricate nature of FPGA programming poses significant challenges in efficiently harnessing their potential. High-Level Synthesis (HLS) has emerged as a promising approach, simplifying FPGA development by automating Register Transfer Level (RTL) code generation from high-level programming models. Yet, existing HLS tools often lack extensibility, hindering optimization and design portability.
Ruifan Xu, Yun Liang 0001
FPGA1
2024 Hestia: An Efficient Cross-Level Debugger for High-Level Synthesis
abstract
High-level synthesis (HLS) offers an opportunity to design hardware at the software level, which automatically trans-forms high-level specifications into RTL designs. However, HLS compilers are often considered complex black-box procedures, lacking transparency for designers and hindering the debugging process. Programmers often rely on simulating the HLS design to comprehend the behavior of the generated hardware. RTL simulation, the prevalent hardware debugging method, is time-consuming and inundates designers with excessive details when applied to HLS designs. Conversely, software-level simulation is fast but does not model hardware-specific details. The debug-ging challenge primarily stems from the semantic gap between software descriptions and RTL implementations. In this paper, we present Hestia, an efficient cross-level debugger enabling debugging HLS designs at different abstraction levels. Hestia provides a multi-level interpreter, aiding in debugging various issues in the HLS procedure with less hardware details and lower time costs. With an equivalent mapping across different levels, Hestia facilitates bug identifi-cation and localization, providing breakpoints and stepping at multiple granularities. We demonstrate the effectiveness of Hestia from three aspects: simulation efficiency, debugging capability, and scalability. Experimental results show that Hestia achieves significant simulation speedup compared to RTL simulators and prior work. The experiment of a case study also illustrates how Hestia helps find and localize bugs easily.
Ruifan Xu, Yibo Lin, Runsheng Wang, Ru Huang 0001, Yun Liang 0001
MICRO1
2022 HECTOR: A Multi-Level Intermediate Representation for Hardware Synthesis Methodologies
abstract
Hardware synthesis requires a complicated process to generate synthesizable register transfer level (RTL) code. High-level synthesis tools can automatically transform a high-level description into hardware design, while hardware generators adopt domain specific languages and synthesis flows for specific applications. The implementation of these tools generally requires substantial engineering efforts due to RTL's weak expressivity and low level of abstraction. Furthermore, different synthesis tools adopt different levels of intermediate representations (IR) and transformations. A unified IR obviously is a good way to lower the engineering cost and get competitive hardware design rapidly by exploring different synthesis methodologies.
Ruifan Xu, Youwei Xiao, Yun Liang 0001
ICCAD1