Yuyang Zou

dblp:243/3320 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
5since 2021 · last 2026
0009-0007-6185-5001ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 FESTAL: Dataflow Accelerator Synthesis Framework with Graph-Based Fusion for FPGA
abstract
High-Level Synthesis (HLS) provides a promising approach to design hardware at the software level. However, recent research efforts primarily focus on computational optimization while assuming perfect memory system. As a result, issues such as limited on-chip buffer capacity and high-latency off-chip memory access frequently become the performance bottlenecks. Dataflow architectures address this by enabling parallel task execution with direct on-chip communication, reducing the need for external memory access. However, dataflow implementation presents significant challenges, such as determining inter-task communication, and balancing compute and memory resources. A comprehensive modeling approach is necessary to fully leverage the benefits of dataflow for enhanced hardware performance. In this paper, we present Festal, a holistic FPGA synthesis framework that automatically generates efficient dataflow accelerators. Festal introduces a novel graph-based algorithm that systematically explores task fusion opportunities, optimizing inter-task communication patterns entirely on-chip and thereby reducing the need for off-chip memory access. By explicitly modeling memory constraints, the framework achieves a critical balance between computational workload and memory resources. Built on the MLIR infrastructure, Festal models memory management and streaming channels during code generation, providing an efficient solution for dataflow designs. Experimental results show that Festal achieves an average speedup of $2.06 \times$ on standard benchmark suites, outperforming the state-of-theart synthesis framework. For real-world applications, Festal demonstrates performance comparable to custom FPGA accelerators, underscoring its practical effectiveness.
Ruifan Xu, Yuyang Zou, Yun Liang 0001
ASP-DAC2
2026 Finding Reusable Instructions via E-Graph Anti-Unification
abstract
Domain-specific accelerators provide an increasingly valuable source of performance for diverse applications. Custom instructions that trigger the execution of dedicated hardware units or accelerators for common application functions become key building blocks in modern computing systems, balancing performance and cost effectiveness. RISC-V, the open and extensible instruction set architecture, is increasingly popularizing this trend. However, exploring custom instructions for an application domain remains challenging. Existing automated approaches suffer from poor reusability and limited performance. They can only identify or merge syntactically similar, scalar instruction sequences while missing semantically equivalent patterns.
Youwei Xiao, Chenyun Yin, Yitian Sun, Yuyang Zou, Yun Liang 0001
ASPLOS (2)4
2025 Clay: High-level ASIP Framework for Flexible Microarchitecture-Aware Instruction Customization
abstract
Application-specific instruction-set processors (ASIPs) pro-vide energy-efficient acceleration for embedded systems and IoT devices. The free and open RISC-V ISA promotes open-source ASIP solutions to accelerate diverse application domains. Existing ASIP tools generate hardware and software artifacts from high-level architecture description languages (ADLs), however, they only support the in-pipeline coupling strategy on specific processors. As a result, they suffer from two critical limitations: they restrict instruction extensions to stateless behavior, preventing hardware implementation of efficient control flow like loops, and they impose rigid microarchitectural constraints that limit register file and memory interactions. These restrictions create a fundamental bottleneck in application acceleration and prevent the efficient deployment of custom instructions across different processors.We introduce Clay, an open-source high-level ASIP framework that overcomes these limitations. Clay introduces a unified instruction extension interface that abstracts different coupling strategies as microarchitecture-agnostic actions and microarchitectural attributes. Clay ADL (CADL) combines the interface actions and high-level syntax to describe general instruction behavior, which can be stateful. We further propose a microarchitecture-aware synthesis flow that selects the best coupling strategy for each custom instruction and schedules the optimal implementation with microarchitectural attributes modeled as constraints. Our evaluation of diverse workloads demonstrates that Clay delivers substantial performance improvements across two RISC-V processors, our custom Clay-core and the open-source Rocket-core.
Weijie Peng, Youwei Xiao, Yuyang Zou, Zizhang Luo, Yun Liang 0001
ICCAD3
2025 Invited Paper: APS: Open-Source Hardware-Software Co-Design Framework for Agile Processor Specialization
abstract
APS is an open-source framework for agile hardware-software co-design of domain-specific processors. It provides both hardware synthesis and compiler infrastructure to facilitate the development of instruction extensions (ISAXs) for application acceleration. The framework proposes a unified instruction extension interface for seamless integration with diverse RISC-V SoC ecosystems. Based on the unified interface, APS introduces a cross-level architecture description language (CADL) for comprehensive instruction behavior specification, which is translated into a dynamic pipeline architecture through its synthesis flow. Besides, APS’s compiler infrastructure introduces a pattern-matching engine for the automated utilization of ISAXs in general programs. It also incorporates bitwidth-aware vectorization that leverages operand bitwidth information to reduce the overhead of calling ISAXs. We conduct case studies across multiple workloads, including cryptography, machine learning, and digital signal processing. With fewer than 175 lines of ISAX description, APS achieves 2.29× to 14.99× speedup for each case study, demonstrating APS’s practical productivity and acceleration capability. Overall, APS offers a complete, end-to-end methodology that significantly reduces the development cycle of ISAXs, making agile processor specialization practical to the research and open-source hardware communities.
Youwei Xiao, Yuyang Zou, Yitian Sun, Chenyun Yin, Ruifan Xu, Renze Chen, Yun Liang 0001
ICCAD2
2024 Optimization of Post-disaster Road Network Repair Strategy Considering Road Recovery Level
abstract
Existing studies on post-disaster road network repair strategies have ignored the impact of different levels of road damage and recovery on the efficiency of network repair. To solve this issue, this study integrates the construction material distribution (CMD) with the repair crew scheduling and routing problem (RCSRP), and determine the level of road recovery through the CMD. Then, a bi-level optimization model is proposed with network performance resilience and recovery speed resilience as the optimization objectives. A two-stage optimization algorithm (TSOA) composed of a genetic algorithm with an improved coding method (ICM-GA) and the Frank-Wolfe algorithm (FW) is then employed to solve this model. Finally, the effectiveness of the model and algorithm is validated through simulation experiments. The results indicate that, under given material and time constraints, the optimal repair strategy proposed in this study outperforms the repair strategy without considering road recovery level by 17.51% and 5.42% in terms of network performance resilience and recovery speed resilience, respectively. This demonstrates the positive significance of considering road recovery level in formulating road network repair strategies. Besides, this strategy can be applied to optimize the configuration of workstation count for different-scale networks.
Chen Mu, Shumei Liu, Jiapei Wang, Yuyang Zou
SMC5