EDBT 2026 Demo / reviewers in the wild / expert
Youwei Xiao
dblp:336/0763
· DBLP profile ↗
8ranked-venue papers
4as first author
8since 2021 · last 2026
0000-0002-3636-3685ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 4 first-author · 8 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Finding Reusable Instructions via E-Graph Anti-UnificationabstractDomain-specific accelerators provide an increasingly valuable source of performance for diverse applications. Custom instructions that trigger the execution of dedicated hardware units or accelerators for common application functions become key building blocks in modern computing systems, balancing performance and cost effectiveness. RISC-V, the open and extensible instruction set architecture, is increasingly popularizing this trend. However, exploring custom instructions for an application domain remains challenging. Existing automated approaches suffer from poor reusability and limited performance. They can only identify or merge syntactically similar, scalar instruction sequences while missing semantically equivalent patterns. Youwei Xiao, Chenyun Yin, Yitian Sun, Yuyang Zou, Yun Liang 0001 |
ASPLOS (2) | 1 |
| 2025 | Cayman: Custom Accelerator Generation with Control Flow and Data Access OptimizationabstractCustom accelerators enhance System-on-Chips’ performance through hardware specialization. High-level synthesis (HLS) can automatically synthesize accelerators for given kernels but requires manual selection and extraction of kernels from applications. This paper proposes Cayman, the first end-to-end framework to synthesize high-performance custom accelerators with both control flow and data access optimization. Cayman automatically selects kernels for hardware acceleration based on a hierarchical program representation, which captures kernel candidates with general control flows. Besides, Cayman optimizes accelerators with specialized processor-accelerator interfaces for data access acceleration. Cayman further introduces a novel accelerator merging mechanism to synthesize reusable accelerators. Experiments on various benchmarks demonstrate that Cayman outperforms two state-of-the-art frameworks by $8.0 \times$ and $14.4 \times$. Youwei Xiao, Fan Cui, Zizhang Luo, Weijie Peng, Yun Liang 0001 |
DAC | 1 |
| 2025 | An Empirical Comparision of LLM-based Hardware Design and High-level SynthesisabstractField-Programmable Gate Arrays (FPGAs) are increasingly used for accelerating diverse applications due to their reconfigurability and ability to implement custom hardware architectures. However, programming FPGAs remains challenging, traditionally relying on low-level Hardware Description Languages (HDLs) like Verilog, which are intricate and time-consuming. High-Level Synthesis (HLS) tools, such as Vitis HLS, have emerged to address these issues by allowing hardware functionality description in high-level languages like C/C++, but they come with their own limitations, including less efficient hardware implementations, delay overhead caused by conservative scheduling strategies, and unpredictable solutions due to semantic differences between software and hardware. Fan Cui, Youwei Xiao, Kexing Zhou, Yun Liang 0001 |
FPGA | 2 |
| 2025 | Clay: High-level ASIP Framework for Flexible Microarchitecture-Aware Instruction CustomizationabstractApplication-specific instruction-set processors (ASIPs) pro-vide energy-efficient acceleration for embedded systems and IoT devices. The free and open RISC-V ISA promotes open-source ASIP solutions to accelerate diverse application domains. Existing ASIP tools generate hardware and software artifacts from high-level architecture description languages (ADLs), however, they only support the in-pipeline coupling strategy on specific processors. As a result, they suffer from two critical limitations: they restrict instruction extensions to stateless behavior, preventing hardware implementation of efficient control flow like loops, and they impose rigid microarchitectural constraints that limit register file and memory interactions. These restrictions create a fundamental bottleneck in application acceleration and prevent the efficient deployment of custom instructions across different processors.We introduce Clay, an open-source high-level ASIP framework that overcomes these limitations. Clay introduces a unified instruction extension interface that abstracts different coupling strategies as microarchitecture-agnostic actions and microarchitectural attributes. Clay ADL (CADL) combines the interface actions and high-level syntax to describe general instruction behavior, which can be stateful. We further propose a microarchitecture-aware synthesis flow that selects the best coupling strategy for each custom instruction and schedules the optimal implementation with microarchitectural attributes modeled as constraints. Our evaluation of diverse workloads demonstrates that Clay delivers substantial performance improvements across two RISC-V processors, our custom Clay-core and the open-source Rocket-core. Weijie Peng, Youwei Xiao, Yuyang Zou, Zizhang Luo, Yun Liang 0001 |
ICCAD | 2 |
| 2025 | Invited Paper: APS: Open-Source Hardware-Software Co-Design Framework for Agile Processor SpecializationabstractAPS is an open-source framework for agile hardware-software co-design of domain-specific processors. It provides both hardware synthesis and compiler infrastructure to facilitate the development of instruction extensions (ISAXs) for application acceleration. The framework proposes a unified instruction extension interface for seamless integration with diverse RISC-V SoC ecosystems. Based on the unified interface, APS introduces a cross-level architecture description language (CADL) for comprehensive instruction behavior specification, which is translated into a dynamic pipeline architecture through its synthesis flow. Besides, APS’s compiler infrastructure introduces a pattern-matching engine for the automated utilization of ISAXs in general programs. It also incorporates bitwidth-aware vectorization that leverages operand bitwidth information to reduce the overhead of calling ISAXs. We conduct case studies across multiple workloads, including cryptography, machine learning, and digital signal processing. With fewer than 175 lines of ISAX description, APS achieves 2.29× to 14.99× speedup for each case study, demonstrating APS’s practical productivity and acceleration capability. Overall, APS offers a complete, end-to-end methodology that significantly reduces the development cycle of ISAXs, making agile processor specialization practical to the research and open-source hardware communities. Youwei Xiao, Yuyang Zou, Yitian Sun, Chenyun Yin, Ruifan Xu, Renze Chen, Yun Liang 0001 |
ICCAD | 1 |
| 2024 | Cement: Streamlining FPGA Hardware Design with Cycle-Deterministic eHDL and SynthesisabstractField-programmable gate arrays (FPGAs) provide opportunities for adopting cutting-edge microarchitectural technologies to accelerate emerging applications. However, it remains challenging to program FPGAs. On one hand, hardware description languages (HDLs), although lauded for their ability to provide circuit representations that closely mimic the inherent hardware structures, have been criticized for their inherent shortcomings, including low-level programming and poor productivity. On the other hand, high-level synthesis (HLS) attempts to raise the abstraction level of hardware design to the software domain. However, it often results in unpredictable solutions due to semantic difference between software and hardware. Furthermore, domain-specific languages (DSLs) tailored for FPGA programming have their own set of limitations, particularly in terms of expressiveness and flexibility. In this work, we introduce a novel hardware design framework named Cement \xspace, which encompasses the embedded HDL (eHDL) CmtHDL \xspace and the compiler CmtC \xspace, providing a better programming framework for FPGA. CmtHDL \xspace introduces event-based procedural specification alongside RTL description, empowering designers to describe hardware productively at a higher level of abstraction while maintaining cycle-deterministic behavior. CmtC \xspace provides a comprehensive compilation workflow that includes analyzing the timing behavior of the hardware and conducting synthesis to yield solutions with anticipated performance for FPGAs. Experiments show that Cement \xspace provides comparable productivity, but offers 1.41\texttimes-3.49\texttimes\xspace speedup, and saves 23%-82% resources compared to existing HLS or DSL tools. The practical significance of Cement \xspace is further validated through a case study of designing real-world FPGA-based accelerators. Youwei Xiao, Zizhang Luo, Kexing Zhou, Yun Liang 0001 |
FPGA | 1 |
| 2024 | OriGen: Enhancing RTL Code Generation with Code-to-Code Augmentation and Self-ReflectionabstractRecent studies have demonstrated the significant potential of Large Language Models (LLMs) in generating Register Transfer Level (RTL) code, with notable advancements showcased by commercial models such as GPT-4 and Claude3-Opus. However, these proprietary LLMs often raise concerns regarding privacy and security. While open-source LLMs offer solutions to these concerns, they typically underperform commercial models in RTL code generation tasks, primarily due to the scarcity of high-quality open-source RTL datasets. To address this challenge, we introduce OriGen, a fully open-source framework that incorporates self-reflection capabilities and a novel dataset augmentation methodology for generating high-quality, large-scale RTL code. Our approach employs a code-to-code augmentation technique to enhance the quality of open-source RTL code datasets. Furthermore, OriGen can rectify syntactic errors through a self-reflection process that leverages compiler feedback. Fan Cui, Chenyang Yin, Kexing Zhou, Youwei Xiao, Guangyu Sun 0003, Qiang Xu 0001, Qipeng Guo, Yun Liang 0001, Xingcheng Zhang, Demin Song, Dahua Lin |
ICCAD | 4 |
| 2022 | HECTOR: A Multi-Level Intermediate Representation for Hardware Synthesis MethodologiesabstractHardware synthesis requires a complicated process to generate synthesizable register transfer level (RTL) code. High-level synthesis tools can automatically transform a high-level description into hardware design, while hardware generators adopt domain specific languages and synthesis flows for specific applications. The implementation of these tools generally requires substantial engineering efforts due to RTL's weak expressivity and low level of abstraction. Furthermore, different synthesis tools adopt different levels of intermediate representations (IR) and transformations. A unified IR obviously is a good way to lower the engineering cost and get competitive hardware design rapidly by exploring different synthesis methodologies. Ruifan Xu, Youwei Xiao, Yun Liang 0001 |
ICCAD | 2 |