VLDB 2026 Research / reviewers in the wild / expert
Zizhang Luo
dblp:259/4093
· DBLP profile ↗
11ranked-venue papers
3as first author
10since 2021 · last 2025
0000-0002-7276-2317ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 2 first-author · 10 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Cayman: Custom Accelerator Generation with Control Flow and Data Access OptimizationabstractCustom accelerators enhance System-on-Chips’ performance through hardware specialization. High-level synthesis (HLS) can automatically synthesize accelerators for given kernels but requires manual selection and extraction of kernels from applications. This paper proposes Cayman, the first end-to-end framework to synthesize high-performance custom accelerators with both control flow and data access optimization. Cayman automatically selects kernels for hardware acceleration based on a hierarchical program representation, which captures kernel candidates with general control flows. Besides, Cayman optimizes accelerators with specialized processor-accelerator interfaces for data access acceleration. Cayman further introduces a novel accelerator merging mechanism to synthesize reusable accelerators. Experiments on various benchmarks demonstrate that Cayman outperforms two state-of-the-art frameworks by $8.0 \times$ and $14.4 \times$. Youwei Xiao, Fan Cui, Zizhang Luo, Weijie Peng, Yun Liang 0001 |
DAC | 3 |
| 2025 | Clay: High-level ASIP Framework for Flexible Microarchitecture-Aware Instruction CustomizationabstractApplication-specific instruction-set processors (ASIPs) pro-vide energy-efficient acceleration for embedded systems and IoT devices. The free and open RISC-V ISA promotes open-source ASIP solutions to accelerate diverse application domains. Existing ASIP tools generate hardware and software artifacts from high-level architecture description languages (ADLs), however, they only support the in-pipeline coupling strategy on specific processors. As a result, they suffer from two critical limitations: they restrict instruction extensions to stateless behavior, preventing hardware implementation of efficient control flow like loops, and they impose rigid microarchitectural constraints that limit register file and memory interactions. These restrictions create a fundamental bottleneck in application acceleration and prevent the efficient deployment of custom instructions across different processors.We introduce Clay, an open-source high-level ASIP framework that overcomes these limitations. Clay introduces a unified instruction extension interface that abstracts different coupling strategies as microarchitecture-agnostic actions and microarchitectural attributes. Clay ADL (CADL) combines the interface actions and high-level syntax to describe general instruction behavior, which can be stateful. We further propose a microarchitecture-aware synthesis flow that selects the best coupling strategy for each custom instruction and schedules the optimal implementation with microarchitectural attributes modeled as constraints. Our evaluation of diverse workloads demonstrates that Clay delivers substantial performance improvements across two RISC-V processors, our custom Clay-core and the open-source Rocket-core. Weijie Peng, Youwei Xiao, Yuyang Zou, Zizhang Luo, Yun Liang 0001 |
ICCAD | 4 |
| 2024 | Cement: Streamlining FPGA Hardware Design with Cycle-Deterministic eHDL and SynthesisabstractField-programmable gate arrays (FPGAs) provide opportunities for adopting cutting-edge microarchitectural technologies to accelerate emerging applications. However, it remains challenging to program FPGAs. On one hand, hardware description languages (HDLs), although lauded for their ability to provide circuit representations that closely mimic the inherent hardware structures, have been criticized for their inherent shortcomings, including low-level programming and poor productivity. On the other hand, high-level synthesis (HLS) attempts to raise the abstraction level of hardware design to the software domain. However, it often results in unpredictable solutions due to semantic difference between software and hardware. Furthermore, domain-specific languages (DSLs) tailored for FPGA programming have their own set of limitations, particularly in terms of expressiveness and flexibility. In this work, we introduce a novel hardware design framework named Cement \xspace, which encompasses the embedded HDL (eHDL) CmtHDL \xspace and the compiler CmtC \xspace, providing a better programming framework for FPGA. CmtHDL \xspace introduces event-based procedural specification alongside RTL description, empowering designers to describe hardware productively at a higher level of abstraction while maintaining cycle-deterministic behavior. CmtC \xspace provides a comprehensive compilation workflow that includes analyzing the timing behavior of the hardware and conducting synthesis to yield solutions with anticipated performance for FPGAs. Experiments show that Cement \xspace provides comparable productivity, but offers 1.41\texttimes-3.49\texttimes\xspace speedup, and saves 23%-82% resources compared to existing HLS or DSL tools. The practical significance of Cement \xspace is further validated through a case study of designing real-world FPGA-based accelerators. Youwei Xiao, Zizhang Luo, Kexing Zhou, Yun Liang 0001 |
FPGA | 2 |
| 2024 | Rubick: A Unified Infrastructure for Analyzing, Exploring, and Implementing Spatial Architectures via Dataflow DecompositionabstractThe fast-growing tensor applications expose tremendous dataflow alternatives when implemented on spatial architectures that feature large PE arrays and abundant interconnection resources. Prior works develop various notations and performance models for dataflows. Though these notations are very useful for understanding the reuse, bandwidth, and performance of dataflows, they do not define the underlying hardware implementation. Due to the semantic gap, analysis based on these notations cannot capture the detailed architectural features between different dataflows, leading to inefficient design space exploration and suboptimal designs. To address these issues, we propose Rubick, a unified infrastructure for analyzing, exploring, and implementing spatial architectures. The main innovation of Rubick is it decomposes the dataflow into two low-level intermediate representations: access entry and data layout. Access entry specifies how data enter into the PE arrays from memory, while data layout specifies how data are arranged and accessed. These two representations allow us to infer the hardware implementation details such as PE interconnection and memory structure, which are amenable for structural analysis and systematic exploration. Based on this decomposition analysis, Rubick provides opportunities for micro-architecture optimization and efficient design space exploration. Our experiments demonstrate that Rubick can reduce 82.4% of wire resources with only a 2.7% latency increase by optimizing access entry IR, and achieve 70.8% memory overhead reduction by optimizing data layout IR. Rubick also accelerates the DSE time of dataflows by up to 1.1×105X, saving the time from several days to minutes. The source code of Rubick is publically available on (https://link-omitted-for-blind-review). Liqiang Lu, Zizhang Luo, Size Zheng 0001, Jieming Yin, Jason Cong, Yun Liang 0001, Jianwei Yin |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2023 | Rubick: A Synthesis Framework for Spatial Architectures via Dataflow DecompositionabstractDataflows are critical for spatial architectures designed for tensor applications. Prior works develop various notations and hardware generation frameworks for dataflows. However, due to the semantic gap between notations and low-level details, analysis based on these notations cannot capture the detailed architectural features between different dataflows, so these works failed to provide architectural optimization and efficient design space exploration (DSE) at the same time.We propose Rubick, a synthesis framework for spatial architecture. Rubick decomposes the dataflow into two low-level intermediate representations including access entry and data layout. Access entry specifies how data enter into the PE arrays from memory, while data layout specifies how data are arranged and accessed. Based on this decomposition, Rubick provides efficient DSE and generates optimized hardware. Experiments show that the DSE time is accelerated by up to 1.1×105X and performance on FPGA is improved by 13%. Zizhang Luo, Liqiang Lu, Size Zheng 0001, Jieming Yin, Jason Cong, Jianwei Yin, Yun Liang 0001 |
DAC | 1 |
| 2023 | Calabash: Accelerating Attention Using a Systolic Array Chain on FPGAsabstractIn recent years, attention mechanism has achieved remarkable performance in natural language processing and computer vision applications, at the expense of high computation cost. FPGAs have been demonstrated to be an effective hardware platform for various AI applications. However, the attention mechanism involves complex data dependency, which makes FPGA acceleration difficult. In this paper, we propose Calabash, an FPGA accelerator for attention-based applications. We design a chain of two systolic arrays, applying the same dataflow. Then, we design two scheduling techniques for different matrices to ensure the intermediate matrix can be cached in the on-chip memory. Finally, we develop analytical models for resource utilization estimation, workload balancing, and latency prediction to guide design space exploration. Experiments show that Calabash achieves 1.76 TOP/s, 1.06 TOP/s on Xilinx VU9P and ZCU102 platforms, yielding an average 50.1X and 3.94X energy-efficiency improvement compared with CPU and GPU, respectively. Zizhang Luo, Liqiang Lu, Yicheng Jin, Liancheng Jia, Yun Liang 0001 |
FPL | 1 |
| 2023 | Automatic Generation of Spatial Accelerator for Tensor AlgebraabstractTensor algebra finds applications in various domains including machine learning applications, data analytics and others. Spatial hardware accelerators are widely used to boost the performance of tensor algebra applications. It has a complex hardware architecture and rich design space. Prior approaches based on manual implementation lead to low programming productivity, making it hard to explore the large design space. In this paper, we propose Tensorlib, a framework for generating spatial hardware accelerators for tensor algebra applications. Tensorlib is motivated by the observation that, tensor dataflows can be expressed with linear transformations, and they share common hardware modules which can be reused across different designs. Tensorlib first uses Space-Time Transformation to explore different dataflows, which can compactly represent the hardware dataflow using a transformation matrix. Next, we identify the common structures of different dataflows and build parameterized hardware module templates. Our generation framework can select the needed hardware modules for each dataflow, connect the modules using a specified interconnection pattern, and automatically generate the complete hardware accelerator design. Tensorlib remarkably improves the productivity for the development and optimization of spatial hardware architecture, providing a rich design space with trade-offs in performance, area, and power. Experiments show that Tensorlib can automatically generate hardware designs with different dataflows for a variety of tensor algebra programs. Tensorlib can achieve 318 MHz frequency and 786 GFLOP/s throughput for matrix multiplication kernel on Xilinx VU9P FPGA, which outperforms the state-of-the-art generators. Liancheng Jia, Zizhang Luo, Liqiang Lu, Yun Liang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2021 | TensorLib: A Spatial Accelerator Generation Framework for Tensor AlgebraabstractTensor algebra finds applications in various domains, and these applications, especially when accelerated on spatial hardware accelerators, can deliver high performance and low power. Spatial hardware accelerator exhibits complex design space. Prior approaches based on manual implementation lead to low programming productivity, rendering thorough design space exploration impossible. In this paper, we propose TensorLib, a framework for generating spatial hardware accelerator for tensor algebra applications. TensorLib is motivated by the observation that, different dataflows share common hardware modules, which can be reused across different designs. To build such a framework, TensorLib first uses Space-Time Transformation to explore different dataflows, which can compactly represent the hardware dataflow using a simple transformation matrix. Next, we identify the common structures of different dataflows and build parameterized hardware module templates with Chisel. Our generation framework can select the needed hardware modules for each dataflow, connect the modules using a specified interconnection pattern, and automatically generate the complete hardware accelerator design. TensorLib remarkably improves the productivity for the development and optimization of spatial hardware architecture, providing a rich design space with tradeoffs in performance, area, and power. Experiments show that TensorLib can automatically generate hardware designs with different dataflows and achieve 21% performance improvement on FPGA compared to the state-of-the-arts. Liancheng Jia, Zizhang Luo, Liqiang Lu, Yun Liang 0001 |
DAC | 2 |
| 2021 | TENET: A Framework for Modeling Tensor Dataflow Based on Relation-centric NotationabstractAccelerating tensor applications on spatial architectures provides high performance and energy-efficiency, but requires accurate performance models for evaluating various dataflow alternatives. Such modeling relies on the notation of tensor dataflow and the formulation of performance metrics. Recent proposed compute-centric and data-centric notations describe the dataflow using imperative directives. However, these two notations are less expressive and thus lead to limited optimization opportunities and inaccurate performance models.In this paper, we propose a framework TENET that models hardware dataflow of tensor applications. We start by introducing a relation-centric notation, which formally describes the hardware dataflow for tensor computation. The relation-centric notation specifies the hardware dataflow, PE interconnection, and data assignment in a uniform manner using relations. The relation-centric notation is more expressive than the compute-centric and data-centric notations by using more sophisticated affine transformations. Another advantage of relation-centric notation is that it inherently supports accurate metrics estimation, including data reuse, bandwidth, latency, and energy. TENET computes each performance metric by counting the relations using integer set structures and operators. Overall, TENET achieves 37.4% and 51.4% latency reduction for CONV and GEMM kernels compared with the state-of-the-art data-centric notation by identifying more sophisticated hardware dataflows. Liqiang Lu, Naiqing Guan, Yuyue Wang 0001, Liancheng Jia, Zizhang Luo, Jieming Yin, Jason Cong, Yun Liang 0001 |
ISCA | 5 |
| 2021 | Sanger: A Co-Design Framework for Enabling Sparse Attention using Reconfigurable ArchitectureabstractIn recent years, attention-based models have achieved impressive performance in natural language processing and computer vision applications by effectively capturing contextual knowledge from the entire sequence. However, the attention mechanism inherently contains a large number of redundant connections, imposing a heavy computational burden on model deployment. To this end, sparse attention has emerged as an attractive approach to reduce the computation and memory footprint, which involves the sampled dense-dense matrix multiplication (SDDMM) and sparse-dense matrix multiplication (SpMM) at the same time, thus requiring the hardware to eliminate zero-valued operations effectively. Existing techniques based on irregular sparse patterns or regular but coarse-grained patterns lead to low hardware efficiency or less computation saving. Liqiang Lu, Yicheng Jin, Hangrui Bi, Zizhang Luo, Peng Li 0031, Tao Wang 0004, Yun Liang 0001 |
MICRO | 4 |
| 2020 | Teaching Platform for Network Communication and Protocols Using a Micro: bit Based Wheeled RobotabstractIn this study, we presented a lightweight inverted curriculum\cite1 for teaching the essential details of network communication and protocols to undergraduate students major in computer science. Students are instructed to construct a wireless communication and control system connecting a computer to a wheeled robot using the Micro:bit platform. This platform consists of a micro-controller loaded with a Python interpreter and an additional extension board integrated with motors and sensors. In this study, we describe how the students were instructed to build the system step by step, from establishing a wired connection to implementing a TCP server on the PC-side for wireless control. Students can learn these knowledge through practice, which improves classroom engagement as a consequence. Zizhang Luo, Bohan Yu |
SIGCSE | 1 |