EDBT 2026 Demo / reviewers in the wild / expert
Yi-Hsiang Lai
dblp:145/9456
· DBLP profile ↗
14ranked-venue papers
5as first author
5since 2021 · last 2026
0000-0002-2358-805XORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 5 first-author · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative DecodingabstractModern large language model (LLM) applications exhibit diverse service-level objectives (SLOs), from low-latency requirements in interactive coding assistants to more relaxed constraints in data wrangling tasks. Existing LLM serving systems, which rely on uniform batching and scheduling strategies, often fail to meet these heterogeneous SLOs concurrently. We present AdaServe, the first LLM serving system designed to support efficient multi-SLO serving through SLO-customized speculative decoding. AdaServe formulates multi-SLO serving as a constrained optimization problem and introduces a hardware-aware algorithm that constructs a speculation tree tailored to each request's latency target. It features a speculate-select-verify pipeline that enables fine-grained control over decoding speed while maximizing system throughput. AdaServe further adapts to workload variation by dynamically adjusting speculation parameters. Evaluations across diverse workloads show that AdaServe reduces SLO violations by up to 4.3X and improves goodput by up to 1.9X compared to the best-performing baselines, highlighting its effectiveness in multi-SLO serving. Zikun Li, Zhuofu Chen, Remi Delacourt, Gabriele Oliaro, Qinghan Chen, Shuhuai Lin, April Yang, Zhihao Zhang 0001, Zhuoming Chen, Yi-Hsiang Lai, Xinhao Cheng, Xupeng Miao |
EuroSys | 11 |
| 2024 | Automated Deep Learning Optimization via DSL-Based Source Code TransformationabstractAs deep learning models become increasingly bigger and more complex, it is critical to improve model training and inference efficiency. Though a variety of highly optimized libraries and packages (known as DL kernels) have been developed, it is tedious and time-consuming to figure out which kernel to use, where to use, and how to use them correctly. To address this challenge, we propose an Automated Deep learning OPTimization approach called Adopter. We design a Domain-Specific Language (DSL) to represent DL model architectures and leverage this DSL to specify model transformation rules required to integrate a DL kernel into a model. Given the source code of a DL model and the transformation rules for a set of kernels, Adopter first performs inter-procedural analysis to identify and express the model architecture in our DSL. Then, Adopter performs scope analysis and sub-sequence matching to identify locations in the model architecture where the transformation rules can be applied. Finally, Adopter proposes a synthesis-based code transformation method to apply the transformation rule. We curated a benchmark with 199 models from Hugging Face and a diverse set of DL kernels. We found that, compared to a state-of-the-art automated code transformation technique, Adopter helps improve the precision and recall by 3% and 56%, respectively. An in-depth analysis of 9 models revealed that on average, Adopter improved the training speed by 22.7% while decreasing the GPU memory usage by 10.5%. Minghai Lu, Cody Hao Yu, Yi-Hsiang Lai, Tianyi Zhang 0001 |
ISSTA | 4 |
| 2022 | Accelerator design with decoupled hardware customizations: benefits and challenges: invitedabstractThe past decade has witnessed increasing adoption of high-level synthesis (HLS) to implement specialized hardware accelerators targeting either FPGAs or ASICs. However, current HLS programming models entangle algorithm specifications with hardware customization techniques, which lowers both the productivity and portability of the accelerator design. To tackle this problem, recent efforts such as HeteroCL propose to decouple algorithm definition from essential hardware customization techniques in compute, data type, and memory, increasing productivity, portability, and performance. Debjit Pal, Yi-Hsiang Lai, Shaojie Xiang, Niansong Zhang, Hongzheng Chen, Jeremy Casas, Pasquale Cocchini, Jin Yang 0006, Louis-Noël Pouchet, Zhiru Zhang |
DAC | 2 |
| 2022 | HeteroFlow: An Accelerator Programming Model with Decoupled Data Placement for Software-Defined FPGAsabstractTo achieve high performance with FPGA-equipped heterogeneous compute systems, it is crucial to co-optimize data placement and compute scheduling to maximize data reuse and bandwidth utilization for both on- and off-chip memory accesses. However, optimizing the data placement for FPGA accelerators is a complex task. One must acquire in-depth knowledge of the target FPGA device and its associated memory system in order to apply a set of advanced optimizations. Even with the latest high-level synthesis (HLS) tools, programmers often have to insert many low-level vendor-specific pragmas and substantially restructure the algorithmic code so that the right data are accessed at the right loop level using the right communication schemes. These code changes can significantly compromise the composability and portability of the original program. To address these challenges, we propose HeteroFlow, an FPGA accelerator programming model that decouples the algorithm specification from optimizations related to orchestrating the placement of data across a customized memory hierarchy. Specifically, we introduce a new primitive named .to(), which provides a unified programming interface for specifying data placement optimizations at different levels of granularity: (1) coarse-grained data placement between host and accelerator, (2) medium-grained kernel-level data placement within an accelerator, and (3) fine-grained data placement within a kernel. We build HeteroFlow on top of the open-source HeteroCL DSL and compilation framework. Experimental results on a set of realistic benchmarks show that, programs written in HeteroFlow can match the performance of extensively optimized manual HLS design with much fewer lines of code. Shaojie Xiang, Yi-Hsiang Lai, Hongzheng Chen, Niansong Zhang, Debjit Pal, Zhiru Zhang |
FPGA | 2 |
| 2021 | Programming and Synthesis for Software-defined FPGA Acceleration: Status and Future ProspectsabstractFPGA-based accelerators are increasingly popular across a broad range of applications, because they offer massive parallelism, high energy efficiency, and great flexibility for customizations. However, difficulties in programming and integrating FPGAs have hindered their widespread adoption. Since the mid 2000s, there has been extensive research and development toward making FPGAs accessible to software-inclined developers, besides hardware specialists. Many programming models and automated synthesis tools, such as high-level synthesis, have been proposed to tackle this grand challenge. In this survey, we describe the progression and future prospects of the ongoing journey in significantly improving the software programmability of FPGAs. We first provide a taxonomy of the essential techniques for building a high-performance FPGA accelerator, which requires customizations of the compute engines, memory hierarchy, and data representations. We then summarize a rich spectrum of work on programming abstractions and optimizing compilers that provide different trade-offs between performance and productivity. Finally, we highlight several additional challenges and opportunities that deserve extra attention by the community to bring FPGA-based computing to the masses. Yi-Hsiang Lai, Ecenur Ustun, Shaojie Xiang, Zhenman Fang, Hongbo Rong, Zhiru Zhang |
ACM Trans. Reconfigurable Technol. Syst. | 1 |
| 2020 | SuSy: A Programming Model for Productive Construction of High-Performance Systolic Arrays on FPGAsabstractSystolic algorithms are one of the killer applications on spatial architectures such as FPGAs and CGRAs. However, it requires a tremendous amount of human effort to design and implement a high-performance systolic array for a given algorithm using the traditional RTL-based methodology. On the other hand, existing high-level synthesis (HLS) tools either (1) force the programmers to do "micro-coding" where too many optimizations must be carried out through tedious code restructuring and insertion of vendor-specific pragmas, or (2) give them too little control to influence a push-button compilation flow to achieve high quality of results. Yi-Hsiang Lai, Hongbo Rong, Size Zheng 0001, Xiuping Cui, Yunshan Jia, Jie Wang 0022, Brendan Sullivan, Zhiru Zhang, Yun Liang 0001, Youhui Zhang, Jason Cong, Nithin George, Christopher J. Hughes, Pradeep Dubey |
ICCAD | 1 |
| 2019 | HeteroCL: A Multi-Paradigm Programming Infrastructure for Software-Defined Reconfigurable ComputingabstractWith the pursuit of improving compute performance under strict power constraints, there is an increasing need for deploying applications to heterogeneous hardware architectures with accelerators, such as GPUs and FPGAs. However, although these heterogeneous computing platforms are becoming widely available, they are very difficult to program especially with FPGAs. As a result, the use of such platforms has been limited to a small subset of programmers with specialized hardware knowledge. To tackle this challenge, we introduce HeteroCL, a programming infrastructure composed of a Python-based domain-specific language (DSL) and an FPGA-targeted compilation flow. The HeteroCL DSL provides a clean programming abstraction that decouples algorithm specification from three important types of hardware customization in compute, data types, and memory architectures. HeteroCL further captures the interdependence among these different customization techniques, allowing programmers to explore various performance/area/accuracy trade-offs in a systematic and productive manner. In addition, our framework produces highly efficient hardware implementations for a variety of popular workloads by targeting spatial architecture templates such as systolic arrays and stencil with dataflow architectures. Experimental results show that HeteroCL allows programmers to explore the design space efficiently in both performance and accuracy by combining different types of hardware customization and targeting spatial architectures, while keeping the algorithm code intact. Yi-Hsiang Lai, Yuze Chi, Jie Wang 0022, Cody Hao Yu, Jason Cong, Zhiru Zhang |
FPGA | 1 |
| 2018 | Rosetta: A Realistic High-Level Synthesis Benchmark Suite for Software Programmable FPGAsabstractModern high-level synthesis (HLS) tools greatly reduce the turn-around time of designing and implementing complex FPGA-based accelerators. They also expose various optimization opportunities, which cannot be easily explored at the register-transfer level. With the increasing adoption of the HLS design methodology and continued advances of synthesis optimization, there is a growing need for realistic benchmarks to (1) facilitate comparisons between tools, (2) evaluate and stress-test new synthesis techniques, and (3) establish meaningful performance baselines to track progress of the HLS technology. While several HLS benchmark suites already exist, they are primarily comprised of small textbook-style function kernels, instead of complete and complex applications. To address this limitation, we introduce Rosetta, a realistic benchmark suite for software programmable FPGAs. Designs in Rosetta are fully-developed applications. They are associated with realistic performance constraints, and optimized with advanced features of modern HLS tools. We believe that Rosetta is not only useful for the HLS research community, but can also serve as a set of design tutorials for non-expert HLS users. In this paper we describe the characteristics of our benchmarks and the optimization techniques applied to them. We further report experimental results on an embedded FPGA device as well as a cloud FPGA platform. Udit Gupta 0001, Steve Dai, Ritchie Zhao, Nitish Kumar Srivastava, Hanchen Jin, Joseph Featherston, Yi-Hsiang Lai, Gai Liu, Gustavo Angarita Velasquez, Zhiru Zhang |
FPGA | 8 |
| 2016 | Analytic approaches to the collapse operation and equivalence verification of threshold logic circuitsabstractThreshold logic circuits gain increasing attention due to their feasible realization with emerging technologies and strong bind to neural network applications. In this paper, for logic synthesis we formulate the fundamental operation of collapsing threshold logic gates, not addressed by prior efforts. A necessary and sufficient condition of collapsibility is obtained for linear combination of two threshold logic gates, and an analytic approach is proposed for fast circuit transformation. On the other hand, for equivalence verification we propose a linear time translation from threshold logic circuits to pseudo-Boolean constraints, in contrast to prior exponential translation costs. Experimental results demonstrate the effectiveness of circuit transformation by the collapse operation and the memory efficiency of equivalence verification by our pseudo-Boolean translation. Nian-Ze Lee, Hao-Yuan Kuo, Yi-Hsiang Lai, Jie-Hong Roland Jiang |
ICCAD | 3 |
| 2016 | Scalable Synthesis of PCHB-WCHB Hybrid Quasi-Delay Insensitive CircuitsabstractThe increasing cost paid in clocking integrated circuits and combating timing variations forces designers to rethink asynchronous approaches to system realization. Among various techniques, quasi-delay insensitive design is promising due to its very relaxed timing assumption. Its expensive logic overhead, however, often nullifies its promise of performance and power improvements, and remains a major obstacle on the way of its adoption. To overcome this obstacle, this paper proposes an efficient static performance analysis procedure and a synthesis flow for precharged half buffer and weak-conditioned half buffer circuit optimization. Experimental results demonstrate efficient performance analysis and effective area reduction under pipeline cycle time constraints. Yi-Hsiang Lai, Chi-Chuan Chuang, Jie-Hong Roland Jiang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2015 | Asynchronous QDI Circuit Synthesis from Signal Transition ProtocolsabstractAsynchronous circuits are promising in resolving the emerging issue of process variation and high synchronization power consumption. Among various asynchronous delay models, quasi-delay insensitive (QDI) model is the most robust and yet practical one due to its relaxed timing assumption. However, automatic synthesis of QDI circuits from signal transition graph (STG) protocol specification has not yet been proposed, despite the fact that algorithms synthesizing circuits under other delay models do exist. In this paper we propose the first algorithm synthesizing protocols specified in STGs into QDI circuits by analyzing STG structures without utilizing state graph assignment techniques. Furthermore, an optimization technique is proposed to simplify QDI circuits. In our synthesis algorithm, the state explosion issue is avoided, and restrictions on STGs are relaxed. Case studies on Advanced Microcontroller Bus Architecture (AMBA) and other protocols indicate the feasibility of our method. Bo-Yuan Huang 0001, Yi-Hsiang Lai, Jie-Hong Roland Jiang |
ICCAD | 2 |
| 2015 | A General Framework for Efficient Performance Analysis of Acyclic Asynchronous PipelinesabstractAsynchronous design methodologies gain recent extensive attention due to the variability issues in fabricating nanometer integrated circuits. Prior work on asynchronous pipeline performance analysis mostly focused on full buffer pipelines. To date half buffer performance analysis still lacks a systematic and precise treatment. In this paper, we propose a general framework abstracting four-phase asynchronous protocols and thus uniquely enable efficient performance analysis on various acyclic quasi-delay insensitive (QDI) pipelines (including the well-known pre-charged full buffer (PCFB), pre-charged half buffer (PCHB), weak-conditioned half buffer (WCHB), and null convention logic (NCL)) whose analysis has been challenging, if not impossible. Two approaches, linear programming-based performance analysis (LPA) and static performance analysis (SPA), that were applicable only to restricted set of full-buffer and half-buffer pipelines, respectively, are extended to support the entire set of considered pipelines. Thereby the two approaches can be directly compared for the first time. Experiments show that on average SPA achieve five orders of magnitude speedup over LPA, while LPA may provide 7% to 22% tighter cycle time estimation than SPA. Our results are essential to scalable performance analysis for a comprehensive set of QDI circuits. Yi-Hsiang Lai, Chi-Chuan Chuang, Jie-Hong Roland Jiang |
ICCAD | 1 |
| 2015 | SPOCK: Static Performance Analysis and Deadlock Verification for Efficient Asynchronous Circuit SynthesisabstractPerformance analysis and deadlock verification are two critical issues in asynchronous circuit design, which can be advantageous over the synchronous counterpart in terms of robustness against timing variability, security against side-channel attack, and other benefits. Nevertheless, asynchronous design automation tools are far away from mature. In this paper, we advance the synthesis of quasi-delay insensitive (QDI) circuits of pre-charged half buffer (PCHB) and weak-conditioned half buffer (WCHB) pipelines in three respects. First, static performance analysis (SPA) with linear time complexity is generalized from acyclic to cyclic PCHB and WCHB pipelines. Second, a deadlock verification (DV) algorithm with linear time complexity is proposed for checking PCHB and WCHB pipelines using their four-phase marked graph models. Third, we propose a new simple register circuitry for PCHB and WCHB pipelines that consists of one reset-latch and one buffer-latch and is amenable to circuit minimization. With the above two algorithms, we develop an efficient synthesis flow for buffer-latch minimization while maintaining the system throughput and deadlock-free property. Experimental results show the efficiency of our SPA and DV algorithms and demonstrate the effectiveness of our synthesis method with an average of 37% reduction on the number of buffer-laches. As our SPA and DV algorithms are applicable to arbitrary PCHB and WCHB pipelines and our buffer-latch minimization algorithm is orthogonal to existing synthesis methods such as cut-based technology mapping and slack matching, our methods can be generally useful in the analysis, verification, and synthesis of PCHB and WCHB pipelines. Chun-Hong Shih, Yi-Hsiang Lai, Jie-Hong Roland Jiang |
ICCAD | 2 |
| 2014 | Synthesis of PCHB-WCHB Hybrid Quasi-Delay Insensitive CircuitsabstractThe increasing cost paid in clocking integrated circuits and combating timing variations forces designers to rethink asynchronous approaches to system realization. Among various techniques, quasi-delay-insensitive (QDI) design is promising due to its very relaxed timing assumption. Its expensive logic overhead, however, often nullifies its promise of performance and power improvements, and remains a major obstacle against its adoption. To overcome this obstacle, this paper proposes an efficient static performance analysis procedure and a synthesis flow for precharged half buffer (PCHB) and weak-conditioned half buffer (WCHB) circuit optimization. Experimental results demonstrate efficient performance analysis and effective area reduction under pipeline cycle time constraints. Chi-Chuan Chuang, Yi-Hsiang Lai, Jie-Hong Roland Jiang |
DAC | 2 |