EDBT 2026 Demo / reviewers in the wild / expert
Kuan-Lin Chiu
dblp:09/7869
· DBLP profile ↗
8ranked-venue papers
3as first author
5since 2021 · last 2025
0000-0002-4892-9711ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 1 first-author · 5 since 2021Computer networks · 2 · 2 first-authorSoftware engineering, systems software and programming languages · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Optimization of Wire Pipelining and Channel Parallelism for 2D-Mesh NoC Physical DesignabstractModern systems-on-chip (SoCs) increasingly rely on high-bandwidth networks-on-chip (NoCs) to support communication among their many heterogeneous components. Meanwhile, as technology nodes advance, physical design (PD) has a growing impact on NoC performance, particularly for high-bandwidth NoCs. However, few published works focus on NoC optimization from a PD perspective. In this work, we study the problem of optimizing the PD of 2D-mesh NoCs by focusing on two prominent techniques: wire pipelining and channel parallelism. Our study is based on experimental results obtained from multiple tape-in NoC designs in a 12 nm technology process. We develop models to approximate the power and area effects of different NoC design approaches and analyze the underlying trends. Our findings show that pipelining does not affect die area but increases NoC power consumption by$1.6 \times$compared to increasing parallelism. In contrast, increasing parallelism can result in a NoC area up to$2 \times$larger than one achieving the same bandwidth through pipelining. Building on these insights, we formulate a mathematical optimization problem, which could be solved by optimization solvers to balance the trade-offs between these two techniques. Our study provides a general framework for analyzing NoC physical design trade-offs and optimizing NoC configurations. Pei-Huan Tsai, Maico Cassel, Joseph Zuckerman, Kuan-Lin Chiu, Luca P. Carloni |
ICCD | 4 |
| 2024 | WOLT: Transparent Deployment of ML Workloads on Lightweight Many-Accelerator ArchitecturesabstractMost available frameworks to develop machine learning applications target software deployment on general-purpose processors or GPUs. We propose Wolt, an end-to-end efficient solution to run TFLite application workloads on lightweight system-on-chip (SoC) architectures that feature many fixed-function hardware accelerators. WOltenables the execution of TFLite operations on pre-designed accelerators without requiring any modification of the application source code. As it establishes an interface between the high-level framework and the device drivers of the accelerators, WOLtincludes a resource manager that allows multi-tenant and conflict-free hardware acceleration of many TFLite applications running in parallel on the SoC. We evaluated WOltwith a comprehensive set of FPGA-based experiments by profiling and running 13 different TFLite workloads on a variety of complete SoC prototypes. We designed these prototypes by combining many CVA6 RISC-V processors with multiple accelerators for vector-matrix multiplication and two-dimensional convolution. For these workloads, Woltdelivers up to 23.2 x performance speedup and up to 14.6 x energy-efficiency gains compared to a purely software execution. When running multiple TFLite workloads in parallel, Wolt achieves up to 4 x of additional performance gain compared to basic hardware acceleration, thanks to efficient resource management. Kuan-Lin Chiu, Guy Eichler, Chuan-Tung Lin, Giuseppe Di Guglielmo, Luca P. Carloni |
ICCD | 1 |
| 2023 | PR-ESP: An Open-Source Platform for Design and Programming of Partially Reconfigurable SoCsabstractDespite its presence for more than two decades and its proven benefits in expanding the space of system design, dynamic partial reconfiguration (DPR) is rarely integrated into frameworks and platforms that are used to design complex reconfigurable system-on-chip (SoC) architectures. This is due to the complexity of the DPR FPGA flow as well as the lack of architectural and software runtime support to enable and fully harness DPR. Moreover, as DPR designs involve additional design steps and constraints, they often have a higher FPGA compilation (RTL-to-bitstream) runtime compared to equivalent monolithic designs. In this work, we present PR-ESP, an open-source platform for a system-level design flow of partially reconfigurable FPGA-based SoC architectures targeting embedded applications that are deployed on resource-constrained FPGAs. Our approach is realized by combining SoC design methodologies and tools from the open-source ESP platform with a fully-automated DPR flow that features a novel size-driven technique for parallel FPGA compilation. We also developed a software runtime reconfiguration manager on top of Linux. Finally, we evaluated our proposed platform using the WAMI-App benchmark application on Xilinx VC707. Biruk B. Seyoum, Davide Giri, Kuan-Lin Chiu, Bryce Natter, Luca P. Carloni |
DATE | 3 |
| 2023 | MindCrypt: The Brain as a Random Number Generator for SoC-Based Brain-Computer InterfacesabstractTrue random number generation on resource-constrained devices is challenging due to inherent hardware limitations; these limitations affect the ability to find a reliable source of randomness with high throughput and sufficient entropy. As recent developments in the field of Brain-Computer Interfaces (BCI) suggest a wide range of future applications that require random numbers, we investigate the usability of electrocorticography-based neural data as seeds for random number generation. We develop algorithms that generate random bits from brain data and evaluate the quality of randomness by using the NIST SP 800-22 test suite. We implement the algorithms as hardware random bit generators (RBGs). Then, we integrate these implementations as hardware accelerators in MindCrypt, a heterogeneous System-on-Chip (SoC) that is equipped with a host processor to run BCI applications. In MindCrypt, applications use our RBG accelerators as random number generators (RNGs) and prime number generators. FPGA prototypes of MindCrypt running software applications on a RISC-V processor that invoke our accelerators show improvements of 376x in throughput and 4885x in energy efficiency compared to using state-of-the-art Linux-based RNGs. By transferring random bits with point-to-point (P2P) communication between the RBG accelerators and cryptographic accelerators, we gain 6.1x in performance and 12.4x in energy efficiency compared to direct memory access (DMA). Finally, we explore the efficacy of a partially reconfigurable FPGA implementation of MindCrypt that dynamically optimizes the throughput of random number generation in a resource-constrained BCI SoC. Guy Eichler, Biruk B. Seyoum, Kuan-Lin Chiu, Luca P. Carloni |
ICCD | 3 |
| 2022 | Work-in-Progress: An Open-Source Platform for Design and Programming of Partially Reconfigurable Heterogeneous SoCsabstractDynamic partial reconfiguration (DPR) enables the design and implementation of flexible, scalable and robust adaptive systems. We present an FPGA-based DPR flow for partially reconfigurable heterogeneous SoCs that uses an incremental compilation technique to reduce the total FPGA compilation time. Biruk B. Seyoum, Davide Giri, Kuan-Lin Chiu, Luca P. Carloni |
CASES | 3 |
| 2020 | ESP4ML: Platform-Based Design of Systems-on-Chip for Embedded Machine LearningabstractWe present ESP4ML, an open-source system-level design flow to build and program SoC architectures for embedded applications that require the hardware acceleration of machine learning and signal processing algorithms. We realized ESP4ML by combining two established open-source projects (ESP and HLS4ML) into a new, fully-automated design flow. For the SoC integration of accelerators generated by HLS4ML, we designed a set of new parameterized interface circuits synthesizable with high-level synthesis. For accelerator configuration and management, we developed an embedded software runtime system on top of Linux. With this HW/SW layer, we addressed the challenge of dynamically shaping the data traffic on a network-on-chip to activate and support the reconfigurable pipelines of accelerators that are needed by the application workloads currently running on the SoC. We demonstrate our vertically-integrated contributions with the FPGA-based implementations of complete SoC instances booting Linux and executing computer-vision applications that process images taken from the Google Street View database. Davide Giri, Kuan-Lin Chiu, Giuseppe Di Guglielmo, Paolo Mantovani, Luca P. Carloni |
DATE | 2 |
| 2011 | Cross-layer design vehicle-aided handover scheme in VANETsabstractAbstract The requirement for in‐vehicle passengers to access Internet multimedia services has risen recently. As a consequence, Vehicle Ad hoc NETwork (VANET) has gained much attention, and is regarded as a promising solution for providing in‐vehicle Internet service through inter‐vehicle and infrastructure communication. A new developed wireless network technique, termed WiMAX Mobile Multihop Relay (MMR), provides a good communication framework for a VANET formed from vehicles on high‐speed freeways. Applying MMR WiMAX allows some public transportation vehicles to act as relay vehicles (RVs) to provide Internet access to passenger vehicles. However, the standard handover procedure of mobile or MMR WiMAX suffers long delay due to the lack of information about the next RV. This study presents a cross‐layer fast handover scheme, called vehicular fast handover scheme (VFHS), where the physical layer information is shared with the MAC layer, to reduce the handover delay. The key idea of VFHS is to utilize oncoming side vehicles (OSVs) to accumulate physical and MAC layers information of passing through RVs and broadcast the information to vehicles that are temporarily disconnected, referred to as disconnected vehicles (DVs). A DV can thus perform a rapid handover when it enters the transmission range of one of approaching RVs. The effectiveness of VFHS is verified using ns2 simulations. Simulation results indicate that VFHS significantly decreases handover latency and packet loss. Copyright © 2009 John Wiley & Sons, Ltd. Kuan-Lin Chiu, Ren-Hung Hwang, Yuh-Shyan Chen |
Wirel. Commun. Mob. Comput. | 1 |
| 2009 | A Cross Layer Fast Handover Scheme in VANETabstractThis study presents a cross-layer fast handover scheme for VANET, called vehicular fast handover scheme (VFHS), where the physical layer information is shared with the MAC layer, to reduce the handover delay. The key idea of VFHS is to utilize oncoming side vehicles (OSVs) to collect physical and MAC layers information of passing through RVs and broadcast the information to vehicles that are temporarily disconnected, referred to as broken vehicles (BVs). A BV can thus perform a rapid handover when it enters the transmission range of the approaching RVs. The effectiveness of VFHS is verified using ns2 simulations. Simulation results indicate that VFHS significantly decreases handover latency and packet loss. Kuan-Lin Chiu, Ren-Hung Hwang, Yuh-Shyan Chen |
ICC | 1 |