EDBT 2026 Demo / reviewers in the wild / expert
Yi-Chien Lin
dblp:15/2124
· DBLP profile ↗
14ranked-venue papers
9as first author
11since 2021 · last 2026
0000-0002-1710-1532ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 7 first-author · 10 since 2021Human-computer interaction and ubiquitous computing · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ARGO+: Achieving Multi-Level Scalable GNN Training on Distributed Multi-Core PlatformabstractGraph Neural Networks (GNNs) have emerged as powerful tools for learning from graph-structured data and are widely used in various applications such as traffic prediction, and Electronic Design Automation, among others. However, training GNNs on large-scale graphs with hundreds of millions of nodes and billions of edges is time-consuming, often requiring days or weeks on a single machine. While distributed GNN training across multiple machines has been explored to leverage greater computation and memory resources, existing approaches primarily focus on inter-machine scalability while overlooking intra-machine scalability across multiple cores. As a result, state-of-the-art distributed GNN training frameworks lead to severe resource underutilization and limited training performance in terms of epoch time. This work introduces ARGO+, a novel GNN system designed to achieve multi-level scalability for distributed GNN training by efficiently scaling across both inter- and intra-machine levels. ARGO+ features a two-level graph partitioning strategy, exploiting parallelisms across both graph topology and feature dimensions while minimizing communication overhead and ensuring balanced workloads. In addition, ARGO+ integrates NUMA-aware optimizations to enhance intra-node scalability, addressing inefficiencies in remote-socket data accessing. During runtime, ARGO+ adopts a two-stage parallel training scheme to further reduce communication overheads and instantiates multiple training processes to exploit computation-communication overlapping. We evaluate ARGO+ on a distributed multi-core CPU cluster consisting of 8 machines, each with a dual-socket 80-core Intel Xeon processor. Our results demonstrate that ARGO+ achieves up to 1.89× speedup compared with DistDGL, and 1.53-1.56× speedup compared with state-of-the-art distributed GNN training systems. In addition, ARGO+ is compatible with the Deep Graph Library (DGL), a widely used GNN framework, enabling seamless integration with existing GNN programs. Finally, while ARGO+ adopts various optimizations to improve scalability and training performance, these optimizations do not alter the semantics of the training algorithm; thus, the model accuracy and convergence remain consistent with the original implementation of the algorithm. Yi-Chien Lin, Sameh Gobriel, Nilesh Jain, Viktor Prasanna 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2025 | ELLIE: Energy-Efficient LLM Inference at the Edge Via Prefill-Decode SplittingabstractAs Large Language Models (LLMs) are increasingly deployed for on-device applications, optimizing inference on edge platforms becomes critical. In real-world scenarios, LLM inference must satisfy diverse constraints and user requirements, such as low latency, high energy efficiency, or low Energy-Delay Product (EDP). Most state-of-the-art edge platforms, such as AI PCs and mobile SoCs, integrate heterogeneous processing units, including CPUs, GPUs, and Neural Processing Units (NPUs), each with distinct performance and power characteristics. However, existing approaches often adopt static mapping to a single processing unit (e.g., CPU, GPU, or NPU) and perform optimizations for either latency or energy consumption. This limits their effectiveness in meeting the requirements of diverse application scenarios. Moreover, LLM inference consists of two distinct phases: a highly parallel, compute-intensive Prefill phase, and a sequential, memory-intensive Decode phase. These phases have different computational characteristics, and splitting them across suitable processing units can potentially yield better energy efficiency and EDP than static mapping. However, such improvements may not be realized in all cases, as the actual benefit depends on many factors, including prompt characteristics, the models used, and features of the target hardware. To address these challenges, we propose ellie, a lightweight inference framework for edge heterogeneous platforms that dynamically selects the optimal execution plan based on the usage scenario, LLM, hardware feature, and input prompt. ELLIE builds performance models by regressing latency and power from offline profiling data, and integrates it with a lightweight output token length predictor. At runtime, it estimates the latency and energy of candidate execution plans using the predicted output length and selects an optimized device mapping accordingly. We implement ellie on an Intel AI PC platform with integrated CPU, GPU, and NPU. On average, when optimizing for EDP, ELLIE reduces energy consumption by$1.8 \times$, improves EDP by$1.5 \times$, and achieves latency comparable to GPU-only inference, across diverse LLMs and prompt types. Haoyang Fan, Yi-Chien Lin, Viktor Prasanna 0001 |
ASAP | 2 |
| 2025 | SMART: High-Performance SAR ATR Through Model-Architecture Co-Design on FPGAabstractSynthetic Aperture Radar (SAR) Automatic Target Recognition (ATR) is a fundamental technique in remote-sensing image recognition. SAR ATR systems demand real-time performance and low power consumption, particularly when operating in resource-constrained environments such as small satellites. This paper presents SMART, a model-architecture co-design on FPGA to address the challenges of high latency, large memory footprint, and high power consumption in state-of-the-art SAR ATR models. Model design: We develop a compact CNN with integrated spatial, channel and cross attention mechanisms, optimized through a softmax-free attention approach and quantization for improved computational efficiency and reduced latency. Architecture design: We implement a customized FPGA accelerator with a streaming dataflow architecture using high-level synthesis (HLS) on the Xilinx Alveo U280 FPGA. The design optimizes the dataflow through parameterized hardware kernels for key operations. Experimental results on the widely used MSTAR (99.84%), SynthWakeSAR (94.85%) and GBSAR (99.92%) datasets, demonstrate that our model achieves superior classification accuracy. Our design delivers up to 29× lower latency and 66× higher energy efficiency compared to implementations on state-of-the-art CPU and GPU platforms. Sachini Wickramasinghe, Yi-Chien Lin, Cauligi S. Raghavendra, Viktor Prasanna 0001 |
FCCM | 2 |
| 2025 | Accelerating GNN Inference via Automated Parallel Execution on Edge Heterogeneous PlatformsabstractRecently, Graph Neural Networks (GNN) have been integrated into various local applications, such as local community detection and local code assistant, making edge inference increasingly important. To support diverse workloads, state-of-the-art edge devices have evolved into heterogeneous platforms, integrating components like CPU, GPU, and NPU. To this end, we propose GNX, a novel GNN system that accelerates GNN inference on edge heterogeneous platforms by leveraging all the heterogeneous processing units. Given a GNN model and a heterogeneous platform, GNX automatically generates parallel execution plans, consisting of both data and pipeline parallelism. To reduce the complexity of the design space, GNX converts GNN models into coarse-grained blocks and performs the search at the block level. By leveraging the APIs provided by state-of-the-art heterogeneous frameworks, GNX can flexibly schedule various parallel execution plans and seamlessly adjust the workload across the heterogeneous processing units for load-balanced execution. Our study shows that GNX effectively accelerates three widely-used GNN models on two state-of-the-art edge heterogeneous platforms. Compared with the baseline approach that uses only a single processing unit, GNX achieves up to a 2.57× speedup. Compared with adopting data parallelism and a state-of-the-art scheduler, GNX achieves up to 1.90× and 1.79× speedup, respectively. We also discuss the applicability of and extensions to GNX to support other GNN models. Yi-Chien Lin, Haoyang Fan, Sameh Gobriel, Nilesh Jain, Viktor Prasanna 0001 |
SBAC-PAD | 1 |
| 2024 | A Unified CPU-GPU Protocol for GNN TrainingabstractTraining a Graph Neural Network (GNN) model on large-scale graphs involves a high volume of data communication and computations. While state-of-the-art CPUs and GPUs feature high computing power, the Standard GNN training protocol adopted in existing GNN frameworks cannot efficiently utilize the platform resources. To this end, we propose a novel Unified CPU-GPU protocol that can improve the resource utilization of GNN training on a CPU-GPU platform. The Unified CPU-GPU protocol instantiates multiple GNN training processes in parallel on both the CPU and the GPU. By allocating training processes on the CPU to perform GNN training collaboratively with the GPU, the proposed protocol improves the platform resource utilization and reduces the CPU-GPU data transfer overhead. Since the performance of a CPU and a GPU varies, we develop a novel load balancer that balances the workload dynamically between CPUs and GPUs during runtime. We evaluate our protocol using two representative GNN sampling algorithms, with two widely-used GNN models, on three datasets. Compared with the Standard training protocol adopted in the state-of-the-art GNN frameworks, our protocol effectively improves resource utilization and improves the overall training time. On a platform where the GPU moderately outperforms the CPU, our protocol speeds up GNN training by up to 1.41×. On a platform where the GPU significantly outperforms the CPU, our protocol speeds up GNN training by up to 1.26×. Our protocol is open-sourced and can be seamlessly integrated into state-of-the-art GNN frameworks and accelerate GNN training. Our protocol particularly benefits those with limited GPU access due to its high demand. Yi-Chien Lin, Gangda Deng, Viktor Prasanna 0001 |
CF | 1 |
| 2024 | ARGO: An Auto-Tuning Runtime System for Scalable GNN Training on Multi-Core ProcessorabstractAs Graph Neural Networks (GNNs) become popular, libraries like PyTorch-Geometric (PyG) and Deep Graph Library (DGL) are proposed; these libraries have emerged as the de facto standard for implementing GNNs because they provide graph-oriented APIs and are purposefully designed to manage the inherent sparsity and irregularity in graph structures. However, these libraries show poor scalability on multi-core processors, which under-utilizes the available platform resources and limits the performance. This is because GNN training is a resource-intensive workload with high volume of irregular data accessing, and existing libraries fail to utilize the memory bandwidth efficiently. To address this challenge, we propose ARGO, a novel runtime system for GNN training that offers scalable performance. ARGO exploits multi-processing and core-binding techniques to improve platform resource utilization. We further develop an auto-tuner that searches for the optimal configuration for multi-processing and core-binding. The auto-tuner works automatically, making it completely transparent from the user. Furthermore, the auto-tuner allows ARGO to adapt to various platforms, GNN models, datasets, etc. We evaluate ARGO on two representative GNN models and four widely-used datasets on two platforms. With the proposed autotuner, ARGO is able to select a near-optimal configuration by exploring only 5% of the design space. ARGO speeds up state-of-the-art GNN libraries by up to 5.06× and 4.54× on a four-socket Ice Lake machine with 112 cores and a two-socket Sapphire Rapids machine with 64 cores, respectively. Finally, ARGO can seamlessly integrate into widely-used GNN libraries (e.g., DGL, PyG) with few lines of code and speed up GNN training. Yi-Chien Lin, Sameh Gobriel, Nilesh Jain, Gopi Krishna Jha, Viktor Prasanna 0001 |
IPDPS | 1 |
| 2024 | HitGNN: High-Throughput GNN Training Framework on CPU+Multi-FPGA Heterogeneous PlatformabstractAs the size of real-world graphs increases, training Graph Neural Networks (GNNs) has become time-consuming and requires acceleration. While previous works have demonstrated the potential of utilizing FPGA for accelerating GNN training, few works have been carried out to accelerate GNN training with multiple FPGAs due to the necessity of hardware expertise and substantial development effort. To this end, we propose HitGNN, a framework that enables users to effortlessly map GNN training workloads onto a CPU+Multi-FPGA platform for acceleration. In particular, HitGNN takes the user-defined synchronous GNN training algorithm, GNN model, and platform metadata as input, determines the design parameters based on the platform metadata, and performs hardware mapping onto the CPU+Multi-FPGA platform, automatically. HitGNN consists of the following building blocks: (1) high-level application programming interfaces (APIs) that allow users to specify various synchronous GNN training algorithms and GNN models with only a handful of lines of code; (2) a software generator that generates a host program that performs mini-batch sampling, manages CPU-FPGA communication, and handles workload balancing among the FPGAs; (3) an accelerator generator that generates GNN kernels with optimized datapath and memory organization. We show that existing synchronous GNN training algorithms such as DistDGL and PaGraph can be easily deployed on a CPU+Multi-FPGA platform using our framework, while achieving high training throughput. Compared with the state-of-the-art frameworks that accelerate synchronous GNN training on a multi-GPU platform, HitGNN achieves up to 27.21× bandwidth efficiency, and up to 4.26× speedup using much less compute power and memory bandwidth than GPUs. In addition, HitGNN demonstrates good scalability to 16 FPGAs on a CPU+Multi-FPGA platform. Yi-Chien Lin, Bingyi Zhang, Viktor Prasanna 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2023 | A Framework for Graph Machine Learning on Heterogeneous ArchitectureabstractGraph Machine Learning (Graph ML) has shown great success in many domains such as Electronic Design Automation (EDA) [1], traffic prediction [2], recommendation systems [3], etc. where data is represented as graphs. In real-world scenarios, these domains often involve large-scale graphs with over billions of edges [4]; training Graph ML models on these large datasets using only a single CPU or GPU would take hours or even days [5], which calls for the need for acceleration. To provide more computing power and to efficiently process different types of workloads, state-of-the-art machines feature heterogeneous architecture [6], [7] that consists of multiple CPUs and a variety of accelerators such as GPUs, FPGAs, or AI-specific accelerators [8]–[10]. Heterogeneous architectures have great potential to accelerate Graph ML; however, due to the complexity of such platforms, accelerating Graph ML remains challenging. First, it requires extensive programming to accelerate Graph ML on a heterogeneous architecture [11]. In particular, launching each accelerator requires a different program (e.g., CUDA for GPU, Verilog for FPGA); one also needs to develop a complex host program to orchestrate the task coordination among the CPUs and the accelerators. Second, in order to achieve high performance, the Graph ML kernels on each device (i.e., the CPUs or the accelerators) need to be highly optimized. In addition, training a Graph ML on multiple devices in parallel suffers from workload imbalance and high communication overhead [12]; these issues limit the speedup that can be achieved on the heterogeneous architecture. Finally, it is challenging to build a portable design that can run on various heterogeneous architectures while achieving high performance. Portability is critical for one to build an impactful design that contributes to the research community and the industry as it allows others to easily build upon his or her work. To this end, we propose a novel framework for accelerating Graph ML training on a heterogeneous architecture. The framework aims to achieve high programmability, high performance, and high portability; we introduce how these three objectives are achieved in Section III. Yi-Chien Lin, Viktor Prasanna 0001 |
FCCM | 1 |
| 2023 | HyScale-GNN: A Scalable Hybrid GNN Training System on Single-Node Heterogeneous ArchitectureabstractGraph Neural Networks (GNNs) have shown success in many real-world applications that involve graph-structured data. Most of the existing single-node GNN training systems are capable of training medium-scale graphs with tens of millions of edges; however, scaling them to large-scale graphs with billions of edges remains challenging. In addition, it is challenging to map GNN training algorithms onto a computation node as state-of-the-art machines feature heterogeneous architecture consisting of multiple processors and a variety of accelerators.We propose HyScale-GNN, a novel system to train GNN models on a single-node heterogeneous architecture. HyScale-GNN performs hybrid training which utilizes both the processors and the accelerators to train a model collaboratively. Our system design overcomes the memory size limitation of existing works and is optimized for training GNNs on large-scale graphs. We propose a two-stage data pre-fetching scheme to reduce the communication overhead during GNN training. To improve task mapping efficiency, we propose a dynamic resource management mechanism, which adjusts the workload assignment and resource allocation during runtime. We evaluate HyScale-GNN on a CPU-GPU and a CPU-FPGA heterogeneous architecture. Using several large-scale datasets and two widely-used GNN models, we compare the performance of our design with a multi-GPU baseline implemented in PyTorch-Geometric. The CPU-GPU design and the CPU-FPGA design achieve up to 2.08× speedup and 12.6× speedup, respectively. Compared with the state-of-the-art large-scale multi-node GNN training systems such as P3and DistDGL, our CPU-FPGA design achieves up to 5.27× speedup using a single node. Yi-Chien Lin, Viktor Prasanna 0001 |
IPDPS | 1 |
| 2022 | HP-GNN: Generating High Throughput GNN Training Implementation on CPU-FPGA Heterogeneous PlatformabstractGraph Neural Networks (GNNs) have shown great success in many applications such as recommendation systems, molecular property prediction, traffic prediction, etc. Recently, CPU-FPGA heterogeneous platforms have been used to accelerate many applications by exploiting customizable data path and abundant user-controllable on-chip memory resources of FPGAs. Yet, accelerating and deploying GNN training on such platforms requires not only expertise in hardware design but also substantial development efforts. Yi-Chien Lin, Bingyi Zhang, Viktor Prasanna 0001 |
FPGA | 1 |
| 2021 | A Memory-Efficient Accelerator for DNA Sequence Alignment with Two-Piece Affine Gap TracebacksabstractPreviously, dynamic-programming-based DNA sequence aligners were mostly implemented with a penalty function of the one-piece affine gap model. When aligning sequences with longer gaps, the two-piece affine gap model provides better results at the cost of memory usage, which becomes an issue especially for aligners with memory-hungry traceback capabilities. In this paper, we design a memory-efficient scheme for traceback recording with the two-piece penalty scoring, so that the aligner can be realized on an ASIC. Our design is implemented with TSMC 40nm technology, and the proposed aligner can speed up pairwise alignment by 71× compared to the CPU approach. Jing-Ping Wu, Yi-Chien Lin, Ying-Wei Wu, Shih-Wei Hsieh, Ching-Hsuan Tai, Yi-Chang Lu |
ISCAS | 2 |
| 2011 | Predicting SLA Students' Behavioral Intentions to Use Multimedia Web-Based English Learning Systems
Yi-Chien Lin, RonTung Yeh, Wen-Tung Hung |
ICCE | 1 |
| 2007 | An Automatic Quiz Generation System for English TextabstractIn this study, we design and prototype an automatic quiz generation system (auto-quiz for short) for a given English text to test learner comprehension of text content and English skills. The auto-quiz process parses an English text into a semantic network representation and enhances the semantic network iteratively with intrinsic knowledge, such as English grammar and writing styles, and extrinsic knowledge, such as the word relationship in WordNet and statistics from corpus or search engines. Then, the quiz generation process generates quiz from the text based on learner comprehension skills and according to the learner learning status and needs, such as English proficiency and frequent errors. Li-Chun Sung, Yi-Chien Lin, Meng Chang Chen |
ICALT | 2 |
| 2006 | REDRP: Reactive Energy Decisive Routing Protocol for Wireless Sensor Networks
Ying-Hong Wang, Yi-Chien Lin, Ping-Fang Fu, Chih-Hsiao Tsai |
UIC | 2 |