EDBT 2026 Demo / reviewers in the wild / expert
Heehoon Kim
dblp:202/9987
· DBLP profile ↗
5ranked-venue papers
2as first author
3since 2021 · last 2025
0000-0002-5392-4413ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 3 · 2 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SPipe: Hybrid GPU and CPU Pipeline for Training LLMs under Memory PressureabstractTraining large language models (LLMs) with limited computing resources is challenging because of their immense memory space requirements. In this paper, we specifically focus on the scenarios where we have insufficient aggregate GPU memory to store all model states but explore pipeline parallelism and offloading across all system resources to train the model. In this context, SPipe presents a hybrid GPU and CPU pipelining mechanism that consists of two pipelines: a GPU pipeline to reduce the bubbles in conventional pipeline parallelism and a GPU-CPU pipeline to alleviate data transfer overhead and CPU bottlenecks in offloading data and computation. We evaluate SPipe for training LLMs of various sizes with diverse configurations in practice. The result indicates that SPipe outperforms the state-of-the-art by $1.26 \times$. Junyeol Ryu, Yujin Jeong, Daeyoung Park, Jinpyo Kim, Heehoon Kim, Jaejin Lee |
PACT | 5 |
| 2024 | TCCL: Discovering Better Communication Paths for PCIe GPU ClustersabstractExploiting parallelism to train deep learning models requires GPUs to cooperate through collective communication primitives. While systems like DGX, equipped with proprietary interconnects, have been extensively studied, the systems where GPUs mainly communicate through PCIe have received limited attention. This paper introduces TCCL, a collective communication library designed explicitly for such systems. TCCL has three components: a profiler for multi-transfer performance measurement, a pathfinder to discover optimal communication paths, and a modified runtime of NCCL to utilize the identified paths. The focus is on ring-based collective communication algorithms that apply to popular communication operations in deep learning, such as AllReduce and AllGather. The evaluation results of TCCL on three different PCIe-dependent GPU clusters show that TCCL outperforms (up to ×2.07) the state-of-the-art communication libraries, NCCL and MSCCL. We also evaluate TCCL with DL training workloads with various combinations of parallelism types. Heehoon Kim, Junyeol Ryu, Jaejin Lee |
ASPLOS (3) | 1 |
| 2022 | SnuQS: scaling quantum circuit simulation using storage devicesabstractSince the state-of-the-art quantum computers are still noisy and error-prone, classical simulation of quantum circuits is essential in verifying/calibrating quantum computers and prototyping/debugging complex quantum algorithms. Classical simulation of large quantum systems is challenging due to its exponential increase in space and computation requirements. In this paper, we propose a full-state simulation framework, SnuQS. It exploits storage devices, such as HDDs and NVMe SSDs, to enlarge the available main memory capacity at a small cost. To achieve maximum I/O bandwidth, we propose an overlay-based memory management technique and optimization techniques. We also propose an I/O subsystem architecture that guarantees the maximum bandwidth of each storage device. We evaluate SnuQS on a 64-core CPU and 4-GPU system with 80 2TB HDDs and 10 4TB NVMe SSDs using quantum supremacy and quantum Fourier transform circuits. The experimental result indicates that SnuQS and the proposed I/O subsystem together is an effective and practical solution to scale the full-state simulation of large quantum circuits at about 300X lower cost than the DDR4 DRAM main-memory-only system. Daeyoung Park, Heehoon Kim, Jinpyo Kim, Taehyun Kim 0002, Jaejin Lee |
ICS | 2 |
| 2020 | SOFF: An OpenCL High-Level Synthesis Framework for FPGAsabstractRecently, OpenCL has been emerging as a programming model for energy-efficient FPGA accelerators. However, the state-of-the-art OpenCL frameworks for FPGAs suffer from poor performance and usability. This paper proposes a high-level synthesis framework of OpenCL for FPGAs, called SOFF. It automatically synthesizes a datapath to execute many OpenCL kernel threads in a pipelined manner. It also synthesizes an efficient memory subsystem for the datapath based on the characteristics of OpenCL kernels. Unlike previous high-level synthesis techniques, we propose a formal way to handle variable latency instructions, complex control flows, OpenCL barriers, and atomic operations that appear in real-world OpenCL kernels. SOFF is the first OpenCL framework that correctly compiles and executes all applications in the SPEC ACCEL benchmark suite except three applications that require more FPGA resources than are available. In addition, SOFF achieves the speedup of 1.33 over Intel FPGA SDK for OpenCL without any explicit user annotation or source code modification. Gangwon Jo, Heehoon Kim, Jeesoo Lee, Jaejin Lee |
ISCA | 2 |
| 2017 | Performance analysis of CNN frameworks for GPUsabstractThanks to modern deep learning frameworks that exploit GPUs, convolutional neural networks (CNNs) have been greatly successful in visual recognition tasks. In this paper, we analyze the GPU performance characteristics of five popular deep learning frameworks: Caffe, CNTK, TensorFlow, Theano, and Torch in the perspective of a representative CNN model, AlexNet. Based on the characteristics obtained, we suggest possible optimization methods to increase the efficiency of CNN models built by the frameworks. We also show the GPU performance characteristics of different convolution algorithms each of which uses one of GEMM, direct convolution, FFT, and the Winograd method. We also suggest criteria to choose convolution algorithms for GPUs and methods to build efficient CNN models on GPUs. Since scaling DNNs in a multi-GPU context becomes increasingly important, we also analyze the scalability of the CNN models built by the deep learning frameworks in the multi-GPU context and their overhead. The result indicates that we can increase the speed of training the AlexNet model up to 2X by just changing options provided by the frameworks. Heehoon Kim, Hyoungwook Nam, Wookeun Jung, Jaejin Lee |
ISPASS | 1 |