Jinpyo Kim

dblp:26/364 · DBLP profile ↗
← Back
12ranked-venue papers
3as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 3 first-author · 6 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Early Attack Identification in the Wild
abstract
While characterizing network connections in their early stage is vital for providing timely responses against network threats, existing methods become less attractive due to the requirement of complete connection information (thus unable to make timely identification) or packet payload inspection (there-fore limited to unencrypted packets under no privacy regulation). To this end, this paper takes an approach ofpacket stream analysisreferencing statistical information of packet sequences, requiringneitherpacket inspectionnorcomplete connection information. To enable practical packet stream analysis, there exist several challenges, such asout-of-order packet sequencesintroduced by network dynamics andclass imbalancewith a tiny fraction of attack connections. To overcome these challenges, we design two deep sequence models: (i) abidirectional recurrent structuredesigned for greater resilience to out-of-order packet streams, and (ii) apre-training-enabled sequence-to-sequence structuredesigned for creating consistent representations from unbalanced class distributions using self-supervised learning. We evaluate the presented deep sequence models using real and synthetic network data collections for extensive experimentation. The experimental results support the feasibility of the proposed models outperforming baseline deep learning models, yielding up to 94.8% (F1 score) only with the first five packets (k=5) from the Internet traffic collection containing a substantial fraction of network flows experiencing out-of-order delivery.
Dongeun Lee 0001, Kookjin Lee, Doowon Kim, Jinpyo Kim, Sangman Lee, Jinoh Kim
IEEE Trans. Netw.5
2025 SPipe: Hybrid GPU and CPU Pipeline for Training LLMs under Memory Pressure
abstract
Training large language models (LLMs) with limited computing resources is challenging because of their immense memory space requirements. In this paper, we specifically focus on the scenarios where we have insufficient aggregate GPU memory to store all model states but explore pipeline parallelism and offloading across all system resources to train the model. In this context, SPipe presents a hybrid GPU and CPU pipelining mechanism that consists of two pipelines: a GPU pipeline to reduce the bubbles in conventional pipeline parallelism and a GPU-CPU pipeline to alleviate data transfer overhead and CPU bottlenecks in offloading data and computation. We evaluate SPipe for training LLMs of various sizes with diverse configurations in practice. The result indicates that SPipe outperforms the state-of-the-art by $1.26 \times$.
Junyeol Ryu, Yujin Jeong, Daeyoung Park, Jinpyo Kim, Heehoon Kim, Jaejin Lee
PACT4
2025 SnuSOLVER: Optimizing Sparse Direct Solvers for Heterogeneous Systems
abstract
Achieving scalability in sparse direct solvers is crucial for addressing the complexity of real-world systems.This paper proposes SnuSOLVER, a library for sparse direct solvers.Unlike conventional approaches, which often apply kernel selection heuristically within supernodal or frontal methods, the SnuSOLVER adopts a structured, two-phase execution strategy tailored to the characteristics of each level in the nested dissection hierarchy.It ensures optimal performance across all hierarchical levels of computation.We evaluate SnuSOLVER on an eight-node heterogeneous cluster with AMD CPUs and NVIDIA GPUs using 31 sparse matrices of varying sizes and domains.Experimental results demonstrate that SnuSOLVER outperforms the state-of-the-art solvers Su-perLU_DIST and STRUMPACK and underscore the scalability, efficiency, and adaptability of SnuSOLVER, establishing * He conducted this work as a Ph.D.
Jinpyo Kim, Kyusu Ahn, Hyung Uk Cho, Seungin Baek, Jaejin Lee
ICS3
2025 Integrating spatial and frequency information for Under-Display Camera image restoration
abstract
Abstract Under-Display Camera (UDC) houses a digital camera lens under a display panel. However, UDC introduces complex degradations such as noise, blur, decrease in transmittance, and flare. Despite the remarkable progress, previous research on UDC mainly focuses on eliminating diffraction in the spatial domain and rarely explores its potential in the frequency domain. In this paper, we revisit the UDC degradations in the Fourier space and figure out intrinsic frequency priors that imply the presence of the flares. Based on these observations, we propose SFIM, a novel multi-level deep neural network that efficiently restores UDC-distorted images by integrating local and global (the collective contribution of all points in the image) information. SFIM uses CNNs to capture fine-grained local details and FFT-based models to extract global patterns. The network comprises a spatial domain block (SDB), a frequency domain block (FDB), and an attention-based multi-level integration block (AMIB). Specifically, SDB focuses more on detailed textures such as noise and blur, FDB emphasizes irregular texture loss in extensive areas such as flare, and AMIB employs cross-domain attention to selectively integrate complementary spatial and frequency features across multiple levels, enhancing detail recovery and mitigating irregular degradations like flare. SFIM’s superior performance over state-of-the-art approaches is demonstrated through rigorous quantitative and qualitative assessments. Our source code is publicly available at: https://github.com/mcrl/SFIM .
Kyusu Ahn, Jinpyo Kim, Chanwoo Park, JiSoo Kim, Jaejin Lee
Pattern Anal. Appl.2
2023 DeepUM: Tensor Migration and Prefetching in Unified Memory
abstract
Deep neural networks (DNNs) are continuing to get wider and deeper. As a result, it requires a tremendous amount of GPU memory and computing power. In this paper, we propose a framework called DeepUM that exploits CUDA Unified Memory (UM) to allow GPU memory oversubscription for DNNs. While UM allows memory oversubscription using a page fault mechanism, page migration introduces enormous overhead. DeepUM uses a new correlation prefetching technique to hide the page migration overhead. It is fully automatic and transparent to users. We also propose two optimization techniques to minimize the GPU fault handling time. We evaluate the performance of DeepUM using nine large-scale DNNs from MLPerf, PyTorch examples, and Hugging Face and compare its performance with six state-of-the-art GPU memory swapping approaches. The evaluation result indicates that DeepUM is very effective for GPU memory oversubscription and can handle larger models that other approaches fail to handle.
Jinpyo Kim, Jaejin Lee
ASPLOS (2)2
2022 SnuHPL: high performance LINPACK for heterogeneous GPUs
abstract
These days, it is typical for a large-scale cluster system to have different kinds of GPUs. However, HPL (High-Performance LINPACK), the de-facto standard LINPACK implementation for evaluating the performance of a cluster system, is originally designed to work only for homogeneous CPU-only systems. In this paper, we develop SnuHPL, an optimized HPL for clusters of modern heterogeneous GPUs. To optimize SnuHPL for the heterogeneous GPUs, we design a performance model, a SnuHPL simulator based on the model, and a greedy heuristic algorithm based on the simulator. The algorithm generates the best data distribution for a given cluster configuration by considering computing power, memory capacity, and network performance altogether. We also present a simple technique to increase the energy efficiency of HPL by adjusting the core clock frequency of the GPUs. The evaluation of the data distribution algorithm on small clusters of different GPU combinations shows that it outperforms well-known other data distribution strategies. We show the effectiveness of SnuHPL on a cluster of 1,760 NVIDIA A100-80GB GPUs and 440 A100-40GB GPUs. We also show the effectiveness of the proposed energy optimization technique on a cluster of 144 A100-80GB GPUs.
Jinpyo Kim, Hyungdal Kwon, Jintaek Kang, Seungwook Lee, Jaejin Lee
ICS1
2022 SnuQS: scaling quantum circuit simulation using storage devices
abstract
Since the state-of-the-art quantum computers are still noisy and error-prone, classical simulation of quantum circuits is essential in verifying/calibrating quantum computers and prototyping/debugging complex quantum algorithms. Classical simulation of large quantum systems is challenging due to its exponential increase in space and computation requirements. In this paper, we propose a full-state simulation framework, SnuQS. It exploits storage devices, such as HDDs and NVMe SSDs, to enlarge the available main memory capacity at a small cost. To achieve maximum I/O bandwidth, we propose an overlay-based memory management technique and optimization techniques. We also propose an I/O subsystem architecture that guarantees the maximum bandwidth of each storage device. We evaluate SnuQS on a 64-core CPU and 4-GPU system with 80 2TB HDDs and 10 4TB NVMe SSDs using quantum supremacy and quantum Fourier transform circuits. The experimental result indicates that SnuQS and the proposed I/O subsystem together is an effective and practical solution to scale the full-state simulation of large quantum circuits at about 300X lower cost than the DDR4 DRAM main-memory-only system.
Daeyoung Park, Heehoon Kim, Jinpyo Kim, Taehyun Kim 0002, Jaejin Lee
ICS3
2021 : Near-Storage Accelerator for High-Performance Log Analytics
abstract
This paper presents, a log analytics platform with near-storage accelerators for high-performance, cost- and power-efficient unstructured log processing. offloads log analytics queries to an efficient near-storage FPGA implementation of a token querying engine, which can take advantage of the high internal bandwidth of storage devices within the available chip resource limitations. This engine is flexible enough to handle complex queries including template search based on user-defined tree-based template libraries, as well as concurrent execution of multiple queries. also uses a log-optimized version of a simple, high-throughput compression algorithm in order to further improve the effective bandwidth of backing storage.
Seongyoung Kang, Jiyoung An, Jinpyo Kim, Sang Woo Jun
MICRO3
2014 An end-to-end analysis of file system features on sparse virtual disks
abstract
Software Defined Data Center (SDDC) is now an emerging area drawing considerable attention in enterprise computing. Software-Defined Storage (SDS), as a key element to enable the SDDC concept, is considered one of the most disruptive storage technologies in modern times. SDS introduces a variety of novel features and functionalities thereby changing the traditional view of the storage stack. In VMware's ESXi virtualization platform, several sparse virtual disk formats have been implemented to support critical features for SDS such as virtual machine (VM) snapshots, Fault-Tolerance (FT), thin provisioning and linked clones. Each virtual disk format supports unique features that may incur complex interactions with other layers of the storage stack such as guest file systems and storage devices.
Ruijin Zhou, Sankaran Sivathanu, Jinpyo Kim, Bing Tsai, Tao Li 0006
ICS3
2007 COBRA: An Adaptive Runtime Binary Optimization Framework for Multithreaded Applications
abstract
This paper presents COBRA (continuous binary re-adaptation), a runtime binary optimization framework, for multithreaded applications. It is currently implemented on Itanium 2 based SMP and cc-NUMA systems. Using OpenMP NAS parallel benchmark, we show how COBRA can adoptively choose appropriate optimizations according to observed changing runtime program behavior. Coherent cache misses caused by true/false data sharing often limit the scalability of multithreaded applications. This paper shows that COBRA can significantly improve the performance of some applications parallelized with OpenMP, by reducing the aggressiveness of data prefetching and by using exclusive hints for prefetch instructions. For example, we show that COBRA can improve the performance of OpenMP NAS parallel benchmarks up to 68%, with an average of 17.5% on the SGI Altix cc-NUMA system.
Jinpyo Kim, Wei-Chung Hsu, Pen-Chung Yew
ICPP1
2005 Performance of Runtime Optimization on BLAST
abstract
Optimization of a real world application BLAST is used to demonstrate the limitations of static and profile-guided optimizations and to highlight the potential of runtime optimization systems. We analyze the performance profile of this application to determine performance bottlenecks and evaluate the effect of aggressive compiler optimizations on BLAST. We find that applying common optimizations (e.g. O3) can degrade performance. Profile guided optimizations do not show much improvement across the board, as current implementations do not address critical performance bottlenecks in BLAST. In some cases, these optimizations lower performance significantly due to unexpected secondary effects of aggressive optimizations. We also apply runtime optimization to BLAST using the ADORE framework. ADORE is able to detect performance bottlenecks and deploy optimizations resulting in performance gains up to 58% on some queries using data cache prefetching.
Abhinav Das, Jiwei Lu, Howard Chen 0002, Jinpyo Kim, Pen-Chung Yew, Wei-Chung Hsu, Dong-yuan Chen
CGO4
2005 Dynamic Code Region (DCR) Based Program Phase Tracking and Prediction for Dynamic Optimizations
Jinpyo Kim, Sreekumar V. Kodakara, Wei-Chung Hsu, David J. Lilja, Pen-Chung Yew
HiPEAC1