Kwangsik Shin

dblp:79/8381 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
6since 2021 · last 2026
0009-0000-5144-5673ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 2 first-author · 6 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 AXLE: Coordinated Offloading with Asynchronous Back-Streaming in Computational Memory Systems
Suyeon Lee, Kangkyu Park, Kwangsik Shin, Ada Gavrilovska
ISCA3
2026 Toward Deployable CXL-PNM: The CMM-Ax Prototype and Software Stack
abstract
Processing-Near-Memory (PNM) over Compute Express Link (CXL) has strong architectural appeal. However, most existing CXL-attached PNM prototypes remain inflexible or simulation-based and lack integration with full software stacks, which limits their practical deployment. This work introduces CMM-Ax, a deployable hybrid CXL-PNM system that combines streaming near-memory pipelines with device-side programmability and integrates tightly with the Heterogeneous Memory Software Development Kit (HMSDK). Through HMSDK, CMM-Ax exposes CXL-attached memory via standard Linux allocation paths (e.g., malloc and mmap). This allows existing vector databases and search engines to adopt heterogeneous memory with minimal application-level changes, such as selecting the CMM-Ax FAISS backend. CMM-Ax further provides a comprehensive software stack—FAISS integration, a domain-specific compiler for user-defined operators, and a Kubernetes device plugin supporting multi-tenant slicing—capabilities not demonstrated in prior CXL-PNM. Using an FPGA-based CXL Type-3 prototype, CMM-Ax achieves 4.54× higher throughput and 5.56× lower energy per query than a CXL memory-only baseline on the exact k-nearest neighbor (kNN) search. For approximate nearest neighbor (ANN) inverted-file (IVF) workloads, CMM-Ax sustains bandwidth-proportional efficiency at small batches (batch=1: 37% vs. 12% CPU, 7% GPU). In Kubernetes deployments, CMM-Ax–equipped servers can replace a significantly larger CPU-only cluster at equivalent memory capacity and throughput, reducing node count and system energy.
Kwangsik Shin, KangKyu Park, Joonseop Sim, Thomas Won Ha Choi, Youngpyo Joo, Hoshik Kim
IEEE Trans. Computers1
2024 MTM: Rethinking Memory Profiling and Migration for Multi-Tiered Large Memory
abstract
Multi-terabyte large memory systems are often characterized by more than two memory tiers with different latency and bandwidth. Multi-tiered large memory systems call for rethinking of memory profiling and migration because of the unique problems unseen in the traditional memory systems with smaller capacity and fewer tiers. We develop MTM, an application-transparent Multi-Tiered Memory management framework, based on three principles: (1) connecting the control of profiling overhead with the profiling mechanism for high-quality profiling; (2) building a universal page migration policy on the complex multi-tiered memory for high performance; and (3) introducing huge page awareness. We evaluate MTM using common big-data applications with realistic working sets (hundreds of GB to 1 TB). MTM outperforms seven solutions by up to 42% (17% on average).
Jie Ren 0015, Dong Xu 0024, Junhee Ryu, Kwangsik Shin, Daewoo Kim, Dong Li 0001
EuroSys4
2024 Computational CXL-Memory Solution for Accelerating Memory-Intensive Applications
abstract
CXL interface is the up-to-date technology that enables effective memory expansion by providing a memory-sharing protocol in configuring heterogeneous devices. However, its limited physical bandwidth can be a significant bottleneck for emerging data-intensive applications. In this work, we propose a novel CXL-based memory disaggregation architecture with a real-world prototype demonstration, which overcomes the bandwidth limitation of the CXL interface using near-data processing. The experimental results demonstrate that our design achieves up to 1.9× better performance/power efficiency than the existing CPU system.
Joonseop Sim, Soohong Ahn, Taeyoung Ahn, Seungyong Lee 0005, Myunghyun Rhee, Kwangsik Shin, Donguk Moon, Euiseok Kim, Kyoung Park
HPCA7
2024 Efficient Tensor Offloading for Large Deep-Learning Model Training based on Compute Express Link
abstract
The deep learning models (DL) are becoming bigger, easily beyond the memory capacity of a single accelerator. The recent progress in large DL training utilizes CPU memory as an extension of accelerator memory and offloads tensors to CPU memory to save accelerator memory. This solution transfers tensors between the two memories, creating a major performance bottleneck. We identify two problems during tensor transfers: (1) the coarse-grained tensor transfer creating difficulty in hiding transfer overhead, and (2) the redundant transfer that unnecessarily migrates value-unchanged bytes from CPU to accelerator. We introduce a cache coherence interconnect based on Compute Express Link (CXL) to build a cache coherence domain between CPU memory and accelerator memory. By slightly extending CXL to support an update cache-coherence protocol and avoiding unnecessary data transfers, we reduce training time by $33.7 \%$ (up to $55.4 \%$) without changing model convergence and accuracy, compared with the state-of-the-art work in DeepSpeed [62].
Dong Xu 0024, Kwangsik Shin, Daewoo Kim, Hyeran Jeon, Dong Li 0001
SC3
2024 FlexMem: Adaptive Page Profiling and Migration for Tiered Memory
Dong Xu 0024, Junhee Ryu, Kwangsik Shin, Dong Li 0001
USENIX ATC3
2010 Online Gaming Traffic Generator for Reproducing Gamer Behavior
Kwangsik Shin, Jinhyuk Kim, Kangmin Sohn, Changjoon Park, Sangbang Choi
ICEC1
2008 Task scheduling algorithm using minimized duplications in homogeneous systems
Kwangsik Shin, MyongJin Cha, MunSuck Jang, JinHa Jung, Wanoh Yoon, Sangbang Choi
J. Parallel Distributed Comput.1