Jinin So

dblp:304/5206 · DBLP profile ↗
← Back
9ranked-venue papers
0as first author
9since 2021 · last 2026
0000-0002-7569-3505ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 7 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Clone: A Collaborative Multi-device System for Retrieval-Augmented Generation over CXL
abstract
As vector databases scale in Retrieval-Augmented Generation (RAG), the retrieval phase increasingly bottlenecks end-to-end latency. While Compute Express Link (CXL) offers scalable memory expansion, naïve CXL deployments suffer from intra-device bandwidth saturation and inter-device load imbalance, which collectively hinder system responsiveness.
Seoyoung Ko, Wanju Doh, Eojin Na, Hyunjeong Shim, Sungmin Yun 0001, Jinin So, Yongsuk Kwon, Sang-Soo Park, Si-Dong Roh, Minyong Yoon, Taeksang Song, Eojin Lee, Jung Ho Ahn
ICS6
2026 S-Tiering: A Unified HW/SW Solution for Memory Tiering Based on the Standard CXL Hotness Monitoring Unit
abstract
In this paper, we proposeS-Tiering, a unified hardware and software solution for memory tiering based on the standard CXL Hotness Monitoring Unit (CHMU).S-Tieringconsists of hardware components that comply with CHMU hardware specification defined in the CXL 3.2 Specification, and software components that control hardware components. Based on these components,S-Tieringminimizes the access to CXL memory by properly steering page migration.We evaluate various probabilistic data structure algorithms and adopt a Count-Min Sketch-based Hot Page Tracker that achieves 99% accuracy with only 0.3% tracking buffer overhead compared to assigning a dedicated counter for every 4KB page. We implement hardware components ofS-Tieringon an Field-Programmable Gate Array board and software components ofS-Tieringon Ubuntu 22.04 with Linux-v6.8 kernel. We evaluate the performance impact ofS-Tieringon benchmarks representative of real applications (e.g., High Performance Computing, Graph-processing, In-Memory Database).S-Tieringachieves a performance improvement of up to 193% compared to first-touch allocation and outperforms AutoNUMA memory tiering by 184%p.S-Tieringminimizes the memory access to CXL memory and increases the bandwidth utilization of DDR memory up to ×11.
Seunghak Lee, Wonjae Lee 0001, Hojin Nam, Jehoon Park, Youngshin Park, Junhyeok Im, Jinin So, Raghu Vamsi Krishna Talanki, Praful Ramesh O, Rajeev Verma, Taeksang Song, Wonhwa Shin, Sangjoon Hwang 0001
IEEE Trans. Computers8
2026 Pangaea v2: CXL-Based Disaggregated Memory System Architecture for Cloud-Native Orchestration
abstract
Today’s data centers suffer from CPU and memory resource stranding because they often over-provision resources when deploying servers for worst-case scenarios. This problem gives rise to a disaggregated system architecture allowing each type of resource to be allocated, utilized and freed separately as required. In particular, research on disaggregated memory systems over the past few years has focused primarily on achieving low remote memory access latency over Ethernet, which is known as the RDMA optimization approach.In this paper, we introduce a dynamic rack-scale disaggregated memory system architecture, so called Pangaea v2 using ASIC-CXL H/W and memory orchestration S/W designed to increase the memory utilization of worker nodes between containerized applications execution in a Kubernetes, a major process container platform in the data center. In our evaluation with in-memory database application, disaggregated CXL memory system shows significantly better throughput improved by up to 10.2x/6.7x and 99th tail latency reduced to 96%/93% compared to RDMA with RoCEv2/InfiniBand.
Han Deok Lee, Jehoon Park, Younghyun Lee, Junhyeok Im, Jin Jung, Jinin So, Siamak Tavallaei, Woo Taek Shim, Chin-Hua Chang, Sungwook Ryu, Taeksang Song, Wonhwa Shin, Sangjoon Hwang 0001
IEEE Trans. Computers6
2025 Scalable Processing-Near-Memory for 1M-Token LLM Inference: CXL-Enabled KV-Cache Management Beyond GPU Limits
abstract
The expansion of context windows in large language models (LLMs) to multi-million tokens introduces severe memory and compute bottlenecks, particularly in managing the growing Key-Value (KV) cache. While Compute Express Link (CXL) enables non-eviction frameworks that offload the full KV-cache to scalable external memory, these frameworks still suffer from costly data transfers when recalling non-resident KV tokens to limited GPU memory as context lengths increase. This work proposes scalable Processing-NearMemory (PNM) for 1M-Token LLM Inference, a CXL-enabled KVcache management system that coordinates memory and computation beyond GPU limits. Our design offloads token page selection to a PNM accelerator within CXL memory, eliminating costly recalls and enabling larger GPU batch sizes. We further introduce a hybrid parallelization strategy and a steady-token selection mechanism to enhance compute efficiency and scalability. Implemented atop a state-of-the-art CXL-PNM system, our solution delivers consistent performance gains for LLMs with up to 405B parameters and 1Mtoken contexts. Our PNM-only offloading scheme (PNM-KV) and GPU-PNM hybrid with steady-token execution (PnG-KV) achieve up to $21.9 \times$ throughput improvement, up to $60 \times$ lower energy per token, and up to $7.3 \times$ better total cost efficiency than the baseline, demonstrating that CXL-enabled multi-PNM architectures can serve as a scalable backbone for future long-context LLM inference.
Janghyeon Kim, Hyucksung Kwon, Hyeonggyu Jeong, Sang-Soo Park, Minyong Yoon, Si-Dong Roh, Yongsuk Kwon, Jinin So, Jungwook Choi
PACT10
2025 Accelerating Confidential Recommendation Model Inference With Near-Memory Processing
abstract
Trusted Executing Environments (TEEs) in hardware designs protect program execution from other untrusted software programs in the processor as well as untrusted off-chip hardware components. Meanwhile, Near-Memory Processing (NMP) has shown performance and energy benefits on memory-intensive workloads. Recently, novel memory encryption schemes have been proposed to allow TEEs to leverage the benefits of NMP without requiring trust in the NMP components. In this paper, we present a system design of confidential computing with NMP that can be directly used in Intel SGX, a TEE platform available in commercial processors today. We develop the full software stack and evaluate the results on commercial processors with the emulated AxDIMM, an FPGA-based NMP platform. In our case study on personalized Deep Learning Recommendation Model (DLRM) inference, the proposed confidential computing in NMP achieves up to 1.51× latency reduction and up to 2.57× throughput improvement.
Wenjie Xiong 0001, Liu Ke 0001, Maxim Ostapenko, Yongmin Tai, Yeongon Cho, Joon-Ho Song, Jinin So, Kyungsoo Kim 0003, Yongsuk Kwon, Jin Jung, Byeongho Kim, Shinhaeng Kang, Sukhan Lee 0002, Jeonghyeon Cho, Kyomin Sohn, Xuan Zhang 0001, Hsien-Hsin S. Lee, G. Edward Suh
IEEE Trans. Dependable Secur. Comput.7
2024 An LPDDR-based CXL-PNM Platform for TCO-efficient Inference of Transformer-based Large Language Models
abstract
Transformer-based large language models (LLMs) such as Generative Pre-trained Transformer (GPT) have become popular due to their remarkable performance across diverse applications, including text generation and translation. For LLM training and inference, the GPU has been the predominant accelerator with its pervasive software development ecosystem and powerful computing capability. However, as the size of LLMs keeps increasing for higher performance and/or more complex applications, a single GPU cannot efficiently accelerate LLM training and inference due to its limited memory capacity, which demands frequent transfers of the model parameters needed by the GPU to compute the current layer(s) from the host CPU memory/storage. A GPU appliance may provide enough aggregated memory capacity with multiple GPUs, but it suffers from frequent transfers of intermediate values among GPU devices, each accelerating specific layers of a given LLM. As the frequent transfers of these model parameters and intermediate values are performed over relatively slow device-to-device interconnects such as PCIe or NVLink, they become the key bottleneck for efficient acceleration of LLMs. Focusing on accelerating LLM inference, which is essential for many commercial services, we develop CXL-PNM, a processing near memory (PNM) platform based on the emerging interconnect technology, Compute eXpress Link (CXL). Specifically, we first devise an LPDDR5X-based CXL memory architecture with 512GB of capacity and 1.1TB/s of bandwidth, which boasts 16× larger capacity and 10× higher bandwidth than GDDR6and DDR5-based CXL memory architectures, respectively, under a module form-factor constraint. Second, we design a CXLPNM controller architecture integrated with an LLM inference accelerator, exploiting the unique capabilities of such CXL memory to overcome the disadvantages of competing technologies such as HBM-PIM and AxDIMM. Lastly, we implement a CXLPNM software stack that supports seamless and transparent use of CXL-PNM for Python-based LLM programs. Our evaluation shows that a CXL-PNM appliance with 8 CXL-PNM devices offers 23% lower latency, 31% higher throughput, and 2.8× higher energy efficiency at 30% lower hardware cost than a GPU appliance with 8 GPU devices for an LLM inference service.
Sangsoo Park, Kyungsoo Kim 0003, Jinin So, Jin Jung, Jonggeon Lee, Kyoungwan Woo, Nayeon Kim 0006, Younghyun Lee, Hyungyo Kim, Yongsuk Kwon, Jinhyun Kim, Yeongon Cho, Yongmin Tai, Jeonghyeon Cho, Hoyoung Song, Jung Ho Ahn, Nam Sung Kim
HPCA3
2023 Samsung PIM/PNM for Transfmer Based AI : Energy Efficiency on PIM/PNM Cluster
Jin Hyun Kim, Yuhwan Ro, Jinin So, Sukhan Lee 0002, Shinhaeng Kang, Yeongon Cho, Byeongho Kim, Kyungsoo Kim 0003, Sangsoo Park, Jin-Seong Kim, Sanghoon Cha, Won-Jo Lee, Jin Jung, Jonggeon Lee, Joon-Ho Song, Seungwon Lee 0006, Jeonghyeon Cho, Jaehoon Yu, Kyomin Sohn
HCS3
2022 Improving In-Memory Database Operations with Acceleration DIMM (AxDIMM)
abstract
The significant overhead needed to transfer the data between CPUs and memory devices is one of the hottest issues in many areas of computing, such as database management systems. Disaggregated computing on the memory devices is being highlighted as one promising approach. In this work, we introduce a new near-memory acceleration scheme for in-memory database operations, called Acceleration DIMM (AxDIMM). It behaves like a normal DIMM through the standard DIMM-compatible interface, but has embedded computing units for data-intensive operations. With the minimized data transfer overhead, it reduces CPU resource consumption, relieves the memory bandwidth bottleneck, and boosts energy efficiency. We implement scan operations, one of the most data-intensive database operations, within AxDIMM and compare its performance with SIMD (Single Instruction Multiple Data) implementation on CPU. Our investigation shows that the acceleration achieves 6.8x more throughput than the SIMD implementation.
Donghun Lee 0001, Jinin So, Minseon Ahn, Jong-Geon Lee, Jeonghyeon Cho, Oliver Rebholz, Vishnu Charan Thummala, Ravi Shankar JV, Sachin Suresh Upadhya, Mohammed Ibrahim Khan, Jin Hyun Kim
DaMoN2
2021 Aquabolt-XL: Samsung HBM2-PIM with in-memory processing for ML accelerators and beyond
abstract
Using PIM to overcome memory bottleneck • Although various bandwidth increase methods have been proposed, it is physically impossible to achieve a breakthrough increase. - Limited by # of PCB wires, # of CPU ball, and thermal constraints • PIM has been proposed to improve performance of bandwidth-intensive workloads and improve energy efficiency by reducing computing-memory data movement.
Jin Hyun Kim, Shinhaeng Kang, Sukhan Lee 0002, Woongjae Song, Yuhwan Ro, Seungwon Lee 0006, David Wang 0003, Hyunsung Shin, BengSeng Phuah, Jihyun Choi, Jinin So, Yeongon Cho, Joon-Ho Song, Jangseok Choi, Jeonghyeon Cho, Kyomin Sohn, Young-Soo Sohn, Kwang-Il Park, Nam Sung Kim
HCS12