Si-Dong Roh

dblp:279/1308 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
3since 2021 · last 2026
0000-0001-5961-948XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 3 since 2021
YearPublicationVenuePosition
2026 Clone: A Collaborative Multi-device System for Retrieval-Augmented Generation over CXL
abstract
As vector databases scale in Retrieval-Augmented Generation (RAG), the retrieval phase increasingly bottlenecks end-to-end latency. While Compute Express Link (CXL) offers scalable memory expansion, naïve CXL deployments suffer from intra-device bandwidth saturation and inter-device load imbalance, which collectively hinder system responsiveness.
Seoyoung Ko, Wanju Doh, Eojin Na, Hyunjeong Shim, Sungmin Yun 0001, Jinin So, Yongsuk Kwon, Sang-Soo Park, Si-Dong Roh, Minyong Yoon, Taeksang Song, Eojin Lee, Jung Ho Ahn
ICS9
2025 Scalable Processing-Near-Memory for 1M-Token LLM Inference: CXL-Enabled KV-Cache Management Beyond GPU Limits
abstract
The expansion of context windows in large language models (LLMs) to multi-million tokens introduces severe memory and compute bottlenecks, particularly in managing the growing Key-Value (KV) cache. While Compute Express Link (CXL) enables non-eviction frameworks that offload the full KV-cache to scalable external memory, these frameworks still suffer from costly data transfers when recalling non-resident KV tokens to limited GPU memory as context lengths increase. This work proposes scalable Processing-NearMemory (PNM) for 1M-Token LLM Inference, a CXL-enabled KVcache management system that coordinates memory and computation beyond GPU limits. Our design offloads token page selection to a PNM accelerator within CXL memory, eliminating costly recalls and enabling larger GPU batch sizes. We further introduce a hybrid parallelization strategy and a steady-token selection mechanism to enhance compute efficiency and scalability. Implemented atop a state-of-the-art CXL-PNM system, our solution delivers consistent performance gains for LLMs with up to 405B parameters and 1Mtoken contexts. Our PNM-only offloading scheme (PNM-KV) and GPU-PNM hybrid with steady-token execution (PnG-KV) achieve up to $21.9 \times$ throughput improvement, up to $60 \times$ lower energy per token, and up to $7.3 \times$ better total cost efficiency than the baseline, demonstrating that CXL-enabled multi-PNM architectures can serve as a scalable backbone for future long-context LLM inference.
Janghyeon Kim, Hyucksung Kwon, Hyeonggyu Jeong, Sang-Soo Park, Minyong Yoon, Si-Dong Roh, Yongsuk Kwon, Jinin So, Jungwook Choi
PACT8
2023 Token Merging with Class Importance Score
abstract
Vision Transformers have achieved high performance in computer vision tasks, but their high computational cost and low throughput are weaknesses. Therefore, much research has been done to reduce the size of Vision Transformers. Among them, studies on pruning unnecessary tokens are being actively conducted to reduce the number of tokens used for self-attention computation inside the Vision Transformer. Recently, token merging has been proposed as a new alternative approach. These studies aim to increase throughput with a small accuracy drop by merging similar tokens instead of pruning them. A previous study finds similar tokens using cosine similarity and merges them with a weighted average. However, merging a large number of tokens at once may lead to an accuracy drop because of the underestimating of important information. In this paper, we propose ToMeCIS, a method that merges similar tokens through a weighted average using the class importance score of tokens to reduce the accuracy drop. When ToMeCIS is applied to a pretrained DeiT-S and evaluated on the ImageNet-1k dataset, the throughput is increased by about 50% with an accuracy drop of less than 1% without additional training. In addition, importance scores were evaluated with different metrics to find the best accuracy versus throughput trade-off.
Kwang-Soo Seol, Si-Dong Roh, Ki-Seok Chung
IECON2
2020 Direct Conversion: Accelerating Convolutional Neural Networks Utilizing Sparse Input Activation
abstract
The amount of computation and the number of parameters of neural networks are increasing rapidly as the depth of convolutional neural networks (CNNs) is increasing. Therefore, it is very crucial to reduce both the amount of computation and that of memory usage. The pruning method, which compresses a neural network, has been actively studied. Depending on the layer characteristics, the sparsity level of each layer varies significantly after the pruning is conducted. If weights are sparse, most results of convolution operations will be zeroes. Although several studies have proposed methods to utilize the weight sparsity to avoid carrying out meaningless operations, those studies lack consideration that input activations may also have a high sparsity level. The Rectified Linear Unit (ReLU) function is one of the most popular activation functions because it is simple and yet pretty effective. Due to properties of the ReLU function, it is often observed that the input activation sparsity level is high (up to 85%). Therefore, it is important to consider both the input activation sparsity and the weight one to accelerate CNN to minimize carrying out meaningless computation. In this paper, we propose a new acceleration method called Direct Conversion that considers the weight sparsity under the sparse input activation condition. The Direct Conversion method converts a 3D input tensor directly into a compressed format. This method selectively applies one of two different methods: a method called image to Compressed Sparse Row (im2CSR) when input activations are sparse and weights are dense; the other method called image to Compressed Sparse Overlapped Activations (im2CSOA) when both input activations and weights are sparse. Our experimental results show that Direct Conversion improves the inference speed up to 2.82× compared to the conventional method.
Wonhyuk Lee, Si-Dong Roh, Sangki Park, Ki-Seok Chung
IECON2