Seokwon Kang

dblp:241/4321 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
7since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 3 first-author · 5 since 2021Software engineering, systems software and programming languages · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 2 · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Intelligent Supply Chain for Communication Navigation and Caching in Multi-UAV Wireless Network
abstract
The emergence of unmanned aerial vehicles (UAVs) in wireless communication has opened up new prospects for improving network performance and user experience. This study aims to create a smart supply chain that enables collaborative communication, navigation, and caching in multi-UAV wireless networks. To achieve this, a user-UAV clustering strategy using the Whale Optimization Algorithm (WOA) optimizes resource distribution and network management. Additionally, Multi-Agent Deep Deterministic Policy Gradients (MADDPG) jointly optimize UAV trajectory planning, bandwidth allocation, and caching decisions, enabling UAVs to enhance trajectory, allocate resources, and manage cached content efficiently. Experimental results show a 10% reduction in average response time and a 12% decrease in system energy consumption, demonstrating the efficiency of the proposed approach in improving data delivery and extending UAV lifespan for sustainable network operations.
Seokwon Kang, Md. Shirajum Munir, Choong Seon Hong
IEEE Trans. Netw. Serv. Manag.1
2025 Supporting Register-based Addressing Modes for in-DRAM PIM ISAs
abstract
Processing-in-Memory architecture presents a promising solution to alleviate the data movement bottleneck that arises from transferring data between memory and compute units in traditional processor-centric systems, particularly for DNN applications. However, this architecture introduces two inherent overheads: PIM code offloading and data transferring between CPU and memory. To address these issues, we propose two register-based addressing modes, indexed and base-offset addressing, for DMA descriptor-based in-DRAM PIM ISAs. Our full-system performance evaluation demonstrates that the approach significantly reduces the overheads, resulting in up to 1.94x speedup compared to the baseline PIM, additionally only with $4.65 \%$ area and $8.61 \%$ power consumption.
Seokyoung Kim 0001, Byung Ho Choi, Seokwon Kang, Yongjun Park 0001, Seon Wook Kim
DAC3
2025 PIM-CARE: A Compiler-Assisted Dynamic Resource Allocation Framework for Real-world DRAM PIM
abstract
Processing-In-Memory (PIM) has recently emerged as a promising solution to alleviate the memory bottleneck by integrating computing capabilities into memory chips. Since PIM provides numerous Processing Elements (PEs) and high-bandwidth on-chip data transfers, full utilization of the PEs becomes a critical mission to maximize the performance of PIM applications. However, due to the diverse and complex characteristics of PIM applications, using more resources does not always improve performance. It is therefore important to find the suitable amount of resources to achieve the best performance and to fully utilize the PIM resources.To address this, we introduce PIM-CARE, a framework for dynamic resource allocation across multiple applications with compiler support on real-world PIM systems. PIM-CARE first determines the best amount of PIM resources to allocate for each application. To enable spatial multitasking, the PIM-CARE daemon monitors resource allocation and deallocation requests and estimates total PIM resource utilization at runtime. It then dynamically schedules applications using a priority-based out-of-order policy, considering both available PIM resources and resource requirements for best performance. Evaluation on real-world PIM systems shows that PIM-CARE improves throughput by 5.49x and average turnaround time by 5.71x compared to the baseline.
Inyong Hwang, Donghyeon Kim 0001, Seokwon Kang, Taehyeong Park 0001, Taehoon Kim 0001, Jiwon Seo 0002, Hanjun Kim 0001, Youngsok Kim, Yongjun Park 0001
ICS3
2025 PIM-CCA: An Efficient PIM Architecture with Optimized Integration of Configurable Functional Units
abstract
Processing-in-Memory (PIM) is a promising architecture for alleviating data movement bottlenecks by performing computations closer to memory.However, PIM workloads often encounter computational bottlenecks within the PIM itself.As these workloads become more compute-intensive by leveraging PIM's high internal bandwidth, a small set of hot code regions emerges as the primary performance bottleneck.Unfortunately, increasing the complexity of the PIM processor is difficult due to inherent memory constraints, such as area and power.Therefore, enhancing the computational capability of PIM while maintaining a lightweight design within limited silicon budgets remains highly challenging.In this paper, we propose PIM-CCA, a novel PIM architecture that integrates a Configurable Compute Accelerator (CCA) to mitigate computational bottlenecks with minimal hardware overhead.The CCA-enabled PIM design allows for the flexible configuration of compute logic, enabling acceleration across diverse workloads.The PIM-CCA compiler constructs an instruction-level dataflow graph to identify hot and compute-bound regions and offload them to the CCA.Furthermore, we analyze the interaction between the PIM threading model and resource utilization to derive the optimal thread count for efficient CCA-enabled PIM usage.We implement PIM-CCA in a cycle-accurate simulator based on a commercially available PIM system, and evaluate it using 14 representative benchmarks.The experimental results show that PIM-CCA achieves up to 1.55× performance improvement over baseline PIM systems, with only 0.036% additional area overhead, based on P&R results with limited metal layers.
Jeehyun Kim, Donghyeon Kim 0001, Seokwon Kang, Bongjoon Hyun, Inho Lee 0002, Yongjun Park 0001
MICRO3
2024 ISP Agent: A Generalized In-storage-processing Workload Offloading Framework by Providing Multiple Optimization Opportunities
abstract
As solid-state drives (SSDs) with sufficient computing power have recently become the dominant devices in modern computer systems, in-storage processing (ISP), which processes data within the storage without transferring it to the host memory, is being utilized in various emerging applications. The main challenge of ISP is to deliver storage data to the offloaded workload. This is difficult because of the information gap between the host and storage, the data consistency problem between the host and offloaded workloads, and SSD-specific hardware limitations. Moreover, because the offloaded workloads use internal SSD resources, host I/O performance might be degraded due to resource conflicts. Although several ISP frameworks have been proposed, existing ISP approaches that do not deeply consider the internal SSD behavior are often insufficient to support efficient ISP workload offloading with high programmability. In this article, we propose an ISP agent, a lightweight ISP workload offloading framework for SSD devices. The ISP agent provides I/O and memory interfaces that allow users to run existing function codes on SSDs without major code modifications, and separates the resources for the offloaded workloads from the existing SSD firmware to minimize interference with host I/O processing. The ISP agent also provides further optimization opportunities for the offloaded workload by considering SSD architectures. We have implemented the ISP agent on the OpenSSD Cosmos+ board and evaluated its performance using synthetic benchmarks and a real-world ISP-assisted database checkpointing application. The experimental results demonstrate that the ISP agent enhances host application performance while increasing ISP programmability, and that the optimization opportunities provided by the ISP agent can significantly improve ISP-side performance without compromising host I/O processing.
Seokwon Kang, Jongbin Kim, Gyeongyong Lee, Jeongmyung Lee, Jiwon Seo 0002, Hyungsoo Jung 0001, Yong Ho Song, Yongjun Park 0001
ACM Trans. Archit. Code Optim.1
2023 Orchestrating Large-Scale SpGEMMs using Dynamic Block Distribution and Data Transfer Minimization on Heterogeneous Systems
abstract
Sparse general matrix-matrix multiplication (SpGEMM) is a major kernel in various emerging applications, such as database management systems, deep learning, graph analysis, and recommendation systems. Since SpGEMM requires extensive computation, many SpGEMM techniques have been implemented based on graphics processing units (GPUs) to exploit massive data parallelism completely. However, traditional SpGEMM techniques usually do not fully utilize the GPU because most non-zero elements of the target sparse matrices exist in a few hub nodes, and non-hub nodes barely have non-zero elements. The data-related characteristics (power law) result in a significant degradation in performance because of the load imbalance between the GPU cores and the low utilization of each core. Many attempts have been made through recent implementations to solve this problem using smart pre-/post-processing. However, the net performance hardly improves and sometimes even deteriorates owing to the large overheads. Additionally, non-hub nodes are inherently not suitable for GPU computing, even after optimization. Furthermore, the performance is no longer dominated by kernel execution, but by data transfers such as device-to-host (D2H) data transfers and file I/Os, owing to the rapid growth in the computing power of GPUs and input data size.Therefore, this work proposes a Dynamic Block Distributor (DBD), a novel full-system-level SpGEMM orchestration framework for heterogeneous systems, improving the overall performance by enabling an efficient CPU-GPU collaboration and further minimizing the overhead in data transfer between all the system elements. This framework first divides the target matrix into smaller blocks and then offloads the computation of each block to an appropriate computing unit between a GPU and CPU based on its workload type and the status of resource utilization at runtime. It also minimizes the overhead in data transfer with simple but suitable techniques, such as Row Collecting, I/O Overlapping, and I/O Binding. Our experiments showed that this framework increased the execution latency of SpGEMM, which included both the kernel execution and D2H transfers, by 3.24x on average, and the overall execution time by 2.07x on average, compared to that of the baseline cuSPARSE library.
Taehyeong Park 0001, Seokwon Kang, Myung-Hwan Jang, Sang-Wook Kim, Yongjun Park 0001
ICDE2
2022 Networked SSD: Flash Memory Interconnection Network for High-Bandwidth SSD
abstract
As the flash memory performance increases with more bandwidth, the flash memory channel or the interconnect is becoming a bigger bottleneck to enable high performance SSD system. However, the bandwidth of the flash memory interconnect is not increasing at the same rate as the flash memory. In addition, current flash memory bus is based on dedicated signaling where separate control signals are used for communication between the flash channel controller and the flash memory chip. In this work, we propose to exploit packetized communication to improve the effective flash memory interconnect bandwidth and propose packetized SSD (pSSD) system architecture. We first show how packetized communication can be exploited and the microarchitectural changes required. We then propose the Omnibus topology for flash memory interconnect to enable a packetized network SSD (pnSSD) among the flash memory – a 2D bus-based organization that maintains a “bus” organization for the interconnect while enabling direct communication between the flash memory chips. The pnSSD architecture enables a new type of garbage collection that we refer to as spatial garbage collection that significantly reduces the interference between I/O requests and garbage collection. Our detailed evaluation of pnSSD shows 82% improvement in I/O latency with no garbage collection (GC) while improving I/O latency by 9.71× when GC occurs in parallel with I/O operation, through spatial garbage collection.
Seokwon Kang, Yongjun Park 0001, John Kim 0001
MICRO2
2020 PreScaler: an efficient system-aware precision scaling framework on heterogeneous systems
abstract
Graphics processing units (GPUs) have been commonly utilized to accelerate multiple emerging applications, such as big data processing and machine learning. While GPUs are proven to be effective, approximate computing, to trade off performance with accuracy, is one of the most common solutions for further performance improvement. Precision scaling of originally high-precision values into lower-precision values has recently been the most widely used GPU-side approximation technique, including hardware-level half-precision support. Although several approaches to find optimal mixed-precision configuration of GPU-side kernels have been introduced, total program performance gain is often low because total execution time is the combination of data transfer, type conversion, and kernel execution. As a result, kernel-level scaling may incur high type-conversion overhead of the kernel input/output data. To address this problem, this paper proposes an automatic precision scaling framework called PreScaler that maximizes the program performance at the memory object level by considering whole OpenCL program flows. The main difficulty is that the best configuration cannot be easily predicted due to various application- and system-specific characteristics. PreScaler solves this problem using search space minimization and decision-tree-based search processes. First, it minimizes the number of test configurations based on the information from system inspection and dynamic profiling. Then, it finds the best memory-object level mixed-precision configuration using a decision-tree-based search. PreScaler achieves an average performance gain of 1.33x over the baseline while maintaining the target output quality level.
Seokwon Kang, Kyunghwan Choi, Yongjun Park 0001
CGO1
2020 Optimization of GPU-based Sparse Matrix Multiplication for Large Sparse Networks
abstract
Sparse matrix multiplication (spGEMM) is widely used to analyze the sparse network data, and extract important information based on matrix representation. As it contains a high degree of data parallelism, many efficient implementations using data-parallel programming platforms such as CUDA and OpenCL have been introduced on graphic processing units (GPUs). Several well-known spGEMM techniques, such as cuS- PARSE and CUSP, often do not utilize the GPU resources fully, owing to the load imbalance between threads in the expansion process and high memory contention in the merge process. Furthermore, even though several outer-product-based spGEMM techniques are proposed to solve the load balancing problem on expansion, they still do not utilize the GPU resources fully, because severe computation load variations exist among the multiple thread blocks.To solve these challenges, this paper proposes a new optimization pass called Block Reorganizer, which balances the total computations of each computing unit on target GPUs, based on the outer-product-based expansion process, and reduces the memory pressure during the merge process. For expansion, it first identifies the actual computation amount for each block, and then performs two thread block transformation processes based on their characteristics: 1) B-Splitting to transform a heavy-computation blocks into multiple small blocks and 2) B- Gathering to aggregate multiple small-computation blocks to a larger block. While merging, it improves the overall performance by performing B-Limiting to limit the number of blocks on each computing unit. Experimental results show that it improves the total performance of kernel execution by 1.43x, on an average, when compared to the row-product-based spGEMM, for NVIDIA Titan Xp GPUs on real-world datasets.
Jeongmyung Lee, Seokwon Kang, Yongseung Yu, Yong-Yeon Jo, Sang-Wook Kim, Yongjun Park 0001
ICDE2
2019 GATE: A Generalized Dataflow-level Approximation Tuning Engine For Data Parallel Architectures
abstract
Although approximate computing is widely used, it requires substantial programming effort to find appropriate approximation patterns among multiple pre-defined patterns to achieve a high performance. Therefore, we propose an automatic approximation framework called GATE to uncover hidden opportunities from any data-parallel program regardless of the code pattern or application characteristics using two compiler techniques, namely subgraph-level approximation (SGLA) and approximate thread merge(ATM). GATE also features conservative/aggressive tuning and dynamic calibration to maximize the performance while maintaining the TOQ level during runtime. Our framework achieves an average performance gain of 2.54x over the baseline with minimum accuracy loss.
Seokwon Kang, Yongseung Yu, Yongjun Park 0001
DAC1
2019 A compiler-based approach for GPGPU performance calibration using TLP modulation (WIP paper)
abstract
Modern GPUs are the most successful accelerators as they provide outstanding performance gain by using CUDA or OpenCL programming models. For maximum performance, programmers typically try to maximize the number of thread blocks of target programs, and GPUs also generally attempt to allocate the maximum number of thread blocks to their GPU cores. However, many recent studies have pointed out that simply allocating the maximum number of thread blocks to GPU cores does not always guarantee the best performance, and identifying proper number of thread blocks per GPU core is a major challenge. Despite these studies, most existing architectural techniques cannot be directly applied to current GPU hardware, and the optimal number of thread blocks can vary significantly depending on the target GPU and application characteristics. To solve these problems, this study proposes a just-in-time thread block number adjustment system using CUDA binary modification upon an LLVM compiler framework, referred to as the CTA-Limiter, in order to dynamically maximize GPU performance on real GPUs without reprogramming. The framework gradually reduces the number of concurrent thread blocks of target CUDA workloads using extra shared memory allocation, and compares the execution time with the previous version to automatically identify the optimal number of co-running thread blocks per GPU Core. The results showed meaningful performance improvements, averaging at 30%, 40%, and 44%, in GTX 960, GTX 1050, and GTX 1080 Ti, respectively.
Yongseung Yu, Seokwon Kang, Yongjun Park 0001
LCTES2
2018 Runtime Profiling of OpenCL Workloads Using LLVM-based Code Instrumentation
abstract
GPUs, which are widely used high-performance hardware accelerators in heterogeneous computing, and programming models for architectures such as OpenCL and CUDA, have recently been developed to achieve high productivity. LLVM is an open-source compiler infrastructure that enables low-level optimization through LLVM intermediate representation (LLVM IR) in various programming language environments. In this paper, we propose a fully-automatic Dynamic Profiling framework which performs instruction-level analysis through IR-level code instrumentation for typical OpenCL workload kernels.
Yongseung Yu, Seokwon Kang, Yongjun Park 0001
TENCON2