Sukjin Kim

dblp:01/9050 · DBLP profile ↗
← Back
12ranked-venue papers
2as first author
3since 2021 · last 2026
0009-0003-2478-7215ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 2
YearPublicationVenuePosition
2026 FlexiWalker: Extensible GPU Framework for Efficient Dynamic Random Walks with Runtime Adaptation
abstract
Dynamic random walks are fundamental to various graph analysis applications, offering advantages by adapting to evolving graph properties. Their runtime-dependent transition probabilities break down the pre-computation strategy that underpins most existing CPU and GPU static random walk optimizations. This leaves practitioners suffering from suboptimal frameworks and having to write hand-tuned kernels that do not adapt to workload diversity. To handle this issue, we present FlexiWalker, the first GPU framework that delivers efficient, workload-generic support for dynamic random walks. Our design-space study shows that rejection sampling and reservoir sampling are more suitable than other sampling techniques under massive parallelism. Thus, we devise (i) new high-performance kernels for them that eliminate global reductions, redundant memory accesses, and random-number generation. Given the necessity of choosing the best-fitting sampling strategy at runtime, we adopt (ii) a lightweight first-order cost model that selects the faster kernel per node at runtime. To enhance usability, we introduce (iii) a compile-time component that automatically specializes user-supplied walk logic into optimized building blocks. On various dynamic random walk workloads with real-world graphs, FlexiWalker outperforms the best published CPU/GPU baselines by geometric means of 73.44× and 5.91×, respectively, while successfully executing workloads that prior systems cannot support. We open-source FlexiWalker in https://github.com/AIS-SNU/FlexiWalker.
Seongyeon Park, Jaeyong Song 0002, Changmin Shin 0002, Sukjin Kim, Junguk Hong, Jinho Lee 0001
EuroSys4
2026 LoCaLUT: Harnessing Capacity-Computation Tradeoffs for LUT-Based Inference in DRAM-PIM
abstract
Lookup tables (LUTs) have recently gained attention as an alternative compute mechanism that maps input operands to precomputed results, eliminating the need for arithmetic logic. LUTs not only reduce logic complexity, but also naturally support diverse numerical precisions without requiring separate circuits for each bitwidth—an increasingly important feature in quantized DNNs. This creates a favorable tradeoff in PIM: memory capacity can be used in place of logic to increase computational throughput, aligning well with DRAM-PIM architectures that offer high bandwidth and easily available memory but limited logic density. In this work, we explore this capacity-computation tradeoff in LUT-based PIM designs, where memory capacity is traded for performance by packing multiple MAC operations into a single LUT lookup. Building on this insight, we propose LoCaLUT, a PIM-based design for efficient low-bit quantized DNN inference using operation-packed LUTs. First, we observe that these LUTs contain extensive redundancy and introduce LUT canonicalization, which eliminates duplicate entries to reduce LUT size. Second, we propose reordering LUT, a lightweight auxiliary LUT that remaps weight vectors to their canonical form required by LUT canonicalization with a simple LUT lookup. Third, we propose LUT slice streaming, a novel execution strategy that exploits the DRAM-buffer hierarchy by streaming only relevant LUT columns into the buffer and reusing them across multiple weight vectors. Evaluated on a real system based on UPMEM devices, we demonstrate a geometric mean speedup of$1.82 \times$across various numeric precisions and DNN models. We believe LoCaLUT opens a path toward scalable, low-logic PIM designs tailored for LUT-based DNN inference. Our implementation of LoCaLUT is available at https://github.com/AIS-SNU/LoCaLUT.
Junguk Hong, Changmin Shin 0002, Sukjin Kim, Si Ung Noh, Taehee Kwon 0002, Seongyeon Park, Hanjun Kim 0001, Youngsok Kim, Jinho Lee 0001
HPCA3
2025 PathWeaver: A High-Throughput Multi-GPU System for Graph-Based Approximate Nearest Neighbor Search
Sukjin Kim, Seongyeon Park, Si Ung Noh, Junguk Hong, Taehee Kwon 0002, Hunseong Lim, Jinho Lee 0001
USENIX ATC1
2018 Architectures and algorithms for user customization of CNNs
abstract
In this paper we present a convolutional neural network architecture that supports user customization through incremental transfer learning. The architecture consists of a large basic inference engine and a small augmenting engine. After training the basic inference engine and augmenting engine on a large general dataset, the basic inference engine is fixed. For user customization, only the augmenting engine is re-trained on-device using a small user specific dataset provided by the user. To accelerate the training of the augmenting engine we map this to a coarsegrained reconfigurable array processor. The complete network architecture is evaluated using the Caffe framework, and a C-code equivalent network is implemented and tested on a CGRA processor. Experiments with NIST'19 and our user-specific datasets show an increase in accuracy of the system from 76.3% to 93.2% after user customization. Mapping this code to a CGRA gives us a speed up of 45x and a 49-and 3-fold reduced energy consumption over an ARMv7 processor and a 3-way VLIW processor, respectively, showing the potential of CGRAs as DNN processors.
Barend Harris, Mansureh S. Moghaddam, Duseok Kang, Inpyo Bae, Euiseok Kim, Hyemi Min, Hansu Cho, Sukjin Kim, Bernhard Egger 0002, Soonhoi Ha, Kiyoung Choi
ASP-DAC8
2017 Incremental training of CNNs for user customization: work-in-progress
abstract
This paper presents a convolutional neural network architecture that supports transfer learning for user customization. The architecture consists of a large basic inference engine and a small augmenting engine. Initially, both engines are trained using a large dataset. Only the augmenting engine is tuned to the user-specific dataset. To preserve the accuracy for the original dataset, the novel concept of quality factor is proposed. The final network is evaluated with the Caffe framework, and our own implementation on a coarse-grained reconfigurable array (CGRA) processor. Experiments with MNIST, NIST'19, and our user-specific datasets show the effectiveness of the proposed approach and the potential of CGRAs as DNN processors.
Mansureh S. Moghaddam, Barend Harris, Duseok Kang, Inpyo Bae, Euiseok Kim, Hyemi Min, Hansu Cho, Sukjin Kim, Bernhard Egger 0002, Soonhoi Ha, Kiyoung Choi
CASES8
2017 A space- and energy-efficient code Compression/Decompression technique for coarse-grained reconfigurable architectures
Bernhard Egger 0002, Duseok Kang, Mansureh S. Moghaddam, Youngchul Cho, Yeonbok Lee, Sukjin Kim, Soonhoi Ha, Kiyoung Choi
CGO7
2017 Fast cycle-accurate compile based simulator for reconfigurable processor
abstract
Reconfigurable Processor (RP) provides great flexibility of hardware re-configurability through software solution for high-performance computing. RP is used as a DSP in Samsung DTV and Camera to run the Audio, Video codecs and image quality enhancement algorithms. RP runs in two modes: VLIW (Very Large Instruction Word) and CGRA (Coarse Grain Reconfigurable Array). To minimize the time-to-market of products, application developers of the RP require a fast and cycle-accurate profiling-enabled simulator to verify the functionality of applications and to optimize the hot-spots (performance critical codes). Further, fast simulation is necessary to verify the functionalities for certification of various standards of multimedia (DivX, Dolby, etc.). The state-of-the-art RP simulators are very slow running at 2 MIPS (Million instructions per second) in CGRA mode and 4 MIPS in VLIW mode, running on x86 host processor at 3.4 GHz. We propose FastSim, a fast compile based cycle-accurate simulator for RP, which yields simulation speed of 900 MIPS in CGRA mode and 1600 MIPS in VLIW mode with > 99.5% core cycle accuracy. Thus, FastSim enables ~400x faster simulation when compared to existing RP simulators. FastSim in VLIW mode is 2x faster when compared to the state-of-the-art functional simulators, and 16x faster when compared to cycle-accurate simulators, of the VLIW and RISC processors in the industry. FastSim speed is also comparable to application running time on native x86. The faster simulation speed is achieved with the use of an innovative maximal static analysis in both VLIW and CGRA modes.
Narasinga Rao Miniskar, Raj Narayana Gadde, Young-chul Rams Cho, Sukjin Kim
ISCAS4
2016 Intra mode power saving methodology for CGRA-based reconfigurable processor architectures
abstract
Reconfigurable processors (RP) such as Samsung Reconfigurable Processor (SRP) are best suited for wide range of embedded DSP application domains such as image, audio, video and vision processing. It provides great flexibility of hardware reconfigurability through software solution which provides high-performance computing with low-energy and fast time-to-market. Power consumption of the processor architecture is one of the important considerations for mobile DSP solutions. However, the performance critical code sections of applications mapped to RPs may not utilize all minicores (group of functional units, local register files and their connections) due to low ILP (Instruction Level Parallelism) even after exploring software pipelining by the compiler. The unused minicores can consume considerable amount of power (including power leakage) from configuration memory banks, functional units, local register files, and also from central register files of RP architectures. There are approaches in RP to power-gate the minicores when we switch from VLIW (Very Large Instruction Word) mode of RP operation to CGRA (Coarse Grained Reconfigurable Array) mode (inter-mode power-gating exploration). In this paper, we propose a novel programmer directive based power-gating technique for RP which explores unused resources during VLIW and CGRA mode of execution (intra-mode power-gating). Our approach has shown up to 33% power savings in the SRP CGRA mode and up to 56% power savings in SRP VLIW mode, with additional power-gating circuits that contribute to <;1% increase in die area.
Narasinga Rao Miniskar, Rahul R. Patil, Raj Narayana Gadde, Young-chul Rams Cho, Sukjin Kim, Shi Hwa Lee
ISCAS5
2015 Flexible video processing platform for 8K UHD TV
Sukjin Kim, Young-Hwan Park, Wonchang Lee, Shihwa Lee
Hot Chips Symposium1
2012 Design space exploration and implementation of a high performance and low area Coarse Grained Reconfigurable Processor
abstract
Coarse Grained Reconfigurable Architectures (CGRAs) have played a key role in the area of domain specific processors due to their programmability and runtime reconfigurability. The Coarse Grained Array (CGA) structure enables target designs to achieve high performance, but it is easy to fall into over-design in term of area. Moreover, the network overhead between the function units (FUs) seriously degrades its clock speed. In this paper, we propose a high performance CGRA that facilitates design space exploration (DSE) to reduce these overheads. It employs a concept of building blocks, named mini cores, to mitigate overhead involved in DSE that aims to achieve high clock speed and small area in the target design. The proposed approach reduces the design time more than 100 times compared with previous design. Experimental results show that the implemented architecture reduces logic area by 14.38% and improves clock frequency by 59.34% without performance loss.
Dongkwan Suh, Kiseok Kwon, Sukjin Kim, Soojung Ryu, Jeongwook Kim
FPT3
2011 An instruction-scheduling-aware data partitioning technique for coarse-grained reconfigurable architectures
abstract
In this paper, we propose a data partitioning technique for the memory subsystem that consists of a multi-ported scratchpad memory (SPM) unit and a single-ported data cache in coarse-grained reconfigurable arrays (CGRA) architecture. The embedded reconfigurable processor executes programs by switching between the Non-VLIW and VLIW modes depending on the type of the code region to achieve high performance. The VLIW mode exploits code regions with high ILP that require high memory bandwidth and the Non-VLIW mode exploits those with low ILP that require low memory latency. Our data partitioning technique between the SPM and the data cache is based on data interference graph reduction and profiling information. Given an SPM size, it finds the optimal data partitions by taking the VLIW instruction schedule into consideration. We evaluate our data partitioning technique for the CGRA architecture with three representative multimedia applications.
Choonki Jang, Jaejin Lee, Hee-Seok Kim, Donghoon Yoo, Sukjin Kim, Hongseok Kim, Soojung Ryu
LCTES6
2010 Resource recycling: putting idle resources to work on a composable accelerator
abstract
Mobile computing platforms in the form of smart phones, netbooks, and personal digital assistants have become an integral part of our everyday lives. Moving ahead to the future, mobile multimedia support will become a key differentiating factor for customers. Features such as high-definition audio and video, video conferencing, 3D graphics, and image projection will lead to the adoption of one phone over another. However, in contrast to wireless signal processing which is dominated by vectorizable computation, mobile multimedia applications often contain complex control flow and variable computational requirements. Moreover, data access is more complex where media applications typically operate on multi-dimensional vectors of data rather than single-dimensional vectors with simple strides. To handle these complexities, composable accelerators such as the Polymorphic Pipeline Array, or PPA, present an appealing hardware platform by adding a degree of hardware configurability over existing accelerators. Hardware resources can be both statically as well as dynamically partitioned among executing tasks to maximize execution efficiency. However, an effective compilation framework is essential to partition and assign resources to make intelligent use of the available hardware. In this paper, a compilation framework is introduced that maximizes application throughput with hybrid resource partitioning of a PPA system. Static partitioning handles part of the resource assignment, but this is followed up by dynamic partitioning to identify idle resources and put them to use -- resource recycling. Experimental results show that real-time media applications can take advantage of the static and dynamic configurability of the PPA for increase.
Yongjun Park 0001, Hyunchul Park 0001, Scott A. Mahlke, Sukjin Kim
CASES4