EDBT 2026 Demo / reviewers in the wild / expert
Bogil Kim
dblp:278/6374
· DBLP profile ↗
4ranked-venue papers
1as first author
3since 2021 · last 2025
0000-0002-2332-7933ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | RoTA: Rotational Torus Accelerator for Wear Leveling of Neural Processing ElementsabstractThis paper introduces a reliability-aware neural accelerator design with a wear-leveling solution that balances the utilization of processing elements (PEs). Neural accelerators deploy many PEs to exploit data-level parallelism, but their designs and operations have focused mostly on performance and energy efficiency metrics. Directional dataflows in PE arrays and dimensional misalignment with variable-sized neural layers cause the underutilization of PEs, which is biased to PE locations and gradually accumulated over time. Consequently, the accelerators experience severe usage imbalance between PEs. To resolve the problem, this paper proposes a rotational torus accelerator (RoTA) with an optimized wear-leveling scheme that shuffles PE utilization spaces to eliminate PE usage imbalance. Evaluation results show that RoTA improves lifetime reliability by 1.69x. Taesoo Lim, Jingu Park, Bogil Kim, William J. Song |
DATE | 4 |
| 2023 | NeuroSpector: Systematic Optimization of Dataflow Scheduling in DNN AcceleratorsabstractThis paper presents an optimization framework namedNeuroSpectorthat systematically analyzes the dataflow of deep neural network (DNN) accelerators and rapidly identifies optimal execution methods. The proposed methodology is demonstrated to work effectively with a variety of accelerator architectures and DNN workloads. It has been a baffling challenge to devise scheduling schemes for neural accelerators to maximize energy efficiency and performance. The challenge lies in that hardware specifications associated with multi-dimensional DNN data create an enormous number of possible scheduling options that can be exerted on accelerators. Related work suggested various techniques to solve the challenge encompassing brute-force search of massive solution spaces pruned by user constraints, solving the objective functions of system models, learning-based optimization, etc. However, each suggested technique was devised only for a specific accelerator model. Therefore, we find that they are not adaptively applicable to different accelerators and DNN workloads in that they produce hit-or-miss results with 100.1% greater energy and cycles on average compared to optimal scheduling schemes obtained from fully comprehensive brute-force searches. In contrast, NeuroSpector identifies efficient execution methods for various accelerators and workloads with only 1.5% differences on average to the optimal scheduling solutions. The optimization strategy of NeuroSpector is based on an observation that optimal executions are strongly correlated with minimizing data movements to the lower-level memory hierarchy of accelerators rather than maximizing the utilization of processing elements. Thus, NeuroSpector prioritizes optimizing lower-level components in the accelerator hierarchy, which is proven highly effective for various accelerators and DNN workloads. Chanho Park 0004, Bogil Kim, Sungmin Ryu, William J. Song |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2021 | The Nebula Benchmark Suite: Implications of Lightweight Neural NetworksabstractThis article presents a benchmark suite namedNebulathat implements lightweight neural network benchmarks. Recent neural networks tend to form deeper and sizable networks to enhance accuracy and applicability. However, the massive volume of heavy networks makes them highly challenging to use in conventional research environments such as microarchitecture simulators. We notice that neural network computations are mainly comprised of matrix and vector calculations that repeat on multi-dimensional data encompassing batches, channels, layers, etc. This observation motivates us to develop a variable-sized neural network benchmark suite that provides users with options to select appropriate size of benchmarks for different research purposes or experiment conditions. Inspired by the implementations of well-known benchmarks such as PARSEC and SPLASH suites, Nebula offers various size options from large to small datasets for diverse types of neural networks. The Nebula benchmark suite is comprised of seven representative neural networks built on a C++ framework. The variable-sized benchmarks can be executed i) with acceleration libraries (e.g., BLAS, cuDNN) for faster and realistic application runs or ii) without the external libraries if execution environments do not support them, e.g., microarchitecture simulators. This article presents a methodology to develop the variable-sized neural network benchmarks, and their performance and characteristics are evaluated based on hardware measurements. The results demonstrate that the Nebula benchmarks reduce execution time as much as 25x while preserving similar architectural behaviors as the full-fledged neural networks. Bogil Kim, Chanho Park 0004, William J. Song |
IEEE Trans. Computers | 1 |
| 2020 | Duplo: Lifting Redundant Memory Accesses of Deep Neural Networks for GPU Tensor CoresabstractThis paper introduces a GPU architecture named Duplo that minimizes redundant memory accesses of convolutions in deep neural networks (DNNs). Convolution is one of the fundamental operations used in various classes of DNNs, and it takes the majority of execution time. Various approaches have been proposed to accelerate convolutions via general matrix multiplication (GEMM), Winograd convolution, fast Fourier transform (FFT), etc. Recent introduction of tensor cores in NVIDIA GPUs particularly targets on accelerating neural network computations. A tensor core in a streaming multiprocessor (SM) is a specialized unit dedicated to handling matrix-multiply-and-accumulate (MMA) operations. The underlying operations of tensor cores represent GEMM calculations, and lowering a convolution can effectively exploit the tensor cores by transforming deeply nested convolution loops into matrix multiplication. However, lowering the convolution has a critical drawback since it requires a larger memory space (or workspace) to compute the matrix multiplication, where the expanded workspace inevitably creates multiple duplicates of the same data stored at different memory addresses. The proposed Duplo architecture tackles this challenge by leveraging compile-time information and microarchitectural supports to detect and eliminate redundant memory accesses that repeatedly load the duplicates of data in the workspace matrix. Duplo identifies data duplication based on memory addresses and convolution information generated by a compiler. It uses a load history buffer (LHB) to trace the recent load history of workspace data and their presence in register file. Every load instruction of workspace data refers to the LHB to find if potentially the same copies of data exist in the register file. If data duplicates are found, Duplo simply renames registers and makes them point to the ones containing the same values instead of issuing memory requests to load the same data. Our experiment results show that Duplo improves the performance of DNNs by 29.4% on average and saves 34.1% of energy using tensor cores. Sungwoo Ahn, Yunho Oh, Bogil Kim, Won Woo Ro, William J. Song |
MICRO | 4 |