EDBT 2026 Demo / reviewers in the wild / expert
Minho Ha
dblp:203/8255
· DBLP profile ↗
5ranked-venue papers
1as first author
4since 2021 · last 2026
0000-0002-9940-072XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enhanced CXL Pooled Memory System for Scalable AI via Embedding Access PredictionabstractThe embedding operation, pivotal in modern AI applications such as recommendation systems and natural language processing, transforms high-dimensional sparse data into dense vector representations. However, embedding tables are memory-intensive and pose significant challenges in DRAM-based architectures due to their substantial size. This paper introduces Sage, a scalable architecture for embedding operations in CXL-based pooled memory systems. Sage employs advanced caching and prefetching strategies, leveraging an online clustering algorithm to predict embedding table access patterns, and selectively uses Near-Data Processing (NDP) to mitigate the latency associated with CXL memory access. Our comprehensive evaluation demonstrates that Sage significantly enhances throughput and efficiency, providing a cost-effective solution for large-scale AI models. Our experimental results demonstrate that Sage enhances throughput by 2.84 × as compared to conventional memory management systems. Hoyeon Lee, Minho Ha, Byungil Koh, Jungmin Choi, Yeseong Kim |
DATE | 4 |
| 2023 | Sidekick: Near Data Processing for Clustering Enhanced by Automatic Memory DisaggregationabstractNear Data Processing (NDP) is a promising solution for data mining/analysis techniques, which extract useful information from big data. In this paper, we propose a novel NDP-enabled memory disaggregation system called Sidekick, based on a type-2 CXL device and enhanced by an automated allocation technique for clustering algorithms. The key enabler of our migration technique is to understand clustering workflows in a unit of the program context, which is the function call stack for functions, threads, and memory allocations to drive the automated decision. The proposed technique relates the migrated computation tasks with a series of function calls and performs GA-based optimization to identify the optimal allocation scenario for a target clustering algorithm. In Scikit-learn, a popular machine learning library, we use the genetic algorithm to find the optimal memory allocation policy and the operation offloading policy using the program context. The results show that the proposed technique increases the clustering performance as compared to the case, which only uses disaggregated memory without NDP cores, by up to 92% in terms of execution time, while reducing the majority of remote CXL memory accesses. Minho Ha, Byungil Koh, Kyoung Park, Yeseong Kim |
DAC | 3 |
| 2022 | QuiltNet: efficient deep learning inference on multi-chip accelerators using model partitioningabstractWe have seen many successful deployments of deep learning accelerator designs on different platforms and technologies, e.g., FPGA, ASIC, and Processing In-Memory platforms. However, the size of the deep learning models keeps increasing, making computations a burden on the accelerators. A naive approach to resolve this issue is to design larger accelerators; however, it is not scalable due to high resource requirements, e.g., power consumption and off-chip memory sizes. A promising solution is to utilize multiple accelerators and use them as needed, similar to conventional multiprocessing. For example, for smaller networks, we may use a single accelerator, while we may use multiple accelerators with proper network partitioning for larger networks. However, partitioning DNN models into multiple parts leads to large communication overheads due to inter-layer communications. In this paper, we propose a scalable solution to accelerate DNN models on multiple devices by devising a new model partitioning technique. Our technique transforms a DNN model into layer-wise partitioned models using an autoencoder. Since the autoencoder encodes a tensor output into a smaller dimension, we can split the neural network model into multiple pieces while significantly reducing the communication overhead to pipeline them. Our evaluation results conducted on state-of-the-art deep learning models show that the proposed technique significantly improves performance and energy efficiency. Our solution increases performance and energy efficiency by up to 30.5% and 28.4% with minimal accuracy loss as compared to running the same model on pipelined multi-block accelerators without the autoencoder. Hyukjun Kwon, Seowoo Kim, Minho Ha, Eui-Cheol Lim, Mohsen Imani, Yeseong Kim |
DAC | 5 |
| 2021 | Layerwise Buffer Voltage Scaling for Energy-Efficient Convolutional Neural NetworkabstractIn order to effectively reduce buffer energy consumption, which constitutes a significant part of the total energy consumption in a convolutional neural network (CNN), it is useful to apply different amounts of energy conservation effort to the different levels of a CNN as the buffer energy to total energy usage ratios can differ quite substantially across the layers of a CNN. This article proposes layerwise buffer voltage scaling as an effective technique for reducing buffer access energy. Error-resilience analysis, including interlayer effects, conducted during design-time is used to determine the specific buffer supply voltage to be used for each layer of a CNN. Then these layer-specific buffer supply voltages are used in the CNN for image classification inference. Error injection experiments with three different types of CNN architectures show that, with this technique, the buffer access energy and overall system energy can be reduced by up to 68.41% and 33.68%, respectively, without sacrificing image classification accuracy. Minho Ha, Younghoon Byun, Seungsik Moon, Youngjoo Lee 0002, Sunggu Lee |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2019 | Low-Complexity Dynamic Channel Scaling of Noise-Resilient CNN for Intelligent Edge DevicesabstractIn this paper, we present a novel channel scaling scheme for convolutional neural networks (CNNs), which can improve the recognition accuracy for the practical distorted images without increasing the network complexity. During the training phase, the proposed work first prepares multiple filters under the same CNN architecture by taking account of different noise models and strengths. We then newly introduce an FFT-based noise classifier, which determines the noise property in the received input image by calculating the partial sum of the frequency-domain values. Based on the detected noise class, we dynamically change the filters of each CNN layer to provide the dedicated recognition. Furthermore, we propose a channel scaling technique to reduce the number of active filter parameters if the input data is relatively clean. Experimental results show that the proposed dynamic channel scaling reduces the computational complexity as well as the energy consumption, still providing the acceptable accuracy for intelligent edge devices. Younghoon Byun, Minho Ha, Sunggu Lee, Youngjoo Lee 0002 |
DATE | 2 |