EDBT 2026 Demo / reviewers in the wild / expert
Minseok Lee
dblp:144/4723
· DBLP profile ↗
11ranked-venue papers
2as first author
5since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 2 first-authorDatabases, data management, data science and information retrieval · 4 · 4 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2Software engineering, systems software and programming languages · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Efficient Rotation Compression for Gaussian Splats via Geometry-Based Point Cloud CompressionabstractRecent progress in 3D scene representation, including Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS), has enabled photorealistic rendering with real-time performance. Among these, 3DGS has gained attention due to its capability to represent complex scenes using Gaussian splats. However, this representation introduces significant storage challenges because each splat contains 59 attributes—far more than conventional point clouds. Efficient compression of these attributes, particularly rotation parameters, is essential for practical deployment. Jongseok Lee, Alexander Alshin, Hyejung Hur, Minseok Lee, Hyn-Mook Oh, Jong-Yeul Suh |
DCC | 4 |
| 2025 | Fast video anomaly detection via context-aware shortcut exploration and abnormal feature distance learning
Donghyeong Kim, MyeongAh Cho, Minjung Kim 0002, Minseok Lee, Seungwook Park, Sangyoun Lee |
Pattern Recognit. | 5 |
| 2024 | Embedding Optimization for Training Large-scale Deep Learning Recommendation Systems with EMBarkabstractTraining large-scale deep learning recommendation models (DLRMs) with embedding tables stretching across multiple GPUs in a cluster presents a unique challenge, demanding the efficient scaling of embedding operations that require substantial memory and network bandwidth within a hierarchical network of GPUs. To tackle this bottleneck, we introduce EMBark—a comprehensive solution aimed at enhancing embedding performance and overall DLRM training throughput at scale. EMBark empowers users to create and customize sharding strategies, and features a highly-automated sharding planner, to accelerate diverse model architectures on different cluster configurations. EMBark groups embedding tables, considering their preferred communication compression method to reduce communication overheads effectively. It embraces efficient data-parallel category distribution, combined with topology-aware hierarchical communication, and pipelining support to maximize the DLRM training throughput. Across four representative DLRM variants (DLRM-DCNv2, T180, T200, and T510), EMBark achieves an average end-to-end training throughput speedup of 1.5 × and up to 1.77 × over traditional table-row-wise sharding approaches. Xavier Simmons, Matthias Langer, Minseok Lee, Zehuan Wang 0001 |
RecSys | 8 |
| 2022 | Merlin HugeCTR: GPU-accelerated Recommender System Training and InferenceabstractIn this talk, we introduce Merlin HugeCTR. Merlin HugeCTR is an open source, GPU-accelerated integration framework for click-through rate estimation. It optimizes both training and inference, whilst enabling model training at scale with model-parallel embeddings and data-parallel neural networks. In particular, Merlin HugeCTR combines a high-performance GPU embedding cache with an hierarchical storage architecture, to realize low-latency retrieval of embeddings for online model inference tasks. In the MLPerf v1.0 DLRM model training benchmark, Merlin HugeCTR achieves a speedup of up to 24.6x on a single DGX A100 (8x A100) over PyTorch on 4x4-socket CPU nodes (4x4x28 cores). Merlin HugeCTR can also take advantage of multi-node environments to accelerate training even further. Since late 2021, Merlin HugeCTR additionally features a hierarchical parameter server (HPS) and supports deployment via the NVIDIA Triton server framework, to leverage the computational capabilities of GPUs for high-speed recommendation model inference. Using this HPS, Merlin HugeCTR users can achieve a 5~62x speedup (batch size dependent) for popular recommendation models over CPU baseline implementations, and dramatically reduce their end-to-end inference latency. Zehuan Wang 0001, Yingcan Wei, Minseok Lee, Matthias Langer, Daniel G. Abel, Jianbing Dong, Kunlun Li |
RecSys | 3 |
| 2022 | A GPU-specialized Inference Parameter Server for Large-Scale Deep Recommendation ModelsabstractRecommendation systems are of crucial importance for a variety of modern apps and web services, such as news feeds, social networks, e-commerce, search, etc. To achieve peak prediction accuracy, modern recommendation models combine deep learning with terabyte-scale embedding tables to obtain a fine-grained representation of the underlying data. Traditional inference serving architectures require deploying the whole model to standalone servers, which is infeasible at such massive scale. Yingcan Wei, Matthias Langer, Minseok Lee, Zehuan Wang 0001 |
RecSys | 4 |
| 2017 | An Activity-Embedding Approach for Next-Activity Prediction in a Multi-User Smart SpaceabstractSince the advent of the IoT era, various IoT devices have proliferated, transforming ordinary spaces into smart spaces such as smart home, smart office, and smart building. To provide user-friendly service to people, the majority of previous studies have focused on activity recognition and prediction in singleuser environments such as ambient assisted living (AAL) and activities of daily living (ADL). However, unlike single-user environments, many real world environments are comprised of multiple activities occurring concurrently in a multi-user smart space. In the presence of multiple activities and multiple users, the process of next-activity prediction rarely produces just a single candidate for the next activity. Thus, we present in this paper an approach that generates multiple next- activity candidates. Our approach is motivated by a specific word-embedding algorithm that is typically used for natural language processing (NLP) to map words into a vector space. By using a similar embedding approach in a multi-user smart space, we map activities to vector coordinates in a vector space. After the vectorization, a long short-term memory (LSTM) network can be trained to predict a single vector coordinate for next activity from a group of previously occurring activities; then from this single prediction, multiple next-activity candidates are selected by choosing several vectors near the LSTM's single output. We tested our approach using real data generated from the multi-user smart space testbed on our campus. After the embedding of activities into a vector space, we were able to find semantically meaningful relations between the resultant vectors. In addition, next-activity prediction had a success rate of approximately 82%. Our activity embedding and next- activity prediction method can be utilized together in multi-user smart spaces to develop smart service systems such as a service recommendation system. Younggi Kim, Jihoon An, Minseok Lee, Younghee Lee |
SMARTCOMP | 3 |
| 2016 | iPAWS: Instruction-issue pattern-based adaptive warp scheduling for GPGPUsabstractThread or warp scheduling in GPGPUs has been shown to have a significant impact on overall performance. Recently proposed warp schedulers have been based on a greedy warp scheduler where some warps are prioritized over other warps. However, a single warp scheduling policy does not necessarily provide good performance across all types of workloads; in particular, we show that greedy warp schedulers are not necessarily optimal for workloads with inter-warp locality while a simple round-robin warp scheduler provides better performance. Thus, we argue that instead of single, static warp scheduling, an adaptive warp scheduler that dynamically changes the warp scheduler based on the workload characteristics should be leveraged. In this work, we propose an instruction-issue pattern-based adaptive warp scheduler (iPAWS) that dynamically adapts between a greedy warp scheduler and a fair, round-robin scheduler. We exploit the observation that workloads that favor a greedy warp scheduler will have an instruction-issue pattern that is biased towards some warps while workloads that favor a fair, round-robin warp scheduler will tend to issue instructions across all of the warps. Our evaluations show that iPAWS is able to adapt to the more optimal warp scheduler dynamically and achieve performance that is within a few percent of the statically determined, more optimal warp scheduler. We also show that iPAWS can be extended to other warp schedulers, including the cache-conscious wavefront scheduling (CCWS) and Memory Aware Scheduling and Cache Access Re-execution (MASCAR) to exploit the benefits of other warp schedulers while still providing adaptivity in warp scheduling. Minseok Lee, Gwangsun Kim, John Kim 0001, Woong Seo, Yeongon Cho, Soojung Ryu |
HPCA | 1 |
| 2014 | Energy-efficient scheduling for memory-intensive GPGPU workloadsabstractHigh performance for a GPGPU workload is obtained by maximizing parallelism and fully utilizing the available resources. However, this is not necessarily energy efficient, especially for memory-intensive GPGPU workloads. In this work, we propose Throttle CTA (cooperative-thread array) Scheduling (TCS) where we leverage two type of throttling - throttling the number of actives cores and throttling of warp execution in the cores - to improve energy-efficiency for memory-intensive GPGPU workloads. The algorithm requires the global CTA or thread block scheduler to reduce the number of cores with assigned thread blocks while leveraging the local warp scheduler to throttle memory requests for some of the cores to further reduce power consumption. The proposed TCS scheduling does not require off-line analysis but can be done dynamically during execution. Instead of relying on conventional metrics such as miss-per-kilo-instruction (MPKI), we leverage the memory access latency metric to determine the memory intensity of the workloads. Our evaluations show that TCS reduces energy by up to 48% (38% on average) across different memory-intensive workload while having very little impact on performance for compute-intensive workloads. Seokwoo Song, Minseok Lee, John Kim 0001, Woong Seo, Yeongon Cho, Soojung Ryu |
DATE | 2 |
| 2014 | Improving GPGPU resource utilization through alternative thread block schedulingabstractHigh performance in GPGPU workloads is obtained by maximizing parallelism and fully utilizing the available resources. The thousands of threads are assigned to each core in units of CTA (Cooperative Thread Arrays) or thread blocks - with each thread block consisting of multiple warps or wavefronts. The scheduling of the threads can have significant impact on overall performance. In this work, explore alternative thread block or CTA scheduling; in particular, we exploit the interaction between the thread block scheduler and the warp scheduler to improve performance. We explore two aspects of thread block scheduling - 1) LCS (lazy CTA scheduling) which restricts the maximum number of thread blocks allocated to each core, and 2) BCS (block CTA scheduling) where consecutive thread blocks are assigned to the same core. For LCS, we leverage a greedy warp scheduler to help determine the optimal number of thread blocks by only measuring the number of instructions issued while for BCS, we propose an alternative warp scheduler that is aware of the “block” of CTAs allocated to a core. With LCS and the observation that maximum number of CTAs does not necessary maximize performance, we also propose mixed concurrent kernel execution that enables multiple kernels to be allocated to the same core to maximize resource utilization and improve overall performance. Minseok Lee, Seokwoo Song, Joosik Moon, John Kim 0001, Woong Seo, Yeongon Cho, Soojung Ryu |
HPCA | 1 |
| 2014 | Energy harvesting from anti-corrosion power sourcesabstractThis work presents energy harvesting techniques from low-voltage current used to prevent galvanic corrosion between a metallic structure and a permanent copper/copper sulfate (Cu/CuSO4) reference electrode. Supercapacitors are adopted to compensate for or overcome the limitations of batteries. Then, a boost converter is used to convert the low voltage levels of galvanic corrosion to that needed by the complementary metal oxide semiconductor (CMOS) technologies used for the wireless sensor systems. Experimental results show that our proposed harvesting schemes significantly reduce the overhead of the charging circuitry, which enables nearly full charging of supercapacitors of up to 350F under the low power conditions of 3mW (i.e., 3mA at 1,V). More importantly, our system enables maintenance-free operation of remote-monitoring cathodic protection (RMCP) systems in harsh environments, where sunlight or wind power may be unavailable or unpredictable. Sehwan Kim, Minseok Lee, Pai H. Chou |
ISLPED | 2 |
| 2014 | Multi-GPU System Design with Memory NetworksabstractGPUs are being widely used to accelerate different workloads and multi-GPU systems can provide higher performance with multiple discrete GPUs interconnected together. However, there are two main communication bottlenecks in multi-GPU systems -- accessing remote GPU memory and the communication between GPU and the host CPU. Recent advances in multi-GPU programming, including unified virtual addressing and unified memory from NVIDIA, has made programming simpler but the costly remote memory access still makes multi-GPU programming difficult. In order to overcome the communication limitations, we propose to leverage the memory network based on hybrid memory cubes (HMCs) to simplify multi-GPU memory management and improve programmability. In particular, we propose scalable kernel execution (SKE) where multiple GPUs are viewed as a single virtual GPU as a single kernel can be executed across multiple GPUs without modifying the source code. To fully enable the benefits of SKE, we explore alternative memory network designs in a multi-GPU system. We propose a GPU memory network (GMN) to simplify data sharing between the discrete GPUs while a CPU memory network (CMN) is used to simplify data communication between the host CPU and the discrete GPUs. These two types of networks can be combined to create a unified memory network (UMN) where the communication bottleneck in multi-GPU can be significantly minimized as both the CPU and GPU share the memory network. We evaluate alternative network designs and propose a sliced flattened butterfly topology for the memory network that scales better than previously proposed alternative topologies by removing local HMC channels. In addition, we propose an overlay network organization for unified memory network to minimize the latency for CPU access while providing high bandwidth for the GPUs. We evaluate trade-offs between the different memory network organization and show how UMN significantly reduces the communication bottleneck in multi-GPU systems. Gwangsun Kim, Minseok Lee, Jiyun Jeong, John Kim 0001 |
MICRO | 2 |