EDBT 2026 Demo / reviewers in the wild / expert
Hyunwuk Lee
dblp:290/4149
· DBLP profile ↗
14ranked-venue papers
2as first author
14since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 2 first-author · 11 since 2021Software engineering, systems software and programming languages · 4 · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Slice: A Selective Local Inference Framework with Codec Exploitation for Accelerating Video Super-Resolution
Mingu Jung, Sungbin Kim, Seunghyun Jin 0002, Hyunwuk Lee, Won Woo Ro |
ISCA | 5 |
| 2026 | MXFFP: Microscaling Flexible Floating Point Format for Large-Scale AI Model Acceleration
Sungwoo Kim 0003, Sungbin Kim, Dongho Ha, Hyunwuk Lee, Junsung Kim, Mingu Jung, Murali Annavaram, Won Woo Ro |
ISCA | 4 |
| 2026 | DeSpa: Heterogeneous multi-core accelerators for energy-efficient dense and sparse computation at the tile level in Deep Neural Networks
Hyungjun Jang, Dongho Ha, Hyunwuk Lee, Won Woo Ro |
J. Syst. Archit. | 3 |
| 2025 | CVMAX: Accelerator Architecture with Polar Form Multiplication for Complex-Valued Neural NetworksabstractComplex-Valued Neural Networks (CVNNs) have demonstrated high performance in applications where complex numbers are essential, but suffer from higher computational and memory overheads. Since their target applications often operate in resource-constrained environments, optimizing CVNNs for energy and area efficiency is important for their acceleration. To resolve these challenges, we present CVMAX, a software-hardware co-design for energy and area-efficient CVNN acceleration. CVMAX introduces a specialized quantization technique based on polar form representation and shift quantization. The technique significantly reduces the bit width of CVNNs and computational complexity compared to conventional quantization with rectangular form. Moreover, shift quantization leverages the computational simplicity of multiplication in polar form, reducing the complexity of complex number multiplication. With the quantization technique, we designed a dedicated hardware accelerator that supports CVMAX data and its associated arithmetic operations. In our evaluation, CVMAX achieves a 75% reduction in energy consumption and achieves a $4.44 \times$ speedup compared to conventional accelerators. Hyunwuk Lee, Sungbin Kim, Sungwoo Kim 0003, Won Woo Ro |
DAC | 1 |
| 2025 | Ditto: Accelerating Diffusion Model via Temporal Value SimilarityabstractDiffusion models achieve superior performance in image generation tasks. However, it incurs significant computation overheads due to its iterative structure. To address these overheads, we analyze this iterative structure and observe that adjacent time steps in diffusion models exhibit high value similarity, leading to narrower differences between consecutive time steps. We adapt these characteristics to a quantized diffusion model and reveal that the majority of these differences can be represented with reduced bit-width, and even zero. Based on our observations, we propose the Ditto algorithm, a difference processing algorithm that leverages temporal similarity with quantization to enhance the efficiency of diffusion models. By exploiting the narrower differences and the distributive property of layer operations, it performs full bit-width operations for the initial time step and processes subsequent steps with temporal differences. In addition, Ditto execution flow optimization is designed to mitigate the memory overhead of temporal difference processing, further boosting the efficiency of the Ditto algorithm. We also design the Ditto hardware, a specialized hardware accelerator, fully exploiting the dynamic characteristics of the proposed algorithm. As a result, the Ditto hardware achieves up to $1.5 \times$ speedup and 17.74% energy saving compared to other accelerators. Sungbin Kim, Hyunwuk Lee, Wonho Cho, Mincheol Park, Won Woo Ro |
HPCA | 2 |
| 2025 | Adversarial Purification via Super-Resolution and Diffusion
Mincheol Park, Cheonjun Park, Seungseop Lim, Mijin Koo, Hyunwuk Lee, Won Woo Ro, Suhyun Kim 0001 |
ICCV | 5 |
| 2025 | COSMOS: An LLC Contention Slowdown Model for Heterogeneous Multi-Core SystemsabstractHeterogeneous multi-core systems are increasingly adopted due to their advantages in area efficiency and energy savings. However, existing analytical models often overlook core heterogeneity, leading to lower performance prediction accuracy compared to homogeneous systems. In this paper, we show that even under identical last-level cache (LLC) contention conditions, heterogeneous cores experience different slowdowns. We categorize memory access time into internal and external components based on whether memory requests are served before reaching LLC and analyze how these two types affect application slowdowns. Furthermore, we examine how these components vary with core heterogeneity. Our analysis reveals that differences in cache hierarchies lead to distinct eviction patterns and variable external accesses, producing LLC miss rates that depend on LLC capacity. Additionally, core heterogeneity influences the execution times of both computation and internal memory accesses, which serve as correction factors that modulate the effect of LLC miss rate differences on application slowdown. Based on these insights, we propose COSMOS, an analytical model designed to accurately predict slowdowns caused by LLC contention in heterogeneous multi-core systems. COSMOS profiles the sensitivity of external accesses to LLC capacity, estimates LLC miss rates and average access latency, and aggregates the weighted contributions of all components. COSMOS achieves an average accuracy of 94.71% in performance prediction, significantly outperforming models that overlook internal resources, which achieve average accuracies of 82.76 % and 89.87 %, respectively. Yongju Lee 0003, Jaewon Kwon, Cheolhwan Kim, Enhyeok Jang, Jiwon Lee 0001, Hyunwuk Lee, Won Woo Ro |
ISPASS | 6 |
| 2025 | BitL: A Hybrid Bit-Serial and Parallel Deep Learning Accelerator for Critical Path ReductionabstractAs deep neural networks (DNNs) advance, their computational demands have grown immensely.In this context, previous research introduced bit-wise computation to enhance silicon efficiency, along with skipping unnecessary zero-bit calculations.However, we observe that existing bit-wise approaches miss an opportunity to optimize the critical computation path, as they process groups of values sequentially from the most significant bits (MSBs) to the least significant bits (LSBs).To address this limitation, we propose BitL, a novel bit-wise computing unit designed to minimize the critical path and improve the throughput.BitL dynamically switches between horizontal and vertical data lookups across sub-tiles during Multiply-Accumulate (MAC) operations.Additionally, it presents an innovative optimization technique to maximize the utilization of computing units while switching its lookup direction.Our evaluation demonstrates that BitL delivers up to 1.92× higher throughput compared to a baseline DNN accelerator and achieves a 1.24× improvement over recent zero-bit skipping accelerators.Furthermore, BitL improves energy efficiency by 2.06× on average, with a silicon area overhead of only 5.71%. Seunghyun Lee 0003, Dongho Ha, Sungbin Kim, Sungwoo Kim 0003, Hyunwuk Lee, Won Woo Ro |
MICRO | 5 |
| 2025 | REC: Enhancing fine-grained cache coherence protocol in multi-GPU systems
Gun Ko, Jiwon Lee 0001, Hongju Kal, Hyunwuk Lee, Won Woo Ro |
J. Syst. Archit. | 4 |
| 2024 | AirGun: Adaptive Granularity Quantization for Accelerating Large Language ModelsabstractTransformer-based models have evolved into Large Language Models (LLMs) by increasing model sizes to achieve higher accuracy, but they incur significant computational and memory costs. As quantization is a promising method to mitigate the huge cost of LLMs, the presence of outliers can lead to accuracy drops during quantization. Previous work pointed out LLMs have outliers only in specific input channels of activations. This suggests that per-input channel quantization would be beneficial, but it poses excessive computational overhead without optimization. To address these challenges, we propose a hardware and software co-design that mitigates the overhead of per-input channel quantization. We first propose AirGun, a quantization method that adaptively quantizes LLM modules. We observe that LLMs have high quantization sensitivity only in specific modules. Based on our observation, AirGun applies hardware-efficient per-tensor quantization for non-sensitive modules and per-input channel quantization for sensitive modules. For per-input channel quantization, we introduce early reconstruction and adaptive dyadic numbering, dismissing the overhead while exploiting its advantages. Additionally, we propose the AirGun accelerator that fully utilizes the advantages of AirGun. As a result, the AirGun accelerator achieves a 4.19 × speedup and 63.16 % lower energy consumption compared to the previous LM accelerator while achieving higher accuracy. Sungbin Kim, Hyunwuk Lee, Sungwoo Kim 0003, Cheolhwan Kim, Won Woo Ro |
ICCD | 2 |
| 2024 | GUMSO: Gating Unnecessary On-Chip Memory Slices for Power Optimization on GPUsabstractThe importance of power efficiency in GPUs has grown significantly for data centers, as it directly impacts costs and sustainability. While there have been many works on power optimization for GPUs, their approaches often target applications consuming high parallelism for their entire application sequence. However, graph applications, characterized by their large data size and inherent irregularity in graph structures, present distinct challenges as they suffer from kernel launches with limited parallelism, resulting in under-utilization of GPU hardware components. Even with this under-utilization, last-level cache (LLC) and network-on-chip (NoC) of the GPU are fully activated even with a single thread running, incurring large leakage power. In our analysis, we find out that LLC and NoC contribute up to 39.9% of the total GPU power consumption during graph applications. To address this power inefficiency, we propose GUMSO, an energy-efficient design that enables adaptive power-gating of LLC slices for small kernel executions in GPUs. By managing the utilization of LLC slices, our approach reduces GPU energy consumption by an average of 18.3% across various graph applications with minimum performance overheads. Seunghyun Jin 0002, Hyunwuk Lee, Won Woo Ro |
ISLPED | 2 |
| 2024 | SHREG: Mitigating register redundancy in GPUs
Seunghyun Jin 0002, Hyunwuk Lee, Junsung Kim 0002, Won Woo Ro |
J. Syst. Archit. | 2 |
| 2023 | Early-Adaptor: An Adaptive Framework forProactive UVM Memory ManagementabstractUnified Virtual Memory (UVM) relieves programmers of the burden of memory management between CPU and GPUs. However, the use of UVM can lead to performance degradation due to its on-demand page migration scheme, especially under memory oversubscription. In this research, we conduct various analyses on real hardware, NVIDIA RTX 3090, to examine such performance degradation with an NVIDIA opensource GPU driver. Our analysis shows that the effectiveness of prefetching highly correlates with the relative number of page faults on a group of contiguous pages, which NVIDIA refers to as a Virtual Address Block (VABlock) spanning across a 2MB virtual address range. Also, the risk of page thrashing is determined by the total number of VABlocks that consistently generate page faults during kernel execution. Hence, the performance impact of the prefetch threshold varies across different workloads. These observations indicate that an adaptive prefetching scheme can resolve the performance bottleneck of memory oversubscription. To this end, we propose the Early-Adaptor (EA) framework, which automatically controls the prefetching aggressiveness based on the page fault history. During runtime, the EA framework monitors patterns of page faults in per-VABlock and in a global scope. After analyzing page fault generation rates and the possibility of page thrashing, the EA framework dynamically controls the prefetching aggressiveness by changing the prefetch threshold. The EA framework requires only minor changes to GPU drivers and needs no changes to the GPU hardware. Experiments on real hardware show that when GPU memory is oversubscribed, the EA framework achieves an average speedup of 1. 74x over the conventional GPU prefetcher. Seokjin Go, Hyunwuk Lee, Junsung Kim 0002, Jiwon Lee 0001, Myung Kuk Yoon, Won Woo Ro |
ISPASS | 2 |
| 2023 | Exploiting Inherent Properties of Complex Numbers for Accelerating Complex Valued Neural NetworksabstractSince conventional Deep Neural Networks (DNNs) use real numbers as their data, they are unable to capture the imaginary values and the correlations between real and imaginary values in applications that use complex numbers. To address this limitation, Complex Valued Neural Networks (CVNNs) have been introduced, enabling to capture the context of complex numbers for various applications such as Magnetic Resonance Imaging (MRI), radar, and sensing. CVNNs handle their data with complex numbers and adopt complex number arithmetic to their layer operations, so they exhibit distinct design challenges with real-valued DNNs. The first challenge is the data representation of the complex number, which requires two values for a single data, doubling the total data size of the networks. Moreover, due to the unique operations of the complex-valued layers, CVNNs require a specialized scheduling policy to fully utilize the hardware resources and achieve optimal performance. To mitigate the design challenges, we propose software and hardware co-design techniques that effectively resolves the memory and compute overhead of CVNNs. First, we propose Polar Form Aware Quantization (PAQ) that utilizes the characteristics of the complex number and their unique value distribution on CVNNs. Then, we propose our hardware accelerator that supports PAQ and CVNN operations. Lastly, we design a CVNN-aware scheduling scheme that optimizes the performance and resource utilization of an accelerator by aiming at the special layer operations of CVNN. PAQ achieves 62.5% data compression over CVNNs using FP16 while retaining a similar error with INT8 quantization, and our hardware support PAQ with only 2% area overhead over conventional systolic array architecture. In our evaluation, PAQ hardware with the scheduling scheme achieves a 32% lower latency and 30% lower energy consumption than other accelerators. Hyunwuk Lee, Hyungjun Jang, Sungbin Kim, Sungwoo Kim 0003, Wonho Cho, Won Woo Ro |
MICRO | 1 |