VLDB 2026 Research / reviewers in the wild / expert
Sungwoo Kim 0003
dblp:145/6643-3
· DBLP profile ↗
8ranked-venue papers
1as first author
8since 2021 · last 2026
0009-0009-0106-6590ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 1 first-author · 8 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Reducing Page Faults via Invalidation-Based Mapping Propagation in Multi-GPU Systems
Junsung Kim, Dongho Ha, Sungwoo Kim 0003, Wonho Cho, Sungbin Kim, Yufei Ding, Won Woo Ro |
ISCA | 3 |
| 2026 | MXFFP: Microscaling Flexible Floating Point Format for Large-Scale AI Model Acceleration
Sungwoo Kim 0003, Sungbin Kim, Dongho Ha, Hyunwuk Lee, Junsung Kim, Mingu Jung, Murali Annavaram, Won Woo Ro |
ISCA | 1 |
| 2025 | CVMAX: Accelerator Architecture with Polar Form Multiplication for Complex-Valued Neural NetworksabstractComplex-Valued Neural Networks (CVNNs) have demonstrated high performance in applications where complex numbers are essential, but suffer from higher computational and memory overheads. Since their target applications often operate in resource-constrained environments, optimizing CVNNs for energy and area efficiency is important for their acceleration. To resolve these challenges, we present CVMAX, a software-hardware co-design for energy and area-efficient CVNN acceleration. CVMAX introduces a specialized quantization technique based on polar form representation and shift quantization. The technique significantly reduces the bit width of CVNNs and computational complexity compared to conventional quantization with rectangular form. Moreover, shift quantization leverages the computational simplicity of multiplication in polar form, reducing the complexity of complex number multiplication. With the quantization technique, we designed a dedicated hardware accelerator that supports CVMAX data and its associated arithmetic operations. In our evaluation, CVMAX achieves a 75% reduction in energy consumption and achieves a $4.44 \times$ speedup compared to conventional accelerators. Hyunwuk Lee, Sungbin Kim, Sungwoo Kim 0003, Won Woo Ro |
DAC | 3 |
| 2025 | PIMFY: Eliminating Remote Page Walks in MCM GPUsabstractMulti-Chip Module (MCM) GPUs are suffering from non-uniform memory access (NUMA) due to communication through in-package interconnects between chiplets. As chiplet-to-chiplet communication increases, the performance saturates even with scaled hardware resources in MCM GPUs. Previous works have attempted to mitigate the NUMA caused by data page access between chiplets. However, we observe that page table page access for address translation requests also generates significant NUMA and degrades overall performance of MCM GPUs. In this paper, we analyze the impact of remote page table page access that causes significant challenges in MCM GPUs. Based on our analysis, we propose Page tables In My Front Yard (PIMFY), a technique that eliminates remote page table page access and accelerates page walks, thus improving overall performance in MCM GPUs. Exploiting static address mapping nature during GPU kernel execution, PIMFY replicates page tables onto all chiplets and prevents remote page table access. Our evaluation shows that PIMFY reduces average page walk latency by 27.07%, and enhances overall performance by$1.21 \times$. Junsung Kim 0002, Sungwoo Kim 0003, Seunghyun Jin 0002, Won Woo Ro |
ICCD | 2 |
| 2025 | BitL: A Hybrid Bit-Serial and Parallel Deep Learning Accelerator for Critical Path ReductionabstractAs deep neural networks (DNNs) advance, their computational demands have grown immensely.In this context, previous research introduced bit-wise computation to enhance silicon efficiency, along with skipping unnecessary zero-bit calculations.However, we observe that existing bit-wise approaches miss an opportunity to optimize the critical computation path, as they process groups of values sequentially from the most significant bits (MSBs) to the least significant bits (LSBs).To address this limitation, we propose BitL, a novel bit-wise computing unit designed to minimize the critical path and improve the throughput.BitL dynamically switches between horizontal and vertical data lookups across sub-tiles during Multiply-Accumulate (MAC) operations.Additionally, it presents an innovative optimization technique to maximize the utilization of computing units while switching its lookup direction.Our evaluation demonstrates that BitL delivers up to 1.92× higher throughput compared to a baseline DNN accelerator and achieves a 1.24× improvement over recent zero-bit skipping accelerators.Furthermore, BitL improves energy efficiency by 2.06× on average, with a silicon area overhead of only 5.71%. Seunghyun Lee 0003, Dongho Ha, Sungbin Kim, Sungwoo Kim 0003, Hyunwuk Lee, Won Woo Ro |
MICRO | 4 |
| 2024 | AirGun: Adaptive Granularity Quantization for Accelerating Large Language ModelsabstractTransformer-based models have evolved into Large Language Models (LLMs) by increasing model sizes to achieve higher accuracy, but they incur significant computational and memory costs. As quantization is a promising method to mitigate the huge cost of LLMs, the presence of outliers can lead to accuracy drops during quantization. Previous work pointed out LLMs have outliers only in specific input channels of activations. This suggests that per-input channel quantization would be beneficial, but it poses excessive computational overhead without optimization. To address these challenges, we propose a hardware and software co-design that mitigates the overhead of per-input channel quantization. We first propose AirGun, a quantization method that adaptively quantizes LLM modules. We observe that LLMs have high quantization sensitivity only in specific modules. Based on our observation, AirGun applies hardware-efficient per-tensor quantization for non-sensitive modules and per-input channel quantization for sensitive modules. For per-input channel quantization, we introduce early reconstruction and adaptive dyadic numbering, dismissing the overhead while exploiting its advantages. Additionally, we propose the AirGun accelerator that fully utilizes the advantages of AirGun. As a result, the AirGun accelerator achieves a 4.19 × speedup and 63.16 % lower energy consumption compared to the previous LM accelerator while achieving higher accuracy. Sungbin Kim, Hyunwuk Lee, Sungwoo Kim 0003, Cheolhwan Kim, Won Woo Ro |
ICCD | 3 |
| 2023 | Exploiting Inherent Properties of Complex Numbers for Accelerating Complex Valued Neural NetworksabstractSince conventional Deep Neural Networks (DNNs) use real numbers as their data, they are unable to capture the imaginary values and the correlations between real and imaginary values in applications that use complex numbers. To address this limitation, Complex Valued Neural Networks (CVNNs) have been introduced, enabling to capture the context of complex numbers for various applications such as Magnetic Resonance Imaging (MRI), radar, and sensing. CVNNs handle their data with complex numbers and adopt complex number arithmetic to their layer operations, so they exhibit distinct design challenges with real-valued DNNs. The first challenge is the data representation of the complex number, which requires two values for a single data, doubling the total data size of the networks. Moreover, due to the unique operations of the complex-valued layers, CVNNs require a specialized scheduling policy to fully utilize the hardware resources and achieve optimal performance. To mitigate the design challenges, we propose software and hardware co-design techniques that effectively resolves the memory and compute overhead of CVNNs. First, we propose Polar Form Aware Quantization (PAQ) that utilizes the characteristics of the complex number and their unique value distribution on CVNNs. Then, we propose our hardware accelerator that supports PAQ and CVNN operations. Lastly, we design a CVNN-aware scheduling scheme that optimizes the performance and resource utilization of an accelerator by aiming at the special layer operations of CVNN. PAQ achieves 62.5% data compression over CVNNs using FP16 while retaining a similar error with INT8 quantization, and our hardware support PAQ with only 2% area overhead over conventional systolic array architecture. In our evaluation, PAQ hardware with the scheduling scheme achieves a 32% lower latency and 30% lower energy consumption than other accelerators. Hyunwuk Lee, Hyungjun Jang, Sungbin Kim, Sungwoo Kim 0003, Wonho Cho, Won Woo Ro |
MICRO | 4 |
| 2023 | MAD MAcce: Supporting Multiply-Add Operations for Democratizing Matrix-Multiplication AcceleratorsabstractModern GPUs commonly employ specialized matrix multiplication units (MXUs) to accelerate matrix multiplication, the core computation of deep learning workloads. However, it is challenging to exploit the MXUs for GPGPU applications whose fundamental algorithms do not rely on matrix multiplication. Furthermore, an additional programming effort is necessary to tailor existing code or algorithms using dedicated APIs or libraries to utilize MXUs. Therefore, MXUs are often underutilized even when GPUs hunger for higher throughput. Seunghwan Sung, Sujin Hur, Sungwoo Kim 0003, Dongho Ha, Yunho Oh, Won Woo Ro |
MICRO | 3 |