VLDB 2026 Research / reviewers in the wild / expert
Sukhyun Han
dblp:406/2294
· DBLP profile ↗
3ranked-venue papers
1as first author
3since 2021 · last 2026
0009-0004-5862-5334ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Lupin: Spatial Resource Stealing with Outlier-First Encoding for Mixed-Precision LLM AccelerationabstractLLM inference often exceeds on-chip memory capacity, causing frequent external memory access. Quantization reduces memory cost but loses accuracy due to outliers. Prior mixed-precision accelerators address this issue with encoding schemes, but often result in accuracy degradation for LLMs and pipeline stalls. We present Lupin, an algorithm-architecture co-design with Outlier-First Encoding, which stores outliers in high precision by reallocating less critical normal values. This preserves maximal outlier representation and enables stall-free execution with low-precision MAC units. Experiments show that Lupin maintains accuracy while achieving a 2.02× speedup. Taein Kim, Sukhyun Han, Seongwook Kim, Gwangeun Byeon, Seokin Hong |
DATE | 2 |
| 2025 | Zebra: Leveraging Diagonal Attention Pattern for Vision Transformer AcceleratorabstractVision Transformers (ViTs) have achieved remarkable performance in computer vision, but their computational complexity and challenges in optimizing memory bandwidth limit hardware acceleration. A major bottleneck lies in the self-attention mechanism, which leads to excessive data movement and unnecessary computations despite high input sparsity and low computational demands. To address this challenge, existing transformer accelerators have leveraged sparsity in attention maps. However, their performance gains are limited due to low hardware utilization caused by the irregular distribution of nonzero values in the sparse attention maps. Self-attention often exhibits strong diagonal patterns in the attention map, as the diagonal elements tend to have higher values than others. To exploit this, we introduce Zebra, a hardware accelerator framework optimized for diagonal attention patterns. A core component of Zebra is the Striped Diagonal (SD) pruning technique, which prunes the attention map by preserving only the diagonal elements at runtime. This reduces computational load without requiring offline pre-computation or causing significant accuracy loss. Zebra features a reconfigurable accelerator architecture that supports optimized matrix multiplication method, called Striped Diagonal Matrix Multiplication (SDMM), which computes only the diagonal elements of matrices. With this novel method, Zebra addresses low hardware utilization, a key barrier to leveraging the diagonal patterns. Experimental results demonstrate that Zebra achieves a 57 × speedup over a CPU and 1.7 × over the state-of-the-art ViT accelerator with similar inference accuracy. Sukhyun Han, Seongwook Kim, Gwangeun Byeon, Jihun Yoon, Seokin Hong |
DATE | 1 |
| 2025 | Avalanche: Optimizing Cache Utilization via Matrix Reordering for Sparse Matrix Multiplication AcceleratorabstractSparse Matrix Multiplication (SpMM) is essential in various scientific and engineering applications but poses significant challenges due to irregular memory access patterns.Many hardware accelerators have been proposed to accelerate SpMM.However, they have yet to focus on on-chip memory utilization.In this paper, we highlight the underutilization of the on-chip memory in the SpMM accelerators.Then we propose Avalanche, a novel hardware accelerator that optimally utilizes the on-chip memory to efficiently cache both matrices 𝐵 and 𝐶.Avalanche incorporates three key techniques: Matrix Reordering (Mat-Reorder), Dead-Product Early Eviction (DP-Evict), and Reuse Distance-Aware Matrix Caching (RM-Caching).Mat-Reorder enhances data locality by reordering the columns of matrix 𝐴, ensuring early completion of computations for matrix 𝐶.DP-Evict optimizes on-chip memory usage by promptly evicting fully computed (dead) products from on-chip memory.RM-Caching maximizes data reuse by caching frequently accessed elements of matrix 𝐵 based on their reuse distance.Experimental results demonstrate that Avalanche achieves an average performance improvement of 1.97× compared to the state-of-theart SpMM accelerator, with a chip area of 6.15 mm 2 CCS Concepts• Computer systems organization → Special purpose systems. Gwangeun Byeon, Seongwook Kim, Sukhyun Han, Jinkwon Kim, Prashant J. Nair, Seokin Hong |
ISCA | 4 |