Gwangeun Byeon

dblp:364/7506 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
7since 2021 · last 2026
0000-0002-9029-5898ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 3 first-author · 7 since 2021Software engineering, systems software and programming languages · 4 · 2 first-author · 4 since 2021
YearPublicationVenuePosition
2026 Scrooge: Accelerating Attention Inference in LLMs via Early Termination Mechanism
abstract
Large Language Models (LLMs) have demonstrated remarkable performance in natural language processing and are now widely adopted in diverse applications. However, their significant computation and memory costs severely limit their acceleration. In particular, the self-attention mechanism is a significant bottleneck, as it cannot exploit batch parallelism across prompts, and its memory traffic grows quadratically with sequence length. In this paper, we propose Scrooge, a novel hardware accelerator framework that leverages an attention early termination mechanism, designed to address the inefficiency of self-attention. The self-attention mechanism does not assign equal importance to all tokens. Instead, semantically important tokens consistently receive higher attention scores. Consequently, preserving sufficient attention for a subset of important tokens is often enough to maintain model accuracy, even without computing attention for all tokens. Our key insight is that once sufficient attention has been accumulated, further computation with the remaining tokens only increases complexity without improving accuracy. Scrooge leverages this insight to approximate the attention of the remaining tokens and terminates the attention computation dynamically once it has gathered sufficient attention. With this method, Scrooge reduces both latency and memory traffic while maintaining accuracy. Experimental results show that Scrooge achieves a 1.7× speedup and a 0.47× reduction in memory traffic with negligible accuracy loss.
Gwangeun Byeon, Seongwook Kim, Taein Kim, Seokin Hong
DATE1
2026 Lupin: Spatial Resource Stealing with Outlier-First Encoding for Mixed-Precision LLM Acceleration
abstract
LLM inference often exceeds on-chip memory capacity, causing frequent external memory access. Quantization reduces memory cost but loses accuracy due to outliers. Prior mixed-precision accelerators address this issue with encoding schemes, but often result in accuracy degradation for LLMs and pipeline stalls. We present Lupin, an algorithm-architecture co-design with Outlier-First Encoding, which stores outliers in high precision by reallocating less critical normal values. This preserves maximal outlier representation and enables stall-free execution with low-precision MAC units. Experiments show that Lupin maintains accuracy while achieving a 2.02× speedup.
Taein Kim, Sukhyun Han, Seongwook Kim, Gwangeun Byeon, Seokin Hong
DATE4
2025 Zebra: Leveraging Diagonal Attention Pattern for Vision Transformer Accelerator
abstract
Vision Transformers (ViTs) have achieved remarkable performance in computer vision, but their computational complexity and challenges in optimizing memory bandwidth limit hardware acceleration. A major bottleneck lies in the self-attention mechanism, which leads to excessive data movement and unnecessary computations despite high input sparsity and low computational demands. To address this challenge, existing transformer accelerators have leveraged sparsity in attention maps. However, their performance gains are limited due to low hardware utilization caused by the irregular distribution of nonzero values in the sparse attention maps. Self-attention often exhibits strong diagonal patterns in the attention map, as the diagonal elements tend to have higher values than others. To exploit this, we introduce Zebra, a hardware accelerator framework optimized for diagonal attention patterns. A core component of Zebra is the Striped Diagonal (SD) pruning technique, which prunes the attention map by preserving only the diagonal elements at runtime. This reduces computational load without requiring offline pre-computation or causing significant accuracy loss. Zebra features a reconfigurable accelerator architecture that supports optimized matrix multiplication method, called Striped Diagonal Matrix Multiplication (SDMM), which computes only the diagonal elements of matrices. With this novel method, Zebra addresses low hardware utilization, a key barrier to leveraging the diagonal patterns. Experimental results demonstrate that Zebra achieves a 57 × speedup over a CPU and 1.7 × over the state-of-the-art ViT accelerator with similar inference accuracy.
Sukhyun Han, Seongwook Kim, Gwangeun Byeon, Jihun Yoon, Seokin Hong
DATE3
2025 Avalanche: Optimizing Cache Utilization via Matrix Reordering for Sparse Matrix Multiplication Accelerator
abstract
Sparse Matrix Multiplication (SpMM) is essential in various scientific and engineering applications but poses significant challenges due to irregular memory access patterns.Many hardware accelerators have been proposed to accelerate SpMM.However, they have yet to focus on on-chip memory utilization.In this paper, we highlight the underutilization of the on-chip memory in the SpMM accelerators.Then we propose Avalanche, a novel hardware accelerator that optimally utilizes the on-chip memory to efficiently cache both matrices 𝐵 and 𝐶.Avalanche incorporates three key techniques: Matrix Reordering (Mat-Reorder), Dead-Product Early Eviction (DP-Evict), and Reuse Distance-Aware Matrix Caching (RM-Caching).Mat-Reorder enhances data locality by reordering the columns of matrix 𝐴, ensuring early completion of computations for matrix 𝐶.DP-Evict optimizes on-chip memory usage by promptly evicting fully computed (dead) products from on-chip memory.RM-Caching maximizes data reuse by caching frequently accessed elements of matrix 𝐵 based on their reuse distance.Experimental results demonstrate that Avalanche achieves an average performance improvement of 1.97× compared to the state-of-theart SpMM accelerator, with a chip area of 6.15 mm 2 CCS Concepts• Computer systems organization → Special purpose systems.
Gwangeun Byeon, Seongwook Kim, Sukhyun Han, Jinkwon Kim, Prashant J. Nair, Seokin Hong
ISCA1
2024 A Case for Speculative Address Translation with Rapid Validation for GPUs
abstract
A unified address space is vital for heterogeneous systems as it enables efficient data sharing between CPUs and GPUs. However, GPU address translation faces challenges due to high TLB pressure, particularly with irregular and memory-intensive applications. Compared to an ideal scenario, we observe that address translation overheads cause a slowdown of up to 34.5% in modern heterogeneous systems. This paper introduces Avatar, a novel framework to accelerate address translation in GPUs. Avatar comprises two key components: Contiguity-Aware Speculative Translation (CAST) and In-Cache Validation (CAVA) mechanisms. Avatar identifies the potential for predicting virtual-to-physical address mapping by monitoring contiguous pages that lie in both virtual and physical address spaces. Leveraging this insight, CAST speculatively translates virtual addresses into physical addresses. This speculative address translation enables immediate data fetching into GPUs while addressing translation occurs in the background, reducing TLB-miss overhead. Unfortunately, modern GPUs lack support for speculative execution, which limits CAST's performance gain. Data fetched from speculated physical addresses is unusable until validation. CAVA addresses this limitation by quickly validating speculated physical addresses. To this end, CAVA embeds page mapping information into each 32B sector of 128B cache lines. Thus, CAVA enables fetching a sector block from memory for a speculated address and rapidly validating the speculative translation using the embedded mapping information. Our experiments show that Avatar achieves a 90.3% (high) speculation accuracy and improves GPU performance by 37.2% (on average).
Junhyeok Park 0001, Osang Kwon, Seongwook Kim, Gwangeun Byeon, Jihun Yoon, Prashant J. Nair, Seokin Hong
MICRO5
2023 SparseFT: Sparsity-aware Fault Tolerance for Reliable CNN Inference on GPUs
abstract
Graphics Processing Units (GPUs), while offering exceptional performance for CNN inference tasks, are susceptible to both transient and permanent hardware faults due to the integration of numerous processing elements and advancements in technology scaling. This paper proposes a novel and cost-effective fault mitigation technique, called Sparsity-aware Fault Tolerance (SparseFT), to ensure reliable CNN inference on GPUs. SparseFT leverages inherent sparsity in the activation maps to detect and correct errors on the processing elements without hardware redundancy. By exploiting the characteristic of dot-products, where multiplications with zero operands are ineffectual, SparseFT dynamically duplicates an effectual computation (i.e., a multiplication with non-zero operands) to the processing element initially assigned to the ineffectual one. It then compares the duplicated computation results to detect errors. Experimental results demonstrate that SparseFT achieves more than 97% error detection coverage with less than 1% performance overhead for the state-of-the-art CNN models.
Gwangeun Byeon, Seungtae Lee, Seongwook Kim, Prashant J. Nair, Seokin Hong
PACT1
2023 Conveyor: Towards Asynchronous Dataflow in Systolic Array to Exploit Unstructured Sparsity
abstract
Systolic array (SA) architecture efficiently offers parallel computation using a simple data movement across processing elements. However, their rigid structure and synchronous dataflow limit flexibility in handling sparse computations, resulting in underutilized resources and suboptimal performance. In this paper, we propose Conveyor-SA, a novel SA-based accelerator architecture leveraging asynchronous dataflow for unstructured sparsity exploitation. Conveyor-SA introduces three core mechanisms: Chunk Propagation for parallel data processing, PE Grouping to accelerate efficiently both sparse and dense CNN models, and Conveyor Queue for load imbalance mitigation. Our experimental results demonstrate that Conveyor-SA achieves an average speedup of 1.68x over the competitors while processing conventional CNN models. In addition, Conveyor-SA delivers 1.42x speedup over state-of-the-art sparse SA architecture while remarkably reducing the chip area requirement by 19.5%.
Seongwook Kim, Gwangeun Byeon, Sihyung Kim, Seokin Hong
ICCD2