Seokin Hong

dblp:78/9350 · DBLP profile ↗
← Back
39ranked-venue papers
7as first author
23since 2021 · last 2026
0000-0001-7842-125XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 38 · 7 first-author · 22 since 2021Software engineering, systems software and programming languages · 11 · 1 first-author · 9 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Scrooge: Accelerating Attention Inference in LLMs via Early Termination Mechanism
abstract
Large Language Models (LLMs) have demonstrated remarkable performance in natural language processing and are now widely adopted in diverse applications. However, their significant computation and memory costs severely limit their acceleration. In particular, the self-attention mechanism is a significant bottleneck, as it cannot exploit batch parallelism across prompts, and its memory traffic grows quadratically with sequence length. In this paper, we propose Scrooge, a novel hardware accelerator framework that leverages an attention early termination mechanism, designed to address the inefficiency of self-attention. The self-attention mechanism does not assign equal importance to all tokens. Instead, semantically important tokens consistently receive higher attention scores. Consequently, preserving sufficient attention for a subset of important tokens is often enough to maintain model accuracy, even without computing attention for all tokens. Our key insight is that once sufficient attention has been accumulated, further computation with the remaining tokens only increases complexity without improving accuracy. Scrooge leverages this insight to approximate the attention of the remaining tokens and terminates the attention computation dynamically once it has gathered sufficient attention. With this method, Scrooge reduces both latency and memory traffic while maintaining accuracy. Experimental results show that Scrooge achieves a 1.7× speedup and a 0.47× reduction in memory traffic with negligible accuracy loss.
Gwangeun Byeon, Seongwook Kim, Taein Kim, Seokin Hong
DATE5
2026 Braid-ZNS: Leveraging Zone Random Write Area for Efficient In-Storage Compression on ZNS SSDs
abstract
Zoned Namespace (ZNS) SSD is an emerging storage solution that reduces device-level garbage collection and write amplification. However, the sequential write constraint of ZNS SSDs poses a challenge to adopting in-storage compression, as data placement rules prevent compressed variable-length data from being packed into optimally sized chunks. In this paper, we propose Braid-ZNS, a novel in-storage compression framework that leverages the Zone Random Write Area (ZRWA) to avoid double reads during data compression on ZNS SSDs. By exploiting the ZRWA to enable temporary in-place updates, Braid-ZNS reorganizes compressed blocks in a size-aware manner and prevents cases where a single logical page is split into multiple fragments. Our evaluation shows Braid-ZNS improved compression efficiency by up to 47.0% and throughput by ×2.24 compared to a state-of-the-art in-storage compression on ZNS SSDs.
Joonseong Hwang, Minjin Park, Seokin Hong
DATE4
2026 Lupin: Spatial Resource Stealing with Outlier-First Encoding for Mixed-Precision LLM Acceleration
abstract
LLM inference often exceeds on-chip memory capacity, causing frequent external memory access. Quantization reduces memory cost but loses accuracy due to outliers. Prior mixed-precision accelerators address this issue with encoding schemes, but often result in accuracy degradation for LLMs and pipeline stalls. We present Lupin, an algorithm-architecture co-design with Outlier-First Encoding, which stores outliers in high precision by reallocating less critical normal values. This preserves maximal outlier representation and enables stall-free execution with low-precision MAC units. Experiments show that Lupin maintains accuracy while achieving a 2.02× speedup.
Taein Kim, Sukhyun Han, Seongwook Kim, Gwangeun Byeon, Seokin Hong
DATE6
2026 Boosting LLC Bandwidth Utilization in GPUs through Adaptive Fine-Grained Data Migration
abstract
Modern server-grade GPUs (e.g., NVIDIA A100) integrate hundreds of cores and tens of memory partitions, providing massive compute capability and memory bandwidth. However, the increased scale amplifies interconnect overhead between cores and memory partitions. To mitigate this, NVIDIA A100 clusters multiple cores and memory partitions into two large groups, thereby simplifying interconnect complexity. Unfortunately, this partitioning introduces a new limitation: remote partition accesses. A core accessing a remote memory partition incurs higher latency and lower bandwidth compared to local accesses.In this paper, we propose a cache-line migration mechanism across partitions to alleviate remote memory access overhead. Our design is motivated by two key observations: (1) conventional GPUs employ limited and often ineffective optimizations for remote access handling, and (2) GPU applications typically exhibit high temporal locality, where a specific partition of cores makes frequent memory accesses for the same data within short time intervals. Leveraging these insights, we dynamically migrate cache-lines to the local partition where the requesting core resides. Experimental results demonstrate that our approach achieves up to 1.24× speedup over the baseline with NVIDIA A100-like replication, highlighting its effectiveness in reducing remote access penalties.
Jihun Yoon, Sungbin Jang, Seokin Hong
DATE3
2025 LibraPIM: Dynamic Load Rebalancing to Maximize Utilization in PIM-Assisted LLM Inference Systems
abstract
Large language models (LLMs) require inference systems that can handle both compute- and memory-intensive workloads. GPUs and NPUs (referred to as xPUs) efficiently process compute-intensive layers, while Processing-In-Memory (PIM) architectures are well suited for memory-bound stages. To exploit this complementary relationship, recent LLM inference systems have adopted heterogeneous architectures integrating xPUs with PIM units. However, this integration poses several challenges. Tight execution dependencies between the two devices limit concurrency, as PIM often must wait for data produced by the xPUs. Moreover, batch size and sequence length affect the computational load on PIM and the accelerator differently, leading to execution imbalance across devices and degrading their utilization. To address these challenges, we propose LibraPIM, a novel PIM framework that orchestrates workload rebalancing and concurrent execution between compute accelerators and in-memory compute units, enabling efficient and scalable LLM inference. LibraPIM addresses the above challenges through two key techniques: Dynamic Batch Offloading (DBO), which adaptively redirects portions of the workload to the underutilized device based on runtime profiling, and Dual-Path Execution (DEX), which enables concurrent PIM operations and memory accesses through sub-bank partitioning with minimal hardware overhead. Together, these techniques improve resource utilization across heterogeneous devices, thereby increasing system throughput in LLM inference. Our experimental results demonstrate that LibraPIM achieves $6.2 \times$ average speedup over the baseline PIM-enabled system and delivers a $2.1 \times$ average speedup compared to a state-of-the-art approach.
Hyeongjun Cho, Yoonho Jang, Hyungi Kim, Seongwook Kim, Keewon Kwon, Gwangsun Kim, Seokin Hong
PACT7
2025 PIMPAL: Accelerating LLM Inference on Edge Devices via In-DRAM Arithmetic Lookup
abstract
Deploying Large Language Models (LLMs) on edge devices poses significant challenges due to their high computational and memory demands. In particular, General MatrixVector Multiplication (GEMV), a key operation in LLM inference, is highly memory-intensive, making it difficult to accelerate using conventional edge computing systems. While Processing-in-memory (PIM) architectures have emerged as a promising solution to this challenge, they often suffer from high area overhead or restricted computational precision. This paper proposes PIMPAL (Processing-In-Memory architecture with Parallel Arithmetic Lookup), a cost-effective PIM architecture leveraging LookUp Table (LUT)-based computation for GEMV acceleration in sLLMs (small LLMs). By replacing traditional arithmetic operations with parallel in-DRAM LUT lookups, PIMPAL significantly reduces area overhead while maintaining high performance. PIMPAL introduces three key innovations: (1) it divides DRAM bank subarrays into compute blocks for parallel LUT processing; (2) it employs Localityaware Compute Mapping (LCM) to reduce row activations by maximizing LUT access locality; and (3) it enables multi-precision computations through a LUT Aggregation (LAG) mechanism that combines results from multiple small LUTs. Experimental results show that PIMPAL achieves up to $17.8 x$ higher performance than previous LUT-based PIM designs and reduces area overhead by $40 \%$ compared to conventional processing unit-based PIM designs.
Yoonho Jang, Hyeongjun Cho, Yesin Ryu, Jungrae Kim, Seokin Hong
DAC5
2025 Zebra: Leveraging Diagonal Attention Pattern for Vision Transformer Accelerator
abstract
Vision Transformers (ViTs) have achieved remarkable performance in computer vision, but their computational complexity and challenges in optimizing memory bandwidth limit hardware acceleration. A major bottleneck lies in the self-attention mechanism, which leads to excessive data movement and unnecessary computations despite high input sparsity and low computational demands. To address this challenge, existing transformer accelerators have leveraged sparsity in attention maps. However, their performance gains are limited due to low hardware utilization caused by the irregular distribution of nonzero values in the sparse attention maps. Self-attention often exhibits strong diagonal patterns in the attention map, as the diagonal elements tend to have higher values than others. To exploit this, we introduce Zebra, a hardware accelerator framework optimized for diagonal attention patterns. A core component of Zebra is the Striped Diagonal (SD) pruning technique, which prunes the attention map by preserving only the diagonal elements at runtime. This reduces computational load without requiring offline pre-computation or causing significant accuracy loss. Zebra features a reconfigurable accelerator architecture that supports optimized matrix multiplication method, called Striped Diagonal Matrix Multiplication (SDMM), which computes only the diagonal elements of matrices. With this novel method, Zebra addresses low hardware utilization, a key barrier to leveraging the diagonal patterns. Experimental results demonstrate that Zebra achieves a 57 × speedup over a CPU and 1.7 × over the state-of-the-art ViT accelerator with similar inference accuracy.
Sukhyun Han, Seongwook Kim, Gwangeun Byeon, Jihun Yoon, Seokin Hong
DATE5
2025 Improving Address Translation in Tagless DRAM Cache by Caching PTE Pages
abstract
This paper proposes a novel caching mechanism for PTE pages to enhance the Tagless DRAM Cache architecture and improve address translation in large on-package DRAM caches. Existing OS-managed DRAM cache architectures have achieved significant performance improvements by focusing on efficient tag management. However, prior studies primarily update PTEs after caching data pages without directly accessing them from the DRAM cache, leading to performance degradation during page walks. To address this, we propose a method to simultaneously cache data pages and PTE pages in the DRAM cache. This approach reduces address translation latency and cache access delays. Additionally, we introduce a shootdown mechanism to maintain PTE and page walk cache consistency in multi-core systems, ensuring all cores access the latest information for shared pages. Experimental results show that caching PTE pages reduces address translation overhead by up to 33.3% compared to traditional OS-managed tagless DRAM caches, improving overall program execution time by an average of 10.5%. This method effectively mitigates bottlenecks caused by address translation.
Osang Kwon, Seokin Hong
DATE3
2025 Buddy ECC: Making Cache Mostly Clean in CXL-Based Memory Systems for Enhanced Error Correction at Low Cost
abstract
As Compute Express Link (CXL) emerges as a key memory interconnect, interest in optimization opportunities and challenges has grown. However, due to the different characteristics of the CXL Memory Module (CMM) compared to traditional DRAM-based Dual In-line Memory Modules (DIMMs), existing optimizations may not be effectively applied. In this paper, we propose an Buddy ECC that leverages the full-duplex nature and features of the CMM to optimize bandwidth, enhance reliability, and reduce area overhead. First, the Proactively Write-back improves bandwidth efficiency by minimizing dirty cachelines in the last-level cache through dead block prediction, proactively identifying and writing back cachelines that are unlikely to be rewritten. Second, the Utilization-aware Policy dynamically monitors the internal bandwidth of the CMM, sending write-back requests only when the module is under low-load rate, thus preventing performance degradation during high traffic. Finally, the Buddy ECC scheme enhances data reliability by separating Error Detection Code (EDC) for clean cachelines and stronger Error Correction Code (ECC) for dirty cachelines. Buddy ECC improved bandwidth utilization by 46%, limited performance degradation to 0.33%, and kept energy consumption increase under 1%.
Junbum Park, Osang Kwon, Sungbin Jang, Seokin Hong
DATE5
2025 SPB: Towards Low-Latency CXL Memory via Speculative Protocol Bypassing
abstract
Compute Express Link (CXL) is an advanced inter-connect standard designed to facilitate high-speed communication between CPUs, accelerators, and memory devices, making it well-suited for data-intensive applications such as machine learning and real-time analytics. Despite its advantages, CXL memory encounters significant latency challenges due to the complex hierarchy of protocol layers, which can adversely impact performance in latency-sensitive scenarios. To address this issue, we introduce the Speculative Protocol Bypassing (SPB) architecture, which aims to minimize latency during read operations by speculatively bypassing several protocol layers of CXL. To achieve this, SPB employs the Snooper mechanism, which extracts essential read commands from the packet data at an early stage, allowing it to bypass multiple protocol layers and reduce memory access time. Additionally, the Hazard Filter (HF) prevents Read-After-Write (RAW) hazards between read and write operations, thereby maintaining data integrity and ensuring system reliability. The SPB architecture effectively optimizes CXL memory access latency, providing a robust solution for high-performance computing environments that require low latency as well as high efficiency. Its minimal hardware overhead makes it a practical and scalable enhancement for future CXL-based memory.
Junbum Park, Sungbin Jang, Wonyoung Lee 0001, Seokin Hong
DATE5
2025 Minimizing Read Disturb via Localized Page Allocation for Modern NAND Flash-Based SSDs
abstract
To meet the increasing demand for higher storage density, modern NAND flash-based SSDs employ Quad-Level Cell (QLC) technology, which stores four bits per memory cell. However, it exacerbates the read disturb issue, where repeated read operations gradually shift the threshold voltages of unselected NAND flash cells. This voltage drift leads to frequent read retries and accelerates device wear, ultimately degrading performance and endurance. To address this issue, we propose Localized Page Allocation (LPA), a novel data mapping scheme that allocates each page to a restricted region of the wordline. LPA assigns four consecutive bits in a single page to a single cell. This approach confines read operations to a limited portion of the bitlines, thereby reducing the number of NAND flash cells exposed to read disturb. However, this localized access increases the complexity of sensing operations. To address this challenge, we introduce Progressive Voltage Convergence (PVC) method that adaptively determines the next sensing voltage for each bitline based on its previous sensing result. This fine-grained control reduces the number of sensing steps during read operations. For additional optimization, we employ lossless compression, which reduces the number of sensing steps for the compressed pages. Experimental evaluations using real-world workloads demonstrate that our design improves I/O performance by up to$\mathbf{1 1 \%}$and extends endurance by 79% compared to the state-of-the-art techniques. Our design requires minor hardware modifications to the conventional NAND flash chips and the SSD controller, which leads to small area and power overhead.
Joonseong Hwang, Minjin Park, Jihun Yoon, Yoonho Jang, Seokin Hong
ICCD6
2025 DDLM: Demand-Aware Dynamic Link Width Management for Energy-Efficient CXL Memory
abstract
Modern data-centric workloads are driving rapid growth in the demand for high-bandwidth, large-capacity memory. Compute Express Link (CXL) has emerged as a key technology for scalable memory expansion through the Peripheral Component Interconnect Express (PCIe)'s high-speed serial interfaces. However, while PCIe 6.0's L0p mode enables lane-level power gating without disrupting traffic flow, it lacks a policy mechanism for dynamically deciding the number of active lanes. Since actual energy savings depend on matching active lanes to runtime bandwidth demand, such a policy is essential. This paper presents DDLM (Demand-Aware Dynamic Link Width Management), a lightweight control framework that improves CXL link energy efficiency by dynamically adjusting link width based on runtime traffic. DDLM integrates two complementary modules: a predictor that predicts bandwidth demand based on short-term temporal locality to set the width ahead of hardware transition latency, and a congestion monitor that detects transient congestion via internal queue inspection. These modules are coordinated by a finite state machine that manages safe link width transitions under physical-layer constraints. We implement DDLM in the Ramulator2 simulator and evaluate it with diverse SPEC CPU workloads. Compared to a fixed-width baseline, DDLM reduces CXL Memory energy by up to 13%, improves utilization by$2.22 \times$, and limits performance loss to under 3%. This work offers a practical path to energy-proportional CXL Memory.
Taejeong Kim, Junbum Park, Seokin Hong
ICCD4
2025 Avalanche: Optimizing Cache Utilization via Matrix Reordering for Sparse Matrix Multiplication Accelerator
abstract
Sparse Matrix Multiplication (SpMM) is essential in various scientific and engineering applications but poses significant challenges due to irregular memory access patterns.Many hardware accelerators have been proposed to accelerate SpMM.However, they have yet to focus on on-chip memory utilization.In this paper, we highlight the underutilization of the on-chip memory in the SpMM accelerators.Then we propose Avalanche, a novel hardware accelerator that optimally utilizes the on-chip memory to efficiently cache both matrices 𝐵 and 𝐶.Avalanche incorporates three key techniques: Matrix Reordering (Mat-Reorder), Dead-Product Early Eviction (DP-Evict), and Reuse Distance-Aware Matrix Caching (RM-Caching).Mat-Reorder enhances data locality by reordering the columns of matrix 𝐴, ensuring early completion of computations for matrix 𝐶.DP-Evict optimizes on-chip memory usage by promptly evicting fully computed (dead) products from on-chip memory.RM-Caching maximizes data reuse by caching frequently accessed elements of matrix 𝐵 based on their reuse distance.Experimental results demonstrate that Avalanche achieves an average performance improvement of 1.97× compared to the state-of-theart SpMM accelerator, with a chip area of 6.15 mm 2 CCS Concepts• Computer systems organization → Special purpose systems.
Gwangeun Byeon, Seongwook Kim, Sukhyun Han, Jinkwon Kim, Prashant J. Nair, Seokin Hong
ISCA8
2025 Leveraging Chiplet-Locality for Efficient Memory Mapping in Multi-Chip Module GPUs
Junhyeok Park 0001, Sungbin Jang, Osang Kwon, Seokin Hong
MICRO5
2025 SoftWalker: Supporting Software Page Table Walk for Irregular GPU Applications
abstract
Address translation has become a significant and growing performance bottleneck in modern GPUs, especially for emerging irregular applications with high TLB miss rates.The limited concurrency of hardware Page Table Walkers (PTWs), due to their small and fixed number, causes severe contention and substantial queueing delays under high translation pressure, which significantly degrades performance.This paper introduces SoftWalker, a novel, scalable, and flexible framework that fundamentally shifts the GPU page table walking from fixed-function hardware to software execution.SoftWalker leverages the GPU's massive thread-level parallelism by dynamically dispatching specialized, lightweight software threads running on GPU cores to handle TLB misses requiring page table walks.In addition, to expand L2 TLB MSHR capacity on demand, SoftWalker incorporates In-TLB MSHRs, a key innovation that repurposes underutilized L2 TLB entries to track outstanding misses when existing MSHRs are saturated.By alleviating MSHR-induced contention, this design preserves the key advantage of highly parallel page table walking in software.SoftWalker enables thousands of concurrent page table walks, significantly reducing PTW-level contention and translation queueing delays.As a result, it achieves an average reduction of 72.8% in page walk latency and delivers an average speedup of 2.24× (3.94× for irregular workloads).
Sungbin Jang, Junhyeok Park 0001, Osang Kwon, Juyoung Seok, Seokin Hong
MICRO7
2024 Rethinking Page Table Structure for Fast Address Translation in GPUs: A Fixed-Size Hashed Page Table
abstract
GPU memory virtualization has become essential for efficient programming, memory management, and address space sharing among computing devices in heterogeneous systems. Conventional GPU virtual memory systems use multi-level Radix Page Tables (RPTs) to store virtual-to-physical address mapping in device (GPU) memory. When a TLB miss occurs, a page table walker accesses each level of the page table sequentially to find the desired mapping. These sequential accesses significantly degrade performance, adding pressure to the GPU memory hierarchy. To make matters worse, recent computing systems now support five-level RPTs, further increasing the number of memory accesses required per page table walk.
Sungbin Jang, Junhyeok Park 0001, Osang Kwon, Seokin Hong
PACT5
2024 Distributed Page Table: Harnessing Physical Memory as an Unbounded Hashed Page Table
abstract
Virtual memory systems rely on the page table, a crucial component that maps virtual addresses to physical addresses (i.e., address translation). While the Radix Page Table (RPT) has traditionally been used for this task, its limitations have become more apparent with the rise of memory-intensive applications. Recently, Hashed Page Tables (HPTs) have been explored as an alternative page table structure to offer faster address translation. However, the HPT introduces its own set of challenges particularly in resizing the page table and allocating contiguous physical memory space for storing the table. To tackle the fundamental problem of the existing HPT designs, this paper introduces Distributed Page Table (DPT), a novel approach that utilizes the physical memory as a huge hashed page table. DPT distributes Page Table Entries (PTEs) across the entire physical memory space, significantly reducing the hash collisions while avoiding the table resizing overheads. When distributing the PTEs across the physical memory, they can be mapped to memory locations already allocated to data pages. This new type of collision, referred to as address collision, may reduce the effectiveness of the DPT. This paper showcases that the DPT can effectively resolve the address collision with three simple yet efficient techniques: Strided Open Addressing (SOA), Collision-Aware Virtual Address Allocation (CVA) and Collided Page Displacement (CPD). Our experimental results demonstrate that DPT achieves average performance improvements of 12.6%, 11.6%, and 8.7% compared to traditional RPT, the latest large-coverage TLB design, and state-of-the-art HPTs, respectively.
Osang Kwon, Junhyeok Park 0001, Sungbin Jang, Byung-Chul Tak, Seokin Hong
MICRO6
2024 A Case for Speculative Address Translation with Rapid Validation for GPUs
abstract
A unified address space is vital for heterogeneous systems as it enables efficient data sharing between CPUs and GPUs. However, GPU address translation faces challenges due to high TLB pressure, particularly with irregular and memory-intensive applications. Compared to an ideal scenario, we observe that address translation overheads cause a slowdown of up to 34.5% in modern heterogeneous systems. This paper introduces Avatar, a novel framework to accelerate address translation in GPUs. Avatar comprises two key components: Contiguity-Aware Speculative Translation (CAST) and In-Cache Validation (CAVA) mechanisms. Avatar identifies the potential for predicting virtual-to-physical address mapping by monitoring contiguous pages that lie in both virtual and physical address spaces. Leveraging this insight, CAST speculatively translates virtual addresses into physical addresses. This speculative address translation enables immediate data fetching into GPUs while addressing translation occurs in the background, reducing TLB-miss overhead. Unfortunately, modern GPUs lack support for speculative execution, which limits CAST's performance gain. Data fetched from speculated physical addresses is unusable until validation. CAVA addresses this limitation by quickly validating speculated physical addresses. To this end, CAVA embeds page mapping information into each 32B sector of 128B cache lines. Thus, CAVA enables fetching a sector block from memory for a speculated address and rapidly validating the speculative translation using the embedded mapping information. Our experiments show that Avatar achieves a 90.3% (high) speculation accuracy and improves GPU performance by 37.2% (on average).
Junhyeok Park 0001, Osang Kwon, Seongwook Kim, Gwangeun Byeon, Jihun Yoon, Prashant J. Nair, Seokin Hong
MICRO8
2023 SparseFT: Sparsity-aware Fault Tolerance for Reliable CNN Inference on GPUs
abstract
Graphics Processing Units (GPUs), while offering exceptional performance for CNN inference tasks, are susceptible to both transient and permanent hardware faults due to the integration of numerous processing elements and advancements in technology scaling. This paper proposes a novel and cost-effective fault mitigation technique, called Sparsity-aware Fault Tolerance (SparseFT), to ensure reliable CNN inference on GPUs. SparseFT leverages inherent sparsity in the activation maps to detect and correct errors on the processing elements without hardware redundancy. By exploiting the characteristic of dot-products, where multiplications with zero operands are ineffectual, SparseFT dynamically duplicates an effectual computation (i.e., a multiplication with non-zero operands) to the processing element initially assigned to the ineffectual one. It then compares the duplicated computation results to detect errors. Experimental results demonstrate that SparseFT achieves more than 97% error detection coverage with less than 1% performance overhead for the state-of-the-art CNN models.
Gwangeun Byeon, Seungtae Lee, Seongwook Kim, Prashant J. Nair, Seokin Hong
PACT6
2023 Facto-CNN: Memory-Efficient CNN Training with Low-rank Tensor Factorization and Lossy Tensor Compression
Seungtae Lee, Jonghwan Ko, Seokin Hong
ACML3
2023 Conveyor: Towards Asynchronous Dataflow in Systolic Array to Exploit Unstructured Sparsity
abstract
Systolic array (SA) architecture efficiently offers parallel computation using a simple data movement across processing elements. However, their rigid structure and synchronous dataflow limit flexibility in handling sparse computations, resulting in underutilized resources and suboptimal performance. In this paper, we propose Conveyor-SA, a novel SA-based accelerator architecture leveraging asynchronous dataflow for unstructured sparsity exploitation. Conveyor-SA introduces three core mechanisms: Chunk Propagation for parallel data processing, PE Grouping to accelerate efficiently both sparse and dense CNN models, and Conveyor Queue for load imbalance mitigation. Our experimental results demonstrate that Conveyor-SA achieves an average speedup of 1.68x over the competitors while processing conventional CNN models. In addition, Conveyor-SA delivers 1.42x speedup over state-of-the-art sparse SA architecture while remarkably reducing the chip area requirement by 19.5%.
Seongwook Kim, Gwangeun Byeon, Sihyung Kim, Seokin Hong
ICCD5
2022 Don't open row: rethinking row buffer policy for improving performance of non-volatile memories
abstract
Among the various NVM technologies, phase-change-memory (PCM) has attracted substantial attention as a candidate to replace the DRAM for next-generation memory. However, the characteristics of PCM cause it to have much longer read and write latencies than DRAM. This paper proposes a Write-Around PCM System that addresses this limitation using two novel schemes: Pseudo-Row Activation and Direct Write. Pseudo-Row Activation provides fast row activation for PCM writes by connecting a target row to bitlines, but it does not fetch the data into the row buffer. With the Direct Write scheme, our system allows for writing operations to update the data even if the target row is in the logically closed state.
Osang Kwon, Seokin Hong
DAC3
2021 CID: Co-Architecting Instruction Cache and Decompression System for Embedded Systems
abstract
Code compression is widely used to reduce the footprint of code memory in cost-sensitive embedded systems. However, despite the small code size, the decompressor and the address translator required to support the code compression incur energy and area overheads. To reduce such overheads while still supporting code compression, we co-architect the instruction cache and decompression system (CID). In CID, each component is placed at the optimal location and the instruction cache is redesigned to recognize the compression state and retain the original address, through the cache division and address space decompression process. As a result of the cache division, the energy consumption and area overheads of the CID instruction cache are reduced. Since the decompressor overhead depends on the code compression technique, we propose a new code compression technique called entropy-based pattern code compression, which reduces overheads of the decompressor. Our experimental results show that the total energy consumption of the instruction cache and decompression system is reduced by up to 29.7 percent and their area is reduced by up to 15.4 percent compared to the post-cache architecture with almost no performance degradation, while achieving an 18.8 percent improvement in the compression ratio compared to the state-of-the-art code compression technique.
Jinkwon Kim, Seokin Hong, Jeongkyu Hong, Soontae Kim
IEEE Trans. Computers2
2020 ADAM: Adaptive Block Placement with Metadata Embedding for Hybrid Caches
abstract
Spin-Transfer Torque Random Access Memory (STT-RAM) is a potential alternative for SRAM-based on-chip caches. STT-RAM offers high density and low leakage power, thereby can be used to build a large capacity last-level caches (LLC). Unfortunately, the write latency of the STT-RAM is significantly longer, and its write energy is considerably higher compared to SRAM. To mitigate these concerns, researchers have proposed hybrid caches that are comprised of SRAM and STT-RAM regions. In such hybrid caches, an intelligent block placement policy is necessary to store as many write-intensive blocks in the SRAM region. This paper proposes an adaptive block placement framework with metadata embedding (ADAM) for hybrid caches. ADAM embeds metadata (i.e., write-intensity) into a cache block when it is evicted from LLC. When a cache block is brought from the main memory, metadata embedded in the block is extracted and used to determine the write-intensity of the block. Our evaluation shows that ADAM can improve performance by 26 % (on average) over a baseline block placement scheme.
Prashant J. Nair, Seokin Hong
ICCD3
2019 Split-CNN: Splitting Window-based Operations in Convolutional Neural Networks for Memory System Optimization
abstract
We present an interdisciplinary study to tackle the memory bottleneck of training deep convolutional neural networks (CNN). Firstly, we introduce Split Convolutional Neural Network (Split-CNN) that is derived from the automatic transformation of the state-of-the-art CNN models. The main distinction between Split-CNN and regular CNN is that Split-CNN splits the input images into small patches and operates on these patches independently before entering later stages of the CNN model. Secondly, we propose a novel heterogeneous memory management system (HMMS) to utilize the memory-friendly properties of Split-CNN. Through experiments, we demonstrate that Split-CNN achieves significantly higher training scalability by dramatically reducing the memory requirements of training algorithms on GPU accelerators. Furthermore, we provide empirical evidence that splitting at randomly chosen boundaries can even result in accuracy gains over baseline CNN due to its regularization effect.
Seokin Hong
ASPLOS2
2019 Touché: Towards Ideal and Efficient Cache Compression By Mitigating Tag Area Overheads
abstract
Compression is seen as a simple technique to increase the effective cache capacity. Unfortunately, compression techniques either incur tag area overheads or restrict cache block placement to only include neighboring addresses. Ideally, we should be able to place compressed cache blocks without any restrictions or overheads.
Seokin Hong, Bülent Abali, Alper Buyuktosunoglu, Michael B. Healy, Prashant J. Nair
MICRO1
2019 Interpage-Based Endurance-Enhancing Lower State Encoding for MLC and TLC Flash Memory Storages
abstract
During the past decade, the endurance of NAND flash memory has severely deteriorated. The maximum number of program and erase cycles has fallen significantly with emerging of multilevel cell (MLC) and triple-level cell (TLC) technology, and scaling down of the cell size. Wear leveling is a general solution used to alleviate this issue; it enables cells to wear down evenly but it cannot actually mitigate the wearing of the cells. Accordingly, techniques are required to minimize the actual cell degradation. This paper proposesendurance-enhancing lower state encoding. The key insight leveraged by the proposed technique is the data pattern-related characteristic of MLC and TLC NAND flash memories, in which the lower the state of the cells, the lower the occurrence of wear out. Thus, our proposed scheme encodes input data to make the cell state as low as possible in consideration of interpage relation. As a result, the wear out of the memory cells can be minimized and their lifetime is improved by 62.7% in a file type and 43.0% in MySQL. Experimental results indicate that our scheme shows better lifetime improvement than other schemes in most cases.
Wonyoung Lee 0001, Mincheol Kang, Seokin Hong, Soontae Kim
IEEE Trans. Very Large Scale Integr. Syst.3
2018 Attaché: Towards Ideal Memory Compression by Mitigating Metadata Bandwidth Overheads
abstract
Memory systems are becoming bandwidth constrained and data compression is seen as a simple technique to increase their effective bandwidth. However, data compressionrequires accessing Metadata which incurs additional bandwidth overheads. Even after using a Metadata-Cache, the bandwidth overheads of Metadata can reduce the benefits of compression. This paper proposes Attaché, a framework that reduces the overheads of Metadata accesses. The Attaché framework consists of two components. The first component, called the Blended Metadata Engine (BLEM), enables data and its Metadata to be accessed together. BLEM incurs additional Metadata accesses only 0.003% times and removes almost all Metadata bandwidth overheads. The second component, called theCompression Pre-dictor(COPR), predicts if the memory block is compressed. TheCOPR predictor uses a fine-grained line-level predictor, a coarse-grained page-level predictor, and a global indicator. This enables Attaché to predict the compressibility of the memory block before sending a memory read request. We implement Attaché on a memory system that uses Sub-Ranking. On average, Attaché achieves 15.3% speedup (ideal 17%) and saves 22% energy consumption (ideal 23%) when compared to a baseline system that does not employ data compression. Attaché is completely hardware-based and uses only 368KB of SRAM.
Seokin Hong, Prashant J. Nair, Bülent Abali, Alper Buyuktosunoglu, Kyu-Hyoun Kim, Michael B. Healy
MICRO1
2017 Partial Row Activation for Low-Power DRAM System
abstract
Owing to increasing demand of faster and larger DRAM system, the DRAM system accounts for a large portion of the total power consumption of computing systems. As memory traffic and DRAM bandwidth grow, the row activation and I/O power consumptions are becoming major contributors to total DRAM power consumption. Thus, reducing row activation and I/O power consumptions has big potential for improving the power and energy efficiency of the computing systems. To this end, we propose a partial row activation scheme for memory writes, in which DRAM is rearchitected to mitigate row overfetching problem of modern DRAMs and to reduce row activation power consumption. In addition, accompanying I/O power consumption in memory writes is also reduced by transferring only a part of cache line data that must be written to partially opened rows. In our proposed scheme, partial rows ranging from a one-eighth row to a full row can be activated to minimize row activation granularity for memory writes and the full bandwidth of the conventional DRAM can be maintained for memory reads. Our partial row activation scheme is shown to reduce total DRAM power consumption by up to 32% and 23% on average, which outperforms previously proposed schemes in DRAM power saving with almost no performance loss.
Yebin Lee, Hyeonggyu Kim, Seokin Hong, Soontae Kim
HPCA3
2016 Designing a Resilient L1 Cache Architecture to Process Variation-Induced Access-Time Failures
abstract
Continuous scaling of process technology increases variations in transistors. The process variations cause large fluctuations in the access times of static random-access memory (SRAM) cells. Caches made of those SRAM cells cannot be accessed within the target clock cycle time, which reduces the yield of processors. Many schemes have been proposed to combat these access time failures in caches. However, these schemes are limited in their coverage and do not scale well at high failure rates. We propose a new level one (L1) cache architecture employing multi-cycle cell access (MCCA) and subarray-level parallel access (SLPA). MCCA eliminates all access-time failures in L1 caches. SLPA minimizes the performance impact of cache bandwidth loss due to MCCA. For further performance improvement, architectural techniques are proposed. Our experimental results show that our proposed L1 cache architecture incurs a performance hit of less than 1.2 percent compared to the conventional cache architecture with no access time failure. Our proposed architecture is not sensitive to access time failure rates and has a low overhead compared with previously proposed competitive schemes.
Seokin Hong, Soontae Kim
IEEE Trans. Computers1
2015 A Low-Cost Mechanism Exploiting Narrow-Width Values for Tolerating Hard Faults in ALU
abstract
Digital circuits are expected to increasingly suffer from more hard faults due to technology scaling. Especially, a single hard fault in ALU (Arithmetic Logic Unit) might lead to a total failure in processors or significantly reduce their performance. To address these increasingly important problems, we propose a novel cost-efficient fault-tolerant mechanism for the ALU, called LIZARD. LIZARD employs two half-word ALUs, instead of a single full-word ALU, to perform computations with concurrent fault detection. When a fault is detected, the two ALUs are partitioned into four quarter-word ALUs. After diagnosing and isolating a faulty quarter-word ALU, LIZARD continues its operation using the remaining ones, which can detect and isolate another fault. Even though LIZARD uses narrow ALUs for computations, it adds negligible performance overhead through exploiting predictability of the results in the arithmetic computations. We also present the architectural modifications when employing LIZARD for scalar as well as superscalar processors. Through comparative evaluation, we demonstrate that LIZARD outperforms other competitive fault-tolerant mechanisms in terms of area, energy consumption, performance and reliability.
Seokin Hong, Soontae Kim
IEEE Trans. Computers1
2015 Ensuring Cache Reliability and Energy Scaling at Near-Threshold Voltage With Macho
abstract
Nanoscale process variations in conventional SRAM cells are known to limit voltage scaling in microprocessor caches. Recently, a number of novel cache architectures have been proposed which substitute faulty words of one cache line with healthy words of others, to tolerate these failures at low voltages. These schemes rely on the fault maps to identify faulty words, inevitably increasing the chip area. Besides, the relationship between word sizes and the cache failure rates is not well studied in these works. In this paper, we analyze the word substitution schemes by employing Fault Tree Model and Collision Graph Model. A novel cache architecture (Macho) is then proposed based on this model. Macho is dynamically reconfigurable and is locally optimized (tailored to local fault density) using two algorithms: 1) a graph coloring algorithm for moderate fault densities and 2) a bipartite matching algorithm to support high fault densities. An adaptive matching algorithm enables on-demand reconfiguration of Macho to concentrate available resources on cache working sets. As a result, voltage scaling down to 400 mV is possible, tolerating bit failure rates reaching 1 percent (one failure in every 100 cells). This near-threshold voltage (NTV) operation achieves 44 percent energy reduction in our simulated system (CPU+DRAM models) with a 1 MB L2 cache.
Tayyeb Mahmood, Seokin Hong, Soontae Kim
IEEE Trans. Computers2
2014 Ternary cache: Three-valued MLC STT-RAM caches
abstract
Spin-transfer torque random access memory (STT-RAM) has become a promising non-volatile memory technology for cache memories. Recently, 2-bit multi-level cell (MLC) STT-RAM has been proposed to enhance data density, but it suffers from low reliability of its read and write operations. In this paper, we propose a novel cache design called Ternary cache. In Ternary cache, a memory cell can store three values (i.e., 0,1,2) while MLC STT-RAM can store four values. In this way, Ternary cache achieves much higher read stability than MLC STT-RAM-based caches. To enhance writability, a write operation is performed with high current and terminated as soon as the data is written. Evaluation results show that Ternary cache achieves the data density benefit of MLC STT-RAM and the reliability benefit of SLC STT-RAM.
Seokin Hong, Jongmin Lee 0002, Soontae Kim
ICCD1
2013 AVICA: an access-time variation insensitive L1 cache architecture
abstract
Ever scaling process technology increases variations in transistors. The process variations cause large fluctuations in the access times of SRAM cells. Caches made of those SRAM cells cannot be accessed within the target clock cycle time, which reduces yield of processors. To combat these access time failures in caches, many schemes have been proposed, which are, however, limited in their coverage and do not scale well at high failure rates. We propose a new L1 cache architecture (AVICA) employing asymmetric pipelining and pseudo multi-banking. Asymmetric pipelining eliminates all access time failures in L1 caches. Pseudo multi-banking minimizes the performance impact of asymmetric pipelining. For further performance improvement, architectural techniques are proposed. Our experimental results show that our proposed L1 cache architecture incurs less than 1% performance hit compared to the conventional cache architecture with no access time failure. Our proposed architecture is not sensitive to access time failure rates and has low overheads compared to the previously proposed competitive schemes.
Seokin Hong, Soontae Kim
DATE1
2013 Skinflint DRAM system: Minimizing DRAM chip writes for low power
abstract
DRAMs are one of the main players of computer system energy consumption due to their large capacities and frequent accesses. Consequently, many schemes have been proposed to reduce DRAM power/energy consumption. Some of them propose new DRAM system and chip organizations, which are effective in reducing power consumption but intrusive. In contrast, we minimize DRAM write accesses at chip level with minimal modification of the conventional DRAM system organization and small addition to caches. When all data going to the same DRAM chips are not modified, the chips are not accessed. Consequently, chips are accessed selectively in our scheme while all chips are accessed simultaneously in the conventional DRAM system. Our chip-based selective DRAM write scheme is shown to reduce DRAM power and energy consumptions by 17% and 14%, respectively, on average. The overheads of our scheme are small in terms of performance, area, and energy consumption.
Yebin Lee, Soontae Kim, Seokin Hong, Jongmin Lee 0002
HPCA3
2013 Macho: A failure model-oriented adaptive cache architecture to enable near-threshold voltage scaling
abstract
Recent interest in CMOS voltage scaling has produced a class of cache architectures which tolerate parametric SRAM failures at low voltage by substituting faulty words of one cache line with healthy words of another line. These caches rely on the fault maps (which grow reciprocally with smaller word sizes) for fault identification. Therefore, the benefits of cache voltage scaling must be rigorously investigated against the cost of their fault map overheads, especially in large caches. This paper reviews the word substitution caches and develops their parametric failure model. Our developed model leads to a non-intrusive and reconfigurable cache (Macho) which can be locally optimized (based on local fault density) by two graph-based algorithms. Specifically, our adaptive matching algorithm increases effective cache capacity by dynamically concentrating healthy cache blocks into active cache sets. Macho enables voltage scaling down to 400mV by tolerating high SRAM-failure rates (≥ 1%) and achieves better energy reduction (44%) than other substitution caches with similar area overheads.
Tayyeb Mahmood, Soontae Kim, Seokin Hong
HPCA3
2011 TLB index-based tagging for cache energy reduction
Jongmin Lee 0002, Seokin Hong, Soontae Kim
ISLPED2
2011 Residue cache: a low-energy low-area L2 cache architecture via compression and partial hits
abstract
L2 cache memories are being adopted in the embedded systems for high performance, which, however, increases energy consumption due to their large sizes. We propose a low-energy low-area L2 cache architecture, which performs as well as the conventional L2 cache architecture with 53% less area and around 40% less energy consumption. This architecture consists of an L2 cache and a small cache called residue cache. L2 and residue cache lines are half sized of the conventional L2 cache lines. Well compressed conventional L2 cache lines are stored only in the L2 cache while other poorly compressed lines are stored in both the L2 and residue caches. Although many conventional L2 cache lines are not fully captured by the residue cache, most accesses to them do not incur misses because not all their words are needed immediately, which are termed as partial hits in this paper. The residue cache architecture consumes much lower energy and area than conventional L2 cache architectures, and can be combined synergistically with other schemes such as the line distillation and ZCA. The residue cache architecture is also shown to perform well on a 4-way superscalar processor typically used in high performance systems.
Soontae Kim, Jongmin Lee 0002, Jesung Kim, Seokin Hong
MICRO4
2010 Lizard: Energy-efficient hard fault detection, diagnosis and isolation in the ALU
abstract
Digital circuits are expected to increasingly suffer from more hard faults due to technology scaling. Especially, a single hard fault in the ALU might lead to a total failure in the embedded systems. In addition, energy efficiency is critical in these systems. To address these increasingly important problems in the ALU, we propose a novel energy-efficient fault-tolerant ALU design called Lizard. Lizard utilizes two 16-bit ALUs to perform 32-bit computations with fault detection and diagnosis. By exploiting predictable operations, fault detection is performed in a single cycle. The 16-bit ALUs can be partitioned into two 8-bit ALUs. When a fault occurs in one of the four 8-bit ALUs, Lizard diagnoses and isolates a faulty 8-bit ALU for itself. After the faulty 8-bit ALU is isolated, Lizard continues its operation using the remaining three 8-bit ALUs, which can detect and isolate another fault. In this way, Lizard can survive faults on at most two sub-ALUs increasing its lifetime and fault tolerance. We conducted comparative evaluations with an unprotected ALU, triple modular redundancy ALU, and quadruple time redundancy ALU in terms of area, energy consumption, performance, and reliability. It is demonstrated that Lizard outperforms other ALU designs in most cases, especially in energy efficiency.
Seokin Hong, Soontae Kim
ICCD1