Yesin Ryu

dblp:72/5513 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
5since 2021 · last 2026
0000-0002-2678-7396ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 3 first-author · 5 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 RowArmor: Efficient and Comprehensive Protection Against DRAM Disturbance Attacks
Minbok Wi, Yoonyul Yoo, Yoojin Kim, Jumin Kim, Yesin Ryu, Saeid Gorgin 0001, Jung Ho Ahn, Jungrae Kim
ASPLOS (2)6
2026 Cerberus: Cross-Layer ECC Co-Design for Robust and Efficient Memory Protection
Junhwan Kim, Yesin Ryu, Saeid Gorgin 0001, Jungrae Kim
ISCA3
2026 DieCARE: Diekill-Correct ECC for HBM Reliability Without Additional Dies
abstract
As High Bandwidth Memory (HBM) continues scaling to address the demands of data-intensive workloads and AI-driven applications, ensuring resilience against increasingly frequent memory faults has become critical. DieCARE introduces a novel memory architecture co-designed with an innovative Error Correcting Code (ECC) scheme to enable die-level fault tolerance without requiring additional dies. It strategically distributes data and ECC check bits across multiple dies, leveraging advanced ECC techniques with flexible symbol layouts to optimize error correction capability, latency, and area efficiency.System-level evaluations demonstrate that DieCARE reduces memory Failure In Time (FIT) rate by 12, 000×, while maintaining an extremely low Silent Data Corruption (SDC) rate. These reliability improvements translate into increased system availability, yielding substantial benefits for large-scale computing systems.
Yesin Ryu, Byungwoo Bang, Hunseong Choi, Yoojin Kim, Hanum Ko, Jungrae Kim
IEEE Trans. Computers1
2025 PIMPAL: Accelerating LLM Inference on Edge Devices via In-DRAM Arithmetic Lookup
abstract
Deploying Large Language Models (LLMs) on edge devices poses significant challenges due to their high computational and memory demands. In particular, General MatrixVector Multiplication (GEMV), a key operation in LLM inference, is highly memory-intensive, making it difficult to accelerate using conventional edge computing systems. While Processing-in-memory (PIM) architectures have emerged as a promising solution to this challenge, they often suffer from high area overhead or restricted computational precision. This paper proposes PIMPAL (Processing-In-Memory architecture with Parallel Arithmetic Lookup), a cost-effective PIM architecture leveraging LookUp Table (LUT)-based computation for GEMV acceleration in sLLMs (small LLMs). By replacing traditional arithmetic operations with parallel in-DRAM LUT lookups, PIMPAL significantly reduces area overhead while maintaining high performance. PIMPAL introduces three key innovations: (1) it divides DRAM bank subarrays into compute blocks for parallel LUT processing; (2) it employs Localityaware Compute Mapping (LCM) to reduce row activations by maximizing LUT access locality; and (3) it enables multi-precision computations through a LUT Aggregation (LAG) mechanism that combines results from multiple small LUTs. Experimental results show that PIMPAL achieves up to $17.8 x$ higher performance than previous LUT-based PIM designs and reduces area overhead by $40 \%$ compared to conventional processing unit-based PIM designs.
Yoonho Jang, Hyeongjun Cho, Yesin Ryu, Jungrae Kim, Seokin Hong
DAC3
2024 Native DRAM Cache: Re-architecting DRAM as a Large-Scale Cache for Data Centers
abstract
Contemporary data center CPUs are experiencing an unprecedented surge in core count. This trend necessitates scrutinized Last-Level Cache (LLC) strategies to accommodate increasing capacity demands. While DRAM offers significant capacity, using it as a cache poses challenges related to latency and energy. This paper introduces Native DRAM Cache (NDC), a novel DRAM architecture specifically designed to operate as a cache. NDC features innovative approaches, such as conducting tag matching and way selection within a DRAM subarray and repurposing existing precharge transistors for tag matching. These innovations facilitate Caching-In-Memory (CIM) and enable NDC to serve as a high-capacity LLC with high set-associativity, low-latency, high-throughput, and low-energy. Our evaluation demonstrates that NDC significantly outperforms state-of-the-art DRAM cache solutions, enhancing performance by $\mathbf{2.8 \%} / \mathbf{52.5 \%} / \mathbf{44.2 \%}$ (up to $8.4 \% / 140.6 \% / 85.5 \%$) in SPEC/NPB/GAP benchmark suites, respectively.
Yesin Ryu, Yoojin Kim, Giyong Jung, Jung Ho Ahn, Jungrae Kim
ISCA1
2008 Clock buffer polarity assignment combined with clock tree generation for power/ground noise minimization
abstract
A new approach to the problem of clock buffer polarity assignment for minimizing power/ground noise on the clock network is presented. The previous approaches solve the assignment problem in two separate steps: (step 1) generating a clock routing tree of minimum total wirelength, satisfying the clock skew constraint and then (step 2) inserting buffering elements with their polarities under the objective of minimizing power/ground noise while satisfying the clock skew constraint. Yet, there is no easy way to predict the result of step 2 during step 1. In our approach, we place the primary importance on the cost of power/ground noise. Consequently, we try to minimize the cost of power/ground noise first and then to construct a clock routing tree later while satisfying the clock skew constraint. Through experimentation using several benchmark circuits, it is shown that this approach is quite effective and produces very good solutions, reducing the power/ground noise by 75% and the peak current by 26% at the expense of 5% wirelength overhead compared to that produced by the conventional approach.
Yesin Ryu
ICCAD1