Yoojin Kim

dblp:220/8239 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
4since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 4 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Artificial intelligence and machine learning · 1
YearPublicationVenuePosition
2026 RowArmor: Efficient and Comprehensive Protection Against DRAM Disturbance Attacks
Minbok Wi, Yoonyul Yoo, Yoojin Kim, Jumin Kim, Yesin Ryu, Saeid Gorgin 0001, Jung Ho Ahn, Jungrae Kim
ASPLOS (2)3
2026 DieCARE: Diekill-Correct ECC for HBM Reliability Without Additional Dies
abstract
As High Bandwidth Memory (HBM) continues scaling to address the demands of data-intensive workloads and AI-driven applications, ensuring resilience against increasingly frequent memory faults has become critical. DieCARE introduces a novel memory architecture co-designed with an innovative Error Correcting Code (ECC) scheme to enable die-level fault tolerance without requiring additional dies. It strategically distributes data and ECC check bits across multiple dies, leveraging advanced ECC techniques with flexible symbol layouts to optimize error correction capability, latency, and area efficiency.System-level evaluations demonstrate that DieCARE reduces memory Failure In Time (FIT) rate by 12, 000×, while maintaining an extremely low Silent Data Corruption (SDC) rate. These reliability improvements translate into increased system availability, yielding substantial benefits for large-scale computing systems.
Yesin Ryu, Byungwoo Bang, Hunseong Choi, Yoojin Kim, Hanum Ko, Jungrae Kim
IEEE Trans. Computers5
2024 Native DRAM Cache: Re-architecting DRAM as a Large-Scale Cache for Data Centers
abstract
Contemporary data center CPUs are experiencing an unprecedented surge in core count. This trend necessitates scrutinized Last-Level Cache (LLC) strategies to accommodate increasing capacity demands. While DRAM offers significant capacity, using it as a cache poses challenges related to latency and energy. This paper introduces Native DRAM Cache (NDC), a novel DRAM architecture specifically designed to operate as a cache. NDC features innovative approaches, such as conducting tag matching and way selection within a DRAM subarray and repurposing existing precharge transistors for tag matching. These innovations facilitate Caching-In-Memory (CIM) and enable NDC to serve as a high-capacity LLC with high set-associativity, low-latency, high-throughput, and low-energy. Our evaluation demonstrates that NDC significantly outperforms state-of-the-art DRAM cache solutions, enhancing performance by $\mathbf{2.8 \%} / \mathbf{52.5 \%} / \mathbf{44.2 \%}$ (up to $8.4 \% / 140.6 \% / 85.5 \%$) in SPEC/NPB/GAP benchmark suites, respectively.
Yesin Ryu, Yoojin Kim, Giyong Jung, Jung Ho Ahn, Jungrae Kim
ISCA2
2021 Sparsity-Aware and Re-configurable NPU Architecture for Samsung Flagship Mobile SoC
abstract
Of late, deep neural networks have become ubiquitous in mobile applications. As mobile devices generally require immediate response while maintaining user privacy, the demand for on-device machine learning technology is on the increase. Nevertheless, mobile devices suffer from restricted hardware resources, whereas deep neural networks involve considerable computation and communication. Therefore, the implementation of a neural-network specialized hardware accelerator, generally called neural processing unit (NPU), has started to gain attention for the mobile application processor (AP). However, NPUs for commercial mobile AP face two challenges that are difficult to realize simultaneously: execution of a wide range of applications and efficient performance.In this paper, we propose a flexible but efficient NPU architecture for a Samsung flagship mobile system-on-chip (SoC). To implement an efficient NPU, we design an energy-efficient inner-product engine that utilizes the input feature map sparsity. We propose a re-configurable MAC array to enhance the flexibility of the proposed NPU, dynamic internal memory port assignment to maximize on-chip memory bandwidth utilization, and efficient architecture to support mixed-precision arithmetic. We implement the proposed NPU using the Samsung 5nm library. Our silicon measurement experiments demonstrate that the proposed NPU achieves 290.7 FPS and 13.6 TOPS/W, when executing an 8-bit quantized Inception-v3 model [1] with a single NPU core. In addition, we analyze the proposed zero-skipping architecture in detail. Finally, we present the findings and lessons learned when implementing the commercial mobile NPU and interesting avenues for future work.
Jun-Woo Jang, Sehwan Lee, Dongyoung Kim, Hyunsun Park, Ali Shafiee, Yeongjae Choi, Channoh Kim, Yoojin Kim, Hyeongseok Yu, Hamzah Abdel-Aziz, Jun-Seok Park, Heonsoo Lee, Myeong Woo Kim, Hanwoong Jung, Heewoo Nam, Dongguen Lim, Seungwon Lee 0006, Joon-Ho Song, Suknam Kwon, Joseph Hassoun, Sukhwan Lim, Changkyu Choi
ISCA8
2019 Graph Neural Network for Music Score Data and Modeling Expressive Piano Performance
abstract
Music score is often handled as one-dimensional sequential data. Unlike words in a text document, notes in music score can be played simultaneously by the polyphonic nature and each of them has its own duration. In this paper, we represent the unique form of musical score using graph neural network and apply it for rendering expressive piano performance from the music score. Specifically, we design the model using note-level gated graph neural network and measure-level hierarchical attention network with bidirectional long short-term memory with an iterative feedback method. In addition, to model different styles of performance for a given input score, we employ a variational auto-encoder. The result of the listening test shows that our proposed model generated more human-like performances compared to a baseline model and a hierarchical attention network model that handles music score as a word-like sequence.
Dasaem Jeong, Taegyun Kwon, Yoojin Kim, Juhan Nam
ICML3