Taeyang Jeong

dblp:219/8048 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2026
0000-0001-9110-6502ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 DH-PIM: Maximizing Computing Unit Utilization in Digital PIM by Dual Half Mode Extension
abstract
Transformer-based Large Language Models (LLMs) rely on both General Matrix-Matrix Multiplication (GEMM) and General Matrix-Vector Multiplication (GEMV) for inference. While existing Processing-in-Memory (PIM) architectures, like HBM-PIM and Newton, efficiently handle GEMV operations, they suffer from low computing unit utilization during GEMM due to frequent input preparation and activation cycles. To address this, we propose Dual Half-mode PIM (DH-PIM), an extension of the Half Division Mode from previous PipePIM. DH-PIM splits sense amplifiers, and computing units into halves, allowing them to execute different operations independently. Dual half mode, which allows both halves of the computing unit to process data broadcast from either half of the sense amplifier, significantly enhances utilization via overlapped and pipelined execution. Our simulations show DH-PIM improves GEMM performance by up to 2.06x in GEMM kernels, with about 1.46x speedup in LLM tasks, compared to HBM-PIM. These performance improvements are achieved with minimal architectural modifications, resulting in low area and energy overhead.
Byoung Jin Kim, Taeyang Jeong, SeongJun Yun, Eui-Young Chung
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2024 PipePIM: Maximizing Computing Unit Utilization in ML-Oriented Digital PIM by Pipelining and Dual Buffering
abstract
A digital processing-in-memory (PIM) that integrates computing units (CUs) with DRAM banks emerges as a promising technique for accelerating matrix–vector multiplication (MV). However, activating and precharging all banks incur significant overheads in a digital PIM based on conventional DRAM, which is limited to activating only a single subarray in a bank. Moreover, a digital PIM utilizes a vector buffer to store and reuse the input vector. This necessitates repeated buffer writes, incurring substantial overhead for large MV. Consequently, these overheads reduce CU utilization in a digital PIM, degrading the performance. To overcome these issues, we propose PipePIM, which maximizes CU utilization in a digital PIM by pipelining and dual buffering. PipePIM consists of two primary schemes: 1) subarray-level pipelining (SAPI) and 2) dual vector buffer. They exploit and extend the features of a multitude of activated subarrays (MASA) introduced by subarray-level parallelism (SALP). SAPI enables a digital PIM to perform activation, precharging, and computation on different subarrays in a pipelined manner. Through SAPI, these operations are overlapped, and activation and precharging overheads are hidden. A dual vector buffer employs two vector buffers and manages them as ping-pong buffering, one for computation and another for buffer write simultaneously. To facilitate it, PipePIM proposes a half-division mode (HDM) enabling independent access to two activated subarrays with marginal area increase. We demonstrate the improvements by PipePIM on the state-of-the-art digital PIMs, Newton and HBM-PIM. Our simulation results indicate that the average speedups of Newton and HBM-PIM on MV are$2.16\times $and$1.74\times $, respectively.
Taeyang Jeong, Eui-Young Chung
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2023 Operand-Oriented Virtual Memory Support for Near-Memory Processing
abstract
Virtual memory support is one of the major challenges of near-memory processing (NMP). Many previous works focused on this issue, but there are practical limitations that conventional CPU hardware or memory allocation schemes should be modified. Another technique uses a specialized page table for NMP to avoid such limitations. However, the previous work proposed NMP-specific page table that has static page table walk latency regardless of data size. This causes unnecessarily long address translation time for relatively small data. In this paper, we propose an operand-oriented technique for virtual memory support. Our scheme does not pre-determine the size of shared space; rather, it allocates shared space depending on the size of operands data for NMP. Then, we significantly reduce page table walk latency by using our flexible page table, which adapts the page table hierarchy to the size of shared spaces. To prove our concept, we implement our scheme in a full-system simulator and an FPGA-based verification platform. We then compared it with CPU's page table and the previous NMP-specific page table. The experimental results show that our technique outperforms page table walk latency by 69.3 percent and 43.8 percent compared to the CPU's page table and the comparison, respectively.
Duheon Choi, Taeyang Jeong, Joonhyeok Yeom, Eui-Young Chung
IEEE Trans. Computers2