EDBT 2026 Demo / reviewers in the wild / expert
Daehyeon Baek
dblp:263/7517
· DBLP profile ↗
5ranked-venue papers
2as first author
4since 2021 · last 2025
0009-0008-0460-3809ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | PAISE: PIM-Accelerated Inference Scheduling Engine for Transformer-based LLMabstractTransformer-based Large Language Models (LLMs) demand significant computational and memory resources due to the autoregressive token generation in decoder blocks. In particular, the attention layer in LLM models has low arithmetic intensity but high memory traffic, thus requiring frequent updates to the KV matrices with each decoder iteration. As a result, LLM inference becomes memory bound, leading to increased latency. To address this, we introduce PAISE, a framework leveraging Processing-In-Memory (PIM) technology to offload memory-intensive tasks. PAISE employs GPU-PIM heterogeneous computing resources to optimize inference operations in transformer-based LLMs. The framework comprises (i) a scheduling algorithm that decides which operations to offload to PIM based on model configuration and PIM hardware specifications and (ii) an enhanced PIM kernel that performs transaction-wise interleave-batched GEMM (General Matrix Multiplication) operations, maximizing data throughput via data layout adjustments. We implemented PAISE on the GPT-2 and Llama2-7B models using an AMD MI100 GPU with HBM-PIM devices. Our evaluations show that offloading the attention layer to PIM reduces execution time by up to 48.3% compared to GPU-only inference, demonstrating PAISE’s significant potential to enhance the efficiency of LLM inference, which could lead to faster and more efficient AI applications. Hyojung Lee, Daehyeon Baek, Jimyoung Son, Jieun Choi, Kihyo Moon, Minsung Jang |
HPCA | 2 |
| 2024 | pSyncPIM: Partially Synchronous Execution of Sparse Matrix Operations for All-Bank PIM ArchitecturesabstractRecent commercial incarnations of processing-in-memory (PIM) maintain the standard DRAM interface and employ the all-bank mode execution to maximize bank-level memory bandwidth. Such a synchronized all-bank PIM control can effectively manage conventional dense matrix-vector operations on evenly distributed matrices across banks with lock-step execution. Sparse matrix processing is another critical computation that can significantly benefit from the PIM architecture, but the current all-bank PIM control cannot support diverging executions due to the random sparsity. To accelerate such sparse matrix applications, this paper proposes a partially synchronous execution on sparse matrix-vector multiplication (SpMV) and sparse triangular matrix-vector solve (SpTRSV), filling the gap between the practical constraint of PIM and the irregular nature of sparse computation. It allows the execution of the processing unit of each bank to diverge in a limited way to manage the irregular execution path of sparse matrix computation. It proposes compaction and distribution policies for the input matrix and vector. In addition to SpMV, this paper identifies SpTRSV is another key kernel, and proposes SpTRSV acceleration on PIM technology. The experimental evaluation shows that the new sparse PIM architecture outperforms NVIDIA Geforce RTX 3080 GPU by $4.43 \times$ speedup for SpMV and $3.53 \times$ speedup for SpTRSV with a similar amount of DRAM bandwidth. Daehyeon Baek, Soojin Hwang, Jaehyuk Huh 0001 |
ISCA | 1 |
| 2024 | Cerberus: Triple Mode Acceleration of Sparse Matrix and Vector MultiplicationabstractThe multiplication of sparse matrix and vector (SpMV) is one of the most widely used kernels in high-performance computing as well as machine learning acceleration for sparse neural networks. The design space of SpMV accelerators has two axes: algorithm and matrix representation. There have been two widely used algorithms and data representations. Two algorithms, scalar multiplication and dot product, can be combined with two sparse data representations, compressed sparse and bitmap formats for the matrix and vector. Although the prior accelerators adopted one of the possible designs, it is yet to be investigated which design is the best one across different hardware resources and workload characteristics. This paper first investigates the impact of design choices with respect to the algorithm and data representation. Our evaluation shows that no single design always outperforms the others across different workloads, but the two best designs (i.e., compressed sparse format and bitmap format with dot product) have complementary performance with trade-offs incurred by the matrix characteristics. Based on the analysis, this study proposes Cerberus, a triple-mode accelerator supporting two sparse operation modes in addition to the base dense mode. To allow such multi-mode operation, it proposes a prediction model based on matrix characteristics under a given hardware configuration, which statically selects the best mode for a given sparse matrix with its dimension and density information. Our experimental results show that Cerberus provides 12.1× performance improvements from a dense-only accelerator, and 1.5× improvements from a fixed best SpMV design. Soojin Hwang, Daehyeon Baek, Jongse Park, Jaehyuk Huh 0001 |
ACM Trans. Archit. Code Optim. | 2 |
| 2021 | InnerSP: A Memory Efficient Sparse Matrix Multiplication Accelerator with Locality-Aware Inner Product ProcessingabstractSparse matrix multiplication is one of the key computational kernels in large-scale data analytics. However, a naive implementation suffers from the overheads of irregular memory accesses due to the representation of sparsity. To mitigate the memory access overheads, recent accelerator designs advocated the outer product processing which minimizes input accesses but generates intermediate products to be merged to the final output matrix. Using real-world sparse matrices, this study first identifies the memory bloating problem of the outer product designs due to the unpredictable intermediate products. Such an unpredictable increase in memory requirement during computation can limit the applicability of accelerators. To address the memory bloating problem, this study revisits an alternative inner product approach, and proposes a new accelerator design called InnerSP. This study shows that nonzero element distributions in real-world sparse matrices have a certain level of locality. Using a smart caching scheme designed for inner product, the locality is effectively exploited with a modest on-chip cache. However, the row-wise inner product relies on on-chip aggregation of intermediate products. Due to uneven sparsity per row, overflows or underflows of the on-chip storage for aggregation can occur. To maximize the parallelism while avoiding costly overflows, the proposed accelerator uses pre-scanning for row splitting and merging. The simulation results show that the performance of InnerSP can exceed or be similar to those of the prior outer product approaches without any memory bloating problem. Daehyeon Baek, Soojin Hwang, Taekyung Heo, Daehoon Kim 0001, Jaehyuk Huh 0001 |
PACT | 1 |
| 2020 | Charge-Aware DRAM Refresh Reduction with Value TransformationabstractAs the memory capacity in a system has been growing, refresh operations consume increasing ratios of the total DRAM power. To reduce the power consumption of such refresh operations, this paper proposes a novel value-aware refresh reduction technique called ZERO - REFRESH which exploits zero values in memory contents. A DRAM cell can retain the discharged state without refresh operations, and ZERO - REFRESH skips refresh operations on rows with all discharged cells. For abundant unallocated memory pages in typical systems, the operating system fills them with zeros to clean the contents. For those idle pages, ZERO - REFRESH can eliminate refresh operations in an OS-transparent way without any new interface to DRAM. However, for allocated memory pages, memory contents may not have many consecutive zero values to match the refresh granularity of DRAM. To increase the frequency of zero values and to arrange them to match the refresh granularity, ZERO - REFRESH transforms the value of memory blocks to the base and delta values, inspired by the prior BDI (Base-Delta-Immediate) compression technique. Once values are converted, bits are transposed to be stored as consecutive discharged bits at the refresh granularity. Such value transformation and rearrangement can make the memory contents friendly to refresh reduction based on discharged cells. The experimental results based on simulation show that the DRAM refresh operations are reduced by 37% on average for a set of benchmark applications, if the entire memory is allocated for the applications. If the memory usage statistics collected from three data center traces are applied, the DRAM refresh operations can be reduced by 46%, 57%, and 83% respectively for the three scenarios. Seikwon Kim, Wonsang Kwak, Changdae Kim 0001, Daehyeon Baek, Jaehyuk Huh 0001 |
HPCA | 4 |